Brand Voice and Editorial Style
Brand Voice Calibration Test
Compare reviewer judgments on representative passages and resolve inconsistent interpretations of voice principles, tone boundaries, and acceptable revisions.
Free editable Markdown · Brand editors, copywriting teams, and quality leads ·
Accessible HTML preview
Blank template
The downloaded file contains the same fields in editable Markdown.
Test setup
- Calibration goal
- [Which decisions need consistency?]
- Voice principles
- [Links and versions]
- Participants
- [Reviewer roles]
- Decision owner
- [Who resolves boundaries?]
- Sample channels
- [Interface, email, page, support, or another]
- Blind-review method
- [How authors and other scores are hidden]
- Test date
- [YYYY-MM-DD]
Sample score
- Sample ID
- [Neutral label]
- Audience and situation
- [Context]
- Reader task
- [What the passage must do]
- Accuracy status
- [Verified / Provided for test / Needs separate review]
- Voice dimension 1
- [1–4 and textual evidence]
- Voice dimension 2
- [1–4 and textual evidence]
- Tone fit
- [1–4 and reason]
- Decision
- [Accept / Revise / Reject]
- Proposed revision
- [Specific change]
- Confidence
- [High / Medium / Low]
Disagreement analysis
- Rating spread
- [Range]
- Different interpretations
- [Summarize fairly]
- Rule ambiguity
- [What guidance failed?]
- Situation ambiguity
- [What context was missing?]
- Risk difference
- [What consequence reviewers weighted differently?]
- Resolved decision
- [Outcome and authority]
- A factual error is not treated as a voice preference.
- Minority reasoning was reviewed before resolution.
- The resolution includes a production-ready example.
Follow-up
- Guide change
- [Rule or anchor revision]
- Example-library entry
- [Link or planned item]
- New test sample
- [Boundary to retest]
- Second-round result
- [Consistency and remaining issue]
- Owner and review date
- [Accountability]
How to use this template
- Define the voice dimensions, production decisions, and rating anchors the exercise will calibrate.
- Prepare varied anonymized samples with enough audience and situation context to judge fairly.
- Collect independent ratings, decisions, textual evidence, and proposed revisions before group discussion.
- Analyze disagreements, resolve consequential boundaries through the authorized owner, and update guidance.
- Repeat with fresh samples and record whether consistency and decision quality improved.
Select samples at the boundaries
Choose real, permitted, or clearly fictional passages representing common channels and difficult situations. Include strong examples, clear failures, and plausible borderline work. Remove confidential data and avoid samples whose answer depends on facts reviewers cannot see. Test one variable at a time when possible: directness, warmth, confidence, technical depth, or humor. A collection of obviously poor passages may produce agreement without revealing how the team interprets actual rules.
Score independently before discussion
Give reviewers the same voice principles, audience context, task, and rating anchors. Ask for a decision and specific textual evidence before anyone hears another view. Separate “accurate and usable” from “voice-aligned”; a passage can sound appropriate while making an unsupported claim. Compare ratings and rationales, not just totals. Disagreement often exposes an ambiguous principle, missing situational rule, cultural assumption, or difference in risk tolerance. Preserve minority reasoning when it identifies a material reader effect.
Turn disagreement into guidance
For every disputed sample, identify the decision owner and resolve what should happen in production: accept, revise, or reject, with an example revision. Update the voice guide, tone matrix, or example library if the existing rule did not support the decision. Do not rewrite scoring anchors merely to match a senior person’s preference. Run a second blind round with new samples to see whether clarification improves consistency. Some variation is healthy; focus on disagreements that create contradictory messages, repeated rework, or reader harm.
See the fields in context
Fictional example: service delay sentence
The service, passage, and reviewer scores are invented.
- Sample: “Good news—your request only needs another three days.”
- Disagreement: Some reviewers scored it warm; others noted that “good news” minimizes an unplanned delay.
- Resolution: Revise to “Your request is delayed by up to three days. You do not need to resubmit it.”
- Guide change: The tone matrix now excludes celebratory framing when the reader bears a delay.
- Retest: A new recovery sample checks whether reviewers apply the boundary consistently.
Frequently asked questions
Should calibration aim for identical scores?
No. Aim for consistent production decisions and shared reasoning on material boundaries. Small rating differences can remain when they do not change what gets published or harm clarity.
Can team members use their own drafts?
Yes, with consent and a psychologically safe process, but anonymization helps discussion focus on text. Fictional boundary samples can reduce defensiveness and isolate one principle more cleanly.
How often should calibration run?
Run it during onboarding, after major guide changes, when repeated review conflicts appear, and at a regular interval suited to publishing volume. Use fresh samples so memory does not masquerade as alignment.
What if the voice guide cannot resolve a sample?
Record the ambiguity, have the authorized owner decide, and update the relevant rule or example. The exercise is working when it reveals that the guidance—not the reviewer—needs improvement.