# Message Testing Scorecard

Compare candidate messages for unaided comprehension, relevance, credibility, differentiation, limitation awareness, recall, and ability to support a suitable action.

## Usage note

Use the scorecard to compare messages under the same research conditions, not to manufacture a winning average. Record participant explanations and misunderstandings beside ratings. Test only claims the organization can support, protect participant data, and do not present small qualitative samples as population estimates.

## How to use this template

1. Define the message decision, audience situation, candidate set, success criteria, and claims that must remain accurate.
2. Standardize presentation, choose an order or randomization plan, and prepare neutral comprehension prompts.
3. Test each message for unaided meaning before asking about preference, appeal, or ratings.
4. Score dimensions with written evidence, recording misunderstandings, limitation awareness, and requested proof.
5. Compare tradeoffs, select revisions rather than an automatic winner, and document what requires retesting.

## Blank template

### Test setup

- **Decision to inform:** [Where the message will be used]
- **Audience and situation:** [Who is deciding what?]
- **Candidate labels:** [A, B, C, without evaluative names]
- **Presentation context:** [Text, page section, ad, email, or another]
- **Order method:** [Randomized, rotated, or fixed with reason]
- **Recruitment criteria:** [Relevant experience]
- **Participant code:** [Non-identifying label]
- **Consent and session date:** [Method and YYYY-MM-DD]

### Candidate observation

- **Candidate:** [A / B / C]
- **Unaided meaning:** [Participant’s explanation]
- **Expected offer or action:** [What do they think happens?]
- **Intended audience:** [Who do they think it serves?]
- **Important words noticed:** [Participant language]
- **Misunderstanding:** [Invented feature, missed condition, or ambiguity]
- **Evidence requested:** [What would make it credible?]
- **Likely next step:** [What would they do?]

### Dimension scorecard

- **Comprehension:** [1–5 and evidence]
- **Relevance:** [1–5 and connection to situation]
- **Credibility:** [1–5 and reason]
- **Differentiation:** [1–5 and compared alternative]
- **Limitation awareness:** [1–5 and condition recalled]
- **Actionability:** [1–5 and expected next step]
- **Recall after delay:** [What remained, if tested]
- [ ] Scores include participant reasoning.
- [ ] The researcher did not correct interpretation before recording it.
- [ ] One severe misunderstanding is not hidden by an average.

### Decision synthesis

- **Strongest candidate by dimension:** [Do not force one overall winner]
- **Material tradeoff:** [What improves while something else weakens?]
- **Unsupported expectation:** [Claim or implication to remove]
- **Revision:** [Specific wording, proof, order, or offer change]
- **Product or policy issue:** [Problem copy cannot solve]
- **Decision owner:** [Name or role]
- **Retest requirement:** [What and with whom]
- **Research limitation:** [What this round cannot establish]
