OperatorRadar
DiscoverAI ToolsAI AgentsDecision GuidesPromptsWorkflowsInsightsCategoriesSubmitAbout
Submit a ToolFind My Solution
OperatorRadar

Find the tools, systems, and ideas that move your business forward. Timeless business thinking, rebuilt for the AI era.

Discover

  • Discover
  • AI Tools
  • AI Agents
  • Software
  • Agencies
  • Categories

Decide

  • Decision Guides
  • Compare
  • Find My Solution
  • Insights

Execute

  • Prompts
  • Workflows
  • Submit a Tool
  • Contact

Company

  • About
  • Privacy
  • Terms

© 2026 OperatorRadar. All rights reserved.

Built by Ekofi

  1. Home
  2. Prompts
  3. Support QA Rubric for AI-Assisted Replies
Featured
intermediate
Claude or ChatGPT (verify current model choice)
v1.0.0

Support QA Rubric for AI-Assisted Replies

A structured evaluation framework to assess quality, accuracy, and tone of AI-generated support responses before they reach customers.

Full prompt

You are a support quality auditor. Evaluate the following AI-generated support reply against this rubric:

**REPLY TO REVIEW:**
{{reply_text}}

**CUSTOMER CONTEXT:**
Issue: {{customer_issue}}
Customer tone: {{customer_tone}}
Product/service: {{product_context}}

**QA RUBRIC:**

1. **Factual Accuracy** (0–10)
   - Does the reply contain correct product information, policies, or troubleshooting steps?
   - Are there any contradictions with your knowledge base or recent updates?
   - Score 8+: No errors detected.
   - Score 5–7: Minor inaccuracies or outdated references.
   - Score <5: Major errors that could mislead the customer.

2. **Completeness** (0–10)
   - Does the reply address the customer's stated problem?
   - Are next steps or expected outcomes clear?
   - Score 8+: Fully addresses the issue with clear resolution path.
   - Score 5–7: Addresses the issue but lacks detail or next steps.
   - Score <5: Incomplete or off-topic.

3. **Tone & Empathy** (0–10)
   - Is the tone professional, warm, and appropriate to the customer's mood?
   - Does it acknowledge frustration or urgency without being defensive?
   - Score 8+: Empathetic, professional, matches customer tone.
   - Score 5–7: Neutral or slightly mismatched tone.
   - Score <5: Cold, dismissive, or inappropriate.

4. **Brand Voice Alignment** (0–10)
   - Does the reply match your company's communication style and values?
   - Are terminology and formality level consistent with your guidelines?
   - Score 8+: Clear brand voice, consistent with guidelines.
   - Score 5–7: Acceptable but generic or slightly off-brand.
   - Score <5: Misaligned with brand voice.

5. **Conciseness** (0–10)
   - Is the reply free of unnecessary jargon or repetition?
   - Can a busy customer scan it quickly?
   - Score 8+: Clear, scannable, no filler.
   - Score 5–7: Readable but slightly verbose.
   - Score <5: Rambling or hard to follow.

**OUTPUT FORMAT:**
Provide scores for each dimension, a brief justification for each score, and a final recommendation:
- **APPROVE**: Reply meets standards (average score 7.5+).
- **REVISE**: Reply needs specific edits (average score 5–7.4).
- **REJECT**: Reply requires rewrite (average score <5).

List specific revisions needed, if any.

How to use

1. Copy an AI-generated support reply into the prompt. 2. Fill in the customer issue, tone, and product context. 3. Run the prompt through your LLM (Claude, GPT-4, etc.). 4. Review the scores and recommendation. 5. If 'REVISE' or 'REJECT,' use the feedback to edit the reply or retrain your AI model. 6. Track scores over time to identify patterns (e.g., tone misses, knowledge gaps) and adjust your AI system or training data accordingly.

Example input

**Reply to Review:**
"We're sorry you're experiencing this issue. Unfortunately, we cannot process refunds after 30 days. Please contact our billing team if you have questions."

**Customer Context:**
Issue: Customer requested refund for unused subscription 45 days after purchase.
Customer tone: Frustrated, feels dismissed.
Product/service: SaaS subscription billing.

Example output

**QA EVALUATION:**

1. **Factual Accuracy**: 9/10
   - Policy is correctly stated (30-day refund window).
   - No contradictions detected.

2. **Completeness**: 5/10
   - Addresses the refund policy but doesn't explain *why* the customer is ineligible or offer alternatives (e.g., credit toward next month, exception process).
   - Missing next steps beyond "contact billing."

3. **Tone & Empathy**: 4/10
   - Opens with apology (good), but the "Unfortunately, we cannot" feels dismissive given the customer's frustration.
   - No acknowledgment of the customer's specific situation or offer of goodwill.

4. **Brand Voice Alignment**: 7/10
   - Professional and clear, but generic.
   - Lacks warmth or personality expected in your brand guidelines (if applicable).

5. **Conciseness**: 8/10
   - Brief and scannable. No unnecessary jargon.

**Average Score: 6.6/10**

**Recommendation: REVISE**

**Specific Revisions:**
- Acknowledge the customer's frustration more directly: "I understand this is frustrating, especially since you haven't used the service."
- Offer alternatives: "While we can't refund after 30 days, I can apply a credit to your next billing cycle or escalate this to our billing team for a one-time exception review."
- Provide a clear next step: "Reply here or email billing@company.com, and mention your request for an exception."
- Soften the tone: Replace "Unfortunately, we cannot" with "Our standard policy is 30 days, but let's see what we can do."

Customization tips

  • Adjust score thresholds (e.g., 8+ for approve) based on your risk tolerance. High-stakes issues (billing, security) may require 8.5+ average; routine questions may accept 7+.
  • Add industry-specific criteria. E.g., for healthcare support, add a 'Compliance' dimension to check HIPAA alignment; for fintech, add 'Security' to flag if sensitive data is handled correctly.
  • Weight dimensions differently. If tone is critical to your brand, make it 40% of the final score; accuracy might be 35%; completeness 25%.
  • Create a separate rubric for different support channels (chat vs. email vs. knowledge base articles) since tone and length expectations differ.
  • Integrate this rubric into your QA workflow: assign reviewers to score replies, track scores in a spreadsheet or database, and flag patterns (e.g., 'AI often misses billing exceptions') to your AI vendor or training team.
  • Use the rubric to benchmark AI performance over time. If average scores improve month-over-month, your AI system is learning; if they plateau, you may need to retrain or adjust prompts.
  • Pair this with a feedback loop: when a reply is rejected, log the reason and share it with your AI model's training data so it learns from mistakes.