
Your rubric is part of the agent's spec, not just a test you run afterward.
"Helpful and accurate" is too vague to grade consistently.
A useful rubric is precise enough that two engineers reach the same verdict on the same output.
Bad: "Don't hallucinate"
Good: "Flag any claim not traceable to provided documents with [UNCERTAIN]."
Same goal. One can be applied consistently. The other cannot.
Writing the rubric is part of designing the agent.
Before it runs, decide what evidence counts as a good answer. Then use those rules to evaluate the output.

English










