Describe acceptable behavior
A support assistant may need to find a policy, explain it accurately, and decline to invent missing information. Those are separate capabilities. Build a rubric that distinguishes correct answers, unsupported claims, appropriate uncertainty, and requests that need human review.
Create a useful evaluation set
Include ordinary questions, ambiguous requests, unavailable information, and examples in each supported language. Keep development examples separate from the final evaluation. For an illustrative bilingual helpdesk, test whether both languages receive the same policy interpretation rather than judging fluency alone.
Repeat and inspect
Generated answers can vary. Record the model version, prompt, retrieved context, and generation settings, then repeat selected difficult cases. Automated grading can help organize results, but sample its judgments against a human rubric and investigate disagreements.
Test the application boundary
Check whether the assistant can expose another customer’s data, follow misleading instructions inside retrieved documents, or perform an action without confirmation. Test timeouts and unavailable dependencies as well as answer quality. Measure latency and cost per completed task alongside quality.
Try this: Start with a small, reviewed collection of representative questions and explicit expected behavior. Expand it using real failure patterns, with private information removed. Keep a release comparison so improvements in one area do not hide regressions elsewhere.