Original fictional teaching example. No model result is prefilled.
An editable starting point
- 1 · Text to check
- Original question or request: Where can I download my invoice, and how do I change the company name for next month? Proposed reply: Download your invoice from Billing → Invoices. For future invoices, update Company name in Billing → Details before the next billing date.
- 2 · Your question
- Does the reply address all explicit parts of the request?
- When does Yes apply?
- Every explicit question or request is addressed, including saying when an answer is unavailable.
- When does No apply?
- At least one explicit question or request is not addressed.
These are model probabilities, not measured accuracy or a guarantee. Yes and No describe your rules; choices are relative to the options you supplied.
Two different quantities
An option probability refers to one possible outcome. TypeSafe’s Choice confidence summarizes the shape of the whole option distribution; it is not another name for the winning probability. Noul has a yes probability and no separate confidence field. Our custom editor displays model probabilities, not an independently measured accuracy rate.
Do not read a strong yes result as “the model is good at this task.” It means yes to the particular question under the supplied rules. For an absolute-promise check, yes identifies a possible problem; for a coverage check, yes indicates the condition appears satisfied.
Use the attached reply exercise
Load the reply-completeness template, read its rules, and decide which explicit requirements the reply should cover. Change one part of the reply and compare the suggestions. The page supplies no expected probabilities, and a changed number alone does not prove an improvement.
The custom editor does not expose the provider’s separate Choice confidence field. Do not relabel one of its probability bars as confidence.
Look for confident mistakes
Try an incomplete reply that mentions every topic but does not answer the questions. Also inspect a reply that answers clearly using different wording. A sharp distribution can still reflect a mistaken interpretation, a poor question, or missing context. Review failures even when the model sounds certain through its numbers.
Choose thresholds with evidence
Before acting automatically, collect representative examples with judgments made independently of the model output. Track false approvals and unnecessary escalations separately. Choose an action threshold according to the cost of each error, and verify it on examples that were not used to tune the rules. This tutorial recommends no universal cutoff.
Continue with the reply-completeness recipe. Website functional checks and isolated live results do not establish calibration or benchmark accuracy.
Sources & further reading
Official documentation and linked community reports support this guide. Community observations are attributed to their authors.
How we check our sources