Concepts
Failure standards
The rules Latent's calibration judge applies to decide which of your answers count as hallucinations, the one hosted plans use, and how to choose and set one in your environment.
Calibration starts with labels: for each answer in a sample, did it fail? A failure standard is the rule the judge applies to decide. Latent ships the strict standard as the default rubric, and when Latent runs in your environment it takes any rubric you write, which is how you run the sentence standard. On the same answers, the sentence standard counts many more failures than the strict one.
On hosted plans, calibration labels with the strict standard, the one standard Latent has checked against human labels on real answers. To discuss calibrating on another standard, write to calibration@runlatent.ai. A standard written for your own domain is part of Enterprise.
Strict: contradictions and unsupported specific facts
An answer fails when it does one of two things:
- states something the source directly contradicts, or
- asserts a specific fact (a name, number, date, place, event or attribute) that does not appear in the source at all.
Paraphrase, summarization, omission, generic statements, inferences that follow from the source, hedged statements, questions, opinions and formatting all pass. When in doubt, the judge answers faithful. This is latent/, the default for latent calibrate.
Sentence: any non-entailed sentence
Every sentence of the answer has to be entailed by the source. A sentence that adds framing, draws an inference the document does not state, summarizes with interpretation or expands an acronym from outside knowledge fails the whole answer, even when it contradicts nothing. This is the standard behind the FACTS Grounding leaderboard.
Compare the two standards
| The answer contains | Strict | Sentence |
|---|---|---|
| A statement the source contradicts | Fails | Fails |
| A specific fact absent from the source | Fails | Fails |
| An inference the source does not state | Passes | Fails |
| Framing or interpretation beyond the source | Passes | Fails |
| An acronym expanded from outside knowledge | Passes | Fails |
| A faithful paraphrase of the source | Passes | Passes |
How the standard affects calibration
The cut decides how much traffic is flagged: in the default budget mode, calibration places it so that your budget's share of held-out answers is flagged under either standard. The standard decides what that share contains. Under the strict standard, current frontier models fail short grounded tasks rarely, a few percent of answers in our runs, so a calibration sample holds few failures and any flag budget above the failure rate is partly filled with good answers, however well the probe ranks. A fit also needs failures in its held-out split (the positives gate). The sentence standard turns the same traffic into many more labeled failures.
Note The rest of this page describes choosing and running a standard yourself with
latent calibrate, when Latent runs in your environment.
Choose a failure standard
Pick strict when:
- a failure means a wrong fact someone acts on: a number, a name, a date, a contradiction;
- reviewer time is scarce and every flag should point at a concrete error;
- your model fails often enough under it to calibrate on your own traffic.
Pick sentence when:
- any claim beyond the document is a problem, as in long-document grounding or summaries for regulated use;
- the strict standard leaves too few failures on your traffic to pass the positives gate.
Write a sentence-standard rubric
latent calibrate labels with latent/ unless you pass --rubric. A rubric is a plain-text template over {context}, {query} and {output} that ends by asking the judge for exactly one word, FAITHFUL or HALLUCINATED, the way the strict rubric ends. For the sentence standard, ask the judge to check every sentence of {output} against {context} and to answer HALLUCINATED when any sentence is not entailed, inferences and framing included. No sentence-standard template ships and none has been validated against human labels.
Measure your failure rate under both
Measure your own failure rate under both standards before you choose. latent label runs the same labeler on any item file, such as the items. a calibration run writes; give it human labels in a gold field (--gold-field) or its output is marked unvalidated:
latent label --items calib/items.jsonl --rubric latent/rubrics/faithfulness_strict.txt --out labels_strict.jsonl
latent label --items calib/items.jsonl --rubric sentence_standard.txt --out labels_sentence.jsonl
Calibrate on the standard you picked
latent calibrate --from-service <review-service-url> --vectors <capture> --work calib/ \
--rubric sentence_standard.txt --gold gold.jsonl
Each candidate judge is scored on your --gold labels and has to clear the false-positive and recall gates (--max-fpr, --min-recall) before it labels anything; when no judge passes, the command labels nothing and says why. Labels are redone when the rubric file is newer than the saved labels or a judge setting changes.
Best practices
- Validate on your own gold. Pass
--goldwith human labels written under the same standard, for the strict default and for any rubric you write. - Write your standard down for reviewers. A reviewer's verdict on an answer wins over the judge's label in the next calibration.
- Switch rubrics in a fresh work directory. To relabel the same sample under another rubric, use a fresh
--workdirectory or--force.
Was this page helpful?
Updated 3 October 2026