ConceptsCalibration

How Latent calibrates on your own answers: the schedule on each hosted plan, what a run does, how many answers it needs, and running one yourself in your environment.

Calibration fits Latent to your traffic. A run takes answers your model actually served, has a judge label a sample as faithful or hallucinated, fits the probe on the activations captured for those answers, sets the cut, and writes a report.

On hosted plans

A calibration run needs 500 scored answers. On Free, Latent runs one when your first 500 answers arrive, and the console shows how many more failures it catches; it is served once you move to Pro. On Pro, you start each run from the console for $25, or add Weekly calibration for $99 a month and Latent runs one every week. Team includes a run every week. Latent's judge labels the sample under the strict failure standard, and the update goes live with nothing to do on your side. Scoring continues on the previous calibration while a run works, and the swap is recorded in the console's change history.

A run takes under an hour for a few hundred to a thousand answers. If your traffic has too few failures to learn from yet, fewer than 30 in the held-out split, the report says so and your threshold is retuned on your traffic instead. The positives gate explains why.

Extra runs on Pro and Team are $25 each for up to 1,000 labeled answers, charged when the run starts. Start one from the console or ask at calibration@runlatent.ai, and it runs the same business day. A run over more than 1,000 answers is billed as one $25 run for each 1,000 answers or part of 1,000, quoted before it starts.

If you turn off answer text for a project, Latent has no text to label, so a calibration run has nothing to work on. Turn answer text back on before a run. The text is then kept for your plan's window, 30 days on Pro and 90 on Team, and deleted when the window ends.

Note The rest of this page describes running a calibration yourself, when Latent runs in your environment.

Run a calibration

export LATENT_REVIEW_TOKEN="$(cat review-token)"
latent calibrate --from-service <review-service-url> --vectors <reservoir-dir> --work calib/

--from-service needs the text store on (LATENT_TEXT_ENABLED, LATENT_GATEWAY_ATTACH_TEXT). With a log of your own:

latent calibrate --responses responses.jsonl --vectors deploy/audit/reservoir \
    --model <served-model-id> --work calib/ --gold gold.jsonl

The vectors come from the reservoir the plugin keeps on its own (LATENT_CAPTURE_RESERVOIR), a bounded sample of scored answers' activations with no text.

What a run does

  1. Pull. With --from-service, stored answers and your reviewers' labels come from the review service; otherwise from the log in --responses.
  2. Items. Truncated answers are counted and left out, and answers with no separable source are not judged.
  3. Label. The judge reads each answer with its whole source under your failure standard. A prompt over the judge budget is left unlabeled and counted; nothing is cut to fit. Reviewer labels win over the judge.
  4. Route (optional, --review-budget). The answers where the probe and the judge disagree most go to a human review queue.
  5. Fit. A held-out split is drawn by prompt, the probe and the out-of-distribution guard are fit on the rest, and the cut is set from the held-out risks (--threshold-mode budget, precision or probability).
  6. Report. report.md and report.json give held-out AUROC, the cut and what it flags (rate, precision, recall, silent misses) with intervals, the learning curve and the positives verdict, with no prompt or answer text.

Every stage writes into --work and resumes from it, so a crashed judge run never pays twice.

How many answers you need

Count failures: what decides whether a fit is worth shipping is how many labeled failures land in the held-out split.

  • On financial-table QA at a 17.6% held-out failure rate, 250 labeled answers were enough for a usable fit.
  • On healthcare literature QA, a run over 1,000 answers labeled 953 and found 34 failures, a 3.6% rate. A failure rate near 3% needs about five times as many answers for the same number of failures: the run's own estimate is about 5,000 answers, roughly $37 of judge labels.
  • On the reader path, a run mainly buys a cut set on your traffic.

The positives gate

After the fit, latent calibrate counts the held-out answers labeled as failures. The gate is 30 held-out failures, the command's default --min-positives. Below it the artifact is still written, and:

  • it carries verdict.fit_direction: insufficient_positives;
  • the report opens with "do not ship this direction", says to keep the current artifact and requantile its cut, and gives traffic_target, the number of further answers to collect;
  • the learning curve reads "too few failures to read";
  • latent doctor warns on the artifact's fit_verdict, and the console's registry tags it "needs more traffic".

A probe fit on a few dozen failures learns the corpus more than the failure. Requantiling moves only the cut, to the matching quantile of your live risks; the review service's budget_tracking does the same continuously:

latent requantile --events <audit.jsonl> --budget <b> --artifact <it> --out <copy>

What a run costs and how long it takes

Judging is the cost. A run over 1,000 answers took 453 seconds and $6.88 of judge spend on financial QA, and 344 seconds and $7.46 on healthcare QA. --judge-mode batch labels through the Message Batches API at its discount: one run paid $6.22 against $12.26 at list (274.5 seconds for the batch labeling). A batch that makes no progress is canceled and labeled synchronously (--judge-batch-stall-min). With your reviewers' labels in --labels, no judge is called at all.

Swap the new artifact in without a restart

Copy the artifact into the artifacts directory with its report beside it (<stem>.json, which names the file's hash), then make it the model's artifact in the policy, from the console's per-model controls or as an admin over the API:

curl -X PUT "<review-service-url>/policy?model=<served-model-id>" \
  -H "Authorization: Bearer $LATENT_REVIEW_TOKEN" \
  -H "Content-Type: application/json" \
  -d '{"artifact": "<registry path>"}'

The plugin picks it up on its next policy poll, checks it through the integrity gate and installs it between engine steps, so every answer is scored under one artifact. In a chaos test the swap was live within 4 to 6 seconds of the policy change. A load that fails leaves the current artifact serving. Only a change of read layer needs a restart.

Note With budget tracking on, the plugin keeps the previous tracked cut for one policy poll after a swap. The service serves the new artifact's own cut until 50 fresh answers have been scored on it, then re-derives the tracked cut (138 seconds in the rehearsal).

Read the learning curve and recalibrate

The report refits the probe on growing fractions of the training answers and reads each fit on the held-out split. When responses carry task or domain, the curve is read per group too, so a thin task never hides behind the pooled verdict. Run again when:

  • the curve is still rising. More answers of the same kind will help.
  • the curve has plateaued and you have improved the labels. More of the same labels will not move the probe; the levers are label quality (the review queue, --review-budget), the read layer or the domain.
  • the run fell below the positives gate. Keep the current artifact with a requantiled cut until you have collected traffic_target more answers.
  • the drift alert fires, or the flagged rate has moved away from the budget. The console's Drift card shows distance_from_calibration: above 0.15, requantile first; a fit comes after.

Was this page helpful?

Updated 3 October 2026