Concepts
Scored answers
What Latent returns for each finished LLM response (risk score, verdict, cut and audit event), where the riskiest sentence is marked, and how to send answers in-band or through the reader.
A scored answer is one finished LLM response that Latent has read and judged. Latent reads the activations behind the answer (the vectors each layer of the network passes to the next) and applies a probe, a linear classifier fit on labeled answers. On hosted plans, and for models behind an API, Latent re-reads the finished prompt and answer with its own reader model and reads that model's activations. When Latent runs in your environment next to an open-weight model, it can read your model's activations inside your engine as the model writes. Each answer gets a risk score, a verdict against your cut, one event in your audit log and, with the sentence read on, a pointer to its riskiest sentence.
The risk score
risk is the probe's score for the answer, read from the mean of the activations over every answer token. Higher means more likely unsupported by the source the model was given, and the cut applies to it. meta. is the same score mapped to the chance the answer is wrong, by the calibration stored in the artifact. meta., shown beside the verdict, marks an answer whose activations sit far from the calibration data.
Verdicts
| Verdict | Meaning |
|---|---|
pass |
risk is below the cut |
escalate |
risk is at or above the cut |
ood |
only when the policy sets ood_ (or LATENT_ on the plugin): the activations are beyond the guard's cut |
unknown |
the answer was not scored; meta. or meta. says why |
unknown events are stored and kept out of every rate and queue.
How the cut is set
The cut (the threshold) is the risk at or above which an answer is flagged. A calibrated artifact carries its own cut, placed on held-out answers so that your review budget's share of traffic is flagged. The review service can override it per model with a pinned value, with budget tracking (the cut re-set from a trailing window of live risks), or with a cut derived from the artifact's held-out operating curve (cut_mode). Each event records its cut in meta. and the cut's origin in meta..
Switch on sentence localization
A scored answer also says where the risk sits.
- Words.
meta.traceholds the probe's read at every answer token in two series:scores, the answer so far (its last value is the risk), andscores_, each token on its own. The console tints the answer with the per-token series, so color gathers on the words that drove the score. On the reader path,last POST /returns the same two series by default,score token_(the answer so far) andrisk_ running token_risk(each token on its own), withtoken_, one entry per answer token. The console's request page tints the answer with them, and the Insights page's Token risk by position chart averages them over flagged and passed answers.pieces - Sentences. A second probe, fit on labeled sentences, scores each sentence from the same activations with no extra pass. In-band it is opt-in: set
LATENT_with a sentence probe attached to the artifact, and each event addsSTREAM_ SENTENCES=1 meta., whosesentence_ read localised_andindex localised_(span [start, end)character offsets in the decoded answer; the event carries no text) locate the riskiest sentence. On the reader path, every scored answer carries the sentence read (the riskiest sentence's probability under the sentence probe) when the probe is attached, which the hosted deployment does by default. The deep analysis adds counterfactual re-reads and prompt attribution on request. - On demand. A reviewer can open a deep analysis of one stored answer (
POST /), which re-reads it with each sentence removed. Check first the sentence whose removal lowers the risk most.explain/{request_ id}
Read the event and the audit log
{
"request_id": "chatcmpl-...",
"ts": "<unix seconds>",
"model": "<served model id>",
"layer": "<read layer>",
"risk": "<probe score>",
"verdict": "escalate",
"prompt_sha256": "<hex digest>",
"output_sha256": "<hex digest>",
"n_tokens": "<answer tokens>",
"meta": {
"finish": "eos",
"threshold": "<cut applied>",
"threshold_source": "policy",
"artifact_sha256": "<hex digest>",
"risk_calibrated": "<probability>",
"trace": {"kind": "mean", "q": "u8", "scores": "<base64>", "scores_last": "<base64>"}
}
}
The event carries hashes of the token ids and no text (what is stored). The review service appends it to audit., idempotent on request_id, and appends reviewer decisions (confirmed_ or false_, with error-type kinds) as new lines. The log alone shows which artifact scored which output (meta.), when, at which cut, and what a human decided. GET / returns the retained log; archived segments stay as files beside it.
Send answers with the SDK
On hosted plans, the Python SDK sends each finished call to Latent's scoring API, where the reader scores it. The Quickstart sets it up in one line.
Note The next two sections describe sending answers when Latent runs in your environment.
Send answers in-band, inside vLLM
The latent_ plugin loads inside vLLM. In the default tap capture it copies one layer's activations inside the CUDA graph at every decode step and keeps a running mean over the answer on the GPU; at the finish it scores the mean, applies the cut and emits the event. On one A100 serving an 8B model on RAG traffic at a 512-token answer cap, with CUDA graphs and the per-token trace on, the plugin served 5.68 requests per second against 5.68 for plain vLLM.
export LATENT_ARTIFACT=/artifacts/<artifact>.joblib
export LATENT_CAPTURE=tap
export LATENT_STREAM_DEVICE=1
export LATENT_SINK_URL=http://<review-service>/events
export LATENT_SINK_TOKEN=<ingest token>
python -m latent_vllm.doctor
vllm serve <model>
Put the optional latent_ in front of vLLM and the verdict comes back on the response (X-Latent-Risk, X-Latent-Verdict and X-Latent-Action headers, plus a latent field in the body). For flagged non-streaming answers it applies your policy's automatic response before delivery: abstain swaps in fallback text, route re-asks a stronger model, regenerate asks once more, and judge delivers the answer and queues it for the review judge.
Send answers through the reader, in your environment
For a managed API, another engine or a closed model, latent_ scores finished pairs. It holds an open reader model on its own GPU, re-reads the prompt and answer through the reader's chat template, and emits the same event with meta.. Its default artifact is the auditor, fit on human labels of other models' answers. Pairs reach POST / from middleware that runs after your response has gone out (latent_) or from a tail of your request log (python -m latent_).
{"request_id": "chatcmpl-...", "messages": [{"role": "user", "content": "..."}],
"output": "the answer text", "model": "the-served-model", "finish_reason": "stop"}
Against a hosted model (Gemini on Vertex AI, 348 answers, the reader on one A100 reached over a tunnel) the sidecar scored each answer in 150 to 170 ms at the median (p95 under 250 ms). The score lands after delivery, so on this path it drives the review queue and alerts.
Best practices
- Hold answers before delivery. The gateway holds a flagged complete answer automatically, in front of your own vLLM, or in front of an OpenAI-compatible API through the reader. In your code,
runlatent.returns the verdict before you return the answer. Streaming and complete answers covers what can act on a streamed answer.check() - Always send
finish_to the sidecar. Without it the event is markedreason finish_.unknown - Gate on
risk. The cut applies torisk; usemeta.when you need a probability.risk_ calibrated - Act on the whole-answer verdict. A hot token is a pointer for the reviewer; the decision to flag belongs to the whole answer.
Was this page helpful?
Updated 3 October 2026