# Scored answers

What Latent returns for each finished LLM response (risk score, verdict, cut and audit event), where the riskiest sentence is marked, and how to send answers in-band or through the reader.

A scored answer is one finished LLM response that Latent has read and judged. Latent reads the activations behind the answer (the vectors each layer of the network passes to the next) and applies a probe, a linear classifier fit on labeled answers. On hosted plans, and for models behind an API, Latent re-reads the finished prompt and answer with its own reader model and reads that model's activations. When Latent runs in your environment next to an open-weight model, it can read your model's activations inside your engine as the model writes. Each answer gets a risk score, a verdict against your cut, one event in your audit log and, with the sentence read on, a pointer to its riskiest sentence.

## The risk score

`risk` is the probe's score for the answer, read from the mean of the activations over every answer token. Higher means more likely unsupported by the source the model was given, and the cut applies to it. `meta.risk_calibrated` is the same score mapped to the chance the answer is wrong, by the calibration stored in the artifact. `meta.low_confidence`, shown beside the verdict, marks an answer whose activations sit far from the calibration data.

## Verdicts

| Verdict | Meaning |
|---|---|
| `pass` | `risk` is below the cut |
| `escalate` | `risk` is at or above the cut |
| `ood` | only when the policy sets `ood_as_verdict` (or `LATENT_OOD_AS_VERDICT=1` on the plugin): the activations are beyond the guard's cut |
| `unknown` | the answer was not scored; `meta.unscored_reason` or `meta.capture_inert` says why |

`unknown` events are stored and kept out of every rate and queue.

## How the cut is set

The cut (the threshold) is the risk at or above which an answer is flagged. A calibrated artifact carries its own cut, placed on held-out answers so that your review budget's share of traffic is flagged. The review service can override it per model with a pinned value, with budget tracking (the cut re-set from a trailing window of live risks), or with a cut derived from the artifact's held-out operating curve (`cut_mode`). Each event records its cut in `meta.threshold` and the cut's origin in `meta.threshold_source`.

## Switch on sentence localization

A scored answer also says where the risk sits.

- **Words.** `meta.trace` holds the probe's read at every answer token in two series: `scores`, the answer so far (its last value is the risk), and `scores_last`, each token on its own. The console tints the answer with the per-token series, so color gathers on the words that drove the score. On the reader path, `POST /score` returns the same two series by default, `token_risk_running` (the answer so far) and `token_risk` (each token on its own), with `token_pieces`, one entry per answer token. The console's request page tints the answer with them, and the Insights page's Token risk by position chart averages them over flagged and passed answers.
- **Sentences.** A second probe, fit on labeled sentences, scores each sentence from the same activations with no extra pass. In-band it is opt-in: set `LATENT_STREAM_SENTENCES=1` with a sentence probe attached to the artifact, and each event adds `meta.sentence_read`, whose `localised_index` and `localised_span` (`[start, end)` character offsets in the decoded answer; the event carries no text) locate the riskiest sentence. On the reader path, every scored answer carries the sentence read (the riskiest sentence's probability under the sentence probe) when the probe is attached, which the hosted deployment does by default. The deep analysis adds counterfactual re-reads and prompt attribution on request.
- **On demand.** A reviewer can open a deep analysis of one stored answer (`POST /explain/{request_id}`), which re-reads it with each sentence removed. Check first the sentence whose removal lowers the risk most.

## Read the event and the audit log

```json
{
  "request_id": "chatcmpl-...",
  "ts": "<unix seconds>",
  "model": "<served model id>",
  "layer": "<read layer>",
  "risk": "<probe score>",
  "verdict": "escalate",
  "prompt_sha256": "<hex digest>",
  "output_sha256": "<hex digest>",
  "n_tokens": "<answer tokens>",
  "meta": {
    "finish": "eos",
    "threshold": "<cut applied>",
    "threshold_source": "policy",
    "artifact_sha256": "<hex digest>",
    "risk_calibrated": "<probability>",
    "trace": {"kind": "mean", "q": "u8", "scores": "<base64>", "scores_last": "<base64>"}
  }
}
```

The event carries hashes of the token ids and no text ([what is stored](/docs/concepts/data)). The review service appends it to `audit.jsonl`, idempotent on `request_id`, and appends reviewer decisions (`confirmed_failure` or `false_alarm`, with error-type `kinds`) as new lines. The log alone shows which artifact scored which output (`meta.artifact_sha256`), when, at which cut, and what a human decided. `GET /export.jsonl` returns the retained log; archived segments stay as files beside it.

## Send answers with the SDK

On hosted plans, the Python SDK sends each finished call to Latent's scoring API, where the reader scores it. The [Quickstart](/docs/quickstart) sets it up in one line.

> Note: The next two sections describe sending answers when Latent runs in your environment.

## Send answers in-band, inside vLLM

The `latent_vllm` plugin loads inside vLLM. In the default `tap` capture it copies one layer's activations inside the CUDA graph at every decode step and keeps a running mean over the answer on the GPU; at the finish it scores the mean, applies the cut and emits the event. On one A100 serving an 8B model on RAG traffic at a 512-token answer cap, with CUDA graphs and the per-token trace on, the plugin served 5.68 requests per second against 5.68 for plain vLLM.

```bash
export LATENT_ARTIFACT=/artifacts/<artifact>.joblib
export LATENT_CAPTURE=tap
export LATENT_STREAM_DEVICE=1
export LATENT_SINK_URL=http://<review-service>/events
export LATENT_SINK_TOKEN=<ingest token>
python -m latent_vllm.doctor
vllm serve <model>
```

Put the optional `latent_gateway` in front of vLLM and the verdict comes back on the response (`X-Latent-Risk`, `X-Latent-Verdict` and `X-Latent-Action` headers, plus a `latent` field in the body). For flagged non-streaming answers it applies your policy's automatic response before delivery: `abstain` swaps in fallback text, `route` re-asks a stronger model, `regenerate` asks once more, and `judge` delivers the answer and queues it for the review judge.

## Send answers through the reader, in your environment

For a managed API, another engine or a closed model, `latent_sidecar` scores finished pairs. It holds an open reader model on its own GPU, re-reads the prompt and answer through the reader's chat template, and emits the same event with `meta.capture: "sidecar"`. Its default artifact is the auditor, fit on human labels of other models' answers. Pairs reach `POST /score` from middleware that runs after your response has gone out (`latent_sidecar.tap.FireAndForget`) or from a tail of your request log (`python -m latent_sidecar.tap --follow`).

```json
{"request_id": "chatcmpl-...", "messages": [{"role": "user", "content": "..."}],
 "output": "the answer text", "model": "the-served-model", "finish_reason": "stop"}
```

Against a hosted model (Gemini on Vertex AI, 348 answers, the reader on one A100 reached over a tunnel) the sidecar scored each answer in 150 to 170 ms at the median (p95 under 250 ms). The score lands after delivery, so on this path it drives the review queue and alerts.

## Best practices

- **Hold answers before delivery.** The gateway holds a flagged complete answer automatically, in front of your own vLLM, or in front of an OpenAI-compatible API through the reader. In your code, `runlatent.check()` returns the verdict before you return the answer. [Streaming and complete answers](/docs/concepts/streaming) covers what can act on a streamed answer.
- **Always send `finish_reason`** to the sidecar. Without it the event is marked `finish_unknown`.
- **Gate on `risk`.** The cut applies to `risk`; use `meta.risk_calibrated` when you need a probability.
- **Act on the whole-answer verdict.** A hot token is a pointer for the reviewer; the decision to flag belongs to the whole answer.

---
Docs index: https://runlatent.ai/llms.txt
