# Streaming and complete answers

How Latent scores a streamed answer and a complete one, and what can act on the verdict in each case, hosted and in your environment.

An answer reaches your user in one of two shapes. A complete answer comes back in one response after the model has finished writing. A streamed answer (`stream: true`) arrives a few tokens at a time while the model writes. Latent scores both. The shape decides when the verdict exists compared with what your user has already seen, and so what can act on it.

## At a glance

| | Complete answer | Streamed answer |
|---|---|---|
| Hosted | Scored when it is finished. With an automatic response on, the SDK holds, replaces, reroutes or regenerates a flagged answer before delivery (OpenAI, Anthropic and Gemini clients, one answer per call), or your code acts on the verdict. | Scored when the stream ends. Flagged answers go to your review queue. |
| In your environment | The gateway holds, replaces, reroutes or regenerates a flagged answer before delivery, with no code in your app. | Read at every token while it streams. Early stop can end a risky answer mid-stream. |

## A complete answer

### Hosted

With any provider, you can act on the verdict in your code. `runlatent.check()` scores the answer and returns its verdict before you return it:

```python
r = runlatent.check(prompt, answer, context=docs)
if r.flagged:
    answer = SAFE_FALLBACK
```

`context` takes your source documents as one string. If Latent cannot score the answer in time, `r.scored` is `False`, and you choose whether to deliver the answer or hold it. With `runlatent.auto_instrument()` on, pass the provider's response id (`request_id=resp.id`) so the check pairs with the captured answer and it appears once in the console; an unpaired check is scored, and counted, twice. From another language, send the pair to `POST /score` directly.

The SDK's one-line setup, `runlatent.auto_instrument()`, sends each answer after your app has it, on a background thread, so your user never waits on Latent. Those answers are scored, flagged and queued for review. Turn on an automatic response in the console and the same setup applies it before the call returns, to each complete, single-answer call made through the OpenAI, Anthropic or Google GenAI SDK: those calls then wait for the read, up to 500 ms by default, and a read that misses under `abstain` holds the answer. See [Act on a flagged answer](/docs/concepts/act-on-a-flagged-answer).

### In your environment

Put the gateway in front of your vLLM and it acts on a flagged answer before delivery, with no code in your app. It holds each response until the plugin's verdict arrives, up to 500 ms (`LATENT_GATEWAY_WAIT_MS`), then applies your policy's response:

| Action | What your user gets |
|---|---|
| `abstain` | Your fallback text in place of the answer |
| `route` | The answer from a stronger model you configure (`LATENT_GATEWAY_ROUTE_UPSTREAM`), asked the same request. Applies to every flagged answer, an `ood` verdict included. |
| `regenerate` | A second answer from your model, asked the same request once more |
| `judge` | The answer unchanged. The pair is queued for a written review when the gateway has a reviewer token (`LATENT_GATEWAY_JUDGE_TOKEN`), the text store is on and the review service's explain judge is set up. Without the token the decision is only recorded (`judge_unconfigured`); without the text or the judge, the queued review fails and is counted (`judge_failed`). |

On our GPU check (one A100 serving Llama 3.1 8B, 2026-09-19), all 1,040 complete answers across two runs had their verdict inside the 500 ms window. Watch `wait_ms` and `unmatched_timeout` on the gateway's `/health` for your own traffic.

Every scored response carries the verdict in a `latent` field in the body and in the `X-Latent-Risk`, `X-Latent-Verdict` and `X-Latent-Action` headers. `LATENT_GATEWAY_CLIENT_HEADERS=0` hides both from your end users. A rerouted or regenerated answer reports `latent.verdict` as `routed` or `regenerated` with `risk: null`, since the probe never read it, and the replaced answer's score under `latent.original_risk`.

Under `abstain`, an answer whose verdict does not arrive in time is withheld too, so a monitoring outage never lets an unscored answer through. Under every other action, it is delivered unchanged and marked `verdict: unknown`. The gateway acts on requests for one answer (`n: 1`).

## A streamed answer

A token your user has seen cannot be taken back. On a stream, the verdict flags the answer and queues it for review once the stream ends. To cut a risky answer short, Latent has to stop it while it is being written, which needs Latent inside your model.

### Hosted

The reader scores a streamed answer once the stream ends, with a per-token read of the whole answer. The SDK sends a stream when your app has read it to the end. If your app closes a healthy stream early (`close()`, or leaving its `with` block without an error), the text so far is sent, marked `finish_reason: "abort"` and `client.stream_complete: false`. A stream your app stops reading without closing it, one that fails partway, and an early exit from OpenAI's `.stream()` helpers are not sent. Flagged answers go to your review queue and count toward your [alert rules](/docs/integrations/alerts).

To hold a streamed answer before your user sees it, read the whole stream on your server first, then check it with `runlatent.check()` before you show it.

### In your environment

The plugin reads your model as it writes, so it has a risk at every token of a streamed answer, and the event keeps that per-token trace.

**Early stop.** With early stop on, the plugin ends an answer mid-stream once its risk stays high: by default, when the running-mean risk of the answer so far is at or above your cut on two consecutive per-token reads, never counting the first 16 tokens. Your user keeps the text already sent, and the stream's last chunk ends the answer with a choice that reads:

```json
{"index": 0, "delta": {}, "finish_reason": "length", "stop_reason": "latent_early_stop"}
```

An answer that reaches its natural end on the same step keeps `finish_reason: "stop"`, with the same `stop_reason`. A stop string that fires on the same token wins and sets its own `stop_reason`. Check `stop_reason` to tell a stopped answer from one that reached `max_tokens`, and replace or hide it in your interface. Early stop works on complete answers too: the answer comes back shorter, with the same `stop_reason`.

Early stop needs `LATENT_STREAM_DEVICE=1`, the default `tap` capture, an artifact with a `last` or `mean` probe, and a single GPU (tensor parallel 1). Capture-only and `mean+last` artifacts cannot stop an answer. Turn it on with `LATENT_EARLY_STOP=1`, or with a policy whose action is `early_stop`.

```bash
export LATENT_CAPTURE=tap   # the default
export LATENT_STREAM_DEVICE=1
export LATENT_EARLY_STOP=1
# Optional: LATENT_EARLY_STOP_MIN_TOKENS (default 16), LATENT_EARLY_STOP_PERSIST (default 2),
# LATENT_EARLY_STOP_THRESHOLD (default: your cut)
```

**The gateway on a stream.** The gateway passes a stream through as it arrives. After the last token it adds one chunk with the verdict, just before `[DONE]`. Abridged, it reads:

```text
data: {"id":"chatcmpl-…","object":"chat.completion.chunk","choices":[],"latent":{"risk":0.83,"verdict":"escalate","action":"none","stream":true,"enforcement":"observational","unsupported_shape":"stream","reason":"stream_not_enforced:abstain"}}

data: [DONE]
```

`reason` names the action your policy would have taken on a complete answer. Streams carry no `X-Latent-*` headers, because headers leave before the verdict exists, and with `LATENT_GATEWAY_CLIENT_HEADERS=0` there is no verdict chunk either. If your model's stream fails partway, the gateway ends it with an error chunk and no `[DONE]`, so a client never mistakes a cut-off answer for a whole one.

While your policy changes answers (`abstain`, `regenerate`, or `route` with a stronger model configured), the gateway answers a `stream: true` request with a 400 that names the fix (`stream: false`), so streaming can never get around your policy. Set `LATENT_GATEWAY_UNSUPPORTED_SHAPE=pass` to forward streams anyway, scored and flagged as above. Requests for several answers (`n` above 1) follow the same rule.

## Choose a setup

- **Every flagged complete answer held before a user sees it.** Turn on an automatic response in the console: the SDK applies it in your app, and in your environment the gateway does in front of your own vLLM. You can also check each answer with `runlatent.check()` before you return it.
- **Streaming with a hard stop.** Run Latent in your environment with early stop on. If the gateway also enforces a policy, set `LATENT_GATEWAY_UNSUPPORTED_SHAPE=pass` so streams reach your model.
- **Streaming on a hosted plan.** Stream as usual and review what is flagged. For the flows where one wrong answer costs the most, read the stream on your server and check it before you show it.

---
Docs index: https://runlatent.ai/llms.txt
