Concepts
Streaming and complete answers
How Latent scores a streamed answer and a complete one, and what can act on the verdict in each case, hosted and in your environment.
An answer reaches your user in one of two shapes. A complete answer comes back in one response after the model has finished writing. A streamed answer (stream: true) arrives a few tokens at a time while the model writes. Latent scores both. The shape decides when the verdict exists compared with what your user has already seen, and so what can act on it.
At a glance
| Complete answer | Streamed answer | |
|---|---|---|
| Hosted | Scored when it is finished. With an automatic response on, the SDK holds, replaces, reroutes or regenerates a flagged answer before delivery (OpenAI, Anthropic and Gemini clients, one answer per call), or your code acts on the verdict. | Scored when the stream ends. Flagged answers go to your review queue. |
| In your environment | The gateway holds, replaces, reroutes or regenerates a flagged answer before delivery, with no code in your app. | Read at every token while it streams. Early stop can end a risky answer mid-stream. |
A complete answer
Hosted
With any provider, you can act on the verdict in your code. runlatent. scores the answer and returns its verdict before you return it:
r = runlatent.check(prompt, answer, context=docs)
if r.flagged:
answer = SAFE_FALLBACK
context takes your source documents as one string. If Latent cannot score the answer in time, r.scored is False, and you choose whether to deliver the answer or hold it. With runlatent. on, pass the provider's response id (request_) so the check pairs with the captured answer and it appears once in the console; an unpaired check is scored, and counted, twice. From another language, send the pair to POST / directly.
The SDK's one-line setup, runlatent., sends each answer after your app has it, on a background thread, so your user never waits on Latent. Those answers are scored, flagged and queued for review. Turn on an automatic response in the console and the same setup applies it before the call returns, to each complete, single-answer call made through the OpenAI, Anthropic or Google GenAI SDK: those calls then wait for the read, up to 500 ms by default, and a read that misses under abstain holds the answer. See Act on a flagged answer.
In your environment
Put the gateway in front of your vLLM and it acts on a flagged answer before delivery, with no code in your app. It holds each response until the plugin's verdict arrives, up to 500 ms (LATENT_), then applies your policy's response:
| Action | What your user gets |
|---|---|
abstain |
Your fallback text in place of the answer |
route |
The answer from a stronger model you configure (LATENT_), asked the same request. Applies to every flagged answer, an ood verdict included. |
regenerate |
A second answer from your model, asked the same request once more |
judge |
The answer unchanged. The pair is queued for a written review when the gateway has a reviewer token (LATENT_), the text store is on and the review service's explain judge is set up. Without the token the decision is only recorded (judge_); without the text or the judge, the queued review fails and is counted (judge_). |
On our GPU check (one A100 serving Llama 3.1 8B, 2026-09-19), all 1,040 complete answers across two runs had their verdict inside the 500 ms window. Watch wait_ms and unmatched_ on the gateway's /health for your own traffic.
Every scored response carries the verdict in a latent field in the body and in the X-Latent-Risk, X-Latent-Verdict and X-Latent-Action headers. LATENT_ hides both from your end users. A rerouted or regenerated answer reports latent. as routed or regenerated with risk: null, since the probe never read it, and the replaced answer's score under latent..
Under abstain, an answer whose verdict does not arrive in time is withheld too, so a monitoring outage never lets an unscored answer through. Under every other action, it is delivered unchanged and marked verdict: unknown. The gateway acts on requests for one answer (n: 1).
A streamed answer
A token your user has seen cannot be taken back. On a stream, the verdict flags the answer and queues it for review once the stream ends. To cut a risky answer short, Latent has to stop it while it is being written, which needs Latent inside your model.
Hosted
The reader scores a streamed answer once the stream ends, with a per-token read of the whole answer. The SDK sends a stream when your app has read it to the end. If your app closes a healthy stream early (close(), or leaving its with block without an error), the text so far is sent, marked finish_ and client.. A stream your app stops reading without closing it, one that fails partway, and an early exit from OpenAI's .stream() helpers are not sent. Flagged answers go to your review queue and count toward your alert rules.
To hold a streamed answer before your user sees it, read the whole stream on your server first, then check it with runlatent. before you show it.
In your environment
The plugin reads your model as it writes, so it has a risk at every token of a streamed answer, and the event keeps that per-token trace.
Early stop. With early stop on, the plugin ends an answer mid-stream once its risk stays high: by default, when the running-mean risk of the answer so far is at or above your cut on two consecutive per-token reads, never counting the first 16 tokens. Your user keeps the text already sent, and the stream's last chunk ends the answer with a choice that reads:
{"index": 0, "delta": {}, "finish_reason": "length", "stop_reason": "latent_early_stop"}
An answer that reaches its natural end on the same step keeps finish_, with the same stop_. A stop string that fires on the same token wins and sets its own stop_. Check stop_ to tell a stopped answer from one that reached max_tokens, and replace or hide it in your interface. Early stop works on complete answers too: the answer comes back shorter, with the same stop_.
Early stop needs LATENT_, the default tap capture, an artifact with a last or mean probe, and a single GPU (tensor parallel 1). Capture-only and mean+last artifacts cannot stop an answer. Turn it on with LATENT_, or with a policy whose action is early_stop.
export LATENT_CAPTURE=tap # the default
export LATENT_STREAM_DEVICE=1
export LATENT_EARLY_STOP=1
# Optional: LATENT_EARLY_STOP_MIN_TOKENS (default 16), LATENT_EARLY_STOP_PERSIST (default 2),
# LATENT_EARLY_STOP_THRESHOLD (default: your cut)
The gateway on a stream. The gateway passes a stream through as it arrives. After the last token it adds one chunk with the verdict, just before [DONE]. Abridged, it reads:
data: {"id":"chatcmpl-…","object":"chat.completion.chunk","choices":[],"latent":{"risk":0.83,"verdict":"escalate","action":"none","stream":true,"enforcement":"observational","unsupported_shape":"stream","reason":"stream_not_enforced:abstain"}}
data: [DONE]
reason names the action your policy would have taken on a complete answer. Streams carry no X-Latent-* headers, because headers leave before the verdict exists, and with LATENT_ there is no verdict chunk either. If your model's stream fails partway, the gateway ends it with an error chunk and no [DONE], so a client never mistakes a cut-off answer for a whole one.
While your policy changes answers (abstain, regenerate, or route with a stronger model configured), the gateway answers a stream: true request with a 400 that names the fix (stream: false), so streaming can never get around your policy. Set LATENT_ to forward streams anyway, scored and flagged as above. Requests for several answers (n above 1) follow the same rule.
Choose a setup
- Every flagged complete answer held before a user sees it. Turn on an automatic response in the console: the SDK applies it in your app, and in your environment the gateway does in front of your own vLLM. You can also check each answer with
runlatent.before you return it.check() - Streaming with a hard stop. Run Latent in your environment with early stop on. If the gateway also enforces a policy, set
LATENT_so streams reach your model.GATEWAY_ UNSUPPORTED_ SHAPE=pass - Streaming on a hosted plan. Stream as usual and review what is flagged. For the flows where one wrong answer costs the most, read the stream on your server and check it before you show it.
Was this page helpful?
Updated 3 October 2026