ConceptsAct on a flagged answer

What can happen to a flagged answer after Latent scores it, where each action runs (the gateway, the plugin, your code), how to set it on the Policy page, and which setup fits your deployment.

Scoring finds the answer: every finished response gets a risk and a verdict against your cut, and every scored answer, flagged or not, is written to your audit log. Acting is what happens next, and where it happens decides what is possible. A component that sees the HTTP response can replace, reroute or regenerate a flagged answer before your user reads it. A component inside the model can stop an answer while it is still being written. Code in your app can hold an answer it has already received. The review service alone can queue a flag for a human and alert on the rate; it never touches a response.

The actions

Your policy names one automatic response for every flagged answer (response.action). The table reads left to right: what the user gets, what Latent writes down, where the action runs, and what has to be true for it to run.

Action What the user sees What is recorded Where it runs Preconditions
abstain Your fallback text in place of the answer. Tool calls, logprobs and reasoning fields are removed from the choice, so the flagged output does not leak beside the replacement. latent.action: abstain, X-Latent-Action: abstain, and the original answer's hash as latent.replaced_output_sha256, so the audit row joins. The audit event is untouched: it was scored and stored before the replacement. The gateway, or the SDK in your app (hosted) A complete (non-streaming) answer with n: 1, and either the gateway in front of your vLLM or, on the hosted service, auto_instrument() with an OpenAI, Anthropic or Gemini client.
route The answer from a stronger model you configure, asked the same request. latent.action: route, latent.verdict: routed, latent.risk: null (the probe never read the routed answer), the withheld answer's score as latent.original_risk and latent.original_verdict, plus latent.routed_to and latent.routed_id. A failure or timeout delivers the original answer as route_failed, with its own risk. The gateway, or the SDK in your app (hosted) As above, plus LATENT_GATEWAY_ROUTE_UPSTREAM behind the gateway. In your app (hosted), the SDK re-asks the policy's response.route_model on the same provider. Applies to every flagged answer, an ood verdict included; only the legacy overflow: route_to_stronger_model passes an ood verdict through (route_not_for_ood). Without a route configured the answer passes through (route_unconfigured).
regenerate A second answer from the same model, asked the same request once more. latent.verdict: regenerated, latent.risk: null, latent.original_risk, latent.regenerated_id, latent.replaced_output_sha256. The second answer is scored as its own request, as a second event on the service; this response does not wait for that score. A failure delivers the original answer as regenerate_failed. The gateway, or the SDK in your app (hosted) A complete answer with n: 1, behind the gateway or through auto_instrument() on the hosted service. One extra generation per flagged answer.
judge The answer, unchanged. latent.action: judge. The pair is queued for the review service's written analysis about 2 seconds after the response left, and the judge's verdict lands on the request's record in the console. Without a reviewer token on the gateway the decision is recorded as judge_unconfigured. The gateway or the SDK, then the review service Behind the gateway: LATENT_GATEWAY_JUDGE_TOKEN, the text store on (LATENT_TEXT_ENABLED=1 on the service, LATENT_GATEWAY_ATTACH_TEXT=1 on the gateway), and the review service's explain judge set up (LATENT_EXPLAIN_SIDECAR_URL, with LATENT_EXPLAIN_JUDGE=1 and an Anthropic key for the written analysis). On the hosted service the SDK queues it; it needs the project's text store and counts toward the plan's daily judge analyses. The judge never changes an answer.
early_stop A shorter answer: the model stops writing once its risk has stayed above the cut. The last choice reads finish_reason: "length" with stop_reason: "latent_early_stop". The text already streamed stays on screen. The event's meta.finish: "early_stop" and meta.early_stop (the decode position, the score and the rule). The gateway, if present, records latent.reason: early_stop_is_the_plugins and takes no action of its own. The plugin, inside vLLM Latent in your environment: the default tap capture, LATENT_STREAM_DEVICE=1, an artifact with a last or mean probe, tensor parallel 1. Off by default. It is the plugin's own rule, and the only action that works on a stream.
none The answer, unchanged. The flag, the verdict and the cut on the event. With the review queue on, the row is stamped tier: review and queued; with it off, tier: auto and counted. Behind the gateway, latent.action: none with latent.tier. The review service Nothing. This is the default.

Three rules hold across the table.

  • A stream carries the verdict and no action. A token your user has seen cannot be taken back, so the gateway passes a stream: true response through untouched and appends one chunk with the verdict before [DONE], with reason: stream_not_enforced:<action> and enforcement: observational. While your policy changes answers (abstain, regenerate, or route with a route configured), the gateway answers a stream request with a 400 that names the fix (stream: false) rather than letting streaming get around the policy; LATENT_GATEWAY_UNSUPPORTED_SHAPE=pass forwards streams anyway, marked. Requests for several answers (n above 1) follow the same rule: the decision is recorded for the first sample and not applied. Streaming and complete answers has the detail.
  • Early stop is the plugin's rule. It runs inside your model with its own threshold and serving flags. The gateway never stops a stream.
  • A routed or regenerated answer is unscored, and says so. verdict: routed or regenerated with risk: null; the withheld answer's score travels as original_risk. In our measured run, routing an 8B model's flagged answers to a 70B fixed 33% of them: routing buys a better answer some of the time and is not a fix for the flag.

Older policy documents name the same actions as action: safe_fallback (the abstain replacement) and the over-budget values hold_and_alert, route_to_stronger_model and safe_fallback_above_threshold. They still work as legacy aliases when response.action is none; the names above are current.

Set it in the console

Everything in the table is set on the Policy page, per project or per model. The page leads with Monitoring mode, four presets over the review fields and the queue switch, saved with the one Save button:

Preset What it writes What stops or refuses output
Review queue Queue on, a length finish goes to review, no automatic response. Humans review flagged answers within the budget. Nothing
Alerts only Queue off, nothing routed, no automatic response. Flags are recorded and counted, and alert rules fire. Nothing
Act automatically Queue off, and the Automatic response below does the work. Pressing it picks a response this deployment enforces, when there is one. early_stop under tap capture with the device stream; abstain, route or regenerate behind the gateway, or in your app through the SDK on the hosted service. The judge never changes an answer.
Custom Keeps the current values and opens Advanced. Set each field yourself. Per the fields

The response select names the actions as Stop generation early, Send to the judge, Route to a stronger model, Regenerate the answer and Abstain with the fallback text; None reads "review queue only" with the queue on and "record the flag only" with it off.

How the console knows what is enforced. It reads the newest scored answer in scope. An answer captured in-band with the device stream unlocks early stop; an answer that came through the gateway (every event the gateway forwards is stamped meta.gateway: true), or that the SDK enforced in your app (meta.enforced_by: "sdk"), unlocks abstain, route and regenerate. In your environment, where nothing is enforceable, the preset says why: "Nothing scored yet; the options appear after the first answer", or "Acting on output needs the gateway or the in-band plugin with the device stream; the newest answer shows neither". On the hosted service the preset is enforced by the SDK in your app (auto_instrument(), runlatent 0.3.0).

The flag threshold is one cut: every answer at or above it is flagged, whatever you do next. It is derived, never stored. The mode row offers Auto (the review budget's share of traffic, lowered to reviewer capacity when one is set), F-beta, Cost, Flag rate and Pinned, and under the slider the page reads off the artifact's held-out curve what the cut means: the flagged share, precision, failures caught and false actions per 1,000. With the queue off, the budget row reads Flag rate: the same number, now defining what is flagged rather than what is reviewed.

The review queue switch is policy.review.enabled. On, every flagged answer is stamped tier: review and sorted by score for your reviewers. Off, flagged rows are stamped tier: auto with tier_basis: review_disabled: out of the queue, still flagged for the rates, the series, /metrics and the alert rules. A flag is a flag whether or not it was queued. Every change is a line in the policy history.

In your app (hosted). While the automatic response is abstain, regenerate, route or judge, auto_instrument() reads each complete, single-answer call before returning it, waiting up to the policy's response.wait_ms (500 ms by default). A read that misses under abstain holds the answer with the fallback text (latent.action: hold); under the other actions the answer is released with verdict: unknown. auto_instrument(enforce=False) turns enforcement off.

What is recorded, and what is enforced. The policy document stores every field so your intent is on the audit log now, and the runtime honors a subset. GET /policy.enforced is the machine-readable table. The console marks the rest: the over-budget list shows the values this deployment enforces and puts the others behind Show all with "N options not enforced in this deployment"; the response list shows every action, and on the hosted service lists early stop as "Available in your VPC today."; a stored response the deployment no longer enforces reads Stored: ...; not enforced by the current capture (...); nothing is queued for an early stop; for the other actions it reads Stored: ...; needs the gateway in front of the model, which the newest answer does not show; nothing is queued in your environment, and Stored: ...; needs the SDK's enforce mode in your app, which the newest answer does not show; nothing is queued on the hosted service, rather than a promise; and in your environment the Operating curve's limits line says that the automatic response applies to non-streaming single-sample requests today, streams and parallel samples are recorded only.

The gateway

The gateway is a reverse proxy that sits between your clients and your model's OpenAI-compatible endpoint. It is the one component in the stack that sees the response, so it is where abstain, route, regenerate and judge become real. In front of your own vLLM it joins each response to the event the plugin scored for it (the plugin posts its events through the gateway, and the gateway forwards them to the review service), puts the verdict on the response in-band, and applies your policy before the answer leaves. One Python process, no GPU, no torch.

# side by side (try it): gateway on ${GATEWAY_PORT:-8001}, vLLM still on ${VLLM_PORT}
docker compose -f deploy/docker-compose.yml --profile gateway up -d --build
# enforcing: the gateway takes ${VLLM_PORT}, vLLM is unpublished, plugin -> gateway -> service
docker compose -f deploy/docker-compose.yml -f deploy/docker-compose.gateway.yml --profile gateway up -d --build

Three variables do most of the work: LATENT_GATEWAY_UPSTREAM (your vLLM's OpenAI server), LATENT_SINK_URL (the review service's /events, which also tells the gateway where to poll the policy) and LATENT_INGEST_TOKEN (one shared secret from the plugin to the gateway to the service). Two more shape the actions: LATENT_GATEWAY_FALLBACK_TEXT is what abstain delivers (default: "I can't answer that reliably right now; a human will follow up."), and LATENT_GATEWAY_ROUTE_UPSTREAM is the stronger endpoint route re-asks, with LATENT_GATEWAY_ROUTE_MODEL and LATENT_GATEWAY_ROUTE_TOKEN when it needs them and a 30 second cap (LATENT_GATEWAY_ROUTE_TIMEOUT_S); the client's own credential is never forwarded there. The gateway needs a project key from your console (LATENT_GATEWAY_SCORE_TOKEN), validated at start and hourly; without one it refuses to start rather than run with enforcement silently off. An in-VPC deployment against your own review service may opt out with LATENT_GATEWAY_REQUIRE_KEY=0 and LATENT_GATEWAY_IN_VPC=1 together.

Every complete response carries the verdict twice: in the headers (X-Latent-Risk, X-Latent-Verdict, X-Latent-Policy-Version, X-Latent-Action, and X-Latent-Reason when the verdict is unknown) and as a latent field on the completion body, which OpenAI SDKs keep as an unknown field (response.model_extra["latent"] in Python). Abridged:

"latent": {
  "risk": 0.83, "verdict": "escalate",
  "action": "abstain", "reason": "policy.response",
  "tier": "auto", "flag_cut": 0.62, "cut_basis": "risk",
  "policy_version": 3, "policy_source": "model",
  "replaced_output_sha256": "..."
}

LATENT_GATEWAY_CLIENT_HEADERS=0 hides both from your end users; enforcement, the alert webhook and the audit records are unchanged.

The 500 ms wait. The response and the event race. The gateway holds a complete response until its event arrives or LATENT_GATEWAY_WAIT_MS (default 500) has elapsed, a budget chosen as the largest value a chat client will not notice against a multi-second generation. In our GPU check (one A100 serving Llama 3.1 8B, 2026-09-19), all 1,040 complete answers across two runs had their verdict inside the window, and at 32 concurrent requests the proxy hop cost nothing measurable: 300 prompts in 52 s through the gateway against 51 s direct. Watch wait_ms and unmatched_timeout on the gateway's /health for your own traffic.

Fail-closed under abstain. LATENT_GATEWAY_FAIL_CLOSED is auto by default: while the policy withholds flagged answers (abstain, or the legacy safe_fallback, hold_and_alert and safe_fallback_above_threshold), an answer whose verdict does not arrive in time is withheld too, with the fallback text and latent.action: hold, because an unscored answer released under a withholding policy is the failure that policy exists to prevent. Under every other action a miss is released unchanged and marked verdict: unknown, reason: unmatched_timeout. Both are counted on /health and each gets one audit record on the review service, so no fail-open is silent. 1 and 0 pin the behavior on or off.

Stop mid-answer

Early stop is the one action that works on a stream, and it runs inside your model. With the plugin's device stream on, every decode step reads back each running request's running-mean risk, the probe's score of the mean activation over the answer so far. A request whose running mean is at or above the cut for 2 consecutive reads, never counting the first 16 tokens (LATENT_EARLY_STOP_PERSIST, LATENT_EARLY_STOP_MIN_TOKENS), is finished on the spot instead of decoding to max_tokens.

What the client sees is vLLM's own length finish with Latent's reason beside it, so OpenAI-compatible clients already handle it as a truncated answer:

{"index": 0, "delta": {}, "finish_reason": "length", "stop_reason": "latent_early_stop"}

Key on stop_reason to tell a stopped answer from one that hit max_tokens, and replace or hide it in your interface; the plugin cannot recall tokens already sent. An answer that reaches its natural end on the same step keeps finish_reason: "stop" with the same stop_reason. The audit event reads meta.finish: "early_stop" with the decision under meta.early_stop. Turn it on with LATENT_EARLY_STOP=1 or a policy whose response is early_stop; it needs the default tap capture, LATENT_STREAM_DEVICE=1 and a last or mean artifact. Under a last artifact the shipped cut was fit for the last token alone, so set LATENT_EARLY_STOP_THRESHOLD or fit a mean artifact.

Measured, with conditions. Served on an A100 with Llama 3.1 8B, a mean artifact at threshold 0.761, 300 held-out prompts, async scheduling: the running-mean rule stopped 17% of requests (51 of 300) at precision 0.80 against Opus labels of the full answers, recall of bad answers 0.33, 5.8% of good answers killed, 15.2% of decode tokens saved. The stop is a decision on a partial answer, so its precision belongs to the plugin's own calibrated rule and differs from the full answer's verdict.

The sentence-head signal (opt-in). LATENT_EARLY_STOP_SIGNAL=sentence_head decides at sentence boundaries instead, from the in-band sentence read over the completed sentences so far, and never on a token inside an open sentence. Served on an A100 on 2026-09-24: 1,000 of Llama 3.1 8B Instruct's own long-context RAG answers at temperature 0, 512-token cap, 16 concurrent, with the all-task sentence probe, labeled by the strict Opus rubric. At the shipped requantiled cut (0.9157) the head stopped 8.3% of answers at precision 0.928 [0.865, 0.976], killed 1.3% [0.4, 2.5] of good answers and saved 6.3% of tokens. At a matched 17% stop rate it read 0.857 against the running mean's 0.842 (paired +0.015 [-0.044, +0.077]), and the delivered stopped prefixes were unsupported 84% against 69%. It decides later (median stop index 172 against 92 in the offline simulation), its lag is +2 to +9 delivered tokens, and it never stopped a summary. The running mean stays the default; the head is for deployments that want precision first and accept a later stop, with the head's cut set on their own traffic. The mechanism cost nothing measurable in throughput (6.03 req/s with the head loop running and its cut set so it never fires, against 6.01 off, at 32 concurrent). It needs LATENT_STREAM_SENTENCES=1 and a sentence probe that carries a sentence_head block; the attach refuses otherwise.

In your code

With any provider, your code can hold the answer it already has. runlatent.check() scores one answer synchronously and returns its verdict before you return the answer (runlatent 0.2.0 or later; acheck() is the awaited form):

import runlatent

r = runlatent.check(prompt, answer, context=docs)   # context: your sources, one string
if not r.scored:          # the read did not happen: r.status, r.detail, r.resume say why
    deliver(answer)       # fail open, or answer = SAFE_FALLBACK to fail closed
elif r.flagged:           # verdict escalate or ood: do not deliver as is
    answer = SAFE_FALLBACK
  • r.flagged is true for an escalate or ood verdict. r.passed is true only for a scored pass, and bool(r) is r.passed, so if runlatent.check(...) holds a refusal and an empty answer alike: the fail-closed test in one word. if not r.flagged is the fail-open one.
  • A refusal never raises by default. A 401, 403, 429, a 5xx, a timeout (2 seconds by default) and a connection error all return scored=False with status, the server's detail and a resume sentence, so the fail-open or fail-closed choice is yours and explicit. A reply the SDK cannot read counts as a refusal too (scored=False, verdict unknown). Pass raise_on_error=True for runlatent.CheckError instead; only a missing URL or key raises ConfigurationError regardless.
  • A checked answer appears once in the console, even with auto_instrument() on, when it pairs with the captured read: pass the provider's response id as request_id (or messages= exactly as the provider saw them). An unpaired check is scored, and counted, twice.
  • It uses the destination and key auto_instrument(), configure() or the environment set (LATENT_API_URL, LATENT_API_KEY): the hosted scoring route, or your sidecar in your environment.

The detail reply. check(..., detail=True) asks the scoring route for one more key on the same reply, and the result gains r.sentences (each answer sentence with its character span and its risk under the sentence probe), r.suite and r.suite_fired (every probe-suite head's risk, cut and whether it fired) and r.read (the raw block: the cut and its source, the OOD distance, the finish and token count, the layer and the reader). The default reply is unchanged, and POST /score takes the same detail: true body field from any language. Nothing in the block is stored text: the sentences are your own answer sliced at the read's spans, so it is available on every plan. Next to auto_instrument(), a detail check posts its own read unless it runs before the capture is sent, so pass request_id and check before the capture fires to keep it to one scored answer. Needs runlatent 0.3.0.

Which one?

  • Hosted, with the SDK. Turn on an automatic response; auto_instrument() applies it to each complete answer before the call returns.
  • Hosted, with an API model. Hold in your code with check() before you return the answer; from another language, POST /score.
  • Your own vLLM, in your environment. The plugin's early stop for streams, and the gateway for complete answers (abstain, route, regenerate, judge). If the gateway enforces a policy that changes answers (abstain, regenerate, or route with a route configured) and your clients stream, set LATENT_GATEWAY_UNSUPPORTED_SHAPE=pass so streams reach the model and early stop.
  • Nothing automatic. Keep the Review queue preset: every flagged answer goes to your reviewers sorted by score, a length finish goes to review whatever its risk, alert rules fire, and no response is ever changed. Start here, and turn actions on once the flag rate looks right.

Was this page helpful?

Updated 4 October 2026