Reference
Python SDK
Every option of the runlatent package: setup, per-client wrappers, environment, what is captured, streaming, redaction, counters and limits.
The runlatent package captures every finished call your app makes through the OpenAI, Anthropic or Google GenAI Python SDK and posts the prompt and answer to Latent's scoring API. Until you turn on an automatic response in the console, posting happens after the provider's SDK returns, on a background thread, so your app's call is never slower by a network round trip and never sees a different return value. It never gets an exception from Latent.
Install
pip install runlatent
The package has no runtime dependencies. The provider SDKs are yours: whichever of openai, anthropic and google-genai is importable gets patched, and the rest are skipped.
auto_instrument
import runlatent
runlatent.auto_instrument() # reads LATENT_API_URL and LATENT_API_KEY
Call it once at startup. The patch applies to the SDK classes, so every client in the process is covered, including clients a framework builds for you.
| Argument | Default | Meaning |
|---|---|---|
openai, anthropic, google_ |
True |
Which SDKs to patch |
api_url |
$LATENT_ |
The API base. The SDK posts to <api_ |
api_key |
$LATENT_ |
Sent as Authorization: Bearer <key> |
redact |
None |
A function run on every pair before it is queued. See Redaction |
timeout |
5.0 |
Seconds per post |
max_queue |
10000 |
Pairs waiting to be posted. A full queue drops new pairs and counts them |
enforce |
None |
Apply the console policy (below). False never changes a response |
Calling it twice is safe: nothing is patched twice. A second call with a different URL, key, timeout or queue size switches to the new settings and drains the old queue in the background. With no URL configured anywhere, auto_ logs one warning, patches nothing and returns stats() with enabled: False.
runlatent. restores every original method and wrapped client.
Hold an answer before your user sees it
runlatent. scores one answer and returns its verdict before you return the answer, so your code decides what your user gets. It needs runlatent 0.2.0 or later.
import runlatent
r = runlatent.check(prompt, answer, context=docs)
if r.flagged:
answer = SAFE_FALLBACK
elif not r.scored:
... # Latent could not score it in time: deliver the answer, or hold it
contexttakes your source documents as one string.system,messages,model,timeout(2 seconds by default) andrequest_idare optional.- The result carries
risk,verdictandrequest_id.r.flaggedis true for anescalateoroodverdict, andr.passedfor a scoredpass. check()never raises when Latent refuses or times out. A 401, 403, 429, a 5xx or a timeout returnsscored=Falsewithstatusanddetail, so you choose to fail open or fail closed. Passraise_to get an exception instead.on_ error=True - With no URL or key configured,
check()raisesConfigurationError. acheck()is the async version.- With
auto_on, a checked answer appears once in the console when you pass the provider's response id asinstrument() request_id(for examplerequest_) or the conversation asid=resp. id messages=. Without either, the check is scored on its own, so the answer appears twice and counts twice toward your plan.
Apply your console policy
auto_ also applies the automatic response you set on the console's Policy page. Until you turn one on, nothing changes: every call returns as the provider sent it, and the SDK only observes. With an automatic response on, each complete, single-choice answer is scored before the call returns, within the policy's wait (500 ms by default), and a flagged answer gets the action below. The SDK reads the policy every 60 seconds, so a change in the console reaches your app within a minute, and calls made before the first read are only observed.
| Action | What your app receives |
|---|---|
abstain |
The provider's own response object, with the policy's fallback text in place of the answer. Unless the policy sets one, the text is "I can't answer that reliably right now; a human will follow up." |
route |
The answer from the model set as the policy's route_, on the same provider, through your own client and key. Without one, the answer is delivered unchanged with its verdict |
regenerate |
A second answer from the same model. It is delivered without a second check and scored afterwards as its own row |
judge |
The answer unchanged, queued for the judge's written analysis |
If the verdict does not arrive in time, abstain delivers the fallback text, and the other actions deliver the answer unchanged with the verdict unknown. If a regenerate or route call fails, the original answer is delivered. The policy's fail_ setting can change either rule. Streams and requests for several answers are delivered unchanged, with the verdict recorded. The SDK never raises into your app, and every call returns the provider's own type. It covers OpenAI, Anthropic and Gemini clients. To keep the SDK observing only, pass auto_ or set LATENT_. Needs runlatent 0.3.0 or later (pip install -U runlatent).
Per-client wrappers
To capture one client and leave the rest of the process untouched:
client = runlatent.wrap_openai(OpenAI()) # or AsyncOpenAI()
client = runlatent.wrap_anthropic(Anthropic()) # or AsyncAnthropic()
client = runlatent.wrap_genai(genai.Client()) # sync and .aio
Each returns the client. Wrapping twice does nothing more, and a wrapped client under a patched class is captured once.
Environment
| Variable | Meaning |
|---|---|
LATENT_ |
The API base: https://, or your sidecar's address when Latent runs in your environment |
LATENT_ |
Your API key (lk_...), or the sidecar's token when Latent runs in your environment |
LATENT_, LATENT_ |
Used when the two variables above are unset |
LATENT_ |
Seconds to wait at a normal exit for queued pairs. Default 2.0; 0 turns the wait off |
LATENT_ |
off keeps the SDK observing only, like auto_. An explicit enforce= wins |
LATENT_ |
0 stops the inferred session_ (What Latent receives) |
What is captured
| Provider | Calls | Modes |
|---|---|---|
| OpenAI | Chat Completions and the Responses API, including chat. and responses. |
Sync, async and streaming |
| Anthropic | messages. and messages., including the .stream() manager |
Sync, async and streaming |
| Google GenAI | generate_, generate_, and chats through chats. |
Sync, async and streaming |
It also captures the Groq, Together, Fireworks, Mistral, Cohere, Azure AI Inference, Hugging Face InferenceClient and boto3 Bedrock SDKs.
Each finished call becomes one pair: the request's messages, system prompt included, and the whole answer text. The What Latent receives page lists every field. An answer with no text, such as a tool call only, is counted and not posted. A failed call posts nothing, and its exception reaches your app unchanged.
Streaming
A stream comes back through a thin wrapper that yields the same items and passes every other attribute through. The pair is posted once, at the end:
- Read to the end. The chunks are joined into the full answer, with the final finish reason and token usage.
- Closed early by your app (the stream wrapper's own
close(), or the end of awithblock on it). The partial text is posted withfinish_andreason: "abort" client.. If no text had arrived, nothing is posted (stream_ complete: false stream_).abandoned_ empty - Left unread without closing. Nothing is posted; the stream is counted as
stream_. Leaving OpenAI'sunfinished chat.orcompletions. stream() responses.helper early counts here too: itsstream() withexit closes the HTTP response and leaves the stream itself unread. - Failed mid-way, for example an overloaded error, or an exception your app raised inside the
withblock. Nothing is posted; the stream is counted asstream_.errored
With n greater than 1, only the first choice or candidate is posted.
Finish reasons
Each provider's finish reason is mapped to one vocabulary, so the console's review rules read the same way for every model. The provider's own word is kept alongside it.
| Posted | OpenAI Chat | OpenAI Responses | Anthropic | Gemini |
|---|---|---|---|---|
stop |
stop |
completed |
end_turn, stop_ |
STOP |
length |
length |
incomplete:max_ |
max_tokens, model_ |
MAX_TOKENS |
tool_calls |
tool_calls, function_ |
tool_use |
MALFORMED_ |
|
content_ |
content_ |
incomplete:content_ |
refusal |
SAFETY, RECITATION, BLOCKLIST, PROHIBITED_, SPII, IMAGE_, LANGUAGE |
abort |
your app closed the stream early | cancelled, in_, queued |
||
| passed through | pause_turn |
OTHER as other, FINISH_ as unknown |
A Responses answer with status failed and a Gemini prompt blocked before any answer are not posted: there is no answer to score (skipped_, skipped_). Any other word passes through lower-cased, cut to 32 characters. A missing finish reason is left unset, and the service marks it finish_.
Under a new project's default review rule (finish_), an answer cut off at the length limit goes to the review queue whatever its score. A deployment whose stored policy predates the rule reads [] until an admin sets it.
Redaction
By default the SDK posts the text exactly as your app sent and received it. Masking changes what Latent reads, so mask only if you must, and mask the prompt and the answer with the same table of names:
def redact(body):
table = my_entities(body["messages"]) # names from the source, one table
body["messages"] = [dict(m, content=mask(m["content"], table)) for m in body["messages"]]
body["output"] = mask(body["output"], table) # the same table on the answer
return body # or None to skip this pair
runlatent.auto_instrument(redact=redact)
The hook sees the full body and can rewrite it or return None to drop it. If the hook raises, or returns anything other than a dict or None, the pair is dropped and counted (redact_), and the unmasked pair never leaves the process.
With an automatic response of abstain on, an answer whose pair your hook drops or fails on is held, and your app gets the fallback text.
Counters and shutdown
runlatent.stats()
# {"enabled": True, "captured": 120, "posted": 118, "failed": 0, "dropped": 2, "queued": 0,
# "skipped_empty_output": 3, "multimodal_parts_dropped": 7, "redact_skipped": 1,
# "errors": {...}, ...}
runlatent.flush(timeout=5.0) # in your shutdown path
capturedcounts pairs assembled.posted,failed,droppedandqueueddescribe delivery.droppedincludes pairs captured while no destination was set.skipped_*andstream_*count calls that were not posted, and why.errorscounts failures by stage and exception class. None of them ever reached your app. Only the class is logged, never the message, so answer text cannot leak into your logs through an error.
A retriable failure (a 5xx, a 429 or a network error) is tried up to three times, 0.5 seconds and then 2 seconds apart. A 4xx other than 429, such as a 401 for a bad key, or a redirect is final after one attempt. After a pair exhausts its attempts, the destination is held down for 5 seconds, and pairs in that window are lost without an attempt and counted. There is no on-disk spill: a pair that cannot be delivered is lost and counted.
Limits
chat.,completions. parse() responses.,parse() beta.*, and the Assistants and Realtime APIs are not captured.- A pair over 1 MiB is refused (413) and counted as failed. A Free key scores up to 10 answers a second; above that the API answers 429
rate_.limited - Pass arguments by keyword. A positional
messagesis not read. with_andraw_ response. create() with_are not captured; they are counted asstreaming_ response. create() skipped_.raw_ response - The stream wrapper is not an instance of the SDK's
Streamclass, soisinstancechecks on it fail. Attribute access works. max_queuecounts pairs, behind one posting thread. 10,000 queued pairs with large retrieval prompts can hold about a gigabyte in memory, so lower it on memory-tight hosts.- A
kill -9oros._exit()loses queued pairs. Callrunlatent.in your shutdown path when that matters.flush()
Note Gemini thinking models count their thought tokens against
max_. A small budget produces empty answers, which are not posted, and truncated answers, which are posted asoutput_ tokens lengthand, under a new project's default review rule, routed to review. Give thinking models a budget that covers the thinking, or turn thinking off.
Note OpenAI's
chat.can raisecompletions. stream() LengthFinishReasonErrorafter the stream has finished. The pair is still posted, aslength, and your app still gets the exception.
Was this page helpful?
Updated 3 October 2026