ReferencePython SDK

Every option of the runlatent package: setup, per-client wrappers, environment, what is captured, streaming, redaction, counters and limits.

The runlatent package captures every finished call your app makes through the OpenAI, Anthropic or Google GenAI Python SDK and posts the prompt and answer to Latent's scoring API. Until you turn on an automatic response in the console, posting happens after the provider's SDK returns, on a background thread, so your app's call is never slower by a network round trip and never sees a different return value. It never gets an exception from Latent.

Install

pip install runlatent

The package has no runtime dependencies. The provider SDKs are yours: whichever of openai, anthropic and google-genai is importable gets patched, and the rest are skipped.

auto_instrument

import runlatent

runlatent.auto_instrument()   # reads LATENT_API_URL and LATENT_API_KEY

Call it once at startup. The patch applies to the SDK classes, so every client in the process is covered, including clients a framework builds for you.

Argument Default Meaning
openai, anthropic, google_genai True Which SDKs to patch
api_url $LATENT_API_URL The API base. The SDK posts to <api_url>/score
api_key $LATENT_API_KEY Sent as Authorization: Bearer <key>
redact None A function run on every pair before it is queued. See Redaction
timeout 5.0 Seconds per post
max_queue 10000 Pairs waiting to be posted. A full queue drops new pairs and counts them
enforce None Apply the console policy (below). False never changes a response

Calling it twice is safe: nothing is patched twice. A second call with a different URL, key, timeout or queue size switches to the new settings and drains the old queue in the background. With no URL configured anywhere, auto_instrument() logs one warning, patches nothing and returns stats() with enabled: False.

runlatent.uninstrument() restores every original method and wrapped client.

Hold an answer before your user sees it

runlatent.check() scores one answer and returns its verdict before you return the answer, so your code decides what your user gets. It needs runlatent 0.2.0 or later.

import runlatent

r = runlatent.check(prompt, answer, context=docs)
if r.flagged:
    answer = SAFE_FALLBACK
elif not r.scored:
    ...  # Latent could not score it in time: deliver the answer, or hold it
  • context takes your source documents as one string. system, messages, model, timeout (2 seconds by default) and request_id are optional.
  • The result carries risk, verdict and request_id. r.flagged is true for an escalate or ood verdict, and r.passed for a scored pass.
  • check() never raises when Latent refuses or times out. A 401, 403, 429, a 5xx or a timeout returns scored=False with status and detail, so you choose to fail open or fail closed. Pass raise_on_error=True to get an exception instead.
  • With no URL or key configured, check() raises ConfigurationError.
  • acheck() is the async version.
  • With auto_instrument() on, a checked answer appears once in the console when you pass the provider's response id as request_id (for example request_id=resp.id) or the conversation as messages=. Without either, the check is scored on its own, so the answer appears twice and counts twice toward your plan.

Apply your console policy

auto_instrument() also applies the automatic response you set on the console's Policy page. Until you turn one on, nothing changes: every call returns as the provider sent it, and the SDK only observes. With an automatic response on, each complete, single-choice answer is scored before the call returns, within the policy's wait (500 ms by default), and a flagged answer gets the action below. The SDK reads the policy every 60 seconds, so a change in the console reaches your app within a minute, and calls made before the first read are only observed.

Action What your app receives
abstain The provider's own response object, with the policy's fallback text in place of the answer. Unless the policy sets one, the text is "I can't answer that reliably right now; a human will follow up."
route The answer from the model set as the policy's route_model, on the same provider, through your own client and key. Without one, the answer is delivered unchanged with its verdict
regenerate A second answer from the same model. It is delivered without a second check and scored afterwards as its own row
judge The answer unchanged, queued for the judge's written analysis

If the verdict does not arrive in time, abstain delivers the fallback text, and the other actions deliver the answer unchanged with the verdict unknown. If a regenerate or route call fails, the original answer is delivered. The policy's fail_closed setting can change either rule. Streams and requests for several answers are delivered unchanged, with the verdict recorded. The SDK never raises into your app, and every call returns the provider's own type. It covers OpenAI, Anthropic and Gemini clients. To keep the SDK observing only, pass auto_instrument(enforce=False) or set LATENT_SDK_ENFORCE=off. Needs runlatent 0.3.0 or later (pip install -U runlatent).

Per-client wrappers

To capture one client and leave the rest of the process untouched:

client = runlatent.wrap_openai(OpenAI())          # or AsyncOpenAI()
client = runlatent.wrap_anthropic(Anthropic())    # or AsyncAnthropic()
client = runlatent.wrap_genai(genai.Client())     # sync and .aio

Each returns the client. Wrapping twice does nothing more, and a wrapped client under a patched class is captured once.

Environment

Variable Meaning
LATENT_API_URL The API base: https://app.runlatent.ai, or your sidecar's address when Latent runs in your environment
LATENT_API_KEY Your API key (lk_...), or the sidecar's token when Latent runs in your environment
LATENT_SIDECAR_URL, LATENT_SIDECAR_TOKEN Used when the two variables above are unset
LATENT_SDK_EXIT_FLUSH_S Seconds to wait at a normal exit for queued pairs. Default 2.0; 0 turns the wait off
LATENT_SDK_ENFORCE off keeps the SDK observing only, like auto_instrument(enforce=False). An explicit enforce= wins
LATENT_INFER_SESSION 0 stops the inferred session_fingerprint (What Latent receives)

What is captured

Provider Calls Modes
OpenAI Chat Completions and the Responses API, including chat.completions.stream() and responses.stream() Sync, async and streaming
Anthropic messages.create and messages.stream, including the .stream() manager Sync, async and streaming
Google GenAI generate_content, generate_content_stream, and chats through chats.send_message Sync, async and streaming

It also captures the Groq, Together, Fireworks, Mistral, Cohere, Azure AI Inference, Hugging Face InferenceClient and boto3 Bedrock SDKs.

Each finished call becomes one pair: the request's messages, system prompt included, and the whole answer text. The What Latent receives page lists every field. An answer with no text, such as a tool call only, is counted and not posted. A failed call posts nothing, and its exception reaches your app unchanged.

Streaming

A stream comes back through a thin wrapper that yields the same items and passes every other attribute through. The pair is posted once, at the end:

  • Read to the end. The chunks are joined into the full answer, with the final finish reason and token usage.
  • Closed early by your app (the stream wrapper's own close(), or the end of a with block on it). The partial text is posted with finish_reason: "abort" and client.stream_complete: false. If no text had arrived, nothing is posted (stream_abandoned_empty).
  • Left unread without closing. Nothing is posted; the stream is counted as stream_unfinished. Leaving OpenAI's chat.completions.stream() or responses.stream() helper early counts here too: its with exit closes the HTTP response and leaves the stream itself unread.
  • Failed mid-way, for example an overloaded error, or an exception your app raised inside the with block. Nothing is posted; the stream is counted as stream_errored.

With n greater than 1, only the first choice or candidate is posted.

Finish reasons

Each provider's finish reason is mapped to one vocabulary, so the console's review rules read the same way for every model. The provider's own word is kept alongside it.

Posted OpenAI Chat OpenAI Responses Anthropic Gemini
stop stop completed end_turn, stop_sequence STOP
length length incomplete:max_output_tokens max_tokens, model_context_window_exceeded MAX_TOKENS
tool_calls tool_calls, function_call tool_use MALFORMED_FUNCTION_CALL
content_filter content_filter incomplete:content_filter refusal SAFETY, RECITATION, BLOCKLIST, PROHIBITED_CONTENT, SPII, IMAGE_SAFETY, LANGUAGE
abort your app closed the stream early cancelled, in_progress, queued
passed through pause_turn OTHER as other, FINISH_REASON_UNSPECIFIED as unknown

A Responses answer with status failed and a Gemini prompt blocked before any answer are not posted: there is no answer to score (skipped_failed_response, skipped_empty_output). Any other word passes through lower-cased, cut to 32 characters. A missing finish reason is left unset, and the service marks it finish_unknown.

Under a new project's default review rule (finish_to_review: ["length"]), an answer cut off at the length limit goes to the review queue whatever its score. A deployment whose stored policy predates the rule reads [] until an admin sets it.

Redaction

By default the SDK posts the text exactly as your app sent and received it. Masking changes what Latent reads, so mask only if you must, and mask the prompt and the answer with the same table of names:

def redact(body):
    table = my_entities(body["messages"])        # names from the source, one table
    body["messages"] = [dict(m, content=mask(m["content"], table)) for m in body["messages"]]
    body["output"] = mask(body["output"], table)  # the same table on the answer
    return body                                   # or None to skip this pair

runlatent.auto_instrument(redact=redact)

The hook sees the full body and can rewrite it or return None to drop it. If the hook raises, or returns anything other than a dict or None, the pair is dropped and counted (redact_failed), and the unmasked pair never leaves the process.

With an automatic response of abstain on, an answer whose pair your hook drops or fails on is held, and your app gets the fallback text.

Counters and shutdown

runlatent.stats()
# {"enabled": True, "captured": 120, "posted": 118, "failed": 0, "dropped": 2, "queued": 0,
#  "skipped_empty_output": 3, "multimodal_parts_dropped": 7, "redact_skipped": 1,
#  "errors": {...}, ...}

runlatent.flush(timeout=5.0)   # in your shutdown path
  • captured counts pairs assembled. posted, failed, dropped and queued describe delivery.
  • dropped includes pairs captured while no destination was set.
  • skipped_* and stream_* count calls that were not posted, and why.
  • errors counts failures by stage and exception class. None of them ever reached your app. Only the class is logged, never the message, so answer text cannot leak into your logs through an error.

A retriable failure (a 5xx, a 429 or a network error) is tried up to three times, 0.5 seconds and then 2 seconds apart. A 4xx other than 429, such as a 401 for a bad key, or a redirect is final after one attempt. After a pair exhausts its attempts, the destination is held down for 5 seconds, and pairs in that window are lost without an attempt and counted. There is no on-disk spill: a pair that cannot be delivered is lost and counted.

Limits

  • chat.completions.parse(), responses.parse(), beta.*, and the Assistants and Realtime APIs are not captured.
  • A pair over 1 MiB is refused (413) and counted as failed. A Free key scores up to 10 answers a second; above that the API answers 429 rate_limited.
  • Pass arguments by keyword. A positional messages is not read.
  • with_raw_response.create() and with_streaming_response.create() are not captured; they are counted as skipped_raw_response.
  • The stream wrapper is not an instance of the SDK's Stream class, so isinstance checks on it fail. Attribute access works.
  • max_queue counts pairs, behind one posting thread. 10,000 queued pairs with large retrieval prompts can hold about a gigabyte in memory, so lower it on memory-tight hosts.
  • A kill -9 or os._exit() loses queued pairs. Call runlatent.flush() in your shutdown path when that matters.

Note Gemini thinking models count their thought tokens against max_output_tokens. A small budget produces empty answers, which are not posted, and truncated answers, which are posted as length and, under a new project's default review rule, routed to review. Give thinking models a budget that covers the thinking, or turn thinking off.

Note OpenAI's chat.completions.stream() can raise LengthFinishReasonError after the stream has finished. The pair is still posted, as length, and your app still gets the exception.

Was this page helpful?

Updated 3 October 2026