# Python SDK

Every option of the runlatent package: setup, per-client wrappers, environment, what is captured, streaming, redaction, counters and limits.

The `runlatent` package captures every finished call your app makes through the OpenAI, Anthropic or Google GenAI Python SDK and posts the prompt and answer to Latent's scoring API. Until you turn on an automatic response in the console, posting happens after the provider's SDK returns, on a background thread, so your app's call is never slower by a network round trip and never sees a different return value. It never gets an exception from Latent.

## Install

```bash
pip install runlatent
```

The package has no runtime dependencies. The provider SDKs are yours: whichever of `openai`, `anthropic` and `google-genai` is importable gets patched, and the rest are skipped.

## auto_instrument

```python
import runlatent

runlatent.auto_instrument()   # reads LATENT_API_URL and LATENT_API_KEY
```

Call it once at startup. The patch applies to the SDK classes, so every client in the process is covered, including clients a framework builds for you.

| Argument | Default | Meaning |
|---|---|---|
| `openai`, `anthropic`, `google_genai` | `True` | Which SDKs to patch |
| `api_url` | `$LATENT_API_URL` | The API base. The SDK posts to `<api_url>/score` |
| `api_key` | `$LATENT_API_KEY` | Sent as `Authorization: Bearer <key>` |
| `redact` | `None` | A function run on every pair before it is queued. See [Redaction](#redaction) |
| `timeout` | `5.0` | Seconds per post |
| `max_queue` | `10000` | Pairs waiting to be posted. A full queue drops new pairs and counts them |
| `enforce` | `None` | Apply the console policy ([below](#apply-your-console-policy)). `False` never changes a response |

Calling it twice is safe: nothing is patched twice. A second call with a different URL, key, timeout or queue size switches to the new settings and drains the old queue in the background. With no URL configured anywhere, `auto_instrument()` logs one warning, patches nothing and returns `stats()` with `enabled: False`.

`runlatent.uninstrument()` restores every original method and wrapped client.

## Hold an answer before your user sees it

`runlatent.check()` scores one answer and returns its verdict before you return the answer, so your code decides what your user gets. It needs runlatent 0.2.0 or later.

```python
import runlatent

r = runlatent.check(prompt, answer, context=docs)
if r.flagged:
    answer = SAFE_FALLBACK
elif not r.scored:
    ...  # Latent could not score it in time: deliver the answer, or hold it
```

- `context` takes your source documents as one string. `system`, `messages`, `model`, `timeout` (2 seconds by default) and `request_id` are optional.
- The result carries `risk`, `verdict` and `request_id`. `r.flagged` is true for an `escalate` or `ood` verdict, and `r.passed` for a scored `pass`.
- `check()` never raises when Latent refuses or times out. A 401, 403, 429, a 5xx or a timeout returns `scored=False` with `status` and `detail`, so you choose to fail open or fail closed. Pass `raise_on_error=True` to get an exception instead.
- With no URL or key configured, `check()` raises `ConfigurationError`.
- `acheck()` is the async version.
- With `auto_instrument()` on, a checked answer appears once in the console when you pass the provider's response id as `request_id` (for example `request_id=resp.id`) or the conversation as `messages=`. Without either, the check is scored on its own, so the answer appears twice and counts twice toward your plan.

## Apply your console policy

`auto_instrument()` also applies the automatic response you set on the console's Policy page. Until you turn one on, nothing changes: every call returns as the provider sent it, and the SDK only observes. With an automatic response on, each complete, single-choice answer is scored before the call returns, within the policy's wait (500 ms by default), and a flagged answer gets the action below. The SDK reads the policy every 60 seconds, so a change in the console reaches your app within a minute, and calls made before the first read are only observed.

| Action | What your app receives |
|---|---|
| `abstain` | The provider's own response object, with the policy's fallback text in place of the answer. Unless the policy sets one, the text is "I can't answer that reliably right now; a human will follow up." |
| `route` | The answer from the model set as the policy's `route_model`, on the same provider, through your own client and key. Without one, the answer is delivered unchanged with its verdict |
| `regenerate` | A second answer from the same model. It is delivered without a second check and scored afterwards as its own row |
| `judge` | The answer unchanged, queued for the judge's written analysis |

If the verdict does not arrive in time, `abstain` delivers the fallback text, and the other actions deliver the answer unchanged with the verdict `unknown`. If a regenerate or route call fails, the original answer is delivered. The policy's `fail_closed` setting can change either rule. Streams and requests for several answers are delivered unchanged, with the verdict recorded. The SDK never raises into your app, and every call returns the provider's own type. It covers OpenAI, Anthropic and Gemini clients. To keep the SDK observing only, pass `auto_instrument(enforce=False)` or set `LATENT_SDK_ENFORCE=off`. Needs runlatent 0.3.0 or later (`pip install -U runlatent`).

## Per-client wrappers

To capture one client and leave the rest of the process untouched:

```python
client = runlatent.wrap_openai(OpenAI())          # or AsyncOpenAI()
client = runlatent.wrap_anthropic(Anthropic())    # or AsyncAnthropic()
client = runlatent.wrap_genai(genai.Client())     # sync and .aio
```

Each returns the client. Wrapping twice does nothing more, and a wrapped client under a patched class is captured once.

## Environment

| Variable | Meaning |
|---|---|
| `LATENT_API_URL` | The API base: `https://app.runlatent.ai`, or your sidecar's address when Latent runs in your environment |
| `LATENT_API_KEY` | Your API key (`lk_...`), or the sidecar's token when Latent runs in your environment |
| `LATENT_SIDECAR_URL`, `LATENT_SIDECAR_TOKEN` | Used when the two variables above are unset |
| `LATENT_SDK_EXIT_FLUSH_S` | Seconds to wait at a normal exit for queued pairs. Default `2.0`; `0` turns the wait off |
| `LATENT_SDK_ENFORCE` | `off` keeps the SDK observing only, like `auto_instrument(enforce=False)`. An explicit `enforce=` wins |
| `LATENT_INFER_SESSION` | `0` stops the inferred `session_fingerprint` ([What Latent receives](/docs/sdk/what-latent-receives)) |

## What is captured

| Provider | Calls | Modes |
|---|---|---|
| OpenAI | Chat Completions and the Responses API, including `chat.completions.stream()` and `responses.stream()` | Sync, async and streaming |
| Anthropic | `messages.create` and `messages.stream`, including the `.stream()` manager | Sync, async and streaming |
| Google GenAI | `generate_content`, `generate_content_stream`, and chats through `chats.send_message` | Sync, async and streaming |

It also captures the Groq, Together, Fireworks, Mistral, Cohere, Azure AI Inference, Hugging Face `InferenceClient` and boto3 Bedrock SDKs.

Each finished call becomes one pair: the request's messages, system prompt included, and the whole answer text. The [What Latent receives](/docs/sdk/what-latent-receives) page lists every field. An answer with no text, such as a tool call only, is counted and not posted. A failed call posts nothing, and its exception reaches your app unchanged.

## Streaming

A stream comes back through a thin wrapper that yields the same items and passes every other attribute through. The pair is posted once, at the end:

- **Read to the end.** The chunks are joined into the full answer, with the final finish reason and token usage.
- **Closed early by your app** (the stream wrapper's own `close()`, or the end of a `with` block on it). The partial text is posted with `finish_reason: "abort"` and `client.stream_complete: false`. If no text had arrived, nothing is posted (`stream_abandoned_empty`).
- **Left unread** without closing. Nothing is posted; the stream is counted as `stream_unfinished`. Leaving OpenAI's `chat.completions.stream()` or `responses.stream()` helper early counts here too: its `with` exit closes the HTTP response and leaves the stream itself unread.
- **Failed mid-way,** for example an overloaded error, or an exception your app raised inside the `with` block. Nothing is posted; the stream is counted as `stream_errored`.

With `n` greater than 1, only the first choice or candidate is posted.

## Finish reasons

Each provider's finish reason is mapped to one vocabulary, so the console's review rules read the same way for every model. The provider's own word is kept alongside it.

| Posted | OpenAI Chat | OpenAI Responses | Anthropic | Gemini |
|---|---|---|---|---|
| `stop` | `stop` | `completed` | `end_turn`, `stop_sequence` | `STOP` |
| `length` | `length` | `incomplete:max_output_tokens` | `max_tokens`, `model_context_window_exceeded` | `MAX_TOKENS` |
| `tool_calls` | `tool_calls`, `function_call` | | `tool_use` | `MALFORMED_FUNCTION_CALL` |
| `content_filter` | `content_filter` | `incomplete:content_filter` | `refusal` | `SAFETY`, `RECITATION`, `BLOCKLIST`, `PROHIBITED_CONTENT`, `SPII`, `IMAGE_SAFETY`, `LANGUAGE` |
| `abort` | your app closed the stream early | `cancelled`, `in_progress`, `queued` | | |
| passed through | | | `pause_turn` | `OTHER` as `other`, `FINISH_REASON_UNSPECIFIED` as `unknown` |

A Responses answer with status `failed` and a Gemini prompt blocked before any answer are not posted: there is no answer to score (`skipped_failed_response`, `skipped_empty_output`). Any other word passes through lower-cased, cut to 32 characters. A missing finish reason is left unset, and the service marks it `finish_unknown`.

Under a new project's default review rule (`finish_to_review: ["length"]`), an answer cut off at the length limit goes to the review queue whatever its score. A deployment whose stored policy predates the rule reads `[]` until an admin sets it.

## Redaction

By default the SDK posts the text exactly as your app sent and received it. Masking changes what Latent reads, so mask only if you must, and mask the prompt and the answer with the same table of names:

```python
def redact(body):
    table = my_entities(body["messages"])        # names from the source, one table
    body["messages"] = [dict(m, content=mask(m["content"], table)) for m in body["messages"]]
    body["output"] = mask(body["output"], table)  # the same table on the answer
    return body                                   # or None to skip this pair

runlatent.auto_instrument(redact=redact)
```

The hook sees the full body and can rewrite it or return `None` to drop it. If the hook raises, or returns anything other than a dict or `None`, the pair is dropped and counted (`redact_failed`), and the unmasked pair never leaves the process.

With an automatic response of `abstain` on, an answer whose pair your hook drops or fails on is held, and your app gets the fallback text.

## Counters and shutdown

```python
runlatent.stats()
# {"enabled": True, "captured": 120, "posted": 118, "failed": 0, "dropped": 2, "queued": 0,
#  "skipped_empty_output": 3, "multimodal_parts_dropped": 7, "redact_skipped": 1,
#  "errors": {...}, ...}

runlatent.flush(timeout=5.0)   # in your shutdown path
```

- `captured` counts pairs assembled. `posted`, `failed`, `dropped` and `queued` describe delivery.
- `dropped` includes pairs captured while no destination was set.
- `skipped_*` and `stream_*` count calls that were not posted, and why.
- `errors` counts failures by stage and exception class. None of them ever reached your app. Only the class is logged, never the message, so answer text cannot leak into your logs through an error.

A retriable failure (a 5xx, a 429 or a network error) is tried up to three times, 0.5 seconds and then 2 seconds apart. A 4xx other than 429, such as a 401 for a bad key, or a redirect is final after one attempt. After a pair exhausts its attempts, the destination is held down for 5 seconds, and pairs in that window are lost without an attempt and counted. There is no on-disk spill: a pair that cannot be delivered is lost and counted.

## Limits

- `chat.completions.parse()`, `responses.parse()`, `beta.*`, and the Assistants and Realtime APIs are not captured.
- A pair over 1 MiB is refused (413) and counted as failed. A Free key scores up to 10 answers a second; above that the API answers 429 `rate_limited`.
- Pass arguments by keyword. A positional `messages` is not read.
- `with_raw_response.create()` and `with_streaming_response.create()` are not captured; they are counted as `skipped_raw_response`.
- The stream wrapper is not an instance of the SDK's `Stream` class, so `isinstance` checks on it fail. Attribute access works.
- `max_queue` counts pairs, behind one posting thread. 10,000 queued pairs with large retrieval prompts can hold about a gigabyte in memory, so lower it on memory-tight hosts.
- A `kill -9` or `os._exit()` loses queued pairs. Call `runlatent.flush()` in your shutdown path when that matters.

> Note: Gemini thinking models count their thought tokens against `max_output_tokens`. A small budget produces empty answers, which are not posted, and truncated answers, which are posted as `length` and, under a new project's default review rule, routed to review. Give thinking models a budget that covers the thinking, or turn thinking off.

> Note: OpenAI's `chat.completions.stream()` can raise `LengthFinishReasonError` after the stream has finished. The pair is still posted, as `length`, and your app still gets the exception.

---
Docs index: https://runlatent.ai/llms.txt
