# Quickstart

Install the Python SDK, add one line at startup, and see your first scored answer.

Latent scores every answer your app gets from OpenAI, Anthropic or Google Gemini. The Python SDK captures each finished call and sends the prompt and the answer to Latent's scoring API. Your app's calls never get an exception from Latent, and until you turn on an automatic response in the console they are never slowed by a network round trip and never see a different result.

## Before you start

- A Latent API key. Sign in to the [console](https://app.runlatent.ai/setup) with your email and password, an email link, Google or GitHub. After your first sign-in the console shows your project's key once, starting with `lk_`, with the install line filled in. To make another later, open Settings > Project > API keys. A key lasts one year, and your admins get an email 14 days before it expires.
- A Python app that calls a model through the `openai`, `anthropic` or `google-genai` Python SDK.

## Set up with your coding agent

Paste this into Claude Code, Codex, Cursor or another coding agent working in your project:

```text
Add Latent hallucination scoring to this project.
Read https://runlatent.ai/docs/quickstart and https://runlatent.ai/docs/sdk before changing any code.
1. Install the runlatent package from PyPI (pip install runlatent) and add it to the project's dependencies.
2. Read LATENT_API_URL and LATENT_API_KEY from the project's environment configuration. Add both to the example env file if there is one, with LATENT_API_URL=https://app.runlatent.ai. Do not write a real key into the code.
3. In the app's entry point, import runlatent and call runlatent.auto_instrument() once at startup, before any OpenAI, Anthropic or Google GenAI client is created, including an openai client pointed at another endpoint by base_url.
4. If LATENT_API_KEY is not set, stop here and ask me to create a key at https://app.runlatent.ai and add it to the project's environment myself. Never ask for the key in this chat.
5. Run the app once so it makes one model call, then call runlatent.flush() and print runlatent.stats(). Confirm that "posted" is at least 1 and "failed" is 0.
Change nothing else.
```

The agent's change is three lines and two environment variables. The [coding agent guide](/docs/coding-agent) explains each step.

## Set up by hand

### 1. Install the SDK

```bash
pip install runlatent
```

The package has no dependencies of its own. It patches whichever of `openai`, `anthropic` and `google-genai` your app already uses.

### 2. Set your API key

```bash
export LATENT_API_URL=https://app.runlatent.ai
export LATENT_API_KEY=lk_...
```

Nothing is sent while neither `LATENT_API_URL` nor its fallback `LATENT_SIDECAR_URL` is set: `auto_instrument()` logs one warning and patches nothing. So you can ship the line below everywhere and turn Latent on per environment.

### 3. Add one line at startup

```python
import runlatent

runlatent.auto_instrument()
```

That covers every OpenAI, Anthropic and Google GenAI client in the process, including clients a framework builds for you. Your model calls stay exactly as they are:

```python
from openai import OpenAI

client = OpenAI()
answer = client.chat.completions.create(
    model="gpt-4.1",
    messages=[{"role": "user", "content": "What is our refund window?"}],
)
```

### 4. Verify

Run the app so it makes at least one model call. Posting runs on a background thread, so flush the queue before you read the SDK's counters:

```python
runlatent.flush()
print(runlatent.stats())
# {"enabled": True, "captured": 1, "posted": 1, "failed": 0, "dropped": 0, ...}
```

`posted` counts the answers delivered to Latent. Each posted answer appears in your project's [console](https://app.runlatent.ai) with its risk score.

For a short script that exits right after its last call, the SDK waits up to 2 seconds at exit for queued answers. In a server, call `runlatent.flush()` in your shutdown path.

## Troubleshooting

**`enabled` is `False`.** Neither `LATENT_API_URL` nor `LATENT_SIDECAR_URL` is set in the process that makes the calls.

**`captured` grows but `posted` stays at 0.** A rejected or unreachable post increments `failed`, and the SDK logs one warning per status every 60 seconds naming the URL and the HTTP status. A 401 means the key is unknown, revoked or expired: create a new one in the console. A 4xx other than 429 is final after one attempt. `stats()["errors"]` is a different channel: it counts failures inside your process by stage, such as `capture:`, `redact:` and `stream_feed:`.

**`patched` is `[]`.** No supported SDK could be imported in this process, and one warning is logged.

**`failed` grows after your plan's included answers.** Free scores up to 5,000 answers a month. Pro and Team keep scoring past their included volume with usage billing, which is on from the upgrade, and pause there if an admin turns it off. Above the limit, the scoring API answers 429 `limit_reached`, and its `reason` says which limit: Free's monthly cap, the included volume, or your usage-billing spend limit. The SDK counts each pair as failed after three tries.

**Nothing is captured.** Pass arguments by keyword, as the provider SDKs require. `with_raw_response`, `with_streaming_response` and OpenAI's `parse()` are not captured; see [Python SDK limits](/docs/sdk#limits).

> Note: Gemini thinking models count their thought tokens against `max_output_tokens`. A small budget produces empty answers, which are not posted, and truncated answers, which are posted as `length` and, under a new project's default review rule, routed to review. Give thinking models a budget that covers the thinking, or turn thinking off, before you read the review queue as a quality signal.

> Note: OpenAI's `chat.completions.stream()` can raise `LengthFinishReasonError` after the stream has finished. The answer is still sent to Latent, marked as cut off at the length limit, and your app still gets the exception.

## Next steps

- [Hold an answer before your user sees it](/docs/sdk#hold-an-answer-before-your-user-sees-it): one call returns the verdict before you return the answer.
- [Python SDK reference](/docs/sdk): every option, per-client wrappers, streaming and shutdown.
- [What Latent receives](/docs/sdk/what-latent-receives): the exact fields sent for each call, and what is never sent.
- [Calibration](/docs/concepts/calibration): fit Latent to your traffic once answers are flowing.

---
Docs index: https://runlatent.ai/llms.txt
