Get startedQuickstart

Install the Python SDK, add one line at startup, and see your first scored answer.

Latent scores every answer your app gets from OpenAI, Anthropic or Google Gemini. The Python SDK captures each finished call and sends the prompt and the answer to Latent's scoring API. Your app's calls never get an exception from Latent, and until you turn on an automatic response in the console they are never slowed by a network round trip and never see a different result.

Before you start

  • A Latent API key. Sign in to the console with your email and password, an email link, Google or GitHub. After your first sign-in the console shows your project's key once, starting with lk_, with the install line filled in. To make another later, open Settings > Project > API keys. A key lasts one year, and your admins get an email 14 days before it expires.
  • A Python app that calls a model through the openai, anthropic or google-genai Python SDK.

Set up with your coding agent

Paste this into Claude Code, Codex, Cursor or another coding agent working in your project:

Add Latent hallucination scoring to this project.
Read https://runlatent.ai/docs/quickstart and https://runlatent.ai/docs/sdk before changing any code.
1. Install the runlatent package from PyPI (pip install runlatent) and add it to the project's dependencies.
2. Read LATENT_API_URL and LATENT_API_KEY from the project's environment configuration. Add both to the example env file if there is one, with LATENT_API_URL=https://app.runlatent.ai. Do not write a real key into the code.
3. In the app's entry point, import runlatent and call runlatent.auto_instrument() once at startup, before any OpenAI, Anthropic or Google GenAI client is created, including an openai client pointed at another endpoint by base_url.
4. If LATENT_API_KEY is not set, stop here and ask me to create a key at https://app.runlatent.ai and add it to the project's environment myself. Never ask for the key in this chat.
5. Run the app once so it makes one model call, then call runlatent.flush() and print runlatent.stats(). Confirm that "posted" is at least 1 and "failed" is 0.
Change nothing else.

The agent's change is three lines and two environment variables. The coding agent guide explains each step.

Set up by hand

1. Install the SDK

pip install runlatent

The package has no dependencies of its own. It patches whichever of openai, anthropic and google-genai your app already uses.

2. Set your API key

export LATENT_API_URL=https://app.runlatent.ai
export LATENT_API_KEY=lk_...

Nothing is sent while neither LATENT_API_URL nor its fallback LATENT_SIDECAR_URL is set: auto_instrument() logs one warning and patches nothing. So you can ship the line below everywhere and turn Latent on per environment.

3. Add one line at startup

import runlatent

runlatent.auto_instrument()

That covers every OpenAI, Anthropic and Google GenAI client in the process, including clients a framework builds for you. Your model calls stay exactly as they are:

from openai import OpenAI

client = OpenAI()
answer = client.chat.completions.create(
    model="gpt-4.1",
    messages=[{"role": "user", "content": "What is our refund window?"}],
)

4. Verify

Run the app so it makes at least one model call. Posting runs on a background thread, so flush the queue before you read the SDK's counters:

runlatent.flush()
print(runlatent.stats())
# {"enabled": True, "captured": 1, "posted": 1, "failed": 0, "dropped": 0, ...}

posted counts the answers delivered to Latent. Each posted answer appears in your project's console with its risk score.

For a short script that exits right after its last call, the SDK waits up to 2 seconds at exit for queued answers. In a server, call runlatent.flush() in your shutdown path.

Troubleshooting

enabled is False. Neither LATENT_API_URL nor LATENT_SIDECAR_URL is set in the process that makes the calls.

captured grows but posted stays at 0. A rejected or unreachable post increments failed, and the SDK logs one warning per status every 60 seconds naming the URL and the HTTP status. A 401 means the key is unknown, revoked or expired: create a new one in the console. A 4xx other than 429 is final after one attempt. stats()["errors"] is a different channel: it counts failures inside your process by stage, such as capture:, redact: and stream_feed:.

patched is []. No supported SDK could be imported in this process, and one warning is logged.

failed grows after your plan's included answers. Free scores up to 5,000 answers a month. Pro and Team keep scoring past their included volume with usage billing, which is on from the upgrade, and pause there if an admin turns it off. Above the limit, the scoring API answers 429 limit_reached, and its reason says which limit: Free's monthly cap, the included volume, or your usage-billing spend limit. The SDK counts each pair as failed after three tries.

Nothing is captured. Pass arguments by keyword, as the provider SDKs require. with_raw_response, with_streaming_response and OpenAI's parse() are not captured; see Python SDK limits.

Note Gemini thinking models count their thought tokens against max_output_tokens. A small budget produces empty answers, which are not posted, and truncated answers, which are posted as length and, under a new project's default review rule, routed to review. Give thinking models a budget that covers the thinking, or turn thinking off, before you read the review queue as a quality signal.

Note OpenAI's chat.completions.stream() can raise LengthFinishReasonError after the stream has finished. The answer is still sent to Latent, marked as cut off at the length limit, and your app still gets the exception.

Next steps

Was this page helpful?

Updated 3 October 2026