Latent

Stop hallucinations before they reach your customers.

Latent reads your model's internal state as it writes. Made-up answers are held before a customer sees them.

Backed by Y Combinator

Watch Latent catch
a made-up answer.

As your model writes each word, Latent reads what is happening inside it. When the answer is finished, it sends it where your policy says. Pick an answer and a policy, and watch it run.

Pick an answer
When Latent flags it
Customer asks

Latent is reading·
Your model
The answer

Customer sees
Your stronger model
Standing by
Review queue
Nothing waiting

 

Why not have another model check every answer?

You can. It is slow, it is expensive, and every answer has to leave your environment to be checked. Latent reads the model that wrote the answer, so it checks output validity in real time.

Answers checked in the same time103× faster
A judge model0 checked
Latent, API models0 checked
Latent, your own modelas it writes

On your own model, Latent reads each answer while it is written, so the check adds no measurable time.

Cost per 1,000 answers530× cheaper
A judge model$10.61
Latent, API models$0.02
Latent, your own modelno extra cost
Want a judge anyway?

Let Latent choose which answers the judge reads. The judge bill falls tenfold, and more of what gets flagged is really wrong.

10×smaller judge bill
$1.06 against $10.61 per 1,000
93%of flags correct
against 82% judging everything

A layer the rest of your stack can't see.

Judges, guardrails and dashboards look at what your model said. Latent reads the model while it writes, and it works alongside all of them if your team wants to keep them.

Works with what you already run

Latent sits at runtime, between the input and the answer, where the rest of your stack can't see. Keep your judge, guardrails and traces on the output. Latent feeds them through a Prometheus endpoint, OpenTelemetry trace ids and labeled answers for fine-tuning.

See why it flagged

Each flag points to the sentence and the words that drove it, the sentence in your source the answer leaned on most, and the kind of failure, such as an unsupported number, a wrong date or an invented name. Your reviewers know where to look.

Alerts where your team already works

Slack, PagerDuty, email or a webhook, when flags spike or stop, confirmed accuracy drops, scoring slows, the review queue backs up or your traffic drifts.

From one line of code to a working review queue.

  1. 01Add the SDK

    One line at startup sends each finished answer from your OpenAI, Anthropic or Gemini calls to Latent. Your calls go to your provider exactly as before.

  2. 02Collect your first answers

    Latent scores every answer from the first call, and keeps a sample of what your model actually served for calibration.

  3. 03Calibrate on your traffic

    Latent runs calibration for you. A judge labels the sample against your sources, Latent tunes its threshold and its read to your traffic, and the update goes live with nothing to do on your side.

  4. 04Review what it flags

    Flagged answers wait in a queue with the words that gave them away. Your reviewers confirm the failure or clear it, and each decision is kept as a labeled example.

Latent
Finance assistant
pip install runlatent ✔ Successfully installed runlatent-0.3.0export LATENT_API_URL=https://app.runlatent.aiexport LATENT_API_KEY=lk_••••••••••••# app.py, once at startupimport runlatentrunlatent.auto_instrument()# after the first model callrunlatent.flush(); print(runlatent.stats()){"enabled": True, "captured": 1, "posted": 1, "failed": 0, ...}
Off-threadposts after each call returns, so no network round trip on your request path
Fail-openSDK errors are counted in stats() and never raised into your app
Zero depspure Python, covering sync, async and streaming calls
Sample from a finance assistant0 of 1,000 answers

    Latent keeps this sample's text for calibration under your retention settings.

    calibration run · 1,000 labeled answers · run by Latent

    0%of the answers Latent flags are really wrong, checking other models' answers
    Checked at random0%
    Flagged by Latent0%
    Under an hourfor the whole run
    Judge includedno key of yours needed
    Nothingto do on your side to go live

    Latent's reading, word by word

    Every review makes it better.

    Each decision your reviewers make is kept as a labeled example. Latent learns from them, and the more answers your team validates, the better the data you post-train your model on.

    FlaggedLatent holds an answer
    ReviewedConfirmed or released
    LearnedKept as a labeled example
    Flag accuracyHow often a flagged answer is really wrong
    At install
    Training data for your modelValidated answers, ready to export
    Recalibrating live, no restart

    Latent recalibrates on your reviewers' decisions while it runs, with no restart. Every validated answer can also be exported as fine-tuning or preference data for your model.

    In your VPC, or hosted by us.

    Security and deployment detail
    Hosted by Latent
    Your appwith the runlatent SDK
    Latent readerscores each finished answer
    Review consolequeue, alerts and history
    Your model serves exactly as it does today. After each call returns, the SDK sends the prompt and answer to Latent's reader, which scores every token of the finished answer. In your VPC, Latent reads your model itself, word by word.
    Your model

    Serves as it does today.

    Latent

    Keeps what your retention settings allow.

    You decide what leaves

    In your VPC, nothing leaves it, and the only outbound calls are ones you configure yourself: the model download, an alert webhook, a stronger model you name. Hosted by Latent, prompts and answers are sent to Latent's cloud for scoring, under the retention and redaction settings you choose.

    Your weights stay yours

    Latent never receives or copies model weights, and calibration does not need them.

    A record of every answer

    Each scored request writes a 1.7 KB append-only record with hashes of the prompt and answer, never the text.

    Questions teams ask first.

    Anything else, ask us. Get in touch

    We call our model through an API. What do we need?

    The runlatent Python SDK and an API key, when Latent hosts it. In your VPC, one GPU for Latent's reader model, the smallest of which fits in about 10 GB. Either way, the reader scores each finished answer and returns its verdict in a fraction of a second. Turn on an automatic response in the console, and the SDK in your app holds, replaces or reroutes a flagged answer before your user sees it. It covers OpenAI, Anthropic and Gemini clients; with any other provider, one call in your code does the same.

    Which models does it work with?

    Any model, including closed-weight models behind an API, which Latent checks with its own reader model. With open-weight models served on vLLM in your VPC, Latent reads your own model's activations per token as it writes, and can act on an answer mid-stream.

    What hardware does it need?

    None when Latent hosts it. In your VPC, it runs on the GPU already serving your model: the plugin adds 16 MiB of GPU memory at idle and at most 58 MiB at peak, whatever the prompt length. Every serving number on this site was measured on NVIDIA A100 40 GB or H100 GPUs.

    Does it change what my customers see?

    Only when you tell it to. Out of the box Latent scores and records every answer. To hold, replace or reroute a flagged answer, turn on an automatic response in the console, and the SDK applies it in your app, or Latent's gateway in front of your model does. You can also act on the verdict in your code.

    How good is it at finding wrong answers?

    Reviewers working from Latent's flags find 2.7 to 4.2 times more wrong answers than checking at random, and when it checks other models' answers, 96 in 100 of the answers it flags are really wrong. Calibrating on your own traffic is what moves you to the top of that range.

    How long does calibration take?

    When Latent hosts it, Latent runs calibration for you, judge included. A run takes under an hour for a few hundred to a thousand answers, and the update goes live with nothing to do on your side. In your VPC, one run over about 1,000 answers takes five to eight minutes and $6 to $8 of judge calls on your own key, and goes live at the plugin's next policy check, within 30 seconds by default, with no restart.

    Can it run air-gapped?

    Yes. After the model download on first start you can block all egress from the host, and the stack keeps running.

    What happens when our traffic changes?

    A drift alert watches the scores. It stays silent on steady traffic and fires within 120 to 181 answers when a shift matters, so you recalibrate when you need to.