Latent

Stop hallucinations before they reach your customers.

Latent reads your model's internal state as it writes and holds the answers it is making up before a customer sees them.

Backed by Y Combinator

Watch Latent catch
a made-up answer.

As your model writes each word, Latent reads what is happening inside it. When the answer is finished, it sends it where your policy says. Pick an answer and a policy, and watch it run.

Pick an answer
When Latent flags it
Customer asks

Latent is reading·
Your model
The answer

Customer sees
Your stronger model
Standing by
Review queue
Nothing waiting

 

Why not have another model check every answer?

You can. It is slow, it is expensive, and every answer has to leave your environment to be checked. Latent reads the model that wrote the answer, so it checks output validity in real time.

Answers checked in the same time103× faster
A judge model0 checked
Latent0 checked

One trip is one answer checked: 3.4 seconds for a judge, 33 ms for Latent.

Cost per 1,000 answers530× cheaper
A judge model$10.61
Latent$0.02
Want a judge anyway?

Let Latent choose which answers the judge reads. The judge bill falls tenfold, and more of what gets flagged is really wrong.

10×smaller judge bill
$1.06 against $10.61 per 1,000
93%of flags correct
against 82% judging everything

A layer the rest of your stack can't see.

Judges, guardrails and dashboards look at what your model said. Latent reads the model while it writes, and it works alongside all of them if your team wants to keep them.

Works with what you already run

Latent sits at runtime, between the input and the answer, where the rest of your stack can't see. Keep your judge, guardrails and traces on the output. Latent feeds them through a Prometheus endpoint, OpenTelemetry trace ids and labelled answers for fine-tuning.

See why it flagged

Each flag points to the sentence and the words that drove it, the sentence in your source the answer leaned on most, and the kind of failure, such as an unsupported number, a wrong date or an invented name. Your reviewers know where to look.

Alerts where your team already works

Slack, PagerDuty, email or a webhook, when flags spike or stop, confirmed accuracy drops, scoring slows, the review queue backs up or your traffic drifts.

From install to a working review queue.

  1. 01Install next to your inference engine, or in front of your API calls

    Latent joins your serving stack. A preflight checks the box, and Latent reads every answer from then on. Your prompts and answers stay on your hardware.

  2. 02Collect a sample

    Latent keeps a sample of what your model actually served. For calibration, an opt-in judge labels it against your sources.

  3. 03Calibrate on your traffic

    Latent studies that sample and learns what a wrong answer looks like in your domain. The run takes minutes, and the update goes live without a restart.

  4. 04Review what it holds

    Held answers wait in a queue. Your reviewers release them or confirm the failure, and every decision lands in the audit log.

Latent
Llama-3.1-8B-Instruct
docker compose -f deploy/docker-compose.yml up -d ✔ Container deploy-service-1  Started ✔ Container deploy-vllm-1     Startedlatent doctor  ✓ plugin_entry_point  latent -> latent_vllm.plugin:register  ✓ gpu                 1 GPU (NVIDIA A100-SXM4-40GB)  ✓ LATENT_CAPTURE      tap (in-graph capture; CUDA graphs stay on)  ✓ sink_reachable      http://service:8080/events  ✓ artifact            Llama-3.1-8B-Instruct_rag_mean.joblib  OK to servedocker compose logs vllm | grep 'latent_vllm ACTIVE'latent_vllm ACTIVE: model=meta-llama/Llama-3.1-8B-Instruct scoring=on capture=tap stream=on
14.56 vs 14.66requests per second, with and without Latent
+16 MiBof GPU memory at idle
Nothingleaves your network while serving
Sample from a finance assistant0 of 1,000 answers

    The sample stays on your box: at most 5,000 scored requests per calibration, kept by the plugin itself.

    latent calibrate --from-service http://localhost:8080 --work calib/

    0%of the answers Latent flags are really wrong
    Checked at random0%
    Flagged by Latent0%
    Under 8 minfor the whole run
    $6.88for all the labelling
    Under 6 sto go live, with no restart

    Latent's reading, word by word

    Every review makes it better.

    Each decision your reviewers make is kept as a labelled example. Latent learns from them, and the more answers your team validates, the better the data you post-train your model on.

    FlaggedLatent holds an answer
    ReviewedConfirmed or released
    LearnedKept as a labelled example
    Flag accuracyHow often a flagged answer is really wrong
    At install
    Training data for your modelValidated answers, ready to export
    Recalibrating live, no restart

    Latent recalibrates on your reviewers' decisions while it runs, with no restart. Every validated answer can also be exported as fine-tuning or preference data for your model.

    It runs inside your environment.

    Security and deployment detail
    Your VPC or your own hardware
    Your app
    Your inference enginewith the Latent plugin
    Review servicequeue and audit log
    On the GPUs that already serve your model. Latent reads the model itself, word by word.
    Latent

    Nothing is sent to Latent. No prompts, answers, activations, weights or telemetry.

    Nothing goes to Latent

    The only outbound calls are ones you configure yourself: the model download, an alert webhook, a stronger model you name.

    Your weights stay yours

    Latent never receives or copies model weights, and calibration does not need them.

    A record of every answer

    Each scored request writes a 1.7 KB append-only record with hashes of the prompt and answer, never the text.

    Questions teams ask first.

    Anything else, ask us. Get in touch

    We use a frontier model API. What do we need?

    One GPU in your environment for Latent's reader model, the smallest of which fits in about 10 GB, and Latent's gateway in front of your API calls. Latent can hold, replace or reroute complete answers. Answers you stream to the customer as they are written are scored and flagged, but they cannot be held.

    Which models does it work with?

    Any model, including closed-weight models behind an API, which Latent checks with its own reader model. With open-weight models served on vLLM, Latent reads your model directly, which adds per-token risk and the ability to act while the answer is being written. It is measured in serving on Llama-3.1-8B-Instruct and Hermes-3, and it runs beside SGLang too.

    What hardware does it need?

    The GPU already serving your model. The plugin adds 16 MiB of GPU memory at idle and at most 58 MiB at peak, whatever the prompt length. Every number on this site was measured on NVIDIA A100 40 GB or H100 GPUs.

    Does it change what my customers see?

    Only when you tell it to. Out of the box Latent scores and records every answer. Holding, replacing or rerouting flagged answers is a policy you switch on in the review console.

    How good is it at finding wrong answers?

    Reviewers working from Latent's flags find 2.7 to 4.2 times more wrong answers than checking at random, and when it checks other models' answers, 96 in 100 of the answers it flags are really wrong. Calibrating on your own traffic is what moves you to the top of that range.

    How long does calibration take?

    Five to eight minutes for one run over about 1,000 answers, for $6 to $8 of judge calls on your own key. The new calibration goes live within six seconds, with no restart.

    Can it run air-gapped?

    Yes. After the model download on first start you can block all egress from the host, and the stack keeps running.

    What happens when our traffic changes?

    A drift alert watches the scores. It stays silent on steady traffic and fires within 120 to 181 answers when a shift matters, so you recalibrate when you need to.