Blog · News · October 4, 2026 · Vedant Gaur · 3 min read
Latent is live
Starting today, any team can sign up for Latent, add one line of Python and see which of its model's answers are made up. Free to start.
Today we're opening Latent to everyone. Sign up, add one line to your app, and every answer your model gives gets a score for how likely it is to be made up. The risky ones land in a review queue, with the sentence that gave them away. The Free plan covers 5,000 answers a month, and you can start without talking to us.
Why we built it
Every team shipping an LLM product has a version of the same story. A customer screenshots an answer that sounds right and isn't: a refund policy that doesn't exist, a number that appears nowhere in the document, a person the model invented. By the time anyone notices, it's in a support ticket, or on social media.
The standard fix is to have a second model judge every answer. That works, but it's slow and expensive. Judging every answer with a frontier model costs about $10.61 per 1,000 answers and takes around three seconds each, so most teams judge a small sample and hope the rest are fine.
Latent reads the internal state of a model as the answer passes through it: the activations behind each word. Published research from Anthropic and others shows that this is where failures show up, often when the text gives nothing away. We wrote about that work in Why internal state. When Latent checks other models' answers, 96 in 100 of the answers it flags are really wrong.
What you get today
- A score on every answer, read token by token. Until you turn on an automatic response, scoring happens after your app already has its answer, so your users never wait on Latent.
- Holds before delivery. Turn on an automatic response in the console, and the SDK holds, replaces or reroutes a flagged answer before your user sees it.
- A review queue, sorted by risk. Your team confirms or clears each flag, and every decision becomes a labeled example for the next calibration.
- The reason behind each flag: the riskiest sentence and the words that drove it, the line in your source the answer leaned on most, and the kind of failure, such as an unsupported number or an invented name.
- Calibration on your own traffic. Latent runs it for you, judge included, and the update goes live with nothing to do on your side.
- Alerts in Slack, PagerDuty, email or a webhook when flags spike or stop, accuracy drops or your traffic drifts.
Holds, the calibrated read, the reasons behind flags and the alerts come with Pro and Team.
Start in two minutes
Install the SDK, set your key and add one line at startup:
pip install runlatent
export LATENT_API_URL=https://app.runlatent.ai
export LATENT_API_KEY=lk_...
import runlatent
runlatent.auto_instrument()
That covers every OpenAI, Anthropic and Google GenAI client in the process, sync, async and streaming. Your model calls stay exactly as they are. If you'd rather your coding agent did it, paste the prompt from Set up with your coding agent into Claude Code, Codex or Cursor.
Free covers 5,000 scored answers a month. Pro is $199 a month for 100,000, and Team is $999 a month for 1,000,000. The details are on pricing.
Serving your own model?
If you run an open-weight model on vLLM in your own VPC, Latent can run there too, as a plugin inside your inference engine. It reads your own model's activations while it writes, so every answer carries a live risk score close to word by word, the check adds no measurable time, and nothing leaves your network. It can even stop an answer partway through. Book a call and we'll set it up with you.
What's next
Next up are a proxy that works from any language with no SDK, and scoring the traces you already send to Braintrust, Langfuse or OpenTelemetry.
Tell us what it gets wrong
If Latent flags an answer it shouldn't, or misses one it should, we want to hear about it. Write to founders@runlatent.ai or use the feedback box at the bottom of any docs page. I read every message.
Latent is backed by Y Combinator. Start free, and let us know what you find.