We call our model through an API. What do we need?
The runlatent Python SDK and an API key, when Latent hosts it. In your VPC, one GPU for Latent's reader model, the smallest of which fits in about 10 GB. Either way, the reader scores each finished answer and returns its verdict in a fraction of a second. Turn on an automatic response in the console, and the SDK in your app holds, replaces or reroutes a flagged answer before your user sees it. It covers OpenAI, Anthropic and Gemini clients; with any other provider, one call in your code does the same.
Which models does it work with?
Any model, including closed-weight models behind an API, which Latent checks with its own reader model. With open-weight models served on vLLM in your VPC, Latent reads your own model's activations per token as it writes, and can act on an answer mid-stream.
What hardware does it need?
None when Latent hosts it. In your VPC, it runs on the GPU already serving your model: the plugin adds 16 MiB of GPU memory at idle and at most 58 MiB at peak, whatever the prompt length. Every serving number on this site was measured on NVIDIA A100 40 GB or H100 GPUs.
Does it change what my customers see?
Only when you tell it to. Out of the box Latent scores and records every answer. To hold, replace or reroute a flagged answer, turn on an automatic response in the console, and the SDK applies it in your app, or Latent's gateway in front of your model does. You can also act on the verdict in your code.
How good is it at finding wrong answers?
Reviewers working from Latent's flags find 2.7 to 4.2 times more wrong answers than checking at random, and when it checks other models' answers, 96 in 100 of the answers it flags are really wrong. Calibrating on your own traffic is what moves you to the top of that range.
How long does calibration take?
When Latent hosts it, Latent runs calibration for you, judge included. A run takes under an hour for a few hundred to a thousand answers, and the update goes live with nothing to do on your side. In your VPC, one run over about 1,000 answers takes five to eight minutes and $6 to $8 of judge calls on your own key, and goes live at the plugin's next policy check, within 30 seconds by default, with no restart.
Can it run air-gapped?
Yes. After the model download on first start you can block all egress from the host, and the stack keeps running.
What happens when our traffic changes?
A drift alert watches the scores. It stays silent on steady traffic and fires within 120 to 181 answers when a shift matters, so you recalibrate when you need to.