We use a frontier model API. What do we need?
One GPU in your environment for Latent's reader model, the smallest of which fits in about 10 GB, and Latent's gateway in front of your API calls. Latent can hold, replace or reroute complete answers. Answers you stream to the customer as they are written are scored and flagged, but they cannot be held.
Which models does it work with?
Any model, including closed-weight models behind an API, which Latent checks with its own reader model. With open-weight models served on vLLM, Latent reads your model directly, which adds per-token risk and the ability to act while the answer is being written. It is measured in serving on Llama-3.1-8B-Instruct and Hermes-3, and it runs beside SGLang too.
What hardware does it need?
The GPU already serving your model. The plugin adds 16 MiB of GPU memory at idle and at most 58 MiB at peak, whatever the prompt length. Every number on this site was measured on NVIDIA A100 40 GB or H100 GPUs.
Does it change what my customers see?
Only when you tell it to. Out of the box Latent scores and records every answer. Holding, replacing or rerouting flagged answers is a policy you switch on in the review console.
How good is it at finding wrong answers?
Reviewers working from Latent's flags find 2.7 to 4.2 times more wrong answers than checking at random, and when it checks other models' answers, 96 in 100 of the answers it flags are really wrong. Calibrating on your own traffic is what moves you to the top of that range.
How long does calibration take?
Five to eight minutes for one run over about 1,000 answers, for $6 to $8 of judge calls on your own key. The new calibration goes live within six seconds, with no restart.
Can it run air-gapped?
Yes. After the model download on first start you can block all egress from the host, and the stack keeps running.
What happens when our traffic changes?
A drift alert watches the scores. It stays silent on steady traffic and fires within 120 to 181 answers when a shift matters, so you recalibrate when you need to.