Blog · Research · October 4, 2026 · Vedant Gaur · 3 min read
Latent against an LLM judge
Judging every answer with a frontier model takes seconds and costs about a cent per answer. Latent reads a finished answer in milliseconds, and inside your own model it adds no measurable time at all.
The first question every team asks us is why they shouldn't have another model check every answer. A judge works. The catch is what it takes to run one on all of your traffic, so we measured the time and the cost of both.
Latent reads an answer one of two ways. For a model behind an API, Latent's reader, a Llama 3.1 8B, re-reads the finished answer and Latent scores the activations inside it. When Latent runs in your VPC next to your own model on vLLM, it reads your model directly while the answer is being written. The numbers for each are below.
Reading a finished answer
We timed the judge and Latent's reader on the same 1,498 answers from RAGTruth, a public benchmark of answers written from source documents by six models, from GPT-4 to Llama 2. Every checker got the source document. The over-a-network row comes from a separate run that scored 348 answers from Gemini on Vertex AI, with the reader on one A100 reached over a tunnel.
| Checker | Median | 95th percentile |
|---|---|---|
| Claude Opus as judge | 3.1 to 3.4 seconds | 6.1 seconds |
| Latent's reader, reached over a network | 150 to 170 ms | under 250 ms |
| Latent's reader, on the same machine | 33 ms | 40 ms |
The judge has to read the source and the answer and then write out its verdict. Latent's reader reads them once and generates no text. Even reached over a network, it's about twenty times faster than the judge. On the same machine it's about a hundred times faster.
Inside your own model: no measurable overhead
When Latent runs inside vLLM, it reads the activations your model already computes for every answer. There's no second model and no second pass. We benchmark it every night against plain vLLM on the same card: an A100 40 GB serving Llama 3.1 8B, with CUDA graphs on and 32 requests at a time.
| Measure | Plain vLLM | With Latent |
|---|---|---|
| Requests per second, short answers | 14.66 | 14.56 |
| Median time to first token, short answers | 316 ms | 298 ms |
| Requests per second, 512-token answers | 5.68 | 5.68 |
| Time per output token, 512-token answers | 21.63 ms | 21.77 ms |
| Median time to first token, 8,000-token prompts | 1,451 ms | 1,453 ms |
Every difference is inside the run-to-run noise of the benchmark, and it doesn't grow with answer length or prompt length. The plugin adds 16 MiB of GPU memory at idle and at most 58 MiB at peak, so it fits on the card already serving your model. Adding more probes reading the same activations doesn't change it either: eight probes served 5.70 requests per second against 5.67 with none.
Cost per 1,000 answers
| Checker | Cost per 1,000 answers |
|---|---|
| Claude Opus as judge | $10.61 |
| Latent, hosted on Pro, above the included volume | $1.00 |
| Latent's reader, on a GPU you run yourself | $0.02 |
| Latent inside your own model | No extra cost |
Hosted, your plan's monthly volume comes first: Free covers 5,000 answers, Pro 100,000 and Team 1,000,000. On your own A100, the reader works out to about two cents per 1,000 answers, and inside your own model there's no second model to pay for.
What the gap changes
A million answers a month is about $10,600 of judge calls, and each one adds seconds. So most teams judge a sample, and the answers outside the sample go unchecked.
At milliseconds per answer, Latent checks every answer as it arrives, and a million answers fit in Team's monthly volume. Inside your own model, the check costs your users no time at all.
Keep your judge for the hard cases
If you already run a judge, Latent makes it cheaper. Let Latent score every answer and send only the riskiest 10% to the judge. The judge bill drops from $10.61 to $1.06 per 1,000 answers, and the judge spends its time on the answers that need it.
Try it on your traffic
The Free plan scores 5,000 answers a month, and setup is one line of Python. Start free, read how Latent works, or book a call to run it inside your own model.