Latent

Blog · Research · October 4, 2026 · Vedant Gaur · 3 min read

Latent against an LLM judge

Judging every answer with a frontier model takes seconds and costs about a cent per answer. Latent reads a finished answer in milliseconds, and inside your own model it adds no measurable time at all.

The first question every team asks us is why they shouldn't have another model check every answer. A judge works. The catch is what it takes to run one on all of your traffic, so we measured the time and the cost of both.

Latent reads an answer one of two ways. For a model behind an API, Latent's reader, a Llama 3.1 8B, re-reads the finished answer and Latent scores the activations inside it. When Latent runs in your VPC next to your own model on vLLM, it reads your model directly while the answer is being written. The numbers for each are below.

Reading a finished answer

We timed the judge and Latent's reader on the same 1,498 answers from RAGTruth, a public benchmark of answers written from source documents by six models, from GPT-4 to Llama 2. Every checker got the source document. The over-a-network row comes from a separate run that scored 348 answers from Gemini on Vertex AI, with the reader on one A100 reached over a tunnel.

Checker Median 95th percentile
Claude Opus as judge 3.1 to 3.4 seconds 6.1 seconds
Latent's reader, reached over a network 150 to 170 ms under 250 ms
Latent's reader, on the same machine 33 ms 40 ms

The judge has to read the source and the answer and then write out its verdict. Latent's reader reads them once and generates no text. Even reached over a network, it's about twenty times faster than the judge. On the same machine it's about a hundred times faster.

Inside your own model: no measurable overhead

When Latent runs inside vLLM, it reads the activations your model already computes for every answer. There's no second model and no second pass. We benchmark it every night against plain vLLM on the same card: an A100 40 GB serving Llama 3.1 8B, with CUDA graphs on and 32 requests at a time.

Measure Plain vLLM With Latent
Requests per second, short answers 14.66 14.56
Median time to first token, short answers 316 ms 298 ms
Requests per second, 512-token answers 5.68 5.68
Time per output token, 512-token answers 21.63 ms 21.77 ms
Median time to first token, 8,000-token prompts 1,451 ms 1,453 ms

Every difference is inside the run-to-run noise of the benchmark, and it doesn't grow with answer length or prompt length. The plugin adds 16 MiB of GPU memory at idle and at most 58 MiB at peak, so it fits on the card already serving your model. Adding more probes reading the same activations doesn't change it either: eight probes served 5.70 requests per second against 5.67 with none.

Cost per 1,000 answers

Checker Cost per 1,000 answers
Claude Opus as judge $10.61
Latent, hosted on Pro, above the included volume $1.00
Latent's reader, on a GPU you run yourself $0.02
Latent inside your own model No extra cost

Hosted, your plan's monthly volume comes first: Free covers 5,000 answers, Pro 100,000 and Team 1,000,000. On your own A100, the reader works out to about two cents per 1,000 answers, and inside your own model there's no second model to pay for.

What the gap changes

A million answers a month is about $10,600 of judge calls, and each one adds seconds. So most teams judge a sample, and the answers outside the sample go unchecked.

At milliseconds per answer, Latent checks every answer as it arrives, and a million answers fit in Team's monthly volume. Inside your own model, the check costs your users no time at all.

Keep your judge for the hard cases

If you already run a judge, Latent makes it cheaper. Let Latent score every answer and send only the riskiest 10% to the judge. The judge bill drops from $10.61 to $1.06 per 1,000 answers, and the judge spends its time on the answers that need it.

Try it on your traffic

The Free plan scores 5,000 answers a month, and setup is one line of Python. Start free, read how Latent works, or book a call to run it inside your own model.