Pilot
Run Latent on your own traffic.
We install Latent next to the model you already serve, fit it to your traffic, and your reviewers work its queue on real answers. You finish with a report of your own numbers.
How it goes
- Kickoff
We agree on the model, the review workflow, the criteria you will judge us by, and the pilot's length and price.
- Install
Latent goes in next to your inference engine, or in front of your API calls. It takes minutes, and a preflight checks the setup.
- Calibrate
One run over about 1,000 of your model's own answers fits Latent to your traffic in five to eight minutes, for $6 to $8 of judge calls on your key. It goes live without a restart.
- Review
Flagged answers land in the queue on your host. Your reviewers confirm or release each one, and every decision is logged.
What calibration buys you
A finance assistant, checked on answers about companies the calibration never saw.
What you get
- A report on your traffic, with the criteria you chose at the top: how many flags were really wrong, reviewer time per failure, and any change to your latency.
- The error types on your traffic, such as wrong dates, unsupported numbers and invented names.
- The audit log, one record per answer, in a file that never left your host.
What we need
- A GPU host with Docker, or access to your existing deployment.
- A sample of real traffic, and a little reviewer time.
- A shared Slack channel, and light paperwork: usually an NDA, a security questionnaire and a short evaluation agreement.
What leaves your environment
Nothing, by default, and there is no telemetry. Two optional features call a provider you choose, with your key: judge labels during calibration, and a written explanation when a reviewer asks for one. Everything the stack can call is on the subprocessors page.
If we miss your criteria, you keep the audit log and the calibration.
Book a pilot call