DeploymentDeployment options

Where Latent can run, what it installs, and what leaves your environment in each case.

Latent runs hosted by Latent, in your cloud account, or on your own hardware. In your cloud or on your hardware, you operate it from the repository we share. On Enterprise, nothing it does while serving reaches Latent. With Run it in your cloud on Team, your project's console at Latent receives scores, hashes and metadata.

Choose where Latent runs

Hosted by Latent Your cloud Your hardware
Who operates it Latent You You
What leaves for Latent Prompts and answers, under your retention settings: to Latent's cloud on Free, Pro and Team, or to a dedicated single-tenant instance on Enterprise. Activations are never sent. With Run it in your cloud on Team: scores, hashes and metadata, to your project in Latent's console; answer text only if you turn it on. On Enterprise: nothing. Nothing
How Latent reads The reader The in-band read and the reader The in-band read and the reader

Hosted Enterprise runs on a dedicated single-tenant instance: its own service, database and GPU reader pool, nothing shared with other customers, in a region chosen at signature. The Free, Pro and Team plans are hosted by Latent. Team can add Run it in your cloud, which we set up with you.

Two ways Latent reads an answer

The in-band read runs inside the vLLM that serves your model. A plugin reads your model's internal state while it writes, so every answer carries a live per-token risk, close to per word, while it is written. If you turn early stop on, an answer can be stopped mid-stream; early stop is off by default. The in-band read needs Latent next to your model, so it works when your model is open-weight and served on vLLM (the plugin is verified on vLLM 0.19.1).

The reader scores each finished answer with Latent's own reader model, in one pass after it is written: one risk and one verdict per answer, with the riskiest sentence marked by default. By default every answer carries a per-token read. It works for any model, including models you call through a hosted API, and engines other than vLLM, and it is the only read on a hosted deployment. Your model's activations are never sent anywhere: the reader works from the prompt and the answer text.

What runs in your environment

The default install is two containers on one GPU host:

  • vLLM with the Latent plugin. There is no fork of vLLM. The image is built on your host from the stock vllm/vllm-openai image (v0.19.1, pinned by digest) with the Latent package added, and the plugin loads through vLLM's general_plugins entry point. vLLM's own code is unchanged. The plugin reads one layer and does not modify weights, logits, sampling or outputs. The one thing it can change is to end an answer early, and only when you turn early stop on.
  • The review service. The review queue, the append-only audit log and the console, served from your own infrastructure.

Two components are opt-in:

  • The gateway, a proxy in front of your vLLM or an OpenAI-compatible API that holds, replaces, reroutes or regenerates a flagged complete answer, with no code in your app. In front of vLLM it acts on the in-band read's score. In front of an API it waits for the sidecar's score, so nothing leaves your environment. What it does with a streamed answer is on Streaming and complete answers. By default the gateway checks its project key with Latent when it starts and once an hour. In your environment, set LATENT_GATEWAY_REQUIRE_KEY=0 and LATENT_GATEWAY_IN_VPC=1 and it makes no call to Latent.
  • The sidecar, which re-reads each finished prompt and answer with its own GPU. On a one-GPU host, lower vLLM's share of the card (GPU_MEM_A) first: sharing a card slowed serving by about a fifth in our measurement. Use it for engines other than vLLM, for models behind a hosted API, or when serving must stay exactly as it is.

What leaves your environment

On Enterprise, with the gateway in in-VPC mode, nothing Latent does while serving reaches Latent: prompts, answers and activations stay on your infrastructure, and so do model weights, scores and reviewer decisions. With Run it in your cloud on Team, scores, hashes and metadata go to your project's console at Latent, and answer text only if you turn it on. Two opt-in steps can reach outside your environment, and neither goes to Latent:

  • Calibration labels. Labeling the calibration sample needs a judge. By default that is Anthropic's Claude, called with your own key, so the sample's sources, prompts and answers go to Anthropic under your agreement with them. You can instead point the judge at a model you host, or have your reviewers label the sample, and then nothing leaves.
  • Written explanations. If you turn them on (LATENT_EXPLAIN_JUDGE=1), each explanation a reviewer asks for, and each flagged answer when the gateway's automatic response is Send to the judge, sends that request's source, prompt and answer to Anthropic's API under your key. It runs after the answer is delivered, outside the serving path.

The other outbound calls are the model download from Hugging Face on first start, vLLM's and Hugging Face's own usage reporting until you add the hardening overlay, and the ones you configure, such as an alert webhook or a stronger model the gateway reroutes to. The security page lists every one.

On the hosted plans, prompts and answers are sent to Latent's cloud for scoring, under the retention and redaction settings you choose. See Data and retention for what is stored and for how long.

Getting the install

For Enterprise customers, and Team customers with Run it in your cloud, we share the repository or a release tarball, the install guide and the latent doctor preflight check. The images are built on your host from digest-pinned public bases, and nothing is pulled from Latent. An engineer works through the install with your team. The security page covers the review questions a deployment usually raises: the boundary, the data flow, what is held where, authentication and model risk.

Best practices

  • Run latent doctor before you serve with a new artifact or a new vLLM version.
  • Start observational: score and review without early stop or the gateway, then turn on the actions you want once the flag rate looks right.
  • For calibration on sensitive data, label with a judge you host or with your own reviewers, so nothing leaves your environment.

Note To plan a deployment in your environment, book an architecture walkthrough.

Was this page helpful?

Updated 3 October 2026