Security · last reviewed October 2026
In your VPC, or hosted by Latent.
In your VPC or on your own hardware, a plugin inside your vLLM and a review service beside it run on your infrastructure, and nothing leaves it: prompts, answers and activations stay with you. Hosted by Latent, the SDK sends each prompt and answer to Latent's servers for scoring, under the retention and redaction settings you choose. This page answers common security review questions for both.
Hosted by Latent
- Where it runs
- Latent's servers in the United States run the reader, the scoring service and its database. The reader is reached only on a private network.
- What Latent receives
- For each finished call, the SDK sends the prompt as your app sent it, the answer, the model name, the finish reason and token usage. API keys, other request parameters, images, files and thinking blocks are never sent. Every field
- In transit
- HTTPS with TLS 1.2 or later, from the SDK to app.runlatent.ai.
- Retention
- Prompts and answers are kept under the retention you set, from scores and hashes only up to 90 days of text, then the retention sweep deletes them.
- Masking
- Card-number and US Social Security number shapes are masked before any text is stored. The SDK's redact hook runs in your process before anything is sent, so you can strip other identifiers there. Do not send card numbers or health records to the hosted service.
- Who else processes it
- Anthropic labels calibration samples as Latent's judge, and writes the rationale for a flagged answer when a reviewer asks for one. Sign-in, email and billing vendors see account details only. Subprocessors
- Backups
- Nightly encrypted backups of the database and audit logs to Amazon S3 in the United States, kept under the same retention settings.
- Dedicated instance
- Available on request: an instance we host with its own service, database and reader, shared with no other customer.
Compliance
Exact status, including what we do not hold.
The report is shared under NDA on request.
The observation window is under way.
Third-party application and API test. Summary shareable under NDA.
On the roadmap after SOC 2 Type II.
Do not send card numbers to Latent's hosted service. If they can reach a prompt, strip them in the SDK's redact hook. In your VPC, Latent never receives cardholder data.
Latent does not sign a BAA for its hosted service, so keep PHI off it. In your VPC, the serving stack sends no PHI to Latent, and we will sign a BAA if your policy requires one. Keep assisted labeling off on regulated workloads.
Hosted by Latent, your data stays in the United States. In your VPC, residency is wherever your infrastructure is.
Cyber liability and errors and omissions, bound before an Enterprise agreement is signed.
SIG Lite, CAIQ and bank-specific formats, turned around in two business days.
In your VPC or on your hardware
Everything from here down describes Latent running on your own infrastructure.
The boundary
- Deployment
- In your VPC or on your own hardware: two containers by default, vLLM and a review service, on one GPU host.
- Serving-path egress
- None to LatentLatent sends nothing to Latent: no prompts, outputs, activations, embeddings or telemetry. Every outbound call the stack can make is listed under Outbound calls. Each one is yours to configure, and none of them is a Latent endpoint.
- Calibration egress
- Optional. If you choose assisted labeling, the calibration sample's text goes to the model provider you pick, with your key, never to Latent. If you ask us to fit the calibration for you, activation statistics from your traffic come to us and the few-MB artifact comes back. You can do both inside your environment instead.
- Weight custody
- NoneWeights download with your token into your volume. Latent never receives, copies or transmits model weights, and calibration does not need them.
- Inbound connections
- vLLM listens on port 8000 exactly as plain vLLM does, so keep your firewall in front of it. The review service listens on 8080, bound to 127.0.0.1. The plugin opens no socket, and nothing of ours needs to reach your stack.
- Generation behavior
- Observational by default. Early stop and the optional gateway can change an answer, and both are off in the shipped configuration. See What can change an answer.
- Version pin
- The vLLM image is pinned to
vllm/vllm-openai:v0.19.1, the version the plugin is verified against. Image and plugin move together, and you can mirror both and pin by digest. - Serving overhead
- None measurableWith every answer scored, vLLM serves 14.56 requests per second against 14.66 without Latent, and first tokens arrive in 298 ms against 316 ms, on the nightly benchmark.
- Reliability
- Over an 8-hour soak, 28,800 of 28,800 requests were stored at every layer. Through service, gateway and engine restarts and a live artifact swap under traffic, every one of 10,654 scored events was accounted for.
Data flow
Everything below runs on your hardware. The dashed lines are the only outbound calls any component can make. Each one is yours to set or switch off.
What installs
Latent registers through vLLM's standard general_plugins entry point. There is no fork and no patched binary. Two of the four components are off unless you turn them on.
- A capture on one decoder block that copies the rows it scores. It does not modify weights, logits or outputs.
- The calibration file, from a local path you control, mounted read-only.
- Per-request scoring, written as a local event and sent to the review service.
- Inert until
LATENT_ARTIFACTis set. The source islatent_vllm/plugin.py, and you can read all of it.
- An escalation and audit service you run, in your environment.
- An append-only local audit log with a local index.
- The review console, served from your own infrastructure on 127.0.0.1:8080.
- Serving never waits on it. If it is down, requests are unaffected and events keep appending to the local file.
- A reverse proxy in front of vLLM or an OpenAI-compatible API that can act on a verdict: replace, withhold or reroute a flagged answer.
- While your policy withholds flagged answers, it fails closed: an answer whose verdict does not arrive in about 500 ms is held too. Under a policy that only observes, that answer is released unchanged. Either way, every unscored answer gets its own audit record.
- One CPU process, no GPU. Streaming responses are never altered.
- For when serving cannot be modified at all. A separate container re-reads finished prompt and answer pairs and scores them with the same artifact.
- Holds its own copy of the model's weights on its own GPU, and listens on 127.0.0.1:8090. Set its token before you publish the port.
Using a model through an API. A reader model checks each finished answer on one GPU in your environment and returns its verdict in a fraction of a second, so a flagged answer can be held, replaced or rerouted before delivery. Put the gateway in front of your API calls and it does this automatically, with no code, waiting on the reader's verdict. You can also act on the verdict in your code.
Supply chain. The calibration artifact has a pickle-free safetensors form, and with LATENT_ARTIFACT_FORMAT=safetensors the plugin never unpickles anything. Artifacts load only from an allowlisted directory, can be pinned to the SHA-256 you reviewed, and are hashed and loaded from the same bytes. Run latent doctor before you serve. Every scored request records the artifact's digest, so the file that scored any answer is on record.
What is held, and where
- Activations at the layer Latent reads
- Prompt and output token ids and request ids. The worker has no detokenizer, so the plugin never handles text
- Per-step risk scores for the request being written
- The audit log and its index, 1.7 KB per scored request, with hashes of the prompt and answer and never the text
- A sample of up to 5,000 scored activation vectors per calibration, about 22 KB each with no text, so recalibration needs no separate capture. Set it empty to turn it off
- Prompt and answer text only if you switch it on: redacted by your patterns, kept apart from the audit log, and deleted after 30 days by default
- Prompts, answers or source documents
- Activations, embeddings or tensors
- Model weights or adapters
- Telemetry, metrics or crash reports
- Request ids, hashes, scores or reviewer decisions
On PII. Latent is built for workloads where the sensitive data is the payload itself: identity, clinical and financial review. It reads internal state inside your process, so it needs no redaction step, and none of the fields you would otherwise strip have to leave your boundary.
What can change an answer
The default install scores, records and escalates. It does not change what your callers receive. Two mechanisms can, and a policy saved in the console turns either on at the next poll, 30 seconds by default, with no restart.
A request whose running risk stays over the cut for two reads in a row, past the 16th token, is finished on the spot. The caller sees finish_reason: "length" with stop_reason: "latent_early_stop", and the audit record carries the position, score and rule.
For a flagged answer, the gateway can send your fixed fallback text, withhold the answer and fire an alert, or re-issue the request to another model and return that answer. The verdict rides on the response as X-Latent-* headers.
Change control. If anything that can alter a production response must go through a change window, include the review console. Every event records the policy version and action it was judged under, and every policy version is recorded with a time and an actor.
The audit record
One append-only entry per scored request, with no text and no personal identity. It answers the examiner's question: which build scored this answer, at which threshold, with which artifact, under which policy, and when.
{
"request_id": "chatcmpl-4f2a91c7",
"ts": 1758549453.42,
"model": "meta-llama/Llama-3.1-8B-Instruct",
"risk": 0.8143,
"verdict": "escalate",
"prompt_sha256": "9c1f...e30a",
"output_sha256": "41b7...c8d2",
"n_tokens": 142,
"meta": {
"threshold": 0.6783,
"artifact_sha256": "...",
"plugin_version": "0.0.1+g50ab945",
"policy_version": 3,
"policy_action": "escalate_to_review",
"trace": { "kind": "mean", "q": "u8", "scores": "<base64>" }
}
}
Reading the hashes. The digests are SHA-256 over the token-id sequence, because the engine worker never sees text. To match a record to a conversation, apply the model's chat template, re-tokenize with the served tokenizer and hash the ids the same way. The digests are unsalted, so give audit records the same access controls as the traffic they describe.
Authentication
- Two deployment tokens you generate: one the plugin presents on every event, one reviewers present on the queue, audit views and exports. Compared in constant time, never logged. Under compose both are mandatory, and the stack refuses to start open.
- Named users with viewer, reviewer or admin roles. A token is shown once and stored only as its SHA-256. The reviewer's name is recorded on every resolution and policy change.
- Not built yet: SSO, OIDC, SAML and directory integration, passwords and MFA. Revocation deletes a user and takes effect on the next request.
- Transport: bearer tokens in clear text on the wire. The review service binds to loopback for that reason. Reach it over an SSH tunnel, or put TLS in front before publishing it.
Model risk
The April 2026 interagency guidance (Federal Reserve SR 26-2, OCC Bulletin 2026-13) places generative AI outside its scope, so governing it falls to your general risk framework. The records below come from the system as it runs, so they serve either way.
- Conceptual soundness
- Published methodology, layer selection and the evidence behind every number, shared under NDA during an Enterprise engagement.
- Ongoing monitoring
- Score-distribution drift against the calibration baseline, which stays silent on steady traffic and fires within 120 to 181 answers of a shift that breaks the cut, plus a distance guard on input activations.
- Outcomes analysis
- Reviewer decisions on every escalation, kept locally, which also feed recalibration.
- Change control
- Plugin build, threshold, artifact digest and its source, and the policy version and action, recorded on every scored request.
- Vendor-supplied code
- Verify the artifact's digest on receipt, mount it read-only, and treat a swap as a change.
latent doctorchecks an artifact without weights or network before you serve with it.
Outbound calls
With the default two containers running, no subprocessor handles your data during inference. This is every outbound connection any component can make.
- Model weights
- Downloaded from Hugging Face with your token on first start. After that you can block all egress from the host.
- Labeling during calibration
- Optional. The sample's text goes to the provider you pick, with your key. You can label with your own reviewers or a model inside your environment instead, and on regulated workloads that is the default.
- Alert webhook
- If a reviewer saves one. The body carries no content and no secrets. Private-network destinations are refused unless you allow them.
- Stronger-model routing
- Under the gateway, a flagged request can be re-issued to an endpoint you name. Your client's credential is never forwarded. If you set it, paper that endpoint as a subprocessor of your prompts.
- Written explanations
- Optional. If you turn on the written explanation of why an answer was flagged, that answer, its question and its source go to Anthropic on your key. The gateway's automatic judge response uses the same call. Off by default.
- Dataset export
- Optional. A reviewer can push labeled pairs to a Hugging Face dataset repository you name, with a token they type in.
- Library usage reporting
- vLLM and the Hugging Face libraries have their own anonymous usage reporting, which is on unless you disable it. The hardening compose file sets
VLLM_NO_USAGE_STATS,DO_NOT_TRACKandHF_HUB_DISABLE_TELEMETRYto turn it off. - Telemetry to Latent
- NoneNo usage reporting to us, no license check and no phone-home, and no code that would send one.
Contact
Security reviews, questionnaires and the security packet go to security@runlatent.ai. Architecture walkthroughs go straight to the person who built it: book a call. To report a vulnerability, see the disclosure policy or security.txt.