Latent

Security · last reviewed October 2026

In your VPC, or hosted by Latent.

In your VPC or on your own hardware, a plugin inside your vLLM and a review service beside it run on your infrastructure, and nothing leaves it: prompts, answers and activations stay with you. Hosted by Latent, the SDK sends each prompt and answer to Latent's servers for scoring, under the retention and redaction settings you choose. This page answers common security review questions for both.

In your VPCNothing to Latent
Hosted regionUnited States
Weight custodyNone
SOC 2Type I

Hosted by Latent

Where it runs
Latent's servers in the United States run the reader, the scoring service and its database. The reader is reached only on a private network.
What Latent receives
For each finished call, the SDK sends the prompt as your app sent it, the answer, the model name, the finish reason and token usage. API keys, other request parameters, images, files and thinking blocks are never sent. Every field
In transit
HTTPS with TLS 1.2 or later, from the SDK to app.runlatent.ai.
Retention
Prompts and answers are kept under the retention you set, from scores and hashes only up to 90 days of text, then the retention sweep deletes them.
Masking
Card-number and US Social Security number shapes are masked before any text is stored. The SDK's redact hook runs in your process before anything is sent, so you can strip other identifiers there. Do not send card numbers or health records to the hosted service.
Who else processes it
Anthropic labels calibration samples as Latent's judge, and writes the rationale for a flagged answer when a reviewer asks for one. Sign-in, email and billing vendors see account details only. Subprocessors
Backups
Nightly encrypted backups of the database and audit logs to Amazon S3 in the United States, kept under the same retention settings.
Dedicated instance
Available on request: an instance we host with its own service, database and reader, shared with no other customer.

Compliance

Exact status, including what we do not hold.

SOC 2 Type IComplete

The report is shared under NDA on request.

SOC 2 Type IIIn progress

The observation window is under way.

Penetration testPlanned

Third-party application and API test. Summary shareable under NDA.

ISO 27001Not held

On the roadmap after SOC 2 Type II.

PCI DSSOut of scope

Do not send card numbers to Latent's hosted service. If they can reach a prompt, strip them in the SDK's redact hook. In your VPC, Latent never receives cardholder data.

HIPAA BAAIn your VPC

Latent does not sign a BAA for its hosted service, so keep PHI off it. In your VPC, the serving stack sends no PHI to Latent, and we will sign a BAA if your policy requires one. Keep assisted labeling off on regulated workloads.

Data residencyUnited States or yours

Hosted by Latent, your data stays in the United States. In your VPC, residency is wherever your infrastructure is.

InsuranceAt contracting

Cyber liability and errors and omissions, bound before an Enterprise agreement is signed.

Security questionnairesSupported

SIG Lite, CAIQ and bank-specific formats, turned around in two business days.

In your VPC or on your hardware

Everything from here down describes Latent running on your own infrastructure.

The boundary

Deployment
In your VPC or on your own hardware: two containers by default, vLLM and a review service, on one GPU host.
Serving-path egress
None to LatentLatent sends nothing to Latent: no prompts, outputs, activations, embeddings or telemetry. Every outbound call the stack can make is listed under Outbound calls. Each one is yours to configure, and none of them is a Latent endpoint.
Calibration egress
Optional. If you choose assisted labeling, the calibration sample's text goes to the model provider you pick, with your key, never to Latent. If you ask us to fit the calibration for you, activation statistics from your traffic come to us and the few-MB artifact comes back. You can do both inside your environment instead.
Weight custody
NoneWeights download with your token into your volume. Latent never receives, copies or transmits model weights, and calibration does not need them.
Inbound connections
vLLM listens on port 8000 exactly as plain vLLM does, so keep your firewall in front of it. The review service listens on 8080, bound to 127.0.0.1. The plugin opens no socket, and nothing of ours needs to reach your stack.
Generation behavior
Observational by default. Early stop and the optional gateway can change an answer, and both are off in the shipped configuration. See What can change an answer.
Version pin
The vLLM image is pinned to vllm/vllm-openai:v0.19.1, the version the plugin is verified against. Image and plugin move together, and you can mirror both and pin by digest.
Serving overhead
None measurableWith every answer scored, vLLM serves 14.56 requests per second against 14.66 without Latent, and first tokens arrive in 298 ms against 316 ms, on the nightly benchmark.
Reliability
Over an 8-hour soak, 28,800 of 28,800 requests were stored at every layer. Through service, gateway and engine restarts and a live artifact swap under traffic, every one of 10,654 scored events was accounted for.

Data flow

Everything below runs on your hardware. The dashed lines are the only outbound calls any component can make. Each one is yours to set or switch off.

Your environment
Your app
Gatewayoptional
vLLMwith the Latent plugin
Review servicequeue, audit log, console
Hugging Face, for the model download on first start
Your alert webhook, if a reviewer saves one
Your stronger model, if the gateway reroutes
Your labeling provider, during calibration, if you choose it
Anthropic, for written explanations, if you turn them on
vLLM and Hugging Face usage reporting, until you turn it off
Latent, the gateway's key check at start and hourly, until you set it to in-VPC mode
Latent's console, scores and hashes, with Run it in your cloud on Team

What installs

Latent registers through vLLM's standard general_plugins entry point. There is no fork and no patched binary. Two of the four components are off unless you turn them on.

DefaultIn the inference process
  • A capture on one decoder block that copies the rows it scores. It does not modify weights, logits or outputs.
  • The calibration file, from a local path you control, mounted read-only.
  • Per-request scoring, written as a local event and sent to the review service.
  • Inert until LATENT_ARTIFACT is set. The source is latent_vllm/plugin.py, and you can read all of it.
DefaultThe review service
  • An escalation and audit service you run, in your environment.
  • An append-only local audit log with a local index.
  • The review console, served from your own infrastructure on 127.0.0.1:8080.
  • Serving never waits on it. If it is down, requests are unaffected and events keep appending to the local file.
Opt-inThe gateway
  • A reverse proxy in front of vLLM or an OpenAI-compatible API that can act on a verdict: replace, withhold or reroute a flagged answer.
  • While your policy withholds flagged answers, it fails closed: an answer whose verdict does not arrive in about 500 ms is held too. Under a policy that only observes, that answer is released unchanged. Either way, every unscored answer gets its own audit record.
  • One CPU process, no GPU. Streaming responses are never altered.
Opt-inThe sidecar
  • For when serving cannot be modified at all. A separate container re-reads finished prompt and answer pairs and scores them with the same artifact.
  • Holds its own copy of the model's weights on its own GPU, and listens on 127.0.0.1:8090. Set its token before you publish the port.

Using a model through an API. A reader model checks each finished answer on one GPU in your environment and returns its verdict in a fraction of a second, so a flagged answer can be held, replaced or rerouted before delivery. Put the gateway in front of your API calls and it does this automatically, with no code, waiting on the reader's verdict. You can also act on the verdict in your code.

Supply chain. The calibration artifact has a pickle-free safetensors form, and with LATENT_ARTIFACT_FORMAT=safetensors the plugin never unpickles anything. Artifacts load only from an allowlisted directory, can be pinned to the SHA-256 you reviewed, and are hashed and loaded from the same bytes. Run latent doctor before you serve. Every scored request records the artifact's digest, so the file that scored any answer is on record.

What is held, and where

In memory, in your serving process
  • Activations at the layer Latent reads
  • Prompt and output token ids and request ids. The worker has no detokenizer, so the plugin never handles text
  • Per-step risk scores for the request being written
On disk, on your host
  • The audit log and its index, 1.7 KB per scored request, with hashes of the prompt and answer and never the text
  • A sample of up to 5,000 scored activation vectors per calibration, about 22 KB each with no text, so recalibration needs no separate capture. Set it empty to turn it off
  • Prompt and answer text only if you switch it on: redacted by your patterns, kept apart from the audit log, and deleted after 30 days by default
Never sent to Latent
  • Prompts, answers or source documents
  • Activations, embeddings or tensors
  • Model weights or adapters
  • Telemetry, metrics or crash reports
  • Request ids, hashes, scores or reviewer decisions

On PII. Latent is built for workloads where the sensitive data is the payload itself: identity, clinical and financial review. It reads internal state inside your process, so it needs no redaction step, and none of the fields you would otherwise strip have to leave your boundary.

What can change an answer

The default install scores, records and escalates. It does not change what your callers receive. Two mechanisms can, and a policy saved in the console turns either on at the next poll, 30 seconds by default, with no restart.

Off by defaultEarly stop, inside the engine

A request whose running risk stays over the cut for two reads in a row, past the 16th token, is finished on the spot. The caller sees finish_reason: "length" with stop_reason: "latent_early_stop", and the audit record carries the position, score and rule.

Off by defaultThe gateway, on the response path

For a flagged answer, the gateway can send your fixed fallback text, withhold the answer and fire an alert, or re-issue the request to another model and return that answer. The verdict rides on the response as X-Latent-* headers.

Change control. If anything that can alter a production response must go through a change window, include the review console. Every event records the policy version and action it was judged under, and every policy version is recorded with a time and an actor.

The audit record

One append-only entry per scored request, with no text and no personal identity. It answers the examiner's question: which build scored this answer, at which threshold, with which artifact, under which policy, and when.

{
  "request_id": "chatcmpl-4f2a91c7",
  "ts": 1758549453.42,
  "model": "meta-llama/Llama-3.1-8B-Instruct",
  "risk": 0.8143,
  "verdict": "escalate",
  "prompt_sha256": "9c1f...e30a",
  "output_sha256": "41b7...c8d2",
  "n_tokens": 142,
  "meta": {
    "threshold": 0.6783,
    "artifact_sha256": "...",
    "plugin_version": "0.0.1+g50ab945",
    "policy_version": 3,
    "policy_action": "escalate_to_review",
    "trace": { "kind": "mean", "q": "u8", "scores": "<base64>" }
  }
}

Reading the hashes. The digests are SHA-256 over the token-id sequence, because the engine worker never sees text. To match a record to a conversation, apply the model's chat template, re-tokenize with the served tokenizer and hash the ids the same way. The digests are unsalted, so give audit records the same access controls as the traffic they describe.

Authentication

  • Two deployment tokens you generate: one the plugin presents on every event, one reviewers present on the queue, audit views and exports. Compared in constant time, never logged. Under compose both are mandatory, and the stack refuses to start open.
  • Named users with viewer, reviewer or admin roles. A token is shown once and stored only as its SHA-256. The reviewer's name is recorded on every resolution and policy change.
  • Not built yet: SSO, OIDC, SAML and directory integration, passwords and MFA. Revocation deletes a user and takes effect on the next request.
  • Transport: bearer tokens in clear text on the wire. The review service binds to loopback for that reason. Reach it over an SSH tunnel, or put TLS in front before publishing it.

Model risk

The April 2026 interagency guidance (Federal Reserve SR 26-2, OCC Bulletin 2026-13) places generative AI outside its scope, so governing it falls to your general risk framework. The records below come from the system as it runs, so they serve either way.

Conceptual soundness
Published methodology, layer selection and the evidence behind every number, shared under NDA during an Enterprise engagement.
Ongoing monitoring
Score-distribution drift against the calibration baseline, which stays silent on steady traffic and fires within 120 to 181 answers of a shift that breaks the cut, plus a distance guard on input activations.
Outcomes analysis
Reviewer decisions on every escalation, kept locally, which also feed recalibration.
Change control
Plugin build, threshold, artifact digest and its source, and the policy version and action, recorded on every scored request.
Vendor-supplied code
Verify the artifact's digest on receipt, mount it read-only, and treat a swap as a change. latent doctor checks an artifact without weights or network before you serve with it.

Outbound calls

With the default two containers running, no subprocessor handles your data during inference. This is every outbound connection any component can make.

Model weights
Downloaded from Hugging Face with your token on first start. After that you can block all egress from the host.
Labeling during calibration
Optional. The sample's text goes to the provider you pick, with your key. You can label with your own reviewers or a model inside your environment instead, and on regulated workloads that is the default.
Alert webhook
If a reviewer saves one. The body carries no content and no secrets. Private-network destinations are refused unless you allow them.
Stronger-model routing
Under the gateway, a flagged request can be re-issued to an endpoint you name. Your client's credential is never forwarded. If you set it, paper that endpoint as a subprocessor of your prompts.
Written explanations
Optional. If you turn on the written explanation of why an answer was flagged, that answer, its question and its source go to Anthropic on your key. The gateway's automatic judge response uses the same call. Off by default.
Dataset export
Optional. A reviewer can push labeled pairs to a Hugging Face dataset repository you name, with a token they type in.
Library usage reporting
vLLM and the Hugging Face libraries have their own anonymous usage reporting, which is on unless you disable it. The hardening compose file sets VLLM_NO_USAGE_STATS, DO_NOT_TRACK and HF_HUB_DISABLE_TELEMETRY to turn it off.
Telemetry to Latent
NoneNo usage reporting to us, no license check and no phone-home, and no code that would send one.

Contact

Security reviews, questionnaires and the security packet go to security@runlatent.ai. Architecture walkthroughs go straight to the person who built it: book a call. To report a vulnerability, see the disclosure policy or security.txt.