Concepts
Data and retention
What Latent keeps on each hosted plan and for how long, and when it runs in your environment, how to set retention and redaction and exactly what leaves it.
What Latent keeps depends on where it runs. On hosted plans, Latent keeps your prompts and answers in its own cloud under your plan's retention. When Latent runs in your environment, it stores scores and hashes with no prompt or answer text by default, and serving sends nothing outside your hosts.
On hosted plans
| Plan | What Latent keeps | For how long |
|---|---|---|
| Free | Scores and hashes, plus the text of your first 500 answers for your included first calibration | 14 days. The text 30 days at most |
| Pro | Scores, hashes, and the prompt and answer text | 30 days |
| Team | Scores, hashes, and the prompt and answer text | 90 days |
| Enterprise, dedicated instance | Set in the contract, including zero retention | Set in the contract |
When a window ends, the retention sweep deletes the text. On Pro and Team, an admin can turn off answer text for a project to keep scores and hashes only; a calibration then has no text to run on until text is turned back on (Calibration). Free keeps text only for its first 500 answers. Data is stored on Latent's servers in the United States and reaches them over TLS 1.2 or later.
To keep an identifier out of Latent entirely, strip it in the SDK's redact hook, which runs in your process before anything is sent. Do not send card numbers or health records to a hosted plan.
Export your reviewed answers and their labels as JSONL from the console at any time. On Free, after the first-calibration window, the export carries scores and hashes, without text.
Note The rest of this page describes Latent running in your environment. There, text is stored only when you turn the text store on, and data leaves only through the settings listed at the end of this page, chiefly the judge calls that label answers for calibration.
What Latent stores
| Store | File | What it holds | Kept until |
|---|---|---|---|
| Events | audit. and index.db |
request id, token hashes, risk, verdict, meta with the token trace |
retention. |
| Decisions | resolution lines in audit., policy_ |
reviewer verdicts and every policy change | never swept |
| Text (opt-in) | text.jsonl |
the stored prompt and answer, redacted on write | retention., else LATENT_ |
| Deep analyses | explain. |
a reviewer's re-read of one answer, with its text | the text window, or retention. |
| Activation vectors | the reservoir (LATENT_); an explicit dump (LATENT_) has no size bound, and the whole file is deleted once it has gone unwritten for longer than the vector window |
half-precision vectors keyed by request id, no text | retention., or LATENT_ |
| Usage ledger | usage_ in index.db |
daily counts per model | never swept; a day is frozen two hours after it ends and never recomputed, except by an explicit usage rebuild |
A calibration run writes the answers it labels into its work directory (items., review_) on the host that ran it. Its report.md and report. carry counts, metrics and ids only, so you can share them as they are.
Join events to your logs
prompt_ and output_ hash the token ids, sha256("ids:" + ","., binding each score to the exact tokens the model read and wrote. To join an event to your own logs, match on request_id (the HTTP response id) or re-tokenise and hash the same way. LATENT_ puts an "Open in your logs" link on every queue row, so reviewers reach the text in your own log viewer.
The token trace
meta.trace is the probe's read at each answer token, quantized to a single byte per token: the answer so far (scores) and each token on its own (scores_). Long answers keep every k-th read (stride). It holds numbers only, and the sentence read's localised_ is a pair of character offsets. Turn the trace off with LATENT_ or the policy's trace: false.
Turn on text storage
With LATENT_ unset, the review service refuses text writes and the console marks requests hash-only. Turning the store on (LATENT_ on the service, LATENT_ on the gateway) keeps each pair in text.jsonl, separate from the audit log, readable by reviewers and admins only, redacted before it is written and swept on the text window.
Warning
LATENT_sets the text window, and switching the window off keeps text forever. To store no text, leaveTEXT_ RETENTION_ DAYS LATENT_unset.TEXT_ ENABLED
Set retention windows
One admin-only block in the review service's policy sets every window, and each change is a versioned policy_ line:
{"retention": {"events_days": null, "text_days": null, "explain_days": null, "reservoir_days": null, "archive": true}}
Every window stays null until an admin sets it. A null window keeps events forever, takes the text window from LATENT_, lets analyses follow the text, and leaves vectors to the plugin's LATENT_ (unset keeps them). A sweep runs at start and daily. Aged events leave the index for a compressed segment beside the log (audit.<from>-<to>.), or are dropped with archive: false. A flagged answer still open for review, an answer decided inside the window, and the rows a frozen drift baseline was fit on never move, whatever their age. Preview a sweep from the shell, or as an admin with POST /:
python -m service.app sweep --dry-run
Redact and mask text
LATENT_ (on the service) and LATENT_ (on the gateway, before a pair leaves its host) take comma-separated regexes; every match in every text field becomes [redacted]. A calibration run counts the redacted pairs and notes them in its report.
Latent does not mask for you beyond these patterns. If your logs must be de-identified before a judge sees them, mask identifiers only: people, phone numbers, emails, card and account numbers, URLs and IP addresses. Leave dates and places in the clear, and build one dictionary from the source for the whole record, so a name gets the same stand-in in the source and the answer. A default de-identification tool masks source and answer separately, and the strict judge then calls known-good answers unsupported. A probe calibrated on masked logs serves unmasked traffic, though masking changes which individual answers are flagged.
What leaves your environment
| Setting | What leaves | Where it goes |
|---|---|---|
| Default serving | nothing | |
latent calibrate with a judge |
the calibration sample: each labeled answer's source, question and answer | Anthropic's API, on your ANTHROPIC_ |
latent calibrate --labels |
nothing; your reviewers' labels replace the judge | |
LATENT_ |
one stored pair with its scores, per deep analysis a reviewer opens or the gateway's judge action queues |
Anthropic's API, on the service's key |
| Alert destinations you save | alert payloads | your webhook, Slack, PagerDuty or email relay |
POST / |
the labeled pairs a reviewer exports | a Hugging Face dataset repo, with a token typed into the request |
The calibration judge never receives your general traffic, and truncated answers, answers with no separable source and answers over the judge budget are never sent. Labeling runs synchronously by default. --judge-mode batch goes through Anthropic's Message Batches API, which holds requests longer than a synchronous call; check your Anthropic data-retention terms before using it on sensitive data, or stay on the default --judge-mode sync.
Best practices
- Keep text storage off unless reviewers need it. The log link gives reviewers the text from your own logs, and
latent calibrate --responsescalibrates from your own log;--from-serviceneeds the text store on. - Set a text window whenever text storage is on. Switching the window off keeps text forever.
- Redact on the gateway, so matches are removed before a pair leaves its host.
- Calibrate with
--labelswhen nothing is allowed to leave.
Was this page helpful?
Updated 3 October 2026