# Data and retention

What Latent keeps on each hosted plan and for how long, and when it runs in your environment, how to set retention and redaction and exactly what leaves it.

What Latent keeps depends on where it runs. On hosted plans, Latent keeps your prompts and answers in its own cloud under your plan's retention. When Latent runs in your environment, it stores scores and hashes with no prompt or answer text by default, and serving sends nothing outside your hosts.

## On hosted plans

| Plan | What Latent keeps | For how long |
|---|---|---|
| Free | Scores and hashes, plus the text of your first 500 answers for your included first calibration | 14 days. The text 30 days at most |
| Pro | Scores, hashes, and the prompt and answer text | 30 days |
| Team | Scores, hashes, and the prompt and answer text | 90 days |
| Enterprise, dedicated instance | Set in the contract, including zero retention | Set in the contract |

When a window ends, the retention sweep deletes the text. On Pro and Team, an admin can turn off answer text for a project to keep scores and hashes only; a calibration then has no text to run on until text is turned back on ([Calibration](/docs/concepts/calibration#on-hosted-plans)). Free keeps text only for its first 500 answers. Data is stored on Latent's servers in the United States and reaches them over TLS 1.2 or later.

To keep an identifier out of Latent entirely, strip it in the SDK's [redact hook](/docs/sdk#redaction), which runs in your process before anything is sent. Do not send card numbers or health records to a hosted plan.

Export your reviewed answers and their labels as JSONL from the console at any time. On Free, after the first-calibration window, the export carries scores and hashes, without text.

> Note: The rest of this page describes Latent running in your environment. There, text is stored only when you turn the text store on, and data leaves only through the settings listed at the end of this page, chiefly the judge calls that label answers for calibration.

## What Latent stores

| Store | File | What it holds | Kept until |
|---|---|---|---|
| Events | `audit.jsonl` and `index.db` | request id, token hashes, risk, verdict, `meta` with the token trace | `retention.events_days` |
| Decisions | resolution lines in `audit.jsonl`, `policy_history.jsonl` | reviewer verdicts and every policy change | never swept |
| Text (opt-in) | `text.jsonl` | the stored prompt and answer, redacted on write | `retention.text_days`, else `LATENT_TEXT_RETENTION_DAYS` |
| Deep analyses | `explain.jsonl` | a reviewer's re-read of one answer, with its text | the text window, or `retention.explain_days` |
| Activation vectors | the reservoir (`LATENT_CAPTURE_RESERVOIR`); an explicit dump (`LATENT_CAPTURE_ACTS`) has no size bound, and the whole file is deleted once it has gone unwritten for longer than the vector window | half-precision vectors keyed by request id, no text | `retention.reservoir_days`, or `LATENT_RESERVOIR_RETENTION_DAYS` |
| Usage ledger | `usage_daily` in `index.db` | daily counts per model | never swept; a day is frozen two hours after it ends and never recomputed, except by an explicit `usage rebuild` |

A calibration run writes the answers it labels into its work directory (`items.jsonl`, `review_queue*.jsonl`) on the host that ran it. Its `report.md` and `report.json` carry counts, metrics and ids only, so you can share them as they are.

## Join events to your logs

`prompt_sha256` and `output_sha256` hash the token ids, `sha256("ids:" + ",".join(ids))`, binding each score to the exact tokens the model read and wrote. To join an event to your own logs, match on `request_id` (the HTTP response `id`) or re-tokenise and hash the same way. `LATENT_REVIEW_LINK_TEMPLATE` puts an "Open in your logs" link on every queue row, so reviewers reach the text in your own log viewer.

## The token trace

`meta.trace` is the probe's read at each answer token, quantized to a single byte per token: the answer so far (`scores`) and each token on its own (`scores_last`). Long answers keep every k-th read (`stride`). It holds numbers only, and the sentence read's `localised_span` is a pair of character offsets. Turn the trace off with `LATENT_TRACE` or the policy's `trace: false`.

## Turn on text storage

With `LATENT_TEXT_ENABLED` unset, the review service refuses text writes and the console marks requests hash-only. Turning the store on (`LATENT_TEXT_ENABLED=1` on the service, `LATENT_GATEWAY_ATTACH_TEXT=1` on the gateway) keeps each pair in `text.jsonl`, separate from the audit log, readable by reviewers and admins only, redacted before it is written and swept on the text window.

> Warning: `LATENT_TEXT_RETENTION_DAYS` sets the text window, and switching the window off keeps text forever. To store no text, leave `LATENT_TEXT_ENABLED` unset.

## Set retention windows

One admin-only block in the review service's policy sets every window, and each change is a versioned `policy_history` line:

```json
{"retention": {"events_days": null, "text_days": null, "explain_days": null, "reservoir_days": null, "archive": true}}
```

Every window stays `null` until an admin sets it. A `null` window keeps events forever, takes the text window from `LATENT_TEXT_RETENTION_DAYS`, lets analyses follow the text, and leaves vectors to the plugin's `LATENT_RESERVOIR_RETENTION_DAYS` (unset keeps them). A sweep runs at start and daily. Aged events leave the index for a compressed segment beside the log (`audit.<from>-<to>.jsonl.gz`), or are dropped with `archive: false`. A flagged answer still open for review, an answer decided inside the window, and the rows a frozen drift baseline was fit on never move, whatever their age. Preview a sweep from the shell, or as an admin with `POST /retention/sweep?dry_run=true`:

```bash
python -m service.app sweep --dry-run
```

## Redact and mask text

`LATENT_TEXT_REDACT_PATTERNS` (on the service) and `LATENT_GATEWAY_TEXT_REDACT_PATTERNS` (on the gateway, before a pair leaves its host) take comma-separated regexes; every match in every text field becomes `[redacted]`. A calibration run counts the redacted pairs and notes them in its report.

Latent does not mask for you beyond these patterns. If your logs must be de-identified before a judge sees them, mask identifiers only: people, phone numbers, emails, card and account numbers, URLs and IP addresses. Leave dates and places in the clear, and build one dictionary from the source for the whole record, so a name gets the same stand-in in the source and the answer. A default de-identification tool masks source and answer separately, and the strict judge then calls known-good answers unsupported. A probe calibrated on masked logs serves unmasked traffic, though masking changes which individual answers are flagged.

## What leaves your environment

| Setting | What leaves | Where it goes |
|---|---|---|
| Default serving | nothing | |
| `latent calibrate` with a judge | the calibration sample: each labeled answer's source, question and answer | Anthropic's API, on your `ANTHROPIC_API_KEY` |
| `latent calibrate --labels` | nothing; your reviewers' labels replace the judge | |
| `LATENT_EXPLAIN_JUDGE=1` | one stored pair with its scores, per deep analysis a reviewer opens or the gateway's `judge` action queues | Anthropic's API, on the service's key |
| Alert destinations you save | alert payloads | your webhook, Slack, PagerDuty or email relay |
| `POST /improve/push` | the labeled pairs a reviewer exports | a Hugging Face dataset repo, with a token typed into the request |

The calibration judge never receives your general traffic, and truncated answers, answers with no separable source and answers over the judge budget are never sent. Labeling runs synchronously by default. `--judge-mode batch` goes through Anthropic's Message Batches API, which holds requests longer than a synchronous call; check your Anthropic data-retention terms before using it on sensitive data, or stay on the default `--judge-mode sync`.

## Best practices

- **Keep text storage off unless reviewers need it.** The log link gives reviewers the text from your own logs, and `latent calibrate --responses` calibrates from your own log; `--from-service` needs the text store on.
- **Set a text window whenever text storage is on.** Switching the window off keeps text forever.
- **Redact on the gateway**, so matches are removed before a pair leaves its host.
- **Calibrate with `--labels` when nothing is allowed to leave.**

---
Docs index: https://runlatent.ai/llms.txt
