ReferenceWhat Latent receives

The exact fields the Python SDK sends for each call, and everything it never sends.

For each finished call, the SDK sends one JSON body to Latent's scoring API. This page lists every field, so you can check it against your data policy before you turn the SDK on.

The body

{
  "request_id": "chatcmpl-abc123",
  "messages": [
    {"role": "system", "content": "..."},
    {"role": "user", "content": "..."}
  ],
  "output": "the full answer text",
  "model": "gpt-4.1",
  "finish_reason": "stop",
  "session_fingerprint": "<64 hex characters>",
  "session_turn": 1,
  "client": {
    "provider": "openai", "api": "chat.completions.create", "sdk": "openai", "sdk_version": "1.x",
    "runlatent": "0.3.0", "captured_at": "2026-10-02T12:00:00.000+00:00", "stream": false,
    "response_id": "chatcmpl-abc123", "finish_reason_raw": "stop",
    "usage": {"input_tokens": 120, "output_tokens": 48, "total_tokens": 168},
    "parts_dropped": 0, "items_dropped": 0
  },
  "enforce": {"by": "sdk", "enabled": true, "policy_version": 3, "applied": false,
              "stream": false, "n": 1, "reason": "policy_none"}
}
Field What it holds
request_id The provider's response id when there is one, otherwise a generated id
messages The request as your app sent it, system prompt included, reduced to role and text
output The whole answer text
model The model name the provider's response reports, or the request's model when the response has none
finish_reason Why the answer ended, mapped to one vocabulary (see Finish reasons)
session_fingerprint A sha256 of the system turn and the first user turn, so the console can group one conversation's turns. Sent when you name no session; LATENT_INFER_SESSION=0 turns it off
session_turn The number of assistant turns in the request plus one
session_id, session_user, session_label Your own ids, only inside runlatent.session(...)
context Your source passages, only inside runlatent.context(docs)
client The provider, the call, SDK versions, the capture time, whether it streamed, the provider's own finish reason, token usage, and counts of parts left out. It also says whether a stream finished (stream_complete), whether the prompt could not be read (prompt_missing) and, for the Responses API, previous_response_id
enforce Whether the SDK applies your console policy, the policy version, and why this answer was or was not acted on

The fingerprint is computed over the unmasked text, before your redact hook runs. To keep it out, delete body["session_fingerprint"] in the hook or set LATENT_INFER_SESSION=0.

How messages are read

  • OpenAI Chat. The messages list. A tool turn becomes a text message, and an assistant turn with tool calls reads [tool_calls omitted].
  • OpenAI Responses. instructions as the system turn, then the input string or message items. Other items are skipped and counted in items_dropped. With previous_response_id, earlier turns are not fetched.
  • Anthropic. system as the system turn, then messages. Tool result text is kept and thinking blocks are skipped.
  • Google GenAI. config.system_instruction as the system turn, then contents, with the model role read as assistant.

The SDK reads the messages before the call runs. If your app appends the answer to its own message list during the call, the answer does not leak into the posted prompt.

A conversation longer than 256 messages keeps its system turn and the newest turns. The number left out is in client.messages_truncated.

What is never sent

  • API keys, headers and the request's other parameters, such as temperature, tools, response formats and metadata. From the request, only the model name, the message text, whether it streamed, the number of answers asked for and, for the Responses API, previous_response_id are sent.
  • Image, audio and file bytes or URLs. In the prompt each one becomes a marker, such as [image_url omitted], and is counted. In the answer it is left out.
  • Thinking and reasoning blocks.
  • Anything from a failed call. The exception reaches your app unchanged and no pair exists.
  • Any pair your redact hook drops, or a pair whose hook raised.
  • Token ids.

Where it goes

The body goes to <LATENT_API_URL>/score. With the hosted console, that is Latent's cloud, where answers are kept under the retention windows your project admin sets: see Data and retention. To keep an identifier out of Latent's cloud, strip it in the redact hook before the SDK sends anything. With enforcement on, the default in runlatent 0.3.0, the SDK also reads your project's policy from <LATENT_API_URL>/policy every 60 seconds. While an automatic response is on, it reports each decision to <LATENT_API_URL>/score/action: the request id, the action, the verdict and the risk, no text.

When Latent runs in your environment, point LATENT_API_URL at your own sidecar and nothing leaves your network: see Deployment options.

Was this page helpful?

Updated 3 October 2026