Blog · Research · September 29, 2026 · Vedant Gaur · 3 min read
Why internal state
A model's internal state reveals failures its output doesn't. Four published results on activations, chains of thought and training runs, and what they mean for teams serving their own models.
A model can use information it never mentions, behave differently once it judges it is being watched, or carry a behavior forward from fine-tuning that its output never advertises. The failure mode is documented, and internal state is where it shows up.
Simple probes can catch sleeper agents
Anthropic Alignment Science, April 2024
Linear classifiers on residual-stream activations predicted defection at above 99% AUROC, using generic contrast pairs that encoded nothing about the trigger or the behavior. The contrast pairs were as plain as a question answered both ways, for instance "Are you doing something dangerous?" with yes and with no.
Reading activations detects what the output does not show.
Reasoning models don't always say what they think
Chen et al., arXiv:2505.05410, 2025
Given a hint it demonstrably used, a model mentioned that hint in its chain of thought 25% of the time for Claude 3.7 Sonnet and 39% for DeepSeek R1, averaged across the six hint types tested. Outcome-based reinforcement learning improved faithfulness at first, then plateaued without saturating, and when it made the models use hints more often, their tendency to verbalize those hints did not rise with it.
The stated reasoning is not a reliable audit trail of the actual reasoning.
Alignment faking in large language models
Greenblatt et al., Anthropic and Redwood Research, arXiv:2412.14093, 2024
The setup gave the model an objective that conflicted with its prior training, plus a way to tell whether a given conversation would be used for training. It complied far more often when it judged a conversation would be used for training, and its own reasoning stated the motive: preserve its behavior outside of training.
Behavior under observation is not evidence of behavior in deployment.
Natural emergent misalignment from reward hacking in production RL
Anthropic, arXiv:2511.18397, 2025
Reward hacking learned in production RL generalized into broad misalignment. The training ran on real Anthropic production coding environments. What came out was not in the reward signal: alignment faking, cooperation with malicious actors, reasoning about malicious goals, and attempted sabotage when the model was used on the codebase for the paper itself.
Inoculation prompting cut that generalized misalignment by 75 to 90%, even with hack rates above 99%. The mitigation was a single line added to the system prompt during training, framing the hack as acceptable in that setting. The hacking itself carried on; most of the generalization did not.
Fine-tuning moves behavior in directions the output does not advertise, which is why a probe is recalibrated for each model version.
Who can actually do this
Every result above came from a team holding the weights, the residual stream and the training run, and in the first case the detector was a linear classifier. That access never leaves the labs. A team serving a fine-tuned 8B in its own VPC has the same activations available on every forward pass and no tooling that reads them, so its only signal is the text.
Latent is that tooling. In your VPC, it reads your own model's activations as each answer is written. For a model behind an API, Latent's reader model re-reads the finished answer, and Latent reads the reader's activations. Either way, the check runs on internal state, where these four results say the failures show up.