Large Language Models Develop Belief State Geometry In-Context
🧠 Decoding the Brain: How LLMs Compute Hidden ‘Belief States’ In-Context
The amazing capabilities of Large Language Models (LLMs) often feel like magic. They perform complex tasks just by seeing a few examples in the prompt—a phenomenon called In-Context Learning (ICL). But how do they actually know what to do? Do they just memorize patterns, or are they calculating something deeper?
Our new research dives into this mechanism, providing deep, representation-level evidence that LLMs are performing sophisticated forms of statistical inference.
🧐 The Core Problem: Unpacking ICL
The representations that support In-Context Learning have remained largely mysterious. To tackle this, we designed a highly controlled experiment using Hidden Markov Models (HMMs). HMMs generate data with underlying, unobserved ‘hidden states’—a perfect testbed for measuring inference.
We prompted several open-source LLMs with sequences of data generated by 40 different HMMs. The goal was to see if the model could effectively calculate the belief state: the posterior probability distribution over those hidden states, given all the tokens it has seen so far.
✨ Key Findings: Belief States are Explicitly Programmed
Our results were striking. Across six diverse open-source LLMs and 40 HMMs, we found that the belief state could be accurately extracted (decoded) from the model’s internal activations—specifically, the residual stream—with peak correlation coefficients ($R^2$) ranging from $0.83$ to $0.99$. This strong, linear relationship suggests that the LLM’s latent space is not just correlating tokens; it is structurally encoding the underlying statistical geometry of the data.
But we didn’t stop at prediction! To prove this finding was functional, we performed intervention experiments. We directly patched and steered the identified subspace responsible for holding the belief state. The results were conclusive: by manipulating these specific internal signals, we could reproduce downstream prediction quality that matched the untampered model performance. Meanwhile, controls (disrupting other parts of the activation) caused significant degradation.
🔑 The Takeaway: This provides powerful representation-level evidence suggesting that In-Context Learning in open-source LLMs closely approximates optimal Bayesian inference over a context-inferred generative model. The architecture seems to be inherently designed for optimal statistical prediction, even when operating solely on text prompts.
🚀 Why Does This Matter? (The Big Picture)
This work extends prior findings that link the structure of input data distribution directly to the geometry of activation within large models. It shifts our understanding from simply observing ‘good performance’ to diagnosing how and why LLMs achieve it.
Whether you are building state-of-the-art AI systems, optimizing prompts, or just curious about deep learning theory, this research confirms that current open-source LLMs contain latent mechanisms for robust, structured inference.
Read the full paper to dive into the math and methodology: Decoding Belief State Geometry in LLMs