Appearance
What LLMs Encode Isn't What They Use
The interpretability pipeline has a leak
The standard story goes like this: train a model, probe its hidden states, find the feature, then steer it. Encoded, decodable, actionable. One continuous chain.
Four papers landed in the last few weeks, and each breaks a different link in that chain. A probing audit shows that information you can decode from hidden states often never reaches the output. A study of "thinking" tokens shows they're doing something else entirely: extending the prompt with scratchpad context. A topological analysis finds real structure in the hidden state manifold, but it's structure attention weights can't see. And a theory paper argues the interpretability program has been measuring the wrong parameter from the start.
Taken together, they point at the same conclusion: the map of a model's internal representations is not the territory of its behavior.
Decodable isn't usable
I've run enough probes to know the feeling. You decode a feature cleanly, you write it up, and somewhere between the hidden state and the output the model quietly ignores it. The "Encoded but Not Actionable" audit is that feeling, measured.
The authors probe six frozen decoder-only LLMs on parametric CAD constraints, a controlled testbed because it separates two levels of structure. Local pairwise relations (is this line coincident with that one?) sit below sketch-level constraint status (is the whole sketch fully constrained?). They test four properties on the same hidden states, and the four properties give four different answers.
| Property | What it tests | What they found |
|---|---|---|
| Linear decodability | Can a linear probe read the property from hidden states? | Local relations decode well after pretraining; sketch-level DOF status decodes even from random init |
| Forced-choice generation | Does the model express the property in its output? | Often fails, even where probes succeed |
| Activation-level influence | Does restoring a patched activation change behavior? | Effects vanish at the patched entity position while decodability persists across depth |
| Behavioral steerability | Can mean-difference steering control outputs? | Not reliably |
The first result is a sanity check: pretraining does improve decoding of local geometric relations, and the advantage survives shuffled-order controls, so it's not positional leakage. The sketch-level result is the warning shot. DOF status is already highly decodable from randomly initialized representations. The probe is reading structure that exists without any learned weights. Run the probe alone and you'd credit the model with geometric understanding it never acquired.
Then the divergence. Information that decodes cleanly often never shows up in generation. Activation-restoration effects vanish at the patched position even while decodability persists across depth. Mean-difference steering doesn't reliably move outputs. The pattern held across all six models, so it's not an artifact of one architecture.
Probe accuracy is a ceiling, not a floor. It tells you what the model could use, not what it does use.
Quick Take: Probing tells you what a model knows, not what it will do with that knowledge, and the gap between encoded and expressed is where interpretability tools quietly fail.
Reading "thinking" tokens as cognition is a mistake
I spent a week reading Qwen3.8's thinking traces like a therapist reading session notes. The verbose ones felt like overthinking, the terse ones felt like a model cutting corners. When the same question produced wildly different trace lengths on different runs, I blamed the model. The paper behind the Reddit discussion convinced me I was projecting.
The evidence is hard to argue with. There's no correlation between solution correctness and trace validity: models routinely produce invalid reasoning traces and still land the right answer. Models trained on corrupted or semantically irrelevant traces match or beat models trained on correct traces, especially out of distribution. Reinforcement learning improves accuracy without improving trace validity, and in some cases validity drops while accuracy climbs. Trace length doesn't track problem difficulty at all.
The framing that makes sense of all of it: intermediate tokens extend the prompt. They're scratchpad context the model writes for itself, not a transcript of cognition. That explains the verbose-but-correct pattern. The trace is doing work, but the work is context, not logic.
The Reddit thread around the paper filled with people who'd had the same experience. Verbose traces on trivial questions, terse traces on hard ones, and no relationship between how much the model "thought" and whether it was right. The hardest result to sit with is the corrupted-trace one. Replace the reasoning with gibberish and the model still solves the problem. The reasoning was never doing what it looked like it was doing.
Key Numbers
- 4 properties tested in the CAD audit: decodability, generation, influence, steerability. They diverge.
- 6 frozen decoder-only LLMs audited. The decode-generate gap held across all of them.
- ≤6 semantic hops bounds deep-layer navigation in long-context manifolds.
- ~3 hops for grounded RAG generations. Hallucinations collapse topologically.
- 6 predictions in the phase paper, each testing whether a suppressed meaning stays active.
The hidden state manifold is a small world
The six degrees of separation paper takes a different route: it ignores attention weights entirely. Attention-based interpretability keeps tripping over routing artifacts like attention sinks, so the authors analyze the dynamic geometry of the hidden state manifold directly. They sparsify continuous similarity matrices into unweighted graphs and trace connectivity between disjoint semantic anchors across long contexts.
The result is a sharp topological phase transition. Early syntactic layers are entirely fractured, a pile of disconnected islands. Deep reasoning layers abruptly compress massive conceptual distances into short, connected pathways, strictly bounded by six semantic hops. Six hops means any concept can reach any other concept through a short path, which is what makes multi-hop reasoning over long contexts tractable in the first place.
The practical payoff is hallucination detection. On the RAGognize dataset, factually grounded generations maintain structural integrity with their source context at about three hops. Hallucinations induce severe topological collapse.
Three hops is close enough to source that structural integrity becomes a usable reliability signal, no classifier training required. That's the most immediately deployable result in this batch.
Phase: the parameter interpretability never measured
The most abstract paper of the four, and in some ways the most provocative. "Language Has Two Parameters" argues that co-occurrence statistics give you amplitude, the strength of association between words. Word embeddings and attention weights refine that count, but they sum every writer in the corpus together. The paper claims a second parameter, phase, that signed weights learned from a corpus cannot supply.
Phase lives between meanings, not between words. It determines how coactivated meanings combine, and it can reverse what a meaning contributes while that meaning stays fully present. Irony is the cleanest example: "great job" said sarcastically keeps the meaning "great job" fully active while contributing the opposite. Population averaging deletes phase, because phase is indexed to individuals, dyads, and history. An agent-deindexed corpus can only identify the population marginal state, never any individual or dyadic state.
The uncomfortable implication is aimed directly at the interpretability program. Measuring progress by monosemanticity is optimizing against the wrong target. The coexistence of meanings that monosemanticity treats as a defect is the condition of allusion, irony, and quotation. A model with perfectly monosemantic features couldn't be ironic, because irony requires a meaning to be present and inverted at the same time.
The paper defends the weak claim: interpretation requires a second relational parameter, signed, persistent, and indexed to individuals and dyads. The strong claim, that quantum calculus is necessary rather than convenient notation, rests on an encounter-order constraint not yet derived. The paper is explicit that quantum probability is notation here, not a claim about quantum processes in the brain. Six predictions test the weak claim, and each is runnable on current models: whether a suppressed meaning stays active, whether encounter order changes what a phrase does, whether marking the signal changes interpretation, whether a model given a history is changed by it or only informed about it.
The architecture the theory calls for is concrete: a language model with agent-indexed, phase-bearing semantic states. That's a design constraint, not a metaphor.
Common pitfalls
Treating probe accuracy as behavioral evidence. The CAD audit shows decodability, generation, influence, and steerability diverge. If you publish a probe result without a behavioral intervention, you're reporting a ceiling, not a finding.
Reading reasoning traces as cognition. If you're filtering or rewarding traces based on semantic validity, the corrupted-trace results say you're shaping the wrong signal. Validity doesn't track correctness.
Using attention weights as a proxy for semantic proximity. Attention sinks and routing artifacts make attention a poor map of what the model knows. The topological paper bypasses attention entirely for a reason.
Averaging away the thing you're studying. The phase paper's sharpest methodological point: population averaging deletes history-indexed phase. Pooling activations across contexts, users, or conversation histories can erase the exact signal that determines interpretation.
Treating monosemanticity as the goal. The coexistence of meanings isn't always interference. For irony, allusion, and quotation, it's load-bearing. A feature that lights up for two meanings isn't necessarily broken.
One thing to remember
If the four papers agree on anything, it's this: the hidden state is a warehouse, not a conveyor belt. Structure that a probe can read, structure that reaches the output, and structure that an intervention can control are three different inventories. You have to verify each one separately.
The bottom line
If you're building interpretability tooling, probes, feature attribution, steering vectors, validate every decoded feature against a behavioral intervention. The CAD audit shows decode accuracy alone will overstate your tool by a wide margin.
If you're training or scaffolding agents on reasoning traces, treat intermediate tokens as scratchpad context, not ground-truth reasoning. Don't reward trace validity as a proxy for correctness; the corrupted-trace results show you'd be shaping the wrong signal.
If you're working on RAG or hallucination detection, the topological signature is the most immediately deployable result here: grounded generations hold about three hops, hallucinations collapse. Watch the phase direction too. Agent-indexed, history-bearing semantic states are the next frontier, and within a year expect someone to ship a model that actually implements them.