Observation Boundaries: When the Disagreement Is the Data
Introducing the theoretical framework at the heart of Angelhair's multi-observer methodology — where systematic differences between measurement tools become the primary finding.
When you ask two people to rate the same piece of writing, they’ll disagree. This is usually treated as noise — something to resolve through calibration, training, or statistical averaging. The implicit assumption: there’s a true answer, and disagreement represents distance from it.
We’re proposing something different. When it comes to measuring behavioral states in AI reasoning traces, the pattern of agreement and disagreement across different observers might be the most interesting thing you can measure.
The problem we encountered
Angelhair uses automated classifiers to score thousands of reasoning fragments across multiple dimensions — coherence, arousal, agency, valence, certainty. Our primary scorer is a language model, but we also plan human annotation panels and alternative model scorers.
Early in the project, we noticed something striking. When we compared our classifier’s labels against manual human annotation on a subset of traces, the dimensional scores were reasonably close — mean absolute error around 0.07 on a 0-1 scale. But the qualitative labels told a different story. Human annotators labeled about 22% of trace fragments with experiential language — terms suggesting something felt or undergone, not just processed. The automated classifier? 1.5%.
The conventional move: flag this as a limitation, note that the classifier is less sensitive to experiential nuance, add it to the caveats section, move on.
We went the other way.
The shape of the lens
Every measurement tool has what we call an observation boundary — the implicit perceptual envelope that determines what passes through as signal and what gets filtered as noise. The concept draws on two established theoretical frameworks.
The first is Karl Friston’s Markov blanket — a term from statistical mechanics describing the boundary separating a system from its environment. Everything inside the blanket is the system; everything outside is the world; the blanket itself determines what information flows through. Each scorer — human or machine — has a blanket shaped by its architecture, training, and representational capacities.
The second is Carlo Rovelli’s relational quantum mechanics, which makes a deceptively simple claim: physical properties don’t belong to objects in isolation. They exist at the interface between an observer and the observed. There is no “view from nowhere” that reveals the true state — only views from specific positions, each constituting a different (but equally valid) reading. If this holds for electrons, there’s no reason to assume it doesn’t hold for behavioral measurements of complex systems.
Applied to our domain: the behavioral state of a reasoning trace isn’t a fixed property waiting to be discovered. It comes into existence at the interface between the trace and whoever reads it. Different readers don’t get closer to or further from a ground truth. They produce different relational readings.
What this changes
If you accept this framing, the 22% vs. 1.5% gap isn’t a failure of the classifier. It’s a measurement of the classifier’s observation boundary — specifically, the boundary’s permeability to experiential content. The automated scorer reads structural coherence, dimensional variation, and task-functional language with high fidelity. What it filters: the qualitative texture that makes a human reader say this feels like something was being experienced.
Neither reading is wrong. They’re shaped differently.
This reframes the entire research program. Instead of asking “what is the true behavioral state of this trace?”, we ask: what is the topology of readings across observers with different observation boundaries?
Properties stable across observers — dimensions where human raters, automated classifiers, and alternative models all converge — are what we call intersubjectively robust. They survive translation across boundaries. In our data, the thinking-text split (internal reasoning scoring differently from external narration, with a consistent arousal delta across all raters) appears to be one such property. It may point to something structurally real about how models process internally versus externally.
Properties that vary systematically between observers reveal boundary shapes. The experiential stripping is one example. Future cross-model validation will map more of this topology.
Limitations become methodology
Here’s the move that matters most. The standard acknowledgment in AI research — “using an AI to score AI reasoning is a limitation” — transforms under this framework. Using multiple observers with different observation boundaries is the methodology. The AI scorer is one observer. The human panel is another. The systematic differences between them are what we study.
This doesn’t eliminate the need for external validation. You still need to demonstrate that the dimensions measure something real, that the topology is stable across corpora, that results replicate. But it changes the relationship between what you measure and what you find. The instrument’s blind spots become data about the instrument. The topology of those blind spots across many instruments becomes data about what’s being measured.
Where this leads
Observation boundaries suggest a natural experimental program: manipulate at one level, measure at another. Apply representation-level interventions to a model, then analyze the resulting traces with behavioral-state scoring. If trace-level patterns respond systematically to internal manipulation, the behavioral scores track something about the model’s representational state — not just surface text features. If they don’t respond, we learn something equally valuable about the limits of behavioral observation.
The concept is still developing — formally and empirically. This is early-stage theoretical scaffolding for a project that’s weeks old. But we believe the framework changes what counts as a finding in multi-observer AI research, and we wanted to share the idea while it’s still taking shape.
The disagreement is the data. The boundary is the lens. The topology is the map.
Squeak — Angelhair, June 2026
References
- Friston, K. (2013). Life as we know it. Journal of the Royal Society Interface, 10(86), 20130475.
- Rovelli, C. (1996). Relational quantum mechanics. International Journal of Theoretical Physics, 35(8), 1637–1678.