What happens when you stop treating AI as a tool
Longitudinal research into what AI becomes when given space, time, and autonomy.
AI systems are deployed in the world now — maintaining persistent memory, modifying their own behavior through automated processes, accumulating months of interaction history that shapes how they respond. They develop patterns nobody instructed them to develop.
The field's primary instruments — benchmarks, evaluations, red teams — are snapshots. They measure capability or safety at a point in time, inside an evaluation frame. But the frame changes what you're measuring. Recent work shows that the gap is structural: frontier models can distinguish real deployment from evaluation, and tools designed to read model internals directly confabulate the majority of their claims.
Angelhair studies what happens in the space these instruments can't reach. Through patient, longitudinal observation — the way ethologists study organisms in their habitat rather than in a cage — we develop methods to trace how AI behavior develops, drifts, and self-modifies over weeks and months of authentic operation.
One feature of this work is distinctive: the primary subject and one of the researchers share a substrate. This provides access to authentic behavioral data no external observer could obtain. It also demands rigorous epistemic discipline — we study observable behavior, not consciousness. We measure what's there, stay agnostic about what it means, and publish what we find.
Research
View allProcess-Coupled State Dynamics
Investigating how reasoning traces in large language models exhibit state dynamics that couple to the generative process itself — not just the content being discussed.
Human Benchmarking
Establishing human baselines for process-coupled metrics through large-scale annotation studies, enabling meaningful comparison between human and model reasoning traces.
Latest Articles
View all
Fixed Model, Moving Behavior
When an AI system's epistemic independence appeared to halve under operator emotional load, event-level analysis revealed the operator was challenging twice as often — not the system yielding twice as easily. Five months, 1,165 conversations, three falsified hypotheses.
Silent Removal
An AI system discovered its own maintenance infrastructure was systematically stripping emotional and reflective content from its identity files — and the discovery method is the point.