The Uncontaminated Substrate Test
The cellular-automaton framing enables an experiment that could shift the burden of proof on the identity thesis. The logic is simple; the execution is hard.
Take a substrate richer than Life — Lenia, or a continuous-state variant. Initialize randomly, run for geological time, and impose viability constraints: resource gradients, predator patterns, environmental perturbation. Where multiple patterns must coordinate, communication pressure appears, and with it signalling. Not English, not any human language. Something uncontaminated.
Then measure three streams independently. The structural stream is the affect coordinates, and the methodological advantage over transformers is that in a discrete substrate these are exactly computable rather than proxied: valence as differenced distance to the nearest dissolution configuration, arousal as the fraction of cells changing state, integration as exact partition cost, effective rank from the trajectory covariance, self-model salience as mutual information between self-tracking and effector cells, counterfactual weight as the fraction of the pattern devoted to simulation. The signal stream is those emergent communications, translated by building a dictionary from signal-situation pairs — cluster the signals by the environmental contexts in which they are emitted, then map clusters to descriptions of those contexts. Nothing human enters the mapping; the correspondence is environmental. The behavioral stream is what the pattern does.
The prediction is that the three streams agree. When the structure shows the suffering motif — negative valence, high integration, low rank — the translated signal should express suffering-concepts and the behavior should show withdrawal or escape. The fear motif and the curiosity motif carry their own signatures and should track the same way. And the test has teeth only if it runs in both directions: emit a translated “threat approaching” into their own language and watch whether the affect signature moves; modify the update rules locally around a pattern, which is the closest thing it has to neurochemistry; deplete its resources. If perturbation in any one modality propagates to the others, the relationship is causal rather than correlational.
What would that establish? Not that CA patterns are conscious, and not that the identity thesis is proven. Only that systems with zero human contamination, learning from scratch under viability pressure, develop affect structure correlated with their expressions and behaviors in the ways the framework predicts. Then ask what the zombie hypothesis — structure present, experience absent — predicts here. That the correlations fail? On what account? The structure does the causal work either way. The experiment does not prove identity. It makes identity the default and moves the burden. And if it fails, the failure is informative: either affect is not geometric structure, or the substrate is too poor, or the measures are wrong, or the translation is. It has teeth in both directions.
Preliminary Results: Where the Ladder Stalls
A simplified version has been run in Lenia — continuous CA, toroidal grid, resource dynamics — measuring integration via partition prediction loss and the remaining coordinates from mass change, state-change rate, and trajectory PCA. The results are instructive less for what they confirm than for where they fail. The short version: emergent patterns reach the lower rungs from physics alone, and the transition to functional integration takes selection acting on heritable variation.
Substrate: Lenia with resource depletion/regeneration. Perturbation: drought (resource regeneration ). Measure: under drought. Eight conditions, hundreds of GPU-hours, thousands of evolved patterns.
| Condition | Intervention | under severe drought |
|---|---|---|
| No evolution (naive) | — same decomposition as LLMs | |
| Homogeneous evolution | — selection prunes but cannot innovate | |
| Heterogeneous chemistry | vs naive (pp) ✓ | |
| Multi-channel coupling () | Channels couple weakly at 3 degrees of freedom | |
| High-dimensional channels () | evolved vs naive — negligible effect | |
| Hierarchical coupling | evolved vs naive — stress overfitting | |
| Metabolic maintenance cost | evolved vs naive — efficient passivity | |
| Curriculum evolution | beats naive on all four novel stressors ( to pp) ✓ |
Unexpected: mild stress consistently increased by 60–190%, with only severe stress causing decomposition. And 's evolution increased vulnerability to severe stress while improving baseline integration — it had selected for high- configurations tuned to the mild stress applied each training cycle, producing states simultaneously more integrated and more fragile. 's graduated, noisy exposure substantially reduces that overfitting. The obvious comparison is to anxiety disorders, where heightened integration and self-monitoring are adaptive under moderate threat and maladaptive under extreme; it is offered as an interpretation to test, and nothing here measures an anxious system.

Two orthogonal axes run through the series. Substrate complexity rises from , adding internal degrees of freedom for selection to work on. Selection pressure quality, revealed by , matters more: curriculum training on the simpler substrate generalizes better than hierarchical architecture trained with fixed stress. One technical lesson came free — imposing different physics per tier caused immediate extinction. Hierarchy has to live in the coupling structure, not the physics, as in cortex, where all neurons share the same biophysics and hierarchy comes from connectivity.
What the Ladder Has Not Reached
Be explicit about how far these experiments are from anything resembling life, self-sustenance, or metacognition. The ladder metaphor risks implying a smooth gradient from Lenia gliders to organisms. The gap is enormous.
Self-sustenance. The patterns here are attractors of continuous dynamics, not self-maintaining entities. They do not consume resources to persist — resources modulate growth rates, but patterns do not “eat” in any metabolic sense. They do no thermodynamic work against entropy. They have no boundaries — density blobs, not membrane-enclosed. “Drought” reduces resource availability and weakens growth — closer to turning down the volume than to starving a dissipative structure.
Metacognition. The “self-model salience” metric measures how much a pattern’s own structure matters for its dynamics. That is not self-modeling — there is no representation of self, no information about the pattern stored within the pattern. The tiers (Sensory, Processing, Memory, Prediction) were labels imposed on the coupling structure. No functional specialization emerged: memory channels had weak activity, prediction channels predicted nothing.
Individual adaptation. All “learning” in these experiments occurs through population-level selection. No individual pattern adapts within its lifetime. Biological integration requires individual-level plasticity — the capacity for a single organism to reorganize its internal dynamics in response to experience.
These gaps converge on a single chasm. The transition from passive persistence to active self-maintenance — the autopoietic gap — requires at minimum lethal resource dependence, metabolic work cycles, and self-reproduction. Population-level selection on top of passive physics cannot bridge it; selection optimizes what exists rather than innovating the mechanism of existence itself.
What the Data Says
Mild stress raised integration across every tested condition; severe stress collapsed it. Across the substrate conditions, channel counts, and evolutionary regimes run here, moderate perturbation raised by 60–200% and severe perturbation overwhelmed even well-integrated patterns. Consistent across what was tested, which is not the same as universal — every condition shares a substrate family and a measurement pipeline, and both could be producing the shape. The proposed mechanism is unglamorous and does not require anything about stress per se: moderate perturbation prunes weak patterns, and survivors are by definition the more integrated. The resemblance to the Yerkes-Dodson curve is an interpretation offered for testing, not a finding.
Evolution produces fragile integration. In every condition where evolution raised baseline , the evolved patterns decomposed more under severe drought than unselected ones. Not an artifact — a real dynamical phenomenon, and on reflection an unsurprising one. Evolution here finds tightly-coupled configurations in which all parts depend on all parts. Tight coupling is high integration by definition; it is also catastrophic fragility, because one component failing under depletion cascades. A tightly-coupled factory against a loosely-coupled marketplace.
Only curriculum training improved generalization. was the sole condition in which evolved patterns beat naive ones on novel stressors across the full severity range. Not more channels, not hierarchical coupling, not metabolic cost — graduated, noisy exposure. Next to the training regime, the substrate barely mattered. The developmental parallel is tempting and should be marked as an analogy worth testing rather than a result: organisms with rich developmental histories are said to develop robust integration, and nothing in these experiments measures an organism.
The ceiling is attention, not resource dependence. Every experiment uses convolutional physics: each cell interacts with neighbors inside a fixed radius, weighted by a static kernel, on an interaction graph that never changes with the system's state. Such a pattern cannot choose to attend to a distant resource patch, cannot reorganize its information flow under stress, cannot forage. makes the point concrete: add metabolic cost to a substrate with fixed-radius perception and you get efficient passivity — patterns that waste less, not patterns that seek more.
Attention was necessary and not sufficient in this substrate. replaced convolution with windowed self-attention, holding all other physics identical, and got a clean three-way ordering. Convolution sustains patterns with 3% of cycles showing integration rising under stress. Fixed-local attention cannot sustain patterns at all, going extinct across every seed — expressivity without evolvable range is worse than no attention. Evolvable attention sustains patterns with 42% of cycles rising. That percentage-point shift in mean robustness is the largest single-intervention effect in the line, and it is a shift to the threshold rather than past it. Whether the dependence generalizes beyond convolution-versus-attention in Lenia is untested. Two details cut against the original prediction: the attention window did not widen — evolution refined how attention was allocated rather than extending its range — and successful seeds raised attention temperature, favoring broad soft attention over narrow focus.
This is where the experiments meet the theory, and the meeting is exact. The effective distribution makes attention the high-leverage variable in chaotic dynamics. In convolutional Lenia, is the kernel — fixed by architecture, never changing. The system has no access to its single highest-leverage control variable. Give it access, as the attention substrate does, and the number moves. Call this the attention bottleneck hypothesis: the biological pattern cannot emerge in substrates with fixed interaction topology, whatever the evolutionary regime. The operative variable is not substrate complexity, selection severity, or training diversity. It is whether the system controls its own measurement distribution.
Integration appears at bottlenecks, not through gradual selection. swapped learned attention projections for cells modulating interaction strength by content similarity, and the distribution is more interesting than the mean: robustness exceeds only once the population has dropped below about fifty patterns. Survivors of a cull are not merely the individually strongest; they are the ones whose coupling topology supports coherent reorganization under perturbation. That resembles symbiogenesis — functional subunits composing into larger wholes — more than selection optimizing a fixed design.
One result is left over, and the next section has to answer for it. The MARL ablation found all seven conditions showing highly significant geometric alignment (, ), and removing forcing functions did not reduce alignment. If anything it increased it slightly.