AI research · Evidence guide

AI consciousness:
what can you actually measure?

Fluent self-reports, persistent memory, global broadcast, and high internal scores are four different kinds of evidence. Treating them as one is how engineering demonstrations turn into unsupported consciousness claims.

Christian BucherUpdated July 17, 202611 min readPrimary sources linked
Position: ANIMA is useful as an inspectable cognitive-state engine and research harness. Its behavior and state transitions are measurable. The project does not measure sentience or phenomenal consciousness.

The productive question for developers is not “did the model sound alive?” It is “which property changed, how was it measured, what intervention caused it, and which stronger conclusions remain unsupported?”

A four-level evidence ladder

Every result should be placed on the narrowest level it directly supports. Moving upward requires additional evidence; attractive terminology cannot do that work.

LEVEL 01

Observable behavior

Outputs, task scores, error recovery, self-reports, tool calls, and consistency across trials.

DIRECTLY TESTABLE
LEVEL 02

Architecture

Persistent state, recurrent processing, limited workspace, global availability, self-model, and explicit interventions.

DIRECTLY INSPECTABLE
LEVEL 03

Internal proxies

Integration scores, salience, affect-inspired values, complexity metrics, or theory-derived indicator properties.

DEFINITION-DEPENDENT
LEVEL 04

Subjective experience

Whether there is something it is like to be the system: sentience or phenomenal consciousness.

NOT DIRECTLY MEASURED

A model saying “I feel afraid” is Level 1 behavior. A durable fear-like variable is Level 2 architecture. A composite score derived from it is Level 3. None of those observations alone establishes Level 4.

Self-report is generated data

Language models learn patterns of first-person reporting and respond to context. Self-reports may be inputs to a carefully designed evaluation, but they are not privileged access to an otherwise hidden ground truth.

What consciousness theories contribute

Theories can inspire precise computational properties. They do not make a software system conscious by analogy.

Global Neuronal Workspace Theory links conscious access in brains to selection, amplification, and broad availability of information. An AI developer can implement a capacity-limited workspace and test whether a selected item becomes available to multiple modules. That demonstrates workspace-like information flow, not the biological mechanism and not subjective experience.

Integrated Information Theory makes claims about the intrinsic cause-effect structure of a physical system. A convenient software statistic called phi is not automatically IIT's Φ. Exact and approximate measures have formal assumptions; renaming a coupling score cannot transfer their interpretation.

Predictive processing motivates models that compare predictions with observations and update their state. Prediction-error processing is widespread in machine learning and control systems. Its presence is useful engineering evidence, but not a consciousness verdict.

A 2025 open-science adversarial collaboration directly tested distinctive predictions of IIT and GNWT in 256 human participants using fMRI, MEG, and intracranial EEG. Results challenged important predictions from both accounts rather than producing a simple winner. If leading theories remain under active empirical pressure in humans, software analogies deserve even more care.

Indicators are not a consciousness checklist

A widely discussed 2023 report translated several scientific theories into computational “indicator properties” for assessing AI systems. That is a more disciplined approach than judging conversational style. But indicators are still theory-dependent evidence, not a binary certification test.

ANIMA as an open case study

The maintained ANIMA Kernel is a zero-runtime-dependency Python package for persistent cognitive state. It implements associative recall, limited workspace selection, temporal context, affect-inspired variables, a simplified self-model, JSON persistence, and model-provider bridges.

Historical API names include ConsciousnessState, Phase.CONSCIOUS, Phi, and CQI. They are provenance, not proof. The maintained documentation now defines them as implementation-level state and proxies.

446tests for lifecycle, transitions, memory, persistence, metrics, bridges, and CLI behavior
+0.92%mean CQI delta in the saved single-run neutral-kernel comparison
13.77%CQI change in the saved working-memory ablation
0.00%CQI change in the short temporal ablation — a published null result

The test count is strong implementation evidence. The benchmark is exploratory: a single saved run, no confidence intervals, no preregistration, and no independent replication. The correct conclusion is that specific components changed specific package-defined measurements under that harness.

ClaimEvidence ANIMA hasEvidence still needed
State survives restartsPersistence code and automated round-trip testsWorkload-specific durability testing for production use
Workspace is capacity limitedInspectable implementation and unit testsIndependent behavioral benefit across real agent tasks
Components affect proxy scoresSaved ablation artifactRepeated trials, uncertainty, stronger controls, replication
Agent outcomes improveSmall exploratory neutral-kernel comparisonPredefined external tasks and model/database baselines
System is consciousNoneNo accepted decisive test currently exists

A practical evaluation protocol

  1. Define the boundary. State whether you are evaluating the base model, prompt, memory service, full agent loop, or physical deployment.
  2. Name the property. Replace “more conscious” with persistence accuracy, recall precision, cross-module availability, calibration, or another operational measure.
  3. Intervene. Disable or randomize one component while holding model, prompts, inputs, and sampling settings constant.
  4. Repeat. Report distributions and uncertainty, not one attractive run.
  5. Use external outcomes. Score tasks independently of the internal metric being promoted.
  6. Publish nulls. A component doing nothing is information, not a marketing failure.
  7. Keep the conclusion narrow. Say exactly what changed and stop before the next unsupported rung.

Why the distinction improves products

Honest boundaries are not anti-ambition. They make an unusual project easier to adopt. A developer can use ANIMA's persistence and workspace without accepting a theory of consciousness. A researcher can replace its proxies. A reviewer can reproduce a claim. Clear falsifiability creates more credibility than a dramatic demo that cannot survive inspection.

Build the capability. Label the evidence. Leave the mystery open.

Persistent memory, global availability, self-monitoring, and affect-inspired control signals can be valuable engineering ideas. Their value does not depend on declaring the system alive.

Primary sources

Consciousness in Artificial Intelligence: Insights from the Science of ConsciousnessButlin et al. (2023) · Theory-derived computational indicator properties and an assessment framework. Adversarial testing of global neuronal workspace and integrated information theoriesCogitate Consortium et al., Nature (2025) · Preregistered, multi-method comparison in human participants. An adversarial collaboration protocol for testing contrasting predictions of GNWT and IITMelloni et al., PLOS ONE (2023) · Predictions, pass/fail criteria, and open-science protocol. ANIMA benchmark methodology and limitationsExact local proxy definitions, control construction, ablations, and the limits of the current saved artifact.

Frequently asked questions

Can an AI saying it feels something prove consciousness?

No. It proves that the system generated that text under that context. Treat self-report as behavior to explain, not as privileged ground truth.

Does a global workspace make an AI conscious?

It demonstrates a workspace-like property when selection, limited capacity, and broadcast are operationally verified. Whether such properties are sufficient for experience is disputed.

Does ANIMA measure consciousness?

No. It measures package-defined state, dynamics, and proxies. Its historical names are not validated consciousness measurements.

Can current AI be conscious?

There is no settled scientific answer or accepted decisive test. That uncertainty is a reason for better operational evidence, not permission to promote any output as proof.

What should developers measure now?

Persistence correctness, recall quality, intervention effects, workspace behavior, calibration, task outcomes, robustness, privacy, and failure recovery.

CB
Christian BucherIndependent builder in Vienna · QTool, Soul MCP, Postcondition, ANIMA
Explore ANIMA KernelAudit the source ↗ANIMA vs system prompts