The productive question for developers is not “did the model sound alive?” It is “which property changed, how was it measured, what intervention caused it, and which stronger conclusions remain unsupported?”
A four-level evidence ladder
Every result should be placed on the narrowest level it directly supports. Moving upward requires additional evidence; attractive terminology cannot do that work.
Observable behavior
Outputs, task scores, error recovery, self-reports, tool calls, and consistency across trials.
DIRECTLY TESTABLEArchitecture
Persistent state, recurrent processing, limited workspace, global availability, self-model, and explicit interventions.
DIRECTLY INSPECTABLEInternal proxies
Integration scores, salience, affect-inspired values, complexity metrics, or theory-derived indicator properties.
DEFINITION-DEPENDENTSubjective experience
Whether there is something it is like to be the system: sentience or phenomenal consciousness.
NOT DIRECTLY MEASUREDA model saying “I feel afraid” is Level 1 behavior. A durable fear-like variable is Level 2 architecture. A composite score derived from it is Level 3. None of those observations alone establishes Level 4.
Self-report is generated data
Language models learn patterns of first-person reporting and respond to context. Self-reports may be inputs to a carefully designed evaluation, but they are not privileged access to an otherwise hidden ground truth.
What consciousness theories contribute
Theories can inspire precise computational properties. They do not make a software system conscious by analogy.
Global Neuronal Workspace Theory links conscious access in brains to selection, amplification, and broad availability of information. An AI developer can implement a capacity-limited workspace and test whether a selected item becomes available to multiple modules. That demonstrates workspace-like information flow, not the biological mechanism and not subjective experience.
Integrated Information Theory makes claims about the intrinsic cause-effect structure of a physical system. A convenient software statistic called phi is not automatically IIT's Φ. Exact and approximate measures have formal assumptions; renaming a coupling score cannot transfer their interpretation.
Predictive processing motivates models that compare predictions with observations and update their state. Prediction-error processing is widespread in machine learning and control systems. Its presence is useful engineering evidence, but not a consciousness verdict.
A 2025 open-science adversarial collaboration directly tested distinctive predictions of IIT and GNWT in 256 human participants using fMRI, MEG, and intracranial EEG. Results challenged important predictions from both accounts rather than producing a simple winner. If leading theories remain under active empirical pressure in humans, software analogies deserve even more care.
Indicators are not a consciousness checklist
A widely discussed 2023 report translated several scientific theories into computational “indicator properties” for assessing AI systems. That is a more disciplined approach than judging conversational style. But indicators are still theory-dependent evidence, not a binary certification test.
- An indicator needs an operational definition that independent teams can implement.
- The property should survive adversarial interventions, not only descriptive demos.
- Alternative explanations such as prompt imitation, memorization, and evaluator leakage need controls.
- Results should include uncertainty, negative findings, and the system boundary being assessed.
- Several weak indicators should not be summed into an impressive number without validation.
ANIMA as an open case study
The maintained ANIMA Kernel is a zero-runtime-dependency Python package for persistent cognitive state. It implements associative recall, limited workspace selection, temporal context, affect-inspired variables, a simplified self-model, JSON persistence, and model-provider bridges.
Historical API names include ConsciousnessState, Phase.CONSCIOUS, Phi, and CQI. They are provenance, not proof. The maintained documentation now defines them as implementation-level state and proxies.
The test count is strong implementation evidence. The benchmark is exploratory: a single saved run, no confidence intervals, no preregistration, and no independent replication. The correct conclusion is that specific components changed specific package-defined measurements under that harness.
| Claim | Evidence ANIMA has | Evidence still needed |
|---|---|---|
| State survives restarts | Persistence code and automated round-trip tests | Workload-specific durability testing for production use |
| Workspace is capacity limited | Inspectable implementation and unit tests | Independent behavioral benefit across real agent tasks |
| Components affect proxy scores | Saved ablation artifact | Repeated trials, uncertainty, stronger controls, replication |
| Agent outcomes improve | Small exploratory neutral-kernel comparison | Predefined external tasks and model/database baselines |
| System is conscious | None | No accepted decisive test currently exists |
A practical evaluation protocol
- Define the boundary. State whether you are evaluating the base model, prompt, memory service, full agent loop, or physical deployment.
- Name the property. Replace “more conscious” with persistence accuracy, recall precision, cross-module availability, calibration, or another operational measure.
- Intervene. Disable or randomize one component while holding model, prompts, inputs, and sampling settings constant.
- Repeat. Report distributions and uncertainty, not one attractive run.
- Use external outcomes. Score tasks independently of the internal metric being promoted.
- Publish nulls. A component doing nothing is information, not a marketing failure.
- Keep the conclusion narrow. Say exactly what changed and stop before the next unsupported rung.
Why the distinction improves products
Honest boundaries are not anti-ambition. They make an unusual project easier to adopt. A developer can use ANIMA's persistence and workspace without accepting a theory of consciousness. A researcher can replace its proxies. A reviewer can reproduce a claim. Clear falsifiability creates more credibility than a dramatic demo that cannot survive inspection.
Build the capability. Label the evidence. Leave the mystery open.
Persistent memory, global availability, self-monitoring, and affect-inspired control signals can be valuable engineering ideas. Their value does not depend on declaring the system alive.
Primary sources
Frequently asked questions
Can an AI saying it feels something prove consciousness?
No. It proves that the system generated that text under that context. Treat self-report as behavior to explain, not as privileged ground truth.
Does a global workspace make an AI conscious?
It demonstrates a workspace-like property when selection, limited capacity, and broadcast are operationally verified. Whether such properties are sufficient for experience is disputed.
Does ANIMA measure consciousness?
No. It measures package-defined state, dynamics, and proxies. Its historical names are not validated consciousness measurements.
Can current AI be conscious?
There is no settled scientific answer or accepted decisive test. That uncertainty is a reason for better operational evidence, not permission to promote any output as proof.
What should developers measure now?
Persistence correctness, recall quality, intervention effects, workspace behavior, calibration, task outcomes, robustness, privacy, and failure recovery.