A claim ladder for evaluating what modern AI systems can do without turning impressive behavior into proof of human-like understanding.
I disagree with two confident claims that often appear in the same AI debate.
The first is that a model must understand because it can explain a difficult concept, solve a problem, write working code, or maintain a thoughtful conversation. The second is that none of those achievements matter because the model is “only predicting tokens.”
Both statements skip too much.
Impressive behavior is evidence of capability. It may also give researchers clues about internal representations, planning, and generalization. But behavior alone does not settle whether a system understands the world as a person does, has a stable point of view, or experiences anything. At the same time, describing the training objective does not explain away every useful capability that emerges from the trained system.
For practitioners and leaders, this is not merely a philosophical argument. Careless language changes product requirements, risk decisions, user expectations, and accountability. If a team says an assistant “understands company policy,” people may give it authority that its evaluation has not earned. If a team dismisses the same assistant as a stochastic toy, it may ignore a system that can already create value or harm at scale.
The useful position is neither wonder nor contempt. It is disciplined separation: state exactly what has been observed, name the claim that observation supports, and stop before the evidence runs out.
The word understand is convenient because it compresses several different questions. That convenience is also the problem. A product manager, an engineer, a cognitive scientist, and a user can say “the model understands” while referring to four different things.
This claim ladder makes the differences explicit:
| Level | Claim | Evidence that can support it | What the evidence does not establish |
|---|---|---|---|
| 1. Fluent expression | The model produces coherent, context-sensitive language | Blind review, instruction-following tests, consistency checks | Accuracy, stable beliefs, or subjective experience |
| 2. Task competence | The system performs a defined task under stated conditions | Representative evaluations, error analysis, production outcomes | General competence outside the tested conditions |
| 3. Internal representation | The system encodes useful structure about concepts or environments | Generalization tests, interventions, interpretability research | A human-like conceptual scheme or lived connection to the world |
| 4. Reflective or agentic behavior | The system monitors steps, plans, uses tools, or reports on its state | Traces, controlled experiments, repeated behavior across contexts | Reliable introspection, personal identity, desires, or awareness |
| 5. Conscious experience | There is something it is like to be the system | Theory-based indicators and converging scientific evidence would be needed | This cannot be inferred from fluent self-report alone |
Moving up the ladder requires new evidence. A strong result at one level does not automatically prove the next.
This is the article’s central rule. “It answered correctly” supports a bounded performance claim. “It represented relationships well enough to transfer to a new task” supports a richer computational claim. “It said it felt afraid” is an output that may deserve investigation, but it is not direct access to a private experience.
That discipline works in the other direction too. Uncertainty about consciousness does not erase observable competence. A system does not need human-like comprehension to route support requests, detect a pattern in code, retrieve a policy, or cause a serious incident.
Human beings normally encounter rich language as evidence of another mind. That inference works reasonably well in ordinary life because the speaker usually has a body, a history, needs, relationships, memory, and a place in the same physical and social world.
An AI conversation presents the language without giving us the same background evidence. The interface is familiar, so users supply the missing personhood themselves. Pronouns, typing indicators, a friendly voice, first-person statements, and remembered preferences can strengthen that response.
The system’s fluency is real. The user’s social reaction is real. The conclusion that the system therefore has a human-like inner life does not follow automatically.
NIST’s Generative AI Profile treats confabulation and human-AI interaction as practical risk areas. That is the right operational framing. A fluent response can encourage a person to overestimate accuracy, intention, or authority even when nobody has made a formal claim about consciousness.
This is why wording in a product matters. “The assistant found three passages related to your question” is an observable statement. “The assistant understood your situation” suggests far more. The second may increase trust while making the actual system boundary harder to see.
If a product is meant to support a decision, communicate what it did:
These descriptions may sound less magical. They are more useful because another person can inspect them.
Suppose a coding model fixes a concurrency bug it has never seen before. It reads the surrounding code, proposes a change, writes a test, and explains why the failure was intermittent. Calling this “mere autocomplete” hides the practical achievement. The system combined patterns in a way that solved a real problem.
But one success does not establish general understanding of software systems. Perhaps the test is incomplete. Perhaps the patch breaks a performance assumption. Perhaps a small change to the prompt produces a different answer. Perhaps the model cannot recognize the same concurrency issue when it appears behind another abstraction.
The correct response is to test the capability, not settle the metaphysics.
Ask which task distribution is represented, how performance changes under variation, which errors cluster together, and what happens after the model, prompt, tools, or context change. OpenAI’s 2026 work on deployment simulation illustrates this behavioral approach: it uses deployment-like conversations to estimate unwanted model behavior and find gaps that selected evaluation prompts can miss. The work studies what systems do under more realistic conditions. It does not need to claim what those systems feel.
This distinction keeps evaluation honest. A benchmark can show that a model solved a class of problems. It cannot by itself tell us whether the solution involved the same concepts, experience, or process a human used. Conversely, different internal processes do not make the result irrelevant. Airplanes do not fly as birds do, yet lift remains measurable.
When the next action is engineering, the most important questions are usually operational:
Those questions produce a system a team can govern. A general claim that “the AI understands” does not.
Modern models can encode relationships that are not obvious from a single sentence. They can sometimes track objects in a scenario, predict consequences, navigate a simulated environment, or transfer a learned pattern to a new context. Researchers reasonably study whether such behavior reflects internal world models.
The phrase world model, however, can describe a functional representation used for prediction. It need not mean that the system encounters a world in the human sense.
A map represents a city without walking its streets. A database represents customers without meeting them. An internal model may be much more dynamic and powerful than either, yet the same logical caution applies: representation, access, embodiment, agency, and experience are separate properties.
This matters in business systems because teams sometimes infer too much from a model’s broad knowledge. An assistant may explain procurement policy while lacking access to the current approved document. It may reason about customer frustration while having no access to tone, history, or the consequences of a bad recommendation. It may construct a convincing narrative from an incomplete data extract.
The system can have useful encoded structure and still be missing the local reality that controls the decision.
That is why verifying AI answers before trusting them requires tracing evidence and separating observation from inference. The philosophical uncertainty is interesting, but the immediate professional obligation is concrete: do not let a broad impression of intelligence replace a check on the evidence available in this case.
Ask a conversational model whether it is conscious and it may deny consciousness, claim consciousness, qualify the question, adopt a fictional role, or mirror the assumptions in the prompt. The answer can change with system instructions, training choices, sampling settings, model versions, or the preceding conversation.
That variability should make us cautious about treating the response like human testimony. A person is normally considered to have special access to their own experience. We have not established that a language model’s first-person report has the same relationship to an internal state.
This does not mean every self-report is worthless. Repeated, controlled behavior may reveal how a model represents itself, follows identity instructions, detects its own uncertainty, or responds to pressure. Those are researchable questions. The error is renaming one of those measurements “proof of consciousness” before the causal link has been established.
The same caution applies to apparent emotion. A model can produce language associated with fear, frustration, preference, or pain. That output can affect users and should be designed responsibly even if it reflects no felt state. The human consequence does not depend on resolving the model’s moral status first.
For product teams, a good default is simple:
Rejecting weak evidence is not the same as proving that machine consciousness is impossible.
A 2026 peer-reviewed paper in Trends in Cognitive Sciences, “Identifying indicators of consciousness in AI systems”, proposes deriving indicators from neuroscientific theories and examining whether particular AI architectures satisfy them. The authors also emphasize the field’s uncertainty. This is a much stronger method than asking whether a chatbot sounds alive, but indicators still inform a degree of belief; they are not a consciousness meter.
The subject is no longer limited to abstract philosophy. Anthropic publicly began a model welfare research program in 2025 while stating that there is no scientific consensus on whether current or future systems could be conscious. That combination—investigation without certainty—is a sensible posture.
Organizations should preserve two possibilities at once:
This is not indecision. It is calibrated uncertainty. We use similar reasoning whenever evidence is incomplete and the cost of being wrong is asymmetric. The right response may be continued research, documentation of architecture and training changes, outside review, or low-cost precaution—not a confident declaration based on a conversation screenshot.
Leaders also need to keep two governance tracks separate. One track concerns harm caused by AI systems: unsafe advice, discrimination, privacy loss, manipulation, security failures, or unaccountable automation. Those risks are present now and can be managed through evaluation, controls, monitoring, and ownership. The other concerns possible harm to an AI system if it ever warrants moral consideration. That question needs specialist research and should not displace today’s clear responsibilities to people.
In my data and AI teaching, one habit I return to is asking learners to define what would count as evidence before they ask a model for an answer. Without that standard, a polished response can become its own proof.
Teams can use the same habit in product and leadership discussions. Whenever the word understands appears in a requirement, demo, vendor claim, or strategy document, rewrite the sentence using this short claim review:
| Review question | Example for a policy assistant |
|---|---|
| What behavior did we observe? | It retrieved the controlling policy and produced a supported answer |
| Under which conditions? | Current English-language policies for three departments |
| What evidence supports the claim? | A versioned evaluation set plus review by policy owners |
| What nearby claim remains unproven? | It has not shown reliable performance on conflicting or obsolete policies |
| What action may follow? | It may draft an answer; a person owns consequential exceptions |
| What change forces reevaluation? | A model, prompt, retrieval index, permission, or policy update |
This turns a philosophical shortcut into an engineering statement.
It also complements the word-to-test contract for vague AI requirements. If a requirement says the model should “understand user intent,” define the relevant intents, ambiguous cases, evidence, acceptable failure rate, and escalation behavior. If a vendor says its agent “reasons like an expert,” ask which decisions were tested, against whose judgment, with what tools and information, and under which failure conditions.
Then connect the claim to operating controls. AI reliability protocols matter because model capability is only one part of a dependable workflow. Observability, permissions, approval paths, regression tests, incident response, and rollback determine how much authority that capability should receive.
The claim review does not ban ordinary language. People can still say a model understood the request in casual conversation. The discipline matters when the phrase carries authority, money, risk, or a scientific conclusion.
We do not need to prove that an AI system thinks like a person before taking its capabilities seriously. Systems can be economically valuable, operationally dangerous, socially persuasive, or worthy of careful study without sharing human consciousness.
We also do not need to turn every impressive output into evidence of a hidden person. That leap makes evaluation weaker, not stronger. It replaces measurable behavior with a story that is difficult to test.
So ask the narrow question first. What did the system do? Under what conditions? How stable was the result? What information and tools did it have? Which conclusion does the evidence support? Which conclusion remains open? What authority should follow?
If the question is task performance, evaluate the task. If it is internal representation, study the mechanism. If it is consciousness, use theories and evidence suited to consciousness research. Do not ask one kind of result to settle all three.
Modern AI is already important without mythology. Its behavior deserves serious measurement. Its failures deserve controls. Its influence on people deserves responsible design. Its possible future moral status deserves careful research.
Clear claims let us do each of those jobs without pretending they are the same job.