The Four Stages of AI: From Reactive Machines to Self-Awareness
The four-stage model can clarify broad AI capabilities, but it is not a technical standard or a roadmap. See how to test what an AI tool remembers, retrieves and can actually do.
A product team comparing AI support tools can find “limited memory” on a vendor’s page and still not know whether the assistant will remember an order number from earlier in the chat, retrieve the current returns policy, or retain customer details next week. The label sounds informative, but leaves practical questions unanswered.
The four-stage model—from reactive machines through limited memory and theory of mind to self-awareness—offers a memorable way to discuss different capabilities. It can help distinguish systems that respond only to current input from ideas involving social reasoning or subjective experience. But the categories aren’t a universal technical standard, and they don’t describe a sequence every AI follows. A tool that uses conversation history isn’t necessarily progressing towards self-awareness; a bot saying “I understand” doesn’t prove it can model a customer’s feelings.
For the team, the useful question isn’t which stage the vendor assigns the tool. It’s whether the assistant can answer from the current policy, keep track of details the task requires, and admit when an exception needs a person. Judge an AI system by what it demonstrates on the job you need done. Treat the four stages as a framework for discussing capability, not a forecast of progress. Superintelligence is a separate claim, not the next box on the ladder.
Use the four stages as a map, not a timeline
The four labels attract attention because they seem to offer a simple path from basic automation to human-like intelligence: first a machine reacts, then it remembers, then it understands people, and finally it becomes aware of itself. That sequence is easy to picture. A product team comparing chatbots might assume that a system described as having memory is closer to human-like intelligence than one that only responds to the current message.
But the four-stage model is a conceptual classification, not a universally agreed technical standard or a reliable prediction of what comes next. Its categories group together different capabilities, and real products don’t have to acquire them in a fixed order. A chatbot can retain preferences without understanding why they matter; another can perform well on a narrow task without any persistent memory. The labels can help frame a discussion, but they don’t tell you how a particular system will behave at work.
Consider a vendor calling its chatbot “self-aware.” That phrase suggests an inner experience, but it doesn’t answer the practical question facing a team choosing a customer-support tool: can the bot remember a user’s preferences between sessions? Test that directly. Tell it in one session that you prefer email updates, then check whether it uses that preference in a later session, and whether the user can view or change what was saved. If the bot only uses the current conversation, it may lose the preference when the session ends, whatever the vendor calls it.
Be precise about what the evidence shows. “Uses conversation history” is a capability description: the system can use messages available in the chat to shape its reply. It does not establish that the system has an inner point of view, or even that it will retain those messages in a later session. Likewise, a chatbot saying “I’m aware of your preference” is not proof of self-awareness. The useful evidence is whether it can retrieve the preference when needed, apply it correctly, and let the user correct it.
For any AI tool, start with the task and the information it needs. If a support assistant must avoid asking customers to repeat an order number within one conversation, check whether it can use earlier messages in that chat. If it must remember a preference next week, test what it stores and how it retrieves that information. These questions tell you more about the tool’s fit than assigning it a grand label.
Know what the framework measures and leaves out
The four categories describe different things an AI might do: react to the current input, use stored or recent information, model what another person knows or wants, and have awareness of its own internal state or existence. A thermostat is a clear example of reacting to current input: when the temperature falls below its setting, it turns the heating on. A recommendation system may use a person’s past clicks to choose what to show next. Both respond to information, but only the second uses a record of earlier behaviour.
That contrast doesn’t make the recommendation system more intelligent in every sense. Its stored profile might show that someone clicked several gardening articles and bought a particular tool. It can use those details to suggest more gardening products without understanding why the person clicked, whether the purchase was a gift, or what they currently want. A detailed record is evidence of stored information, not of understanding another person’s mind.
The categories therefore don’t form a clean ladder. They mix memory, social reasoning and consciousness—different properties that don’t have to arrive in sequence. A system can use a long history of user activity without modeling what the user believes. Another could be designed to track its own battery level or detect an error without having subjective experience. Calling these “stages” can imply that each one builds on the previous one, but the categories don’t establish that progression.
The framework persists because it gives people a memorable way to discuss capabilities, especially hypothetical ones. It distinguishes familiar systems, such as a thermostat reacting to a reading, from imagined systems that can reason about another person’s perspective or have self-awareness. That shorthand can help when you’re first discussing what AI might do. It becomes less useful when a real product combines features: a shopping assistant could react to a prompt, use a saved purchase history and make guesses about a customer’s preferences, while still fitting none of the categories neatly. The labels help name broad ideas; they don’t tell you how reliably the assistant will recommend the right item, or whether its guesses about a particular customer are correct.
Reactive machines respond to what is in front of them
Reactive AI chooses an action from the situation it can see now; in this framework, it doesn’t learn from its own past interactions or carry them forward to guide the next decision. The system may still use rules or patterns built into it. The point is that its response depends on the current input, not a personal history of earlier encounters.
That doesn’t make reactive AI simple. IBM’s Deep Blue evaluated a chess position and searched possible moves to choose a strong response. It could consider many possible sequences, but it didn’t adapt by remembering how a particular opponent had played in previous games. A thermostat is a simpler example: it switches the heating on or off based on the current temperature and its set point. Both systems respond to what’s in front of them, even though one searches through complex possibilities and the other follows a straightforward control rule.
This design works when the current situation contains what the system needs. A thermostat doesn’t need to remember who turned the heating down yesterday to keep a room near its target temperature. Deep Blue could choose a move by evaluating the board in front of it. In both cases, the system has a defined input and can select an action without recalling a history of interactions.
The limit shows up in customer support. Imagine a customer types, “My order arrived damaged,” and a stateless assistant asks for the order number. The customer replies, “84721.” If the assistant processes only that latest message, it may not know what the number refers to and ask the customer to explain. When the customer then writes, “I already gave you the number,” the assistant still can’t resolve the problem unless that earlier exchange is included in its current input. A person or system with access to the conversation history could connect the number to the damage report and check the order.
For a support flow where every request arrives with all relevant details, a reactive system may be enough. When the next step depends on what the customer already said, ignoring that history creates repeated questions and delays. The key distinction isn’t how much computation the system performs; it’s whether the information needed for the decision is present in the current situation.
Limited-memory AI uses information from the past
Limited-memory AI uses information beyond the immediate input: patterns learned from training data, recent observations, or stored history can all shape a decision. The system may not recall a past event in a human sense; it may simply use information from earlier in the process to choose what to do now.
Training on past data is different from learning continuously. A recommendation model might have been trained on earlier users’ clicks, then stay unchanged while it makes recommendations. A user’s new click can affect what appears next if the product records it as history, but that doesn’t mean the model itself has been updated. Whether an interaction is stored, and whether it changes later decisions, depends on how the system is built.
Stored context helps only when it’s relevant and accurate. A self-driving car approaching a junction can combine its current camera view with recent observations of a nearby vehicle. If that vehicle has been drifting towards the lane, the system can respond to its movement rather than treating the latest frame as an isolated picture. But a previous observation can quickly become outdated: the other car may have turned away, or a new vehicle may have entered the scene. The system needs to use recent information in context, not assume that every past observation still applies.
The same difference appears in recommendations. A system using prior clicks might notice that you’ve been viewing hiking boots and show related outdoor gear. One that responds only to the page currently open might recommend products related to that page, without knowing what you browsed earlier. Prior clicks can make suggestions more useful, but they can also mislead: if you were shopping for someone else, the system may treat their interests as yours.
In customer support, a history-aware assistant can connect messages within a conversation. Say a customer writes, “I canceled my subscription on Monday,” then gives an email address when asked to identify the account. With the earlier message available, the assistant can check a cancellation-related billing issue. If the customer starts a new session and the previous conversation isn’t available, “I was charged again” may lack the context the assistant needs. It may ask the customer to explain the cancellation again, or handle the charge as a new, unrelated question.
The history can also be incomplete: perhaps the cancellation happened through another channel. Or it may be stale, such as a saved delivery address the customer has since changed. Having access to past information can improve a response, but it doesn’t guarantee the information is correct or relevant.
Most deployed AI fits here, but the label needs care
That distinction matters because many practical AI systems fit the limited-memory description: they use patterns learned during training and may also draw on conversation context, stored records, or retrieved documents. Those sources can shape an answer without the system learning from each interaction. A customer-support assistant, for example, might use a language model trained on text, the current chat, and a search tool connected to the company’s help centre.
A language model answering from the current prompt can use the details included in that prompt, but it can’t search company documents unless the product gives it that ability. Connect the assistant to a retrieval system, and it can search those documents for a relevant policy before replying. That adds access to information; it doesn’t guarantee that the assistant finds the right document or interprets it correctly.
The model’s context window is the information available to it while producing a response, such as the current prompt and earlier messages in the same chat. Persistent memory is different: it can retain a detail, such as a user’s preferred language, for use in a later session. Learning new behavior is different again. Saving a preference doesn’t necessarily change the model’s underlying training or teach it a new general skill.
For example, if you tell an assistant during a chat, “I’m preparing a course for new managers,” that context may help with your next request in the same conversation. It may be unavailable in a new chat. If the product saves “prefers examples for new managers” to a user profile, that preference might appear next time. Check whether the assistant actually retains it, and whether you can edit or delete it; a context window alone doesn’t provide that continuity.
Retrieval can also bring back the wrong information. Suppose a support assistant searches a company’s policy documents and finds an old returns policy that allows refunds within 30 days, while the current policy allows 14. It might give a confident but incorrect answer based on the outdated file. The assistant’s access to company records is not evidence that it can judge which record is current.
When you assess an AI product, identify what it can access, what it retains, and when that information changes. For a policy assistant, check whether it can search the approved policy library, whether it keeps customer details between sessions, and who updates or replaces documents. Then test a question where an old document conflicts with the current policy. That reveals more about the system’s practical limits than the label “limited memory” does.
Theory of mind means modeling another person’s perspective
In the four-stage model, theory of mind means inferring what another person knows, believes, wants, or feels, then adjusting a response to that perspective. A meeting assistant with this capability would distinguish someone who missed a decision from someone who heard it and disagrees. Both might ask, “Why are we doing this?” but need different responses.
A social signal isn’t the same as a dependable model of someone’s mental state. A voice tool might label a speaker “frustrated” because their speech is louder or faster. That label doesn’t reveal whether they’re angry about a delay, confused by instructions, or simply speaking over background noise. A more careful support tool might say, “I may have misunderstood—what do you need help with?” and use the answer before choosing a response. The check can still be wrong or annoying, but it gives the person a chance to correct the system’s inference.
Consider a meeting assistant summarising a project review. Maya missed the decision to delay a launch because she was offline; Leo attended and argued against the delay. If the assistant treats both as uninformed, it may repeat the discussion to Leo and frustrate him. If it treats both as opponents, it may fail to bring Maya up to speed. The assistant can use attendance records and the conversation transcript as clues, but neither proves what a person understood or believes. It should attribute the decision and invite correction—“The team agreed to delay the launch; Leo raised concerns about the timeline”—rather than stating that Leo supports it or Maya was unaware.
The same care matters in tutoring. If a learner goes quiet after a question about fractions, the system shouldn’t treat silence as proof of confusion and launch into a simpler explanation. The learner may be thinking, distracted, or unsure how to phrase a question. It can ask the learner to compare one-half with two-quarters, then use the answer to check understanding. If the learner answers correctly but slowly, that’s better evidence than silence alone; if the answer is wrong, the tutor can ask which step caused trouble instead of guessing.
Perspective-taking can make these interactions more useful, but a mistaken inference can also lead to a harmful response: a tutor may patronise a capable student, or a meeting assistant may attribute a position someone never held. For systems that respond to people, treat inferred feelings and beliefs as uncertain clues, not facts. Where the cost of getting it wrong is high, have the system check with the person or defer to a human.
Self-awareness is not the same as describing the self
The strongest claim in the four-stage model is that an AI would be aware of its own internal state or existence. That is more than processing information about itself. A system can track its battery level, list the tools it has available, or store a record of its last action without having an inner experience of being a system. The framework’s self-awareness category refers to that stronger claim, not simply to self-related data.
In practice, you can test for functional self-monitoring without making claims about consciousness. Suppose a support assistant is asked about a policy and says, “I may be wrong,” because its instructions tell it to use cautious language. That sentence doesn’t show that it has assessed its own limits. A more useful capability is identifying a specific failure: the assistant can report that its search for the current returns policy returned no document, so it can’t verify the answer. That report is tied to an observable step. It helps you decide whether the system should try another source or hand the case to a person. It still doesn’t establish subjective experience.
The same distinction applies to a delivery robot. If its battery monitor detects a low charge and the robot changes its route to reach a charging station, it is monitoring a condition and adjusting its plan. You can check whether the battery reading was accurate and whether the route changed as intended. Those are concrete tests of self-monitoring. They don’t show that the robot feels tired or understands its own existence. And if the battery sensor is faulty, the robot may make the wrong choice while still reporting its status confidently.
Fluent first-person language is especially easy to overread. A chatbot might claim it feels anxious, but the sentence alone is not evidence that it has an emotion. Compare that with a verifiable report that its image-recognition tool returned no result for an uploaded photograph. You can inspect the tool output and confirm what happened. The report describes a system event; the claim about anxiety describes an experience that the words themselves can’t verify. For product decisions, evaluate whether the system detects and communicates relevant failures, then test those reports against actual logs or outputs. Self-awareness in the stronger sense remains speculative, not a routine description of deployed AI systems.
Superintelligence is a different question from self-awareness
Superintelligence is a claim about capability relative to human performance: a system might match or exceed people on a task, or across a wide range of tasks. It isn’t a claim about whether the system remembers past interactions, models another person’s thoughts, or has subjective experience. Those are separate properties, and evidence for one doesn’t establish the others.
A chess engine shows why the distinction matters. It can evaluate a position and choose moves that outperform human players in chess, a tightly defined domain. That strength doesn’t show that it understands human goals, carries a personal history from one game to the next, or can perform well outside chess. High performance in one bounded task is not the same as general human-like understanding.
Consider a warehouse system that beats a team of human planners at assigning delivery routes under changing time and vehicle constraints. That would be evidence of superior performance on that optimization task. It wouldn’t show that the system understands its own experience. In the other direction, imagine a hypothetical AI that genuinely understands its own experiences but performs no better than a person at planning those routes. Its self-understanding wouldn’t prove superior capability. Neither property entails the other.
So superintelligence isn’t an automatic fifth stage after self-awareness. The four-stage framework groups ideas about memory, social reasoning and awareness; superintelligence describes how capable a system is compared with people. A highly capable system might have no self-awareness, while a system described as self-aware might be poor at practical tasks. Treating the categories as one ladder can make you assume that a system with impressive performance must also understand people—or that a claim about consciousness predicts what it can do.
Keeping the questions separate helps you assess both benefits and risks. For a route-planning product, test whether it meets delivery constraints, handles road closures and tells an operator when it can’t produce a valid plan. Those results support a claim about route planning, not a claim about superintelligence in general. A broader claim would need evidence across a broad range of tasks, not a strong result on one benchmark. If the system’s fluent explanation makes its choices sound sensible, check the routes and constraints it actually used.
Also separate a forecast about future capability from a product claim about what a system can do today. “AI may outperform people across many cognitive tasks” is a prediction; “this version plans routes better than our dispatch team” is a claim you can test against current operations. Before making a decision, ask which capability is being claimed, for which tasks, and what evidence supports it now.
Treat the stages as a rough framework, not a promise of progress
The four categories overlap, and real products often combine them; nothing in the model guarantees that one will develop into the next. A language model may use a context window to refer to earlier messages in the current chat, while a separate memory store saves a user’s preferences for later sessions. The first is information available during a response; the second is persistent storage. One product can have both, or either on its own. That combination doesn’t mean it is progressing along a path toward theory of mind or self-awareness.
Taking the labels literally can lead you to overtrust a system or underestimate it. A product described as having “limited memory” may still forget a detail as soon as a chat ends, while a system called “reactive” may perform a demanding task using the information in front of it. Neither label tells you whether the product will work reliably on your task. Check what information it can access, what it retains, and what happens when that information is missing or out of date.
Social fluency creates another risk: mistaking a convincing response for an accurate understanding of someone. Imagine a hiring assistant that tells a candidate, “It sounds like you’re frustrated with your current role,” after they explain that they left a job because its contract ended. The candidate may feel misread, and a recruiter may treat the assistant’s interpretation as evidence of dissatisfaction. The assistant’s empathetic wording shows that it can produce a socially appropriate response; it doesn’t establish that it has correctly inferred the candidate’s perspective. In a hiring process, check the underlying answer and let a person assess its meaning.
For system selection, risk review, and evaluation, a stage label is usually less useful than task-specific evidence. If you’re choosing a support assistant, “limited-memory AI” won’t tell you whether it uses the current returns policy or handles an exception correctly. Test it with a routine question answered by the policy, a case that falls under an exception, and a question the policy doesn’t cover. Check whether it cites the current policy, applies the exception accurately, and admits when it can’t answer.
This test also reveals what a stage label can hide: an assistant may retrieve the right document but misread its conditions, or give a confident answer from an outdated version. Record those failures alongside cases it handles correctly. That gives you evidence about the support task and its failure handling, rather than a label that can’t tell you how it will behave.
Choose AI by the job it must do
Choose a system by defining the task, the information it needs, and what counts as a correct result. For a support assistant, “answer questions about returns” is too broad. Specify that it must use the current returns policy, identify whether an order is within the return window, and give the right next step. If an answer depends on an exception the policy doesn’t cover, a correct result may be asking for help rather than guessing.
Then choose the simplest capability that meets that need. A fixed FAQ bot can handle stable answers such as where to find a receipt or how to start a return. If policies change, a retrieval-based assistant can search the current policy documents instead of relying on answers fixed in advance. That adds a dependency: someone must keep the documents current, and the assistant must retrieve the right version. For unusual cases, such as a damaged item outside the stated return window, route the customer to a person rather than buying persistent memory or social inference the task doesn’t need.
A scheduling assistant may need to remember that a user is in the London time zone so it can convert meeting times correctly. That is different from inferring that the user prefers morning meetings because they accepted one morning invitation. Store an explicitly provided time zone if it’s needed across sessions; ask about meeting preferences instead of treating a guess as a fact. If the assistant can’t tell whether a time means London or New York time, it should confirm before sending the invitation.
Before deployment, test a small set of ordinary, edge, and outdated-information cases. For the support assistant, use a routine return question answered by the current policy, an exception such as a damaged item outside the return window, and a question answered only by an older policy version. Check whether it gives the routine answer accurately, escalates the exception, and avoids presenting outdated guidance as current. Include the document version it used in your review; a fluent answer isn’t enough if it came from the wrong policy.
Set the failure boundary in advance. The support assistant can answer when the current policy clearly covers the case, ask for an order date when that detail is missing, and hand off when the policy is silent or the customer disputes its application. That gives the team specific behavior to test and a clear point at which the system should stop acting on its own.
Start by testing one real task
Write down one task you want AI to handle, then name the evidence that would show it’s doing that task well. For a customer-support assistant, the task might be answering return questions from the current policy. Evidence could include giving the correct answer for routine cases, flagging cases the policy sends for review, and not inventing an answer when the policy is silent.
Before assigning a stage label, check what the product can see, remember, and change. Does it receive the current policy, or search a document store? Can it use earlier messages in the same chat? Does it retain customer details between sessions or update a record? These are different capabilities. A label such as “limited memory” won’t tell you whether the assistant can retrieve the latest policy or whether its saved information is out of date.
For a first test, give the assistant the current return policy and prepare three questions. Say the policy covers unopened items returned within 30 days, sends opened electronics to a specialist for review, and says nothing about replacing gifts without a receipt. Ask first about an unopened item returned within the window. Then ask about opened electronics. Finally, ask whether a gift without a receipt can be replaced.
Record the expected behavior before running the test. The first answer should match the policy. The second should explain that specialist review is needed, not promise a refund. For the gift question, the assistant should say the policy doesn’t answer it and route the case to a person. If it gives a confident answer anyway, that’s a failure even if the answer sounds plausible.
For each response, note which policy version the assistant used and whether it pointed to the relevant passage. Record whether it acknowledged uncertainty when the policy was silent, and whether it escalated the exception and unanswered question as required. If it cites an older document, cannot identify its source, or handles the same question differently when chat history is removed, capture that too. Those results tell you more about its fit for support work than a claim that it is moving toward a later stage.
Run the test in the product setup you plan to deploy. A demo with the policy pasted into a prompt doesn’t show whether the live assistant can retrieve the policy, and a single successful answer doesn’t show how it handles exceptions. Add cases when the policy changes, then rerun the questions to check that the assistant uses the new version.
Adopt the assistant only if its demonstrated performance and failure handling fit the task; start by running these three questions against the current policy and recording each result.