Chapter 4
04How Does AI “Think”?
“A machine’s thinking is not like a human’s; it is advanced pattern processing built on mathematical foundations.”
To apply AI, we must first remove the mystery around how it works. It is often called a magical “black box,” yet it is an engineered system built on understandable logical principles. In this chapter we break the box into its components: the key differences between the terms, how a model builds its “understanding,” and the cycle it follows to make any decision.
Three Nested Circles
Imagine three circles, each inside the other:
- Artificial intelligence: the umbrella: any computer system that mimics human cognitive tasks such as problem-solving, language understanding and decision-making. It is the machine’s “made” intelligence.
- Machine learning: the main way to achieve AI today. Instead of telling the computer “pointed ears and a small mouth mean a cat,” we give it thousands of labelled examples of cats and dogs and it learns to tell them apart.
- Deep learning: the most powerful technique inside ML: multi-layer neural networks that learn highly complex patterns traditional models could not.
The “Layer Cake” Analogy
To understand how a deep-learning model thinks, picture its mind as a layer cake. When it receives an image, it does not grasp it all at once; it processes it through hierarchical layers:


Text in this figure
Top layers · concept: “this is a cat” · Middle layers · parts: eye, ear, nose · Bottom layers · features: edges, lines, curves · image pixels · Figure 7
- Bottom layers: very simple; they detect basic features: edges, lines and curves.
- Middle layers: combine simple features into more complex shapes, learning to recognize an “eye” or an “ear.”
- Top layers: assemble the complex parts; seeing two eyes, two ears and a nose in the right arrangement, they conclude: “this is a cat.”
This “hierarchical learning” is what gives deep learning its power to understand images, sounds and text.
From Cake to “Attention”: How Language Models Read
2026 Update
Modern language models are built on the Transformer architecture. Text is split into small units called “tokens,” and each token becomes a numeric vector carrying its meaning (an embedding). The “attention” mechanism then lets each word “look at” every other word and weigh its relevance; the word “bank” learns from its neighbors whether it means a riverbank or a financial institution. Attention layers stack like the layers of the cake until the model predicts the most probable next token. The span of text a model can “see” at once is called its “context window.”
The Core Thinking Cycle: From Perception to Action
However complex the model, most applied systems follow a four-step cycle. Take a self-driving car:


Text in this figure
Self-driving · car · Perception · camera sees a red light · Inference · red means: stop · Decision · apply the brakes · Action · the car stops · Figure 8
- Perception: gather data from the world with sensors: the camera “sees” a red traffic light.
- Inference: analyze the data to grasp its meaning: the model “infers” that red here means “stop.”
- Decision: choose the best course of action: the system “decides” to apply the brakes.
- Action: execute the decision in the world: the mechanical command is sent and the car stops.
Perceive → infer → decide → act: this cycle governs almost every AI application we meet, from voice assistants to automated trading, and on to the smart agents of Chapter 10.
Lessons Learned
- 1AI is not magic; it is an engineered system built on layers of pattern processing.
- 2Machine learning is “how” AI works, and deep learning is its most powerful tool.
- 3Machine thinking is hierarchical, like a layer cake: from simple to complex.
- 4Language models read through tokens, vectors and attention, and predict the next token.
- 5Every application follows a cycle: perceive, infer, decide, act.
Tip: use ← → to move between sections.

