Skip to main content
Agentic AI architecture · Concept guide

The 7 layers of an AI agent

Every agent — from a simple chatbot with tools to a multi-step autonomous system — is built from the same stack. Here's what's actually inside it, layer by layer.

7 layers 4 memory types 4 planning patterns 2 wrapping systems
SAFETY & GUARDRAILS — WRAPS EVERY LAYER OBSERVABILITY — TRACES EVERY LAYER 1 · Perception Ingests text, images, files, voice, API events 2 · Short-term memory The context window — finite, and cleared each session 3 · Long-term memory Semantic, episodic, procedural, user/entity 4 · RAG pipeline Bridges long-term memory into the context window 5 · Planning ReAct, Chain-of-Thought, Reflection, Tree of Thought 6 · Tools Web search, code execution, APIs, other agents 7 · Observe & decide Checks the goal, then returns an answer or loops again

One reasoning loop, top to bottom — with two systems wrapped around all seven layers rather than sitting inside any one of them.

1

Perception

The agent ingests whatever comes in — text, images, files, voice, API events — and tokenizes it into a form the model can reason over.

2

Short-term memory — the context window

Think of it as the agent's RAM. It holds the system prompt (identity and rules), retrievals pulled from a vector database, recent conversation turns, and the results of the last tool call. Everything the agent "knows right now" lives here — and it's finite.

Layer 3

Long-term memory: four types

Without it, every conversation is a blank slate.

Semantic

Factual knowledge stored as embeddings — Qdrant, Pinecone.

Episodic

Memory of past sessions — "last time you asked this, I…"

Procedural

Saved workflows that have worked before.

User / Entity

Who you are, your preferences, your history.

Layer 4

The RAG pipeline

The bridge between long-term memory and the context window — how an agent answers questions about your private data without hallucinating facts.

Chunk Embed Index Retrieve Inject
Layer 5

Planning: the agent doesn't just react

Four common patterns for how an agent thinks before, during, and after acting.

ReAct

Thought Action Observe

Thought → Action → Observe → repeat.

Chain-of-Thought

Act

Reason step-by-step before acting.

Reflection

Generate Critique Revise

Generate → critique → revise.

Tree of Thought

best path

Explore multiple paths, prune the bad ones.

Layer 6

Tools — the agent's hands

Web search, code execution, database queries, REST APIs, and even calls to other agents.

Web search Code execution Database queries REST APIs Calls to other agents

Model Context Protocol (MCP) is becoming the standard connector for all of this — read the architecture guide.

Layer 7

Observe & decide — the loop closes

After every action, the agent checks one question: did I achieve the goal? If yes, it returns the answer. If no, it loops again — until the task is done or a safety limit is hit.

Plan Act Observe Goal met? No → loop again Yes → return answer
The part most tutorials skip

What wraps around all seven layers

None of this runs in isolation — two systems sit around the whole stack, not inside any one layer of it.

Safety & guardrails

Input and output inspection, sandboxed execution, and human approval before any high-stakes action runs.

Observability

Every step traced, every token logged — so you can debug why the agent got it wrong.

See how this connects to MCP

Tools are one of the seven layers — and MCP is the protocol that's standardizing how agents reach them.

Read the MCP architecture guide