"Let's just slap an agent on top of that."
Same meeting: for one person, it's an LLM call with a system prompt. For another, it's a multi-agent orchestrator with persistent memory and an execution sandbox. Nobody flinches. Everybody nods.
LLM, agent, agentic, long- / medium- / short-term memory, MCP, RAG, A2A… how do you file all of this, and what should I prioritize: buy, build, host, rent, distribute, centralize?
Hence this map: seven layers, two cross-cutting concerns. No ranking, no vendor to recommend to you — just enough to place each term and understand which layer your next problem is coming from.
Be warned, this is not a neutral model. Two biases run through it — treating security and observability as walls rather than as layers, and reading the same map in two opposite ways — and I own both further down. Other maps exist; this one has a point of view.
The map
Let's go back to the shelf. Seven levels stacked up, one for each kind of toy: a spot on each shelf, and the rule that holds everywhere — nothing gets left lying on the floor. Tidying up isn't shoving things in at random; it's knowing which shelf each thing belongs on.
Seven stacked slabs, then, and two walls running their full height alongside them: the cross-cutting concerns. Each slab carries its typical components painted on the floor, color-coded by their nature — a differentiating brick to build, a managed service, infrastructure, an open standard, a user-facing surface, or a still-emerging component.

Why are security and observability cross-cutting rather than layers of their own? Because they don't slip in between two strata: they apply to each one. A permission governs the tool call (layer 4) just as much as the interface surface (layer 7); a trace instruments inference (layer 1) just as much as multi-agent coordination (layer 6). Filing them into a layer demotes them to the rank of a mere module — and you realize too late that they should have been everywhere. The "nothing gets left lying on the floor" rule doesn't live on one shelf: it holds for all of them.
The seven layers
| # | Layer | Responsibility | Typical components |
|---|
| 1 | Compute & inference | Run the models and route requests | Accelerators (GPU/TPU), serving (vLLM, TGI), model gateways, multi-model routing |
| 2 | Models | Provide reasoning/generation capability | Foundation models, fine-tunes, small specialized models (SLM) |
| 3 | Orchestration / runtime | Run the agentic loop and manage state | Act-observe-verify loop, planning, state management, session memory |
| 4 | Tools & integrations | Give the agent the ability to act | Function calling, MCP, connectors, execution sandbox |
| 5 | Memory & knowledge | Provide durable, retrievable context | RAG, vector stores, knowledge graphs, persistent memory |
| 6 | Multi-agent coordination | Make several agents collaborate | Orchestrator/workers, delegation, sub-agents, protocols (A2A) |
| 7 | Interface & surfaces | Expose the agent to humans and systems | Chat, IDE, dashboards, scheduled execution (cron), webhooks |
Where it breaks
The map becomes genuinely useful when you use it as a diagnostic grid. A tidy shelf messes itself up the moment nobody knows which shelf anything belongs on anymore — each symptom below is a toy put back on the wrong level, or left on the floor.
| Symptom | Layer at fault |
|---|
| The bill explodes and nobody knows why | 1 — no gateway, so no per-call metering |
| Switching models forces you to touch business code | 1 — model access not centralized |
| The agent "forgets" in the middle of a long task | 3 — no state after context compaction |
| Every tool costs you a bespoke integration | 4 — no standard protocol |
| The agent did something destructive | A — unbounded blast radius, no approval |
| Impossible to tell whether v2 beats v1 | B — without evals, reliability is an opinion |
| Adding a channel (cron, webhook) = rewrite everything | 7 — agentic logic coupled to the surface |
If several rows ring true, there's a good chance they don't point to the same owner in your organization. And that, often, is the real problem.
Two readings of the same map
If you own the architecture
What you're trying to optimize: coherence, managed debt, legible dependencies.
- Layer 1 — the gateway is the only point that lasts. Centralizing model access behind a single interface decouples everything else from the vendor. Three days of work that save you thirty.
- Layer 3 — this is where the product is decided. What differentiates you isn't the orchestration framework, it's state management: recovery after compaction, step idempotency. The framework, you can swap out; a poorly thought-out state, you pay for.
- Layer 4 — MCP turns "N agents × M tools" into "N + M". And sandboxing stops being negotiable the moment the agent executes code.
- Layer 7 — this is your maturity test. If the agentic logic knows which surface it's running on, every new channel becomes a rewrite. Decoupling the runtime from the surface is exactly what separates a real platform from a demo.
- Cross-cutting — reframe the question. Not "is the agent safe?" but "what's its worst possible action, and who approves it?".
If you own the buying decision
What you're trying to optimize: risk, total cost, dependency, maturity.
| Layer | Build vs Buy | Dominant cost | Lock-in | Maturity |
|---|
| 1. Compute & inference | Buy | Usage (tokens / GPU hour) | High without a gateway | High |
| 2. Models | Buy / Build (fine-tunes) | Price per token, fine-tuning | Medium (portable via gateway) | High |
| 3. Orchestration / runtime | Build or OSS framework | Internal engineering | Medium to high if proprietary | Medium |
| 4. Tools & integrations | Buy + standard (MCP) | Integration & maintenance | Low with a standard protocol | Medium |
| 5. Memory & knowledge | Buy / Build | Storage + vector queries | Medium | Medium-high |
| 6. Multi-agent coordination | Build | Engineering | Low | Low (early) |
| 7. Interface & surfaces | Buy / Build | Per-seat licenses | Medium | High |
Five questions to ask before signing anything:
- Does model access go through a standard gateway?
- Do tools go through an open protocol (MCP)?
- Is memory exportable — and in what format?
- Do traces and costs export to the observability you already have in place?
- Are security and audit native, or sold later as an add-on?
A fuzzy answer on one of the five is fine. Two is a pattern.
What this map doesn't say
What it gives you is modest, but useful: a shared vocabulary. The day the architect says "that's a layer 3 problem" and the buyer understands why it won't be fixed by switching vendors, the map has done its job. The room is never tidied once and for all — but at least everyone finally knows which shelf anything belongs on.
A word on standards, while we're at it. Beneath layer 2, an ISO standard already exists: ISO/IEC 23053 breaks down the anatomy of a machine learning system — model, training, inference, task. It's the normative dictionary for your layer 2, and the edge of layer 1. But it stops at the model: the agentic loop, MCP, multi-agent coordination, the surfaces (layers 3 through 7) have no ISO equivalent. As for the two walls, they fall under other standards in the same family (SC 42) — 42001 for governance, 23894 for risk — not 23053. In other words: the bottom half of the shelf has an official dictionary, the agentic half doesn't yet. That's not a hole in the map, it's why it's useful.
One question it doesn't settle remains: does it hold up against what the CNCF, platform engineering, and the analysts describe on their side? It converges with them on most of the boundaries, diverges on two points that I own — and stumbles on a distinction nobody really makes: the one between the agent you produce and the agent that produces. That's exactly the subject of a second post.