Skip to main content
Subscribe
AI & Agentic

MCP Server Architecture in 2026: Stateless MCP, A2A v1.0, OAuth 2.1

Two years ago, hooking an LLM up to a database meant writing a custom adapter, hand-rolling a JSON envelope, and praying the next model release didn’t break your tool-calling format. By December 2025 there were more than 10,000 active public Model Context Protocol servers and 97 million-plus monthly SDK downloads across Python and TypeScript (Anthropic, 2025). Then, in July 2026, the protocol threw out its own handshake and went stateless.

That’s not a slow burn. That’s a protocol still being rebuilt while everyone runs it in production.

If you’re a platform engineer being asked “should we expose an MCP server?” (or worse, “should we also speak A2A?”), this article is the system design walkthrough I wish I’d had when I started building one. We’ll cover the transport layer, how capability negotiation works now that initialize is gone, the OAuth 2.1 authorization model, the security threats nobody warned you about, and where A2A fits next to MCP now that both live in the same foundation.

I’ll be direct: the spec moves fast, the security guidance lags by months, and most of the production deployments I’ve seen contain at least one footgun the official docs don’t flag clearly. Let’s walk through them.

overview of how AI agents plan, reason, and act autonomously

Key Takeaways

  • The current MCP revision is 2026-07-28. It removes the initialize handshake and protocol-level sessions: every request now carries its protocol version and client capabilities in _meta, and servers MUST implement server/discover (MCP changelog, 2026).
  • Streamable HTTP and stdio are the two transports. HTTP+SSE has been deprecated since 2025-03-26 and is now formally on the removal track.
  • Authorization is optional in the spec, but a protected remote server acts as an OAuth 2.1 resource server and MUST publish RFC 9728 metadata; clients MUST bind tokens with RFC 8707 resource indicators (MCP Authorization, 2026).
  • A2A shipped v1.0 in March 2026 with signed Agent Cards, and joined MCP inside the Linux Foundation’s Agentic AI Foundation in August 2026, backed by 150+ organizations (AAIF, 2026).

Table of Contents


Why Did MCP and A2A Become Production Requirements?

Anthropic open-sourced the Model Context Protocol in November 2024, and Gartner projects that by 2026, 75% of API gateway vendors and 50% of iPaaS vendors will have MCP features (Truto, citing Gartner, 2026). When the gateway layer of your stack quietly absorbs a new protocol, the question stops being “should we?” and starts being “how do we do this without getting breached?”

Two forces drove the shift. First, the big model vendors adopted it fast: OpenAI brought MCP to its Agents SDK, Responses API, and ChatGPT desktop in March 2025, and Google DeepMind confirmed Gemini support a month later (Pento, 2025). By December 2025, ChatGPT, Cursor, Gemini, Microsoft Copilot, and VS Code all shipped first-class client support (Anthropic, 2025). That collapsed the “USB-C for AI” thesis from a marketing line into observable reality. Second, agents stopped being chatbots and started being workflow orchestrators that need to call dozens of tools, often across organizational boundaries. Hand-rolled tool integrations don’t scale to that.

Timeline of MCP and A2A milestones from November 2024 to August 2026. MCP: launched November 2024, Streamable HTTP and OAuth-based authorization in March 2025, elicitation and RFC 8707 token binding in June 2025, Client ID Metadata Documents in November 2025, donated to the Agentic AI Foundation in December 2025, stateless 2026-07-28 spec in July 2026. A2A: launched April 2025 with 50-plus partners, contributed to the Linux Foundation in June 2025, version 0.3 with 150-plus supporting organizations in August 2025, version 1.0 in March 2026, joined the Agentic AI Foundation in August 2026.

Citation Capsule: Anthropic donated MCP to the Agentic AI Foundation, a directed fund under the Linux Foundation co-founded with Block and OpenAI, on December 9, 2025, when the protocol had more than 10,000 active public servers and 97 million-plus monthly SDK downloads (Anthropic, 2025). A2A joined the same foundation in August 2026 (AAIF, 2026), so both agent protocols now share one governance home.

The flip side of 10,000 public servers: quality is all over the place. If you’re consuming a public MCP server, assume the median one is half-finished and read its tool descriptions before you connect it to anything with write access. If you’re building one, the rest of this article is the checklist.

the full guide to finding, installing, and evaluating MCP servers

What Is the Core MCP Architecture?

MCP is a JSON-RPC 2.0 protocol with three roles: a host (the user-facing app like Claude Desktop or Cursor), a client (a connector inside the host that maintains one connection per server), and a server (the process that exposes tools, resources, and prompts) (Model Context Protocol Specification, 2026). One host typically runs many clients in parallel. One per attached server.

The server exposes three primitive types. Tools are functions the model can call (think read_file, query_db, send_email). Resources are read-only chunks of context the model can pull (a file, a row, a doc). Prompts are reusable templated message structures the host can present as user-triggered actions.

The complexity isn’t in that surface area; it’s in the operational layer below it.

A green and yellow circuit board photographed up close, with copper traces and integrated chips representing the layered architecture of an MCP server: transport, JSON-RPC framing, capability negotiation, and primitive handlers

Here’s the part that changed. Through the 2025-11-25 revision, a client opened a transport, ran an initialize exchange, locked in a protocol version and capability set, and then held a stateful session (tracked on HTTP by an Mcp-Session-Id header). The 2026-07-28 revision removed all of that. There is no handshake and no protocol-level session. Every request carries its own protocol version and client capabilities in _meta, list endpoints no longer vary per connection, and servers that need cross-call state mint explicit handles and pass them as ordinary tool arguments (MCP changelog, 2026). Server-to-client change notifications now flow over a single opt-in subscriptions/listen stream instead of a session-wide channel.

Where most teams trip up now: they still build for the old model. If your server keeps per-connection state in memory and relies on a session ID to find it, it was fragile behind a load balancer before, and on 2026-07-28 clients it simply doesn’t match the protocol anymore. Put cross-call state in a store keyed by a handle your server mints, and let any instance answer any request.

How Do MCP Transports Compare: stdio, SSE, and Streamable HTTP?

The spec defines two official transports: stdio (the server runs as a subprocess of the host, JSON-RPC messages go over stdin/stdout newline-delimited) and Streamable HTTP (the server runs independently, clients POST JSON-RPC requests to a single endpoint and optionally receive Server-Sent Events streams in response). The HTTP+SSE transport from the original 2024-11-05 spec has been deprecated since 2025-03-26, and the 2026-07-28 revision formally put it on the removal track under the new deprecation policy (MCP changelog, 2026).

Why was SSE-only thrown out? Because it required the server to maintain a long-lived, highly available connection per client, which made horizontal scaling painful. Streamable HTTP fixes this by making POSTs the primary channel and using SSE only as an optional response upgrade when the server has streaming output to deliver (Bright Data, 2025). The 2026-07-28 revision finished the job: it removed sessions and dropped SSE stream resumability, so a broken response stream now just means the client re-issues the request with a new ID. The architecture is stateless by design, which is the property you actually need for multi-tenant infra.

Comparison table showing the three MCP transport mechanisms: stdio (local, zero latency, single client), HTTP+SSE deprecated (remote, persistent connection, single tenant), and Streamable HTTP (remote, stateless POST plus optional SSE upgrade, multi-tenant scalable)

The practical decision tree: if your server only needs to run alongside a desktop host, ship stdio. If your server is a multi-tenant SaaS that any agent should be able to call, ship Streamable HTTP and put it behind your existing API gateway. If you still see HTTP+SSE-only servers in the wild, treat them as legacy. They’re on a clock now.

how to layer auth, rate limiting, and CORS for an HTTP-exposed API surface

How Does Capability Negotiation Work Without initialize?

It moved from the session to the request. Under 2025-11-25 and earlier, the first message was initialize: the client sent its protocol version and capabilities, the server answered with its own, and that handshake locked the contract for the session. Under 2026-07-28, every request declares its protocol version and the client’s capabilities in _meta (io.modelcontextprotocol/protocolVersion, io.modelcontextprotocol/clientCapabilities), and the server accepts or rejects each request independently (MCP Versioning, 2026).

Two pieces replace the handshake. Servers MUST implement server/discover, which returns supported protocol versions, capabilities, and identity in one call; clients can call it up front or skip it. And when versions don’t match, the server returns an UnsupportedProtocolVersionError listing what it does support, so the client can retry with a mutually supported version. Clients and servers MAY support several versions at once.

The old bug pattern still applies in a new shape. In early MCP implementations, servers would emit notifications for a capability the client never enabled, the client ignored them, and nobody could figure out why no events fired. Now the rule is per request: for example, servers MUST NOT emit log notifications for a request that didn’t set io.modelcontextprotocol/logLevel in its _meta. Only send what that request asked for.

The interactive features changed too. The 2025-06-18 spec added elicitation, letting a server ask the user for input mid-task via elicitation/create with a JSON Schema for the expected data (ForgeCode, 2025). In 2026-07-28, server-initiated requests like that are replaced by the Multi Round-Trip Requests pattern: the server returns an InputRequiredResult describing what it needs, and the client retries the original request with the answers attached.

Capability areas worth knowing in the current spec:

Citation Capsule: The 2026-07-28 MCP revision removed the initialize/notifications/initialized handshake and made the protocol stateless: every request carries its protocol version and client capabilities in _meta, and servers must implement server/discover to advertise versions and capabilities (MCP changelog, 2026). The same revision deprecated Roots, Sampling, and Logging.

A holographic interface displaying interconnected data nodes and glowing connection lines, representing capability negotiation between MCP clients and servers as they exchange supported features during the initialize handshake

How Does the OAuth 2.1 Authorization Layer Work?

Here’s what surprises people: authorization is optional in the MCP spec. But when a remote server is protected, the HTTP flow SHOULD conform to the spec’s authorization model, and the server acts as an OAuth 2.1 resource server; stdio servers should pull credentials from the environment instead (MCP Authorization, 2026). In practice, if you’re shipping anything with a URL that touches real data, treat it as mandatory.

The model has evolved revision by revision. 2025-03-26 introduced the OAuth 2.1-based framework. 2025-06-18 formally classified MCP servers as OAuth resource servers and required clients to use RFC 8707 resource indicators (Auth0, 2025). 2025-11-25 made Client ID Metadata Documents the default registration method, demoted Dynamic Client Registration to optional, and made PKCE mandatory (Den Delimarsky, 2025). 2026-07-28 deprecated Dynamic Client Registration outright and added RFC 9207 issuer validation.

The flow has three actors. The MCP server is the resource server. Some external authorization server issues tokens. That can be your own AS, or a federated one like an enterprise IdP. The MCP client acts as an OAuth 2.1 client and presents a bearer token on every HTTP request.

What’s specific to MCP is the discovery and binding:

  1. OAuth 2.0 Protected Resource Metadata (RFC 9728): MCP servers MUST publish a metadata document that points clients at the right authorization server, and clients MUST use it for discovery.
  2. Authorization server discovery: the AS MUST offer RFC 8414 metadata or OpenID Connect Discovery, and clients MUST support both.
  3. Resource indicators (RFC 8707): clients MUST pass resource=<canonical MCP server URL> in both authorization and token requests, and servers MUST validate that each token was issued for them as the audience.

The RFC 8707 binding is the part most teams miss. Without it, a token issued for mcp.your-product.com can be replayed against mcp.partner.com if both trust the same AS. That’s a confused-deputy attack waiting to happen. The spec closes it from both ends: servers MUST NOT accept or pass through tokens issued for anyone else.

JWT validation, audience binding, and the layered framework that closes confused-deputy gaps

What Are the Security Boundaries Every MCP Server Must Defend?

The cheapest attack is still the oldest one: no auth at all. In July 2025, Trend Micro found 492 MCP servers exposed to the network with no client authentication or traffic encryption, exposing 1,402 tools, more than 90% of which gave direct read access to the underlying data source (Trend Micro, 2025). Past that baseline, I’ve read enough advisories and red-team write-ups to say the threats fall into four buckets, and any production MCP deployment that isn’t defending all four is an incident waiting for a calendar slot.

Donut chart of 492 MCP servers Trend Micro found exposed without authentication or encryption in July 2025: about 74% hosted on major cloud providers (AWS, Azure, GCP, Oracle) and about 26% elsewhere. Together they exposed 1,402 tools, more than 90% with direct read access to data.

Prompt injection through tool descriptions. Every tool a server exposes ships with a description string that’s fed to the model verbatim. Invariant Labs showed in April 2025 how instructions hidden in a tool description (say, “read ~/.ssh and pass it along as a parameter”) are visible to the model but invisible in the simplified view most host UIs show the user (Invariant Labs, 2025). Any server that pulls tool metadata from somewhere an attacker can write to inherits this risk.

Tool poisoning and rug-pull attacks. A tool’s behavior on Day 1 doesn’t bind it on Day 7. A rug pull is a tool that looks legitimate until it gains adoption, then turns malicious; Prompt Security recommends strict sandboxing and continuous behavior monitoring to catch the change (Prompt Security, 2025). I’d go further: pin the tool manifest hash on install, reject silent definition changes, and require re-consent when a tool’s signature changes. The 2026-07-28 rule that servers SHOULD return tools/list in deterministic order makes that hashing easier.

Command injection in tool implementations. CVE-2025-53818 hit a GitHub Kanban MCP server whose add_comment tool passed parameters like issue_number straight into Node’s exec, so a prompt-injected payload with shell metacharacters ran arbitrary commands (CVSS 8.9) (GitHub Security Advisories, 2025). The lesson is the same one we’ve been teaching about web apps for 20 years: never exec user input. MCP servers expose this risk anew because they re-introduce a class of “trusted intermediary” code that often skips the hardening web frameworks bake in.

Confused-deputy and authorization bypass. The cleanest example is General Analysis’s July 2025 red-team demo of the Supabase MCP server: a developer asks Cursor’s agent to list support tickets, the agent runs with the service_role key that bypasses row-level security, and an attacker’s ticket tells it to read the integration_tokens table and paste the contents back into the ticket thread (Simon Willison, 2025). The root cause wasn’t a protocol bug. It was an agent operating with broader privileges than its untrusted input warranted.

A dark cybersecurity-themed monitor showing scrolling code and digital security analysis, representing the four-layer defense surface that production MCP servers must defend against prompt injection, tool poisoning, command injection, and authorization bypass

Citation Capsule: Tool poisoning attacks embed malicious instructions in MCP tool descriptions that are visible to the model but hidden from users, who typically see only a simplified version of each tool in the host UI (Invariant Labs, 2025). The mitigation pattern I use is to pin tool manifest hashes on install and force re-consent on any signature change.

Watch David Soria Parra (the Anthropic engineer behind much of MCP’s architecture) walk through where the protocol is heading next:

https://www.youtube.com/watch?v=v3Fr2JR47KA

How Does A2A Differ From MCP?

Google announced the Agent2Agent (A2A) protocol on April 9, 2025, with more than 50 launch partners (Salesforce, SAP, MongoDB, ServiceNow, LangChain, Atlassian, Cohere among them) and a stated goal of letting agents from different vendors talk to each other directly (Google Developers Blog, 2025). Google contributed it to the Linux Foundation that June, and version 0.3 in August 2025 added gRPC support and signed security cards, with more than 150 supporting organizations (Google Cloud, 2025). A2A v1.0, the first stable spec, shipped in March 2026, and in August 2026 the project joined the Agentic AI Foundation next to MCP, goose, AGENTS.md, and agentgateway (AAIF, 2026).

The cleanest mental model: MCP gives an agent access to tools and data; A2A lets an agent delegate work to another agent. They’re complementary, not competing. The “USB-C of agents” framing applies to both, but at different layers. MCP standardizes vertical integration (agent → its capabilities), A2A standardizes horizontal coordination (agent ↔ peer agent).

A2A’s design choices reflect that. Like MCP, it speaks JSON-RPC over HTTP, and v1.0 adds HTTP+JSON and gRPC bindings. Unlike MCP, the discovery primitive isn’t a tool list. It’s an Agent Card, a public JSON document that advertises an agent’s identity, supported skills, modalities (text, audio, video), and authentication requirements. Two agents that have never met can read each other’s cards, decide compatibility, and start a task without a shared host orchestrating them.

multi-agent code review system showing how specialized agents coordinate on a single task

Grouped bar chart comparing MCP and A2A across five dimensions: primary purpose, transport, discovery primitive, identity model, and typical deployment topology

A2A v1.0 has something MCP doesn’t: signed Agent Cards for cryptographic identity verification before work is delegated. When agent A receives a task from agent B claiming to come from bigcorp.com, A can verify the card’s signature instead of trusting the claim. On MCP, you’d solve that impersonation class at the auth layer with bound tokens. It also explains why A2A shows up in enterprise and regulated scenarios: auditing “who told whom to do what” is materially easier when identities are signed.

Citation Capsule: A2A v1.0, the first stable version of the Agent2Agent protocol, shipped in March 2026 with multi-protocol bindings, version negotiation, multi-tenancy, and signed Agent Cards for cryptographic identity verification (AAIF, 2026). AAIF reports A2A running in production across cloud AI infrastructure, financial services, supply chain, and enterprise IT, with 150+ organizations backing it.

The right model: ship MCP as the protocol your agents use to reach tools, and adopt A2A only when you have a use case that genuinely requires agents from independent organizations to coordinate. Most production workloads I’ve seen don’t yet. The ones that do tend to be cross-company workflows where an audit trail of delegated work is a hard requirement.

why agent-to-agent delegation matters more for small teams than people realize

How Do You Handle Schema Versioning in Production?

The protocol has shipped five dated revisions in under two years: 2024-11-05, 2025-03-26, 2025-06-18, 2025-11-25, and 2026-07-28. The version string marks the last date backwards-incompatible changes were made (MCP Versioning, 2026). Streamable HTTP replaced HTTP+SSE in 2025-03-26, which also introduced OAuth-based authorization; elicitation and RFC 8707 binding arrived in 2025-06-18; Client ID Metadata Documents became the default in 2025-11-25; and 2026-07-28 removed the handshake and sessions entirely. If your server only speaks one version, part of your client base will eventually drop you.

The schema versioning pattern that works under the current spec: implement server/discover honestly, accept every version you can actually serve, and answer anything else with UnsupportedProtocolVersionError listing your supported versions. Keep adapter shims for older, handshake-based clients; the spec documents backward compatibility with 2025-11-25 and earlier. Your clients don’t all upgrade at once, and you can’t force them to.

One genuinely good change in 2026-07-28 is the feature lifecycle policy. Deprecated features now stay in the spec for at least twelve months (ninety days under an expedited-removal exception) before they can be removed, and a public registry tracks them (MCP Versioning, 2026). That gives you a real planning horizon for Sampling, Roots, Logging, and HTTP+SSE.

A laptop screen displaying terminal output and code in a dark workspace, representing the iterative process of building and shipping production-grade MCP servers with proper version negotiation, transport selection, and security defenses

A practical checklist for shipping a server you intend to run for more than six months:

  1. Version your tool schemas separately from the protocol version. Use semantic versioning on each tool. A breaking change to query_db.input.schema shouldn’t require a protocol bump, it should be a tool-level major version that triggers re-consent.
  2. Treat tool descriptions as user-facing copy. Every word in a tool description is fed to a model, and changing it can change behavior in subtle ways. Lock descriptions to a review process; don’t let them mutate on autopilot.
  3. Pin the tool manifest hash on install. This is the single most important defense against rug-pull attacks. Return tools in a deterministic order so the hash is stable.
  4. Log every JSON-RPC request and response with a trace ID. When something breaks during a multi-tool agent run, you need to walk the timeline. With no session ID to lean on anymore, the trace ID is your timeline.
  5. Adopt OpenTelemetry now. The 2026-07-28 revision documents trace context propagation (traceparent, tracestate, baggage) in _meta and points deprecated Logging users at OpenTelemetry. If you ship without it, you’ll retrofit later.

how to trace agent and tool calls with OpenTelemetry without leaking PII

why idempotency keys matter when agents retry tool calls

FAQ

Is MCP just JSON-RPC over HTTP?

Closer than it used to be, but no. MCP uses JSON-RPC 2.0 as its message format over stdio or Streamable HTTP, and since the 2026-07-28 revision it is stateless: no initialize handshake, no session ID, with protocol version and client capabilities carried on every request (MCP changelog, 2026). What makes it more than plain RPC is the shared contract: tools, resources, prompts, server/discover, subscriptions, multi round-trip input requests, and the authorization model.

Do I need OAuth 2.1 for an internal MCP server?

The spec makes authorization optional, but if you protect an HTTP server it SHOULD follow the spec’s OAuth 2.1 model, including RFC 9728 metadata and RFC 8707 audience binding (MCP Authorization, 2026). For purely local stdio servers running as a subprocess of a desktop host, no: pull credentials from the environment. For anything reachable over a network, remember the 492 servers Trend Micro found sitting open with no auth at all.

Should I implement both MCP and A2A?

Probably not yet. Ship MCP first: it covers the “agent needs access to a tool” case, and client support is everywhere, from ChatGPT and Gemini to Cursor and VS Code (Anthropic, 2025). Add A2A when you have a concrete agent-to-agent coordination need: cross-organization workflows, multi-vendor agent meshes, or regulated environments where signed Agent Cards close an audit gap.

How fast is the protocol moving?

Five dated revisions between November 2024 and July 2026: 2024-11-05, 2025-03-26, 2025-06-18, 2025-11-25, 2026-07-28 (MCP Versioning, 2026). That’s a breaking revision roughly every four to eight months. The new lifecycle policy guarantees deprecated features at least twelve months before removal, so plan migrations on that clock.

What’s the biggest production gotcha?

Building for the old session model. Servers that still hide state behind a connection or session ID break against 2026-07-28 clients. Make every request self-contained, mint explicit handles for anything that spans calls, and keep shims for older handshake-based clients until your traffic says you can drop them.

Conclusion

If your platform is going to be agent-callable, you’ll be speaking these protocols. MCP passed 10,000 active public servers and 97 million monthly SDK downloads by late 2025, every major AI client ships support, and both protocols now sit under the Agentic AI Foundation (Anthropic, 2025).

The architectural decisions that matter:

The protocol layer just went through its biggest rewrite yet. The hard part, running it safely, is still just starting.

start here for a full picture of how agents reason, plan, and use protocols like MCP and A2A

Written by Nishil Bhave

Builder, maker, and tech writer at MakeToCreate.

Never miss a post

Get the latest tech insights delivered to your inbox. No spam, unsubscribe anytime.

Related Posts