AI & Machine Learning

MCP vs API: What's Actually Different

The usual answer - APIs are stateless, MCP holds a stateful session - stopped being true on 28 July 2026, when the specification removed protocol-level sessions outright. This guide gives the difference that actually holds: who reads the contract, and when. It covers what the 2026-07-28 revision removed, answers the live "is stateless MCP just an API again" argument honestly, and includes a decider for the fork you really face - build an MCP server, or just write tool definitions.

Mohammed Yaseen
Mohammed Yaseen
Last Updated: · 12 min read
ShareXLinkedIn
MCP vs API: What's Actually Different

Quick Answer: An API is a contract between two pieces of software, read by a developer at build time. MCP is a contract between software and a model, read by the model at runtime. That difference — who reads the contract, and when — is the whole thing. MCP does not replace your API; an MCP server is nearly always a wrapper that calls one. The popular answer, that APIs are stateless while MCP holds a stateful session, stopped being true on 28 July 2026, when the specification removed protocol-level sessions entirely.

If you search this question today, almost every result tells you the same thing: APIs are stateless and MCP maintains stateful sessions that preserve context across tool calls.

That was a reasonable summary a year ago. It is now wrong.

The 2026-07-28 revision of the Model Context Protocol removed protocol-level sessions, removed the Mcp-Session-Id header, removed the initialize handshake, and removed SSE stream resumability. MCP is a stateless request/response protocol now. Every article that leads with the statefulness distinction is describing a version of MCP that no longer exists.

So this piece does two things. It gives you the difference that is actually load-bearing — and it walks through what the July revision changed, because if you are choosing between MCP and a plain API integration this month, you are choosing between things whose shapes moved very recently.

The answer that just expired

Here is what the specification changed on 28 July 2026, in the words of its own changelog: "Remove protocol-level sessions and the Mcp-Session-Id header from the Streamable HTTP transport", and "Make MCP stateless: remove the initialize/notifications/initialized handshake."

The motivation was operational rather than philosophical. Stateful sessions meant session affinity: every request in a conversation had to reach the same server process, or reach shared storage that knew about it. That is a tax on anyone trying to run MCP servers behind an ordinary load balancer. The maintainers describe statelessness as "one of the most highly-requested features from developers who were eager to get better reliability and scalability."

What replaced it: each request now carries its own protocol version and client capabilities in a _meta field (io.modelcontextprotocol/protocolVersion, io.modelcontextprotocol/clientCapabilities). Requests are self-describing, so any server instance can answer any request. Servers that genuinely need state across calls now mint an explicit handle and pass it back as an ordinary tool argument — the state became visible application data instead of invisible protocol machinery.

This matters for your decision because the most commonly cited reason to prefer MCP over a REST integration — "it manages context for you" — was never quite what people thought it meant, and now it is not there at all.

What MCP actually is

MCP is a JSON-RPC 2.0 protocol that lets a client ask a server two questions and then act on the answers: what can you do? and do this specific thing.

A server exposes three kinds of capability. Tools are functions the model can call. Resources are data the model can read. Prompts are reusable templates a user can invoke. Of these, tools are what nearly everyone builds and what nearly every "MCP vs API" question is really about.

The important structural point is what sits underneath. When a model calls a create_issue tool on a Linear MCP server, that server turns around and calls Linear's REST API with a credential the model never sees. The API did not go anywhere. MCP is a uniform, model-readable façade in front of it.

MCP vs API architecture diagram — an API is read by a developer at build time and wired into code, while an MCP server is queried by a client at runtime and calls the same underlying API

That is why "MCP vs API" is a slightly misleading framing, and why the honest version of the question is the one further down this article: MCP server, or just write tool definitions?

MCP vs API: the comparison

REST / GraphQL API MCP (2026-07-28)
Primary consumer A developer writing an integration A language model choosing an action
When the contract is read Build time, by a human Runtime, by a client and model
Discovery Read the docs; redeploy when they change tools/list — the client asks, live
Transport HTTP stdio (local) or Streamable HTTP (remote)
Message format Whatever you designed JSON-RPC 2.0, fixed shape
State Stateless by convention Stateless by specification (since 2026-07-28)
Call shape POST /issues with your schema tools/call with a JSON Schema input
Auth Your scheme; the caller holds the key OAuth 2.x; the model never sees the credential
Error semantics HTTP status codes, your conventions JSON-RPC error codes, spec-allocated ranges
Versioning /v2/, headers, whatever you chose Per-request protocol version in _meta
Who it scales across Consumers you onboard deliberately Any MCP-capable client, no onboarding

The row that does the most work is the last one. An API is something you deliberately expose to consumers you know about. An MCP server is something any MCP-capable client can drive without you having written a line of code for that client. That is the actual trade — reach in exchange for giving up control of the calling pattern.

"If MCP is stateless now, isn't it just an API?"

This is the live argument, not a rhetorical one. InfoQ ran a piece on 12 August 2026 titled almost exactly that, and it is a fair question: strip out the sessions, the handshake and the persistent stream, and what is left is a JSON-RPC endpoint over HTTP. Which is, structurally, an API.

Three things still separate them, and only one of them is technical.

Runtime discovery. With an API, a developer reads documentation and writes an integration; when the API changes, someone edits code. With MCP, the client calls tools/list and gets the current tool set with JSON Schema for each. Add a tool to your server and every connected client can use it without redeploying. That is a genuinely different lifecycle, and it survives statelessness untouched.

A uniform call shape. Every MCP server speaks the same protocol, so one client implementation drives all of them. This is the M×N problem the protocol was built for: without it, every client needs bespoke code for every service. It is standardisation, not capability — but standardisation is why USB-C matters, and it is the honest reason MCP spread.

Standard routing headers. The 2026-07-28 revision requires Mcp-Method and Mcp-Name headers on Streamable HTTP POSTs. This sounds like plumbing, and it is the part infrastructure people care most about — a gateway can now route, rate-limit and audit per individual tool without opening the JSON body. As Cloudflare's Matt Carey put it, quoted in that InfoQ piece: "A gateway, rate limiter or WAF can read those headers and act on them, per method or per tool, using the same primitives it already applies to every other API."

So the fair verdict: stateless MCP is much closer to an API than 2025-era MCP was, and anyone who says otherwise is selling something. What is left is a convention — a strong, widely-implemented, tooling-supported convention for describing capabilities to a model. Conventions are not nothing. HTTP is a convention.

What the July 2026 revision changed

If you are working from a tutorial written before August 2026, this table is the diff. Everything here is from the official changelog.

Change What it means for you
Protocol sessions + Mcp-Session-Id removed No session affinity. Deploy behind any load balancer. Cross-call state becomes an explicit handle you pass as a tool argument
initialize / initialized handshake removed No connection setup step. Version and capabilities ride in _meta on every request
server/discover added Servers MUST implement it. Clients may call it first for version selection
SSE resumability + Last-Event-ID removed A broken stream loses the in-flight request; the client must re-issue it with a new request ID
MRTR replaces server-initiated requests Instead of the server calling back, it returns resultType: "input_required" with inputRequests; the client retries the original call with inputResponses
resultType required on all results "complete" or "input_required". Results from older servers that omit it must be treated as "complete"
Mcp-Method / Mcp-Name headers required Gateways can route and throttle per tool without parsing the body
ttlMs + cacheScope required on list results List responses are now cacheable with an explicit freshness hint, reducing polling
ping, logging/setLevel removed Log level is per-request via _meta
Roots, Sampling, Logging deprecated Still functional, minimum twelve-month window. Don't build new things on them
HTTP+SSE transport deprecated Formally, with a year-long offramp. Use Streamable HTTP
Dynamic Client Registration deprecated Prefer Client ID Metadata Documents

Two of these deserve a flag if you are migrating.

MRTR is the subtle one. Previously, a server that needed user input mid-call sent a request back to the client — which required an open bidirectional channel, which is precisely what statelessness removes. Now the server returns an interim result saying "I need these things", and the client re-issues the original request with the answers attached. Your tool handler has to be written to be re-entered rather than to block. That is a real code change, not a config change.

The Sampling deprecation catches people. Sampling let a server ask the client's model to generate something. The suggested migration is to integrate directly with an LLM provider API instead — which means a server that relied on it now needs its own model credentials and its own budget. If your design assumed you could borrow the client's model, revisit it.

The decision you actually face

Nobody chooses "MCP or REST" for a given service — you almost always end up with both. (If the fork you are actually staring at is server or agent rather than MCP or API, that is a different boundary — see MCP servers vs AI agents.) The real fork is narrower, and it is this: should I build an MCP server, or just define tools directly in my model API calls?

Plain tool definitions mean you describe your functions in the request you send to the model, and your own code executes them when the model asks. No server, no protocol, no transport. For a single application this is dramatically less machinery, and it is the right answer more often than the current enthusiasm suggests.

Decider · nothing is stored

MCP server, or just tool definitions?

"MCP vs API" is not the fork you actually face — you end up with both. The decision that costs something is whether to stand up a server at all, or describe your functions inline in the model request and run them yourself.

Who calls these tools?

When does the tool set change?

How many tools?

Must it work in Claude Desktop, Cursor or similar?

Verdict

Skip it

Discovery is not firing

Discovery value

0 / 13

How much runtime discovery buys you here

Just write tool definitions

Nothing here is discovering anything: you own every caller and the tool list changes when you deploy. Describe the functions in your model request and execute them yourself — less code, one fewer service, and a stack trace that points at your own process. Promote to MCP the day that stops being true.

  • Only your app calls these tools, so nothing needs to discover them — you already know what exists at compile time.
  • A tool set that changes on deploy can be described in the deploy. Discovery adds nothing you are not already doing.
  • A small tool set is easy to hold in one request and easy for a model to choose from.
  • No third-party MCP client in the picture, so protocol compatibility is not buying you anything.
Runtime discovery is the only thing MCP does that inline definitions cannot. Every question above is a proxy for whether anything in your system is genuinely discovering anything — if nothing is, the protocol is overhead for a capability that never fires.

The pattern underneath the widget is straightforward: runtime discovery is the feature you are paying for. If nothing in your system is discovering anything — one app, one model, a tool list that changes when you deploy — you are paying protocol overhead for a capability you never exercise.

When you should not build an MCP server

Four cases where the answer is "don't", drawn from the shape of the problem rather than from taste:

You control both ends and the tool set is small. One app, one model, five tools. Tool definitions in your request are less code and easier to debug. You can always promote to MCP later; the tool schemas transfer.

Your tools are latency-critical and chatty. MCP adds a hop. If a tool is called in a tight loop and the budget is single-digit milliseconds, that hop is a real cost for a benefit — discovery — you are not using.

You need transactional semantics across calls. With sessions gone, there is no protocol-level notion of "these calls belong together". You can pass an explicit handle, but you are building that yourself, and a direct API with a real transaction is a better fit.

You have not worked out the authorisation story. This is the one that bites hardest. An MCP server exposing tools to a model is exposing them to whatever text reaches that model, which includes text an attacker controls. Take identity from the authenticated session, never from a model-supplied argument — the tool signature is your security boundary. We have written about prompt injection and agent security in detail, and it applies with full force the moment you put a tool behind a protocol designed for reach.

Common mistakes

Treating MCP as an API replacement. It is a façade. Your API still needs to exist, still needs versioning, still needs to serve every non-model consumer you have.

Exposing your REST surface one-to-one. Thirty endpoints do not make thirty good tools. Models choose badly from long, undifferentiated lists. Expose the handful of operations that map to things a user would ask for, and write the descriptions for the model rather than for a developer — say when to call a tool, not just what it does.

Building on deprecated features. Roots, Sampling and Logging all have twelve-month clocks running. HTTP+SSE has a year-long offramp. Anything you start today should be Streamable HTTP.

Assuming pre-August-2026 tutorials still apply. Session handling, the initialize handshake and server-initiated requests are all gone or changed. Code written against 2025-11-25 needs work, not just a version bump.

Returning API responses verbatim. A 40-field JSON object burns context and buries the answer. Return what the model needs to decide the next step.

Frequently asked questions

Is MCP replacing REST APIs?

No, and the framing misleads. MCP is a description and invocation layer that typically sits in front of an existing API. A tools/call request arrives at your MCP server, which authenticates with a credential the model never sees, calls your REST or GraphQL endpoint, and reshapes the result for a model to read. Your API continues serving your web app, mobile clients and partners unchanged.

Is MCP stateful or stateless?

Stateless, since the 2026-07-28 specification revision. That release removed protocol-level sessions, the Mcp-Session-Id header, the initialize handshake and SSE stream resumability. Requests are self-describing, carrying protocol version and client capabilities in _meta. Any article describing MCP as session-based is describing a revision earlier than 2026-07-28.

What transports does MCP support?

Two: stdio, for servers running as a local subprocess, and Streamable HTTP, for remote servers. The older HTTP+SSE transport has been deprecated since protocol version 2025-03-26 and was formally reclassified as Deprecated in the 2026-07-28 revision with a year-long offramp. Build on Streamable HTTP.

Do I need an MCP server to use tools with an LLM?

No. Every major model API supports tool definitions passed directly in the request — you describe the function, the model asks you to call it, your code executes it. That is the simpler path and the right one for a single application. MCP is worth its overhead when the same tools must serve clients you do not control, or when the tool set changes independently of your deploys.

Is MCP secure?

The protocol specifies OAuth 2.x authorisation and keeps credentials away from the model, and the 2026-07-28 revision hardened several details — issuer validation per RFC 9207, credentials bound to the issuing authorisation server, application_type required during registration. But the dominant risk is not the transport. It is that a tool exposed to a model is exposed to any text that reaches that model. Derive identity from the session, validate every argument, and treat the tool signature as your security boundary.

Does MCP work with any model?

MCP is model-agnostic — it is a protocol between a client and a server, not between a client and a model. Anthropic introduced it and open-sourced it, and it has since been adopted across the industry. Tier 1 SDKs (TypeScript, Python, Go and C#) shipped 2026-07-28 support on release day, with the Rust SDK in beta.

Conclusion

The difference between MCP and an API is not statefulness, whatever the current top ten results say — that distinction was removed from the specification on 28 July 2026. It is who reads the contract. An API is documentation a developer turns into code. MCP is a description a model reads at runtime and acts on immediately.

That leaves a narrower, more useful question than the one people usually ask. Not "MCP or API" — you will have both — but whether the reach that runtime discovery buys you is worth its overhead for the system you are actually building. If you control both ends and your tool list changes when you deploy, plain tool definitions will serve you better and you can promote later. If your tools need to work in clients you have never heard of, that is exactly what the protocol exists for.

And whichever way you go: check the spec revision date on any MCP tutorial before you follow it. This one moved in July, and most of the internet has not caught up.

If you are building in this space, our free MCP server generator scaffolds a working server from a tool description, and the walkthrough on building an MCP server in Python covers the implementation end to end.

Mohammed Yaseen

Mohammed Yaseen

Founder, SolutionGigs

Mohammed builds AI and data tooling at SolutionGigs, including an MCP server generator and a set of agent-backed products where the tool boundary is the security boundary. LinkedIn →

Try Free MCP Server Generator

Free, no signup — right in your browser.

Try Free MCP Server Generator →
Found this useful? Share it.
ShareXLinkedIn

Comments

0

Join the conversation. Sign in to leave a comment — we'd love to hear your thoughts.