Skip to content

[Server][Streamable HTTP] Concurrent requests in the same session can overwrite queued responses and cause stale/unknown message IDs #275

Description

@tebaly

Describe the bug

Concurrent HTTP requests within the same MCP session can corrupt session-backed protocol state in the PHP SDK.

In Streamable HTTP mode, the server stores MCP protocol state (including outgoing messages and pending requests) inside a shared session payload. The SDK mutates that payload using a read -> modify -> write whole session pattern without any locking or atomic merge semantics.

When multiple requests are processed concurrently for the same Mcp-Session-Id (for example, several parallel tools/call requests from Cursor IDE), one request can overwrite session changes made by another request. In practice, this appears to cause lost outgoing responses / queue entries and leads the client to report stale or unknown message IDs, while the server has actually already produced the response.

This looks like an SDK-level concurrency bug rather than an application bug.

To Reproduce

Steps to reproduce the behavior:

  1. Run an MCP server using the PHP SDK over Streamable HTTP.
  2. Configure it with a shared session store (for example PSR-16 cache via Symfony MCP Bundle, but the same read/modify/write pattern appears to affect file storage too).
  3. Use a client that sends multiple requests concurrently within the same MCP session.
  4. Trigger 2-3 parallel tools/call requests (for example multiple get_record-style read calls).
  5. Observe that some calls succeed, while another response is effectively lost from the session-backed outgoing queue.
  6. On the client side, this can surface as:
    • hanging tool calls,
    • Received a response for an unknown message ID,
    • stale responses after reconnect.

Expected behavior

Concurrent requests for the same MCP session should not overwrite each other's protocol state.

At minimum, the SDK should guarantee safe mutation of session-backed MCP state (outgoing_queue, pending request/response maps, counters, etc.) when multiple HTTP requests for the same session are processed in parallel.

Possible valid fixes could include:

  • locking per MCP session during request processing,
  • atomic session mutation,
  • splitting queue state out of the monolithic session blob,
  • or another concurrency-safe design.

Logs

Client-side symptoms observed in Cursor IDE:

Ignoring stale response (unknown message ID): Received a response for an unknown message ID: {"jsonrpc":"2.0","id":10,"result":{...}}

We also observed cases where multiple parallel get_record calls were started, two completed successfully, and one response payload appeared only as a stale/unknown-message-id response on the client side.

Relevant SDK code showing the read/modify/write pattern:

  1. The PSR-16 session store performs plain get() / set() with no locking:
class Psr16SessionStore implements SessionStoreInterface
{
    public function read(Uuid $id): string|false
    {
        try {
            return $this->cache->get($this->getKey($id), false);
        } catch (\Throwable) {
            return false;
        }
    }

    public function write(Uuid $id, string $data): bool
    {
        try {
            return $this->cache->set($this->getKey($id), $data, $this->ttl);
        } catch (\Throwable) {
            return false;
        }
    }
}
  1. The session object saves the full session payload as one JSON blob:
public function save(): bool
{
    return $this->store->write($this->id, json_encode($this->data, \JSON_THROW_ON_ERROR));
}
  1. Outgoing messages are appended by reading the queue from session state, mutating it in memory, and writing it back later:
private function queueOutgoing(Request|Notification|Response|Error $message, array $context, SessionInterface $session): void
{
    try {
        $encoded = json_encode($message, \JSON_THROW_ON_ERROR);
    } catch (\JsonException $e) {
        $this->logger->error('Failed to encode message to JSON.', [
            'exception' => $e,
        ]);

        return;
    }

    $queue = $session->get(self::SESSION_OUTGOING_QUEUE, []);
    $queue[] = [
        'message' => $encoded,
        'context' => $context,
    ];
    $session->set(self::SESSION_OUTGOING_QUEUE, $queue);
}
  1. The protocol saves the session in a finally block, meaning concurrent requests can each persist their own in-memory snapshot of the same session:
} finally {
    $session->save();
}
  1. Pending request state is mutated the same way:
$pending = $session->get(self::SESSION_PENDING_REQUESTS, []);
$pending[$requestId] = [
    'request_id' => $requestId,
    'timeout' => $timeout,
    'timestamp' => time(),
];
$session->set(self::SESSION_PENDING_REQUESTS, $pending);

This combination strongly suggests a lost-update race condition when two or more requests mutate the same session concurrently.

Additional context

In our integration, this happens systematically when a client sends parallel MCP requests over HTTP within the same session.

A project-level workaround is to serialize all MCP HTTP requests per session with a lock such as mcp-session:{id} around the entire server->run($transport) call. However, that appears to be a mitigation for an SDK concurrency issue, not the ideal long-term fix.

Activity

  1. Chi-teck commented on Apr 14, 2026

    @Chi-teck

    Same issue exists in FileSessionStore.

  2. Chi-teck commented on Apr 14, 2026

    @Chi-teck

    Here is explanation from AI agent:


    The bug is in how Mcp\Server\Protocol uses any session store, not in FileSessionStore itself. Here is the exact sequence:

    1. Session::readData() reads the entire JSON blob from the store on the first get() call and caches it in memory for the rest of that Session instance's lifetime. Each HTTP request gets its own Session instance, so two concurrent requests hold two divergent in-memory snapshots.
    2. Protocol::processInput() dispatches the request, mutates the in-memory session (including appending to _mcp.outgoing_queue via Protocol::queueOutgoing()), then calls $session->save() exactly once. save() serializes the whole blob and writes it back — whichever request writes last clobbers everything the other request wrote, including the other's queued response.
    3. Protocol::consumeOutgoingMessages() then creates yet another fresh Session instance, reads the blob again, and atomically sets the queue to [] and saves. So a request can easily read another request's queued response (and deliver it as its own HTTP response — that is the Received a response for an unknown message ID error), or read an empty queue because the other request already consumed it (and return nothing — that is the timeout).

    Any SessionStoreInterface implementation — file, KV, Redis, whatever — exhibits the same race. The store is opaque: it sees only read(Uuid) / write(Uuid, string) and cannot atomically merge two concurrent blob writes without knowing the schema. The only places to fix it are inside the SDK itself, or outside it by never letting two requests on the same session id run in parallel (a controller-level lock).

  3. tebaly commented on Apr 15, 2026

    @tebaly
    Author

    As I understand it, this isn't a session storage issue. It's a problem with the SDK that describes how MCP uses session storage. There's no atomic data storage, so data gets corrupted. There's no goal in merging data; the goal is not to store this data together, but to store them in separate files.

  4. dathanabhaya-bot commented on May 19, 2026

    @dathanabhaya-bot

    Good

  5. guillaume-sainthillier commented on Sep 25, 2026

    @guillaume-sainthillier
    Contributor

    Another data point, and a proposal for the part that #508 and #500 leave open.

    Setup: mcp/sdk v0.8.1, symfony/mcp-bundle v0.14.0, api-platform/mcp 4.4.1, handshake era, JSON responses, FileSessionStore. The client is an n8n AI Agent (LangChain), which runs tool calls in parallel on one session. A lost response surfaces as MCP error -32001: Request timed out.

    Server Parallel tools/call on one session Result
    FrankenPHP, worker mode (production) 10 1 answered 202 with an empty body
    PHP_CLI_SERVER_WORKERS=8 php -S 3 × 10 1 × 202, 3 × 500 (the 500s are #498)
    Same, with a per-session lock around the controller 50 50 × 200

    The workaround decorates mcp.server.<name>.controller and wraps handle() in LockFactory::createLock('mcp-session-'.$sessionId, ttl: 60)->acquire(true) from symfony/lock. initialize has no session id yet, so it skips the lock.

    Proposal: answer the POST directly, without the queue. #508 filters the queue by request id when it is read. That stops one request from receiving another's response, but it can't bring back a response that a concurrent save() already overwrote. That is the residual loss measured on #508, about 1 pair in 30.

    A response to a request carried by the current POST could go back through the transport the same way the null === $session path already sends it (Protocol::sendResponse()), and createJsonResponse() would return it. The session queue would remain for what can't be answered inline: server-initiated requests and notifications, SSE streams, and handlers suspended in a fiber. Responses would then no longer depend on session writes at all, and #508's id filter would still apply to whatever goes through the queue.

    That alone doesn't cover the other session keys (pending requests, client info). Those would still need an opt-in per-session lock, for example flock in FileSessionStore and a symfony/lock adapter for the PSR-16 store.

    @chr-hertel Would you take this as a PR? Should it build on #508 or replace its filtering? I'm happy to write it, with a deterministic test that interleaves two requests on one store.

  6. chr-hertel commented on Oct 5, 2026

    @chr-hertel
    Member

    Thanks for reporting and looking into this - i just set up a similar to test script like @michalcharvat in #508 (at least i assume) and i would bring in #500 and #508 now, but we really have a gap still and i'd be curious to see you approach here @guillaume-sainthillier, yes - thanks already!

    this is what i was using: https://github.com/chr-hertel/mcp-concurrency-test
    (purely generated)


    edit: changed my mind while reviewing #508 - merged #500 but would prefer your proposal over #508 @guillaume-sainthillier - if you're still up for that :)

  7. added
    ServerIssues & PRs related to the Server component
    P1Significant bug affecting many users, highly requested feature
    on Oct 5, 2026
  8. changed the title [-][Streamable HTTP][Server] Concurrent requests in the same session can overwrite queued responses and cause stale/unknown message IDs[/-] [+][Server][Streamable HTTP] Concurrent requests in the same session can overwrite queued responses and cause stale/unknown message IDs[/+] on Oct 5, 2026
  9. vbcherepanov commented on Oct 6, 2026

    @vbcherepanov

    Agreed, answering the POST directly is the better fix for the response loss, and #508 can't close the residual case on its own.

    I can take the other part guillaume mentioned and that #500 doesn't cover: an opt-in per-session lock, so pending requests and client info aren't lost when two requests save the same session. Rough plan: flock in FileSessionStore, and a small adapter for symfony/lock behind the PSR-16 store. Off by default.

    @guillaume-sainthillier does that stay clear of what you're writing?
    @chr-hertel if that works for you, I'll open it as a separate PR.

    For #508: if the id filter is still useful for what stays in the queue, I'll rebase it on top of guillaume's change and add a TTL so unclaimed responses don't pile up in the session. Otherwise I'll close it.

  10. guillaume-sainthillier commented on Oct 6, 2026

    @guillaume-sainthillier
    Contributor

    Thanks @vbcherepanov — yes, that's clear of #535. It doesn't touch the stores or locking, and lists the per-session lock as a follow-up, so it's yours.

    One thing to know for the lock: #535 makes consumeOutgoingMessages() skip the save when the queue is empty, so a JSON POST now writes the session once (in doProcessInput()) instead of twice.

    On #508: with #535, responses never enter the queue; only server-initiated requests and notifications do, and those aren't matched by id. So there's nothing left for the id filter to take, and no unclaimed responses to expire. I'd close it, but that's your and @chr-hertel's call.

  11. added a commit that references this issue on Oct 8, 2026
    f561a01
  12. added a commit that references this issue on Oct 9, 2026
    6dcd620
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    P1Significant bug affecting many users, highly requested featureServerIssues & PRs related to the Server componentbugSomething isn't working

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions