Skip to content

v4: the 4.0 core, surface and alpha branch (tracking) - #3262

Draft
mgravell wants to merge 815 commits into
mainfrom
v4
Draft

mgravell wants to merge 815 commits into
mainfrom
v4

Conversation

@mgravell

@mgravell mgravell commented Oct 6, 2026 •

Copy link
Copy Markdown
Collaborator

Tracking PR, not for immediate merge. v4 is the long-lived 4.0 branch; this PR exists to show its
diff against main and to collect review. It supersedes #3219, which was the spike this grew from.
When it lands it must be a real merge, not a squash: main is merged into v4 as it moves, so the
final merge stays small. Pending work and open decisions are in design/v4-ledger.md.

What 4.0 is

A new core, with the old one deleted. Message, ResultProcessor, PhysicalBridge,
PhysicalConnection and RedisDatabase are gone. Commands are rendered by an interpolated-string RESP
writer and sent through executors over the RESPite transport. Reconnect, backlog, resubscribe, redirects,
maintenance events, sentinel and logging (shipped event ids preserved) were ported and checked against
the old core before it was deleted.

IDatabase stays, served by the new core. Existing code keeps compiling and working; see
docs/LegacyApi.md. The old API is not deprecated.

A grouped, async-first command surface alongside it:

await db.Strings.SetAsync("k", "v");
using var members = await db.SortedSets.RangeByScoreAsync(key, 0, 10);   // a lease over the reply buffer

Commands are extension members on group types (Strings, Hashes, Keys, ...), so adding one - ours or a
third party's - is never a binary break to an interface. Replies can be leased rather than allocated, and
every call takes a CancellationToken.

Highlights

  • Less work on the IO loop. Commands are formatted on the caller's thread, straight into pooled
    buffers, and replies are parsed on the caller's continuation rather than by the reader. The core's IO
    loop now only frames bytes in and out, so one slow or large reply no longer holds up everyone else's.
  • A much more efficient API for ad-hoc commands and extension libraries (NRedisStack and similar).
    Instead of building object[] argument arrays for Execute/ExecuteAsync and picking through a
    RedisResult, a library writes ctx.SendAsync<T>($"FT.SEARCH {index} {query}") with a typed reply
    handler, getting the same zero-copy formatting, cluster slot routing, key prefixing and caching as the
    built-in commands, and can hang its own command groups off the same targets as extension members.
  • Client-side caching driven by the library itself (CLIENT TRACKING), keyed by the rendered frame,
    with invalidation that cannot race the send.
  • Batches and transactions on the new surface (SER014): db.BeginBatch() / db.BeginTransaction()
    return structs that are ordinary keyspace targets, with explicit ExecuteAsync, discard-on-dispose, and
    conditions via the existing Condition type.
  • Sync without the thread pool: Blocking() contexts send and wait on the calling thread, so a group
    method returns an already-completed ValueTask; the synchronous IDatabase methods are built on it, and
    so can a library's. DedicatedThreads connections connect and handshake synchronously, so on Linux their
    sockets never need a pool thread either.
  • For library authors: one generic accessor per group, Blocking() for sync forms, IDatabaseAsync.AsTask(...)
    for Task forms (async state and observed faults), and context.Multiplexer for connection-level settings
    such as AddLibraryNameSuffix. See docs/Extending.md.
  • Client-side caching (ConfigurationOptions.ClientCache, RESP3): see docs/ClientSideCaching.md.
  • Cluster: routing by CLUSTER SLOTS, batches split per slot, MOVED/ASK handled in the executor.

Compatibility

  • Package and assembly version 4.0; prereleases publish from this branch as 4.0.N-alpha.
  • Newer surfaces are gated behind [Experimental] IDs (SER011, SER012, SER014; see docs/exp/).
  • Additions to the IDatabase family are deliberate and expected in a major version; removals and
    signature changes to shipped members are not.

Status

The full suite runs on Ubuntu (net10.0) and Windows (net10.0, net481). The branch's first run had one
net481 failure (GetServerByKeyMemoization, after a 20s cluster connect), which is being investigated;
known flakes are listed in the ledger.

Pending

Mirrors design/v4-ledger.md, which has the detail; ticked here as items close.

Decisions

Work

  • CI: GetServerByKeyMemoization (RESP3) on Windows net481 - speculative fix in d45eadf, unconfirmed
  • Two-core leftovers: Subscription._subscribed/IsSubscribed/AnySubscribed; ~30 ConnectionsIfCreated?. null-tolerances
  • Subscription re-aim waits for its in-flight send to fail (~2s on Windows)
  • An abandoned connect (timeout) keeps its socket in SYN-SENT until the OS gives up - ConnectAsync is called without the token
  • Stale design notes on the cache (local-write and tracking "not wired", key expiry "invalidates", scans/PFCOUNT "wrong to cache")
  • Retry: stale comments in RetryDatabase.ExecuteAsync
  • Prose still says "the new core" in comments
  • Keep absorbing main by merging (ongoing; the final landing must be a real merge)

Watch (flakes, not yet reproducible)

  • One "nothing inbound" 5s timeout in RespCacheChurnTests on CI (Ubuntu, 2026-10-09)
  • net8.0 cluster tests timing out at 10s under the full suite
  • QueuedResultTests.RetryTransactionFireAndForgetSharesOneCompletedTask (once)
  • MaintenanceOptInServerTests.SequenceIdsAdvanceAndCanBeRepeated fails when its class runs alone
  • Known flakes: TouchIdleTime, RespAggregateTiming, ADeferredWalkNeedsNoStorageAtAll

Backlog (after the alpha)

  • Client-side cache across a geo/active-active failover (no flush on switch, by design today)
  • Alternative client-side-cache invalidation sources, and opt-in/out
  • Unobserved faults on the new ValueTask surface
  • Batch/transaction buffer packing
  • Performance: remaining ideas
  • Replace the method-replaying retry decorators with executor decorators
  • Trusted-callback completion mode (opt-in, if it proves to be where Respire's speed comes from)
  • Non-nullable AppendFormatted(long)/(double) overloads (NRedisStack; needs an overload-resolution sweep)
  • .NET 11 TFM: try runtime async; prototype caller-driven TLS (TlsBufferSession); run the suite and RespFest on 11 early

"Nothing arrived on this connection while we waited" and "plenty arrived and none
of it was ours" are different faults - a stalled socket against a desynchronised
reader - and this core's timeouts gave no way to tell them apart. The shipped core
reports the deltas from `Message.TryGetPhysicalState`; here they come from the
counters the connection stamps on the operation when it takes it, which is exactly
what `OnEnqueued` exists for, so nothing is paid until something times out.

Worded as the shipped core words it - `outbound=0KiB, inbound=0KiB, 1954ms
elapsed, timeout is 1000ms` - because somebody who has been pasting that format
into issues for years should not have to learn a second one.

This is also the first piece of the instrumentation 9b-xiii asked for. That rare
RESP3 wedge reports `qs: 6094` and nothing about whether the socket was still
delivering; with these numbers the next sighting distinguishes the two causes
without a repro.

`MaintenanceRelaxationEvidenceTests.ATimeoutInsideAWindowSaysWhichEventCausedIt`
asserts the deltas are there, and was checked both ways - it fails with the clause
disabled.

Engine flag 3 and 3 across two runs (the ReconnectRetryPolicy pair plus one
rotating each), default suite 11173/0, RESPite 1947/0, build clean.
…rned

The priority is removing the dual engines, so this starts on the dependency that
actually holds the bridges up - and it is not the one the Message inventory
counts.

`ReconfigureAsync` establishes every `ServerEndPoint`'s beliefs by handshaking the
SHIPPED bridge and reading ECHO/INFO/CLUSTER/CONFIG through it. None of that
traffic appears in the inventory, because it goes out through
`WriteDirectOrQueueFireAndForgetAsync` rather than `CheckMessage`. Until the
client's beliefs can come from somewhere else, that bridge must connect, and so
must exist - which makes it the gate on D2.8 and therefore on the retry-policy
pair, on D2.3's mode default, and on every remaining two-socket artefact.

This core already handshakes its own socket and learns the same facts. Now it
publishes them: version and server type to the modelled `ServerEndPoint`, and
`ServerEndPoint.Protocol` answers from this core when the bridge has no answer -
which under the flag is often, since that bridge may never have handshaken while
this core has been speaking RESP3 to the server since the first command. Roles and
selectability already flowed the other way; this is the same shape, one fact
discovered once and told to whoever needs it.

Two guards that are load-bearing rather than defensive. `ServerType` is published
only when it was actually DETERMINED - the type has no "unknown" value, so an
undetermined handshake reports Standalone and publishing that would demote a
cluster on the strength of a question nobody answered. And never over a Sentinel
or a proxy: those are decided by configuration and discovery, and the no-HELLO
fallback is CLUSTER INFO, which can only tell cluster from not-cluster and would
report a sentinel as standalone.

The design notes now carry the whole D2.8 dependency list - discovery,
subscriptions, the IServer long tail, sentinel - and how to sequence it, including
the measurement that says discovery is done: turn the shipped handshake off under
the flag and categorise what breaks, exactly as D2.3 was measured.

Engine flag: two consecutive runs at exactly 2 - the ReconnectRetryPolicy pair and
nothing else. Default suite 11172/0, whole-solution build clean.
…lag pair

Two things, and the second is the more valuable.

**`databases` and `replica-read-only` now come from this core's handshake**, in
`RespHandshake.DiscoverServerConfigAsync`: two `CONFIG GET` reads on an interactive
dial, published to the modelled `ServerEndPoint`. That is the next item on the D2.8
discovery gate - the one dependency that actually holds the shipped bridges up,
because a bridge that exists for discovery is a bridge that connects.

Asked ONLY when nothing has described the server yet, which `Databases == 0` says
exactly: it is the value a `ServerEndPoint` starts at, and both settings are
discovered together on the shipped path. Unconditional asking would cost two round
trips on every dial - per node on a large cluster, which is the cost the lazy design
exists to avoid - and they are server-wide answers, so a second ask learns nothing.
Self-adjusting: while the shipped handshake still discovers, this does nothing; when
it stops, this is what knows. A declined `CONFIG` answers null rather than failing
the dial, because it is restricted on plenty of managed deployments and the client
has a default for every one of these facts.

**And the measurement is BOTH flags, not the engine flag.** Measuring
`SEREDIS_NEW_CORE_ENGINE` alone gives 9 stable failures; every one of them passes
with `SEREDIS_NEW_DATABASE_SURFACE` as well. One mechanism in every case: the surface
flag decides what carries an `IDatabase` command and the engine flag decides what
carries a context one, so with only the engine flag an `IDatabase` write goes out on
the shipped bridge's socket while `IServer` and the caching surface go out on this
core's - and a test whose mover and whose oracle are then on different sockets fails
for a reason that has nothing to do with either core being wrong.
`ServerExecuteDatabaseTests` moves the selection with `IDatabase` and reads it back
with `IServer.Execute("CLIENT", "INFO")`; the tracking and cache-invalidation pairs
register interest on one socket and write on the other. The notes now say so, so the
next reader does not spend a day driving an artefact down.

Both flags: 2 failures - the `ReconnectRetryPolicy` pair and nothing else - 11,179
pass. Default suite 11,177/0. Whole-solution build clean.
…ect wait

Ran the measurement the notes prescribed: `AutoConfigureAsync` early-returns, both
flags set, full suite. Two runs gave 4 failures and then exactly 2 - the
`ReconnectRetryPolicy` pair and nothing else, 11,179 passing, which is the baseline
with the probe burst still running. The two extras in the first run were both known
rotators.

So the burst is no longer load-bearing: `CONFIG GET` x3, `INFO replication`,
`INFO server`, the keyed `SET` role probe, `CLUSTER SLOTS`, `CLUSTER NODES` and the
tie-breaker `GET`. This core's handshake supplies version, mode, role and peers, slot
ranges, `databases` and `replica-read-only`, and publishes all of it.

Three caveats recorded rather than glossed, because "the suite passes" is weak
evidence for a probe whose consumers the suite never exercises: sentinel is skipped
in this environment, the tie-breaker only decides between MULTIPLE primaries, and
`run-id` only matters across a server restart. Those move deliberately, which is why
the burst is still in tree rather than skipped under the flag - the measurement was
to learn where the gate is, not to delete the probes.

And it moved: the gate is no longer BELIEFS, it is the CONNECT WAIT.
`ReconfigureAsync` waits per endpoint on `ServerEndPoint.OnConnectedAsync`, which
only a `PhysicalBridge` completes, so a bridge is still constructed and dialled for
every endpoint and that is now the whole of what holds it up. Satisfying it from this
core is three small pieces - `IsConnected`, completing the pending monitors, and
`Activate` dialling - but none of them can land before the bridge stops dialling,
because until then both cores dial everything and the connection-count tests count
twice. That trio and "stop constructing bridges" are one step, and it is the next one.
…face

The plan read "76 `Message.Create` sites on `IServer`, 462 on sentinel", which makes
the remaining work sound like a year of hand-porting before a bridge can stop being
dialled. Measured it instead, and the shape is much better.

A `Message` can already be rendered without a `PhysicalConnection`: `MessageWriter`
has a `(channelPrefix, map, IBufferWriter<byte>)` constructor and detects a
`RespFrameWriter`, so `Message.WriteTo(in MessageWriter)` produces a frame this core
can send. The coupling that is left is the REPLY -
`SetResultCore(PhysicalConnection, Message, ref RespReader)`.

92 `SetResultCore` overrides; **26 touch `connection` at all**, 19 of them in
`ResultProcessor.cs`. And what those 26 want is a flat list of eight connection
facts - `OnDetailLog`, `BridgeCouldBeNull`, `RecordConnectionFailed`, `SetProtocol`,
`ConnectionId`, `SubscriptionCount`, `Protocol`, `MultiDatabasesOverride` - every one
of which this core's connection knows.

So: an interface over those eight, implemented by `PhysicalConnection` and by
`RespClientConnection`, and `SetResultCore` takes it. 66 overrides change a parameter
type; 26 need reading. One wide mechanical change against 538 hand-ports, and it
makes every remaining `Message` site work through this core at once - which is what
lets the interactive bridge stop being dialled without `IServer` and sentinel being
rewritten first.

Revised sequence in the notes: reply interface, Message executor over this core, then
the connect wait and "stop constructing bridges" together, then subscriptions. The
`IServer` and sentinel tails stop being prerequisites and become cleanup.
The previous note recommended an interface over the eight connection facts
`SetResultCore` wants, on the strength of "92 overrides, only 26 touch it". Counting
overrides missed the cost. `SetResult` runs for every reply on the shipped path, so
making its `connection` an interface puts an indirection in the hot path of the
library as it ships today - and it does so for scaffolding that gets deleted at the
end. It is also one big-bang diff across 92 overrides, with the risk concentrated in
the reply dispatch every command depends on.

The tail wins on every axis but speed. 68 `Message.Create` sites in `RedisServer.cs`
across ~22 commands, each independently verifiable, costing the shipped path nothing,
and - the deciding point - a group method is the DESTINATION rather than scaffolding.
The reply interface would be thrown away; these are kept.

And sentinel does not have to move for the interactive bridge to go: a sentinel is
its own `ServerEndPoint` with its own `ServerType` and connection model, so it can
keep a bridge after ordinary servers stop having one. That takes 462 of the
inventory's messages off the critical path, and leaves the reply interface as the
fallback lever for if sentinel ever does have to move with everything else.

Sequence: the `IServer` tail, then subscriptions, then the connect wait and "stop
constructing bridges for non-sentinel servers" together, with sentinel last.
First batch of D2.8's item (3), the `IServer` tail: `LATENCY RESET`, `LATENCY
HISTORY`, `LATENCY LATEST` and `MEMORY STATS` are now `Diagnostics` group methods,
which already carried both `DOCTOR`s and `MEMORY PURGE`. 68 `Message.Create` sites in
`RedisServer.cs` down to 59, of which 26 are `SENTINEL` and staying.

**The element parses are shared rather than rewritten.**
`LatencyHistoryEntry.TryParseEntry` and `LatencyLatestEntry.TryParseEntry` are now
internal statics called both by the shipped `ArrayResultProcessor` and by the new
handler, so a server's own account of its latency cannot depend on which core asked -
the same argument as `Diagnostics.ParseInfo`, and it keeps
`ResultProcessorUnitTests.Latency` covering both paths. `MEMORY STATS` needed no
handler at all: `RedisResultHandler` already exists for the members that promise a
`RedisResult`, and the reply is an open-ended version-dependent tree that should not
be modelled.

**The risk in a port of this shape is the wire form, not the result.** A `LATENCY
HISTORY` with its arguments transposed fails at the server, but a `LATENCY RESET` that
dropped an event name would quietly reset the wrong set, and no result assertion would
notice. So `RespSurfaceDiagnosticsParityTests` keeps the shipped `Message.Create`
calls - driven for real through `MessageWriter` - as the expected bytes, as the stream
and key parity suites do. It also pins the handlers' array walk and the null case,
because the two `IServer` tests that exercise `LATENCY HISTORY`/`LATEST` against a
real server are `Skip.UnlessLongRunning` and so never run in CI; they do pass, checked
by turning the gate on locally.

One deliberate behaviour change, documented in the test: a null reply now reads as an
empty array rather than null. `IServer` declares these members as returning a
non-nullable array, so that was a null the signature said could not happen, and
`LATENCY` has no null reply to send anyway.

Default suite 11,187/0. Both flags: 2 - the `ReconnectRetryPolicy` pair - 11,188
passing. Whole-solution build clean.
Second batch of D2.8's item (3). `Diagnostics` already carried `SLOWLOG RESET`;
`SLOWLOG GET` now joins it, and `RedisServer.cs` is at 57 `Message.Create` sites from
68, 26 of which are `SENTINEL` and staying.

Same shape as the latency batch. `CommandTrace.ParseArray` is now an internal static
called both by the shipped `CommandTraceProcessor` and by the new handler, so there is
one reading of a server's slow-command log rather than one per core - and it keeps
answering null for a bad element rather than throwing, because that is what let the
shipped processor report an unexpected response for the whole reply rather than for
one entry. The non-positive `count` still omits the argument rather than sending zero,
which is not the same command: `SLOWLOG GET 0` asks for no entries.

**The retry category moved with it.** `CommandRetryCategoryUnitTests` asserted that
`SLOWLOG GET` is categorised as a node-local read, and could only reach that through
the `Message` builder it was deleting. `RespRequest.Flags` carries the category, so
the assertion now lives in `RespSurfaceDiagnosticsParityTests` against the request the
group method actually issues - covering `SLOWLOG GET` both ways, `LATENCY
HISTORY`/`LATEST` and `MEMORY STATS`. The alternative was keeping a dead `Message`
builder alive for a test to look at.

Default suite 11,191/0. Both flags: 2 stable - the `ReconnectRetryPolicy` pair - plus
`MovedUnitTests.MovedToSameEndpoint_BatchCommands_QueuedDuringReconnect` and
`RespClientCacheTests.SweepIfDueHonoursTheInterval`, both of which pass in isolation
and belong in the rotating column. Whole-solution build clean.
Third batch of D2.8's item (3). `Config` already carried `REWRITE` and `RESETSTAT`;
`GET` and `SET` join it, and `RedisServer.cs` is at 53 `Message.Create` sites from 68,
26 of which are `SENTINEL` and staying.

**`CONFIG GET`'s group method is internal**, as `Hashes.GetAllArray` is. An array of
string pairs is the OLD spelling, and what the new surface should offer for `CONFIG
GET` is a separate question from getting `IServer` off the `Message` path; answering
it by accident during a port would be a public shape chosen for a porting
convenience. Its reply reads through `ResultProcessor.StringPairs` - the shipped pair
reader, now typed so a handler with no `PhysicalConnection` can drive it, which the
base's `ParseArray` was already written to allow. Jagged is permitted and then
detected from the bytes, which is what covers both wire shapes: RESP3 answers a map
and RESP2 a flat array, and a setting's value is always a scalar, so the detection has
nothing to misfire on. Pinned three ways in the parity suite.

**`ConfigSet`'s follow-up read is deliberately left on the shipped path.** The point of
that read is not its reply - it is discarded - but what `ResultProcessor.AutoConfigure`
does with it: publish the setting to this `ServerEndPoint`, which is how `databases`,
`timeout` and `replica-read-only` stay true after a caller changes them. There is no
context-surface equivalent yet. The split is harmless because `CONFIG` is
server-global rather than connection state, so reading it back on the other socket
answers the same question - and `RelearnSetting` now says all of that in one place
instead of being an unexplained second `ExecuteSync` in two methods.

`CONFIG GET`'s retry category moved to the parity suite with the rest: connection
rather than server-admin, node-scoped, asserted on the request the group method
issues. The parity file is now `RespSurfaceServerParityTests`, since it covers the
`IServer` tail rather than diagnostics alone.

Default suite 11,199/0. Both flags: exactly 2 - the `ReconnectRetryPolicy` pair -
11,201 passing. Whole-solution build clean.
…and says why

Fourth batch of D2.8's item (3), and the more useful half of it is what did NOT land.

`COMMAND GETKEYS` and `COMMAND LIST` are now `Diagnostics` group methods, alongside
the `COMMAND COUNT` that was already there. `RedisServer.cs` is at 47
`Message.Create` sites from 68; 26 of those are `SENTINEL` and staying, so the real
remainder is 21, now itemised in the notes. Both are internal, like `Config.GetArray`
- `RedisKey[]`/`string[]` is the old spelling, and what the new surface should offer
for a server's opinion about commands this client may not model is a question worth
deciding on its own rather than during a port. `COMMAND LIST` keeps the shipped
"more then one filter is not allowed", because the signature it serves is public.

**`CLIENT LIST` is written, tested, and deliberately not wired up.** Wiring it broke
three `DefaultOptionsTests` under both flags - proved by reverting that one line and
watching them pass - and the cause is not the port. It is the first context-surface
call those tests make, so routing it to the other core makes that core dial its own
socket, and a test asserting "this server has two clients" sees three. That assertion
is right about the shipped library and right again once the bridges stop dialling; it
is only wrong in between. `Diagnostics.ClientListArray` and its parity coverage stay,
and `IServer.ClientList` keeps its `Message` with a remark explaining the wait.

The general rule, now in the notes: **the remaining tail ports are no longer free -
any one that is the first context call on a connection-counting path costs a socket
while two engines exist.** So the tail splits, and the sequence is amended: items
nothing counts sockets on can go now, the rest go WITH the stop-dialling step. Which
makes subscriptions the better next large piece, and "stop constructing bridges"
something to want sooner rather than later. Paying for a port with a redder tree is
the trade to refuse: a tree red for known reasons cannot report unknown ones, which
was the whole argument for two flags.

Default suite 11,205/0. Both flags: the `ReconnectRetryPolicy` pair plus
`RespAggregateTimingTests.ADeferredWalkAgainstMaterialiseThenRead`, which passes in
isolation and is already in the rotating column. CI build clean.
…he Message path

68 `Message.Create` sites in `RedisServer.cs` down to 36, of which 26 are `SENTINEL`
and staying - so the real remainder is 10: `SCAN`, `CLIENT KILL`, `REPLICAOF`/
`SLAVEOF`, `CLUSTER NODES`/`SLOTS`, and the multiplexer's own plumbing.

**`CLIENT LIST` is wired up after all, and the tests absorb the cost.** Holding it
back last time was the wrong call: the direction is removing the `Message` machinery,
and an artefact that exists only because two paths exist is not a reason to keep one
of them. Two shapes of artefact, both handled by saying what is true of the process
as it is actually running:

- **Sockets.** Three `DefaultOptionsTests` count what the server can see, and this
  core dials its own connection for an `IServer` command. `OtherCoreSockets` adds
  this core's connections to the expected count and goes to zero on its own when
  there is no second core; the connection SHAPE those tests exist to check - one
  interactive, plus a subscriber under RESP2, per core - is still asserted.
- **Ordering.** `ConfigTests.ClientLibraryName` relies on a fire-and-forget
  `CLIENT SETINFO` being ordered before the `CLIENT LIST` that reads it back. One
  socket orders them; two do not, so the read could overtake the write. It polls now
  rather than asserting an ordering nothing is offering, and collapses back to a
  single read once there is one socket.

Also ported: `ROLE` (sharing `ResultProcessor.ParseRole`, extracted so both cores read
a server's own account of itself the same way), the `SAVE`/`BGSAVE`/`BGREWRITEAOF`
family (sharing `ResultProcessor.ScalarSays`, because "Background saving started" is
the difference between a save that is happening and a server that said something
else), `SHUTDOWN` - whose success looks like a failure, so the connection-fault
swallowing stays with `IServer`, which knows it asked - and the no-`SCAN` `KEYS`
fallback, which nothing else in the suite sends because the test topology always has
`SCAN`, and which is therefore pinned in the parity suite rather than hoped for.

`RespHandlers` gains `KeyArray` and `StringArray`, because `KEYS`, `COMMAND GETKEYS`,
`COMMAND LIST` and the `SCAN` pages all want "array of scalars, nil reads as empty"
and three copies of it is three chances for one to disagree.

Default suite 11,215/0. Both flags: exactly 2 - the `ReconnectRetryPolicy` pair -
11,217 passing. CI build clean.
Tried to move the single-node subscription onto this core. The outbound half worked -
`PubSub.SubscribeAsync`/`UnsubscribeAsync` with the six spellings, the send sites
routed under the flag, 221 pub/sub tests green twice - and then the rest of the suite
took it apart. The routing is reverted; two real defects it exposed are fixed, and the
mapping it settled is kept and pinned so the next attempt does not re-derive it.

**`SubscriptionEndpoint` trusted an ASSUMPTION about RESP3.** It used the ordinary
connection whenever `KnowOrAssumeResp3` said yes, including when "yes" came from
configuration rather than a negotiated fact. That is safe for every other reader of
that question and not for this one, because subscribing is sticky: under RESP2 a
connection that subscribes enters subscriber mode and refuses everything but
(un)subscribe, `PING` and `QUIT`. Guessing wrong does not cost a socket, it poisons the
ordinary connection for every command after it - 46 failures, reported in the server's
own words as `ERR only [P|S][UN]SUBSCRIBE / PING / QUIT allowed in this context (got:
'PUBLISH')`. Now: an endpoint that already has a subscription socket keeps it;
otherwise the ordinary connection is used only for KNOWN RESP3, and an unknown protocol
opens one of its own. One socket more than needed under RESP3 until something has
handshaken the endpoint, and never wrong - which the alternative is.

**This core did not refuse a command the map had disabled.** `Validate` checked
database, admin mode and primary-only, but not availability: the shipped pipeline
refuses while rendering, where `MessageWriter` throws on an empty mapped name, and a
core that renders its own frames had nowhere making that refusal.
`ConfigTests.ConnectWithSubscribeDisabled` asks for it by name, and the fix applies to
every ported command rather than just `SUBSCRIBE`.

**The blocker, written down so it is not rediscovered: a subscription must record which
core holds it.** `IsConnectedAny` asks `ServerEndPoint.IsSubscriberConnected`, which
describes the bridge's socket; point it at this core and every subscription the SHIPPED
path owns is judged by the wrong one. There is one right answer per subscription and no
field to hold it. It showed up as a lost re-subscription after a RESP3 downgrade - the
in-process server's transcript reads `PUBLISH => :0` followed by three late
`SUBSCRIBE`s. The notes also record the other two things the suite insisted on: `Ping`
travels with the subscription or flushes nothing, and "the server is simply known" is
false wherever a `-MOVED` can move it.

Kept and pinned: the group methods, and the six-spelling mapping checked against
`Subscription.GetSubscriptionMessage` - both the command words and the frames - because
a pattern subscription that silently becomes a literal one is the failure that mapping
prevents.

Default suite 11,222/0. Both flags: the `ReconnectRetryPolicy` pair, plus
`ClusterTests.MovedProfiling` and `NewCorePubSubTests.TestBasicPubSubFireAndForget`,
which pass in isolation and are already in the rotating column. CI build clean.
…thing left

Second attempt at moving the single-node subscription, also reverted, but much further
in - and the notes now carry enough that a third attempt starts from here rather than
from the beginning.

Built and verified working: the ownership field the first attempt proved necessary
(`_onNewCore`, set by whoever actually sent the (un)subscribe, with the liveness
questions asking the owning core and the shipped processor claiming ownership BACK when
a bridge establishes one - a subscription can move between cores, so "who holds it" has
to be re-recorded rather than assumed); a re-subscribe triggered when THIS core's
subscription socket establishes, which nothing else was going to do, since the shipped
core re-subscribes from its own subscription bridge coming up; and the "own socket" rule
tightened to NEGOTIATED RESP2 rather than "not known to be RESP3" - which is what
finally made `Resp3DowngradeTests` pass, because unknown had been treated as "probably
RESP2" and so accepted exactly the subscriptions a later downgrade could re-home.

That took the whole pub/sub, handshake, cluster-sharded, config and default-options set
to two failures, and both are one thing, now diagnosed: an `UNSUBSCRIBE` on a
connection holding no subscriptions gets no reply from the in-process test server -
`RedisClient.Unsubscribe` returns early when `SubscriptionsIfAny` is null, so no
confirmation is sent and the command waits for ever. A real server answers with a count
regardless, which is why nobody had hit it. One test sends that directly; the other
sends it as the DEGRADED ping, because `PING` is not answered on a subscriber connection
by every server and the fallback probe is "unsubscribe from something nobody subscribed
to" - a round trip only where a subscription already exists.

So the subscriber ping has to travel on the connection holding THIS CLIENT's
subscriptions, not merely on a subscription-shaped one, and the ownership field is what
makes that answerable. That closes the loop on `Ping` being inseparable: this is the
precise sense in which it is, and it is the first piece of the next attempt.

No source change: this commit is the map.
Third attempt, and the two that failed first are why the pieces are worth reading
rather than just the diff. Five parts, each one there because a test proved it had to
be.

**Ownership, and it is STICKY.** `Subscription._onNewCore`, set by whoever actually
sent the (un)subscribe, with `IsConnectedAny`/`IsConnectedTo`/
`RemoveDisconnectedEndpoints` asking the owning core and the shipped processor claiming
it back when a bridge establishes one - a subscription MOVES between cores when a
downgrade re-homes it. Without the field, liveness is answered by the wrong core and a
subscription reads as live because a connection it is not on happens to be up, after
which nothing re-subscribes it. Sticky because whether this core would take a FRESH
subscription changes over time - it depends on the negotiated protocol, unknown at
first - so deciding per call let one core subscribe and the other unsubscribe, leaving
a channel subscribed on a connection nobody tracked (`Issue1101Tests`: "expected 0
subscribers, found 1"). A placed subscription now decides where its next command goes.

**Recorded at SEND time for a subscribe**, not on completion: a caller who declines the
outcome pings the subscriber to flush it, and the ping must find this core holding the
subscription or it goes out on the other one's socket and flushes nothing
(`TestBasicPubSubFireAndForget`). An unsubscribe still forgets the endpoint only once
confirmed, because until then deliveries can arrive.

**Re-subscribe when THIS core's subscription socket establishes.** The shipped core does
it from its own subscription bridge coming up; a socket this core brought back had no
equivalent trigger, so subscriptions stayed off until something unrelated asked -
`Resp3DowngradeTests` showed it as `PUBLISH => :0` with the `SUBSCRIBE` arriving after.

**Only a socket whose choice cannot change underneath it**: this endpoint already has a
subscription socket, or its protocol is NEGOTIATED RESP2. Unknown says no. Sharing is
right under RESP3 and is also where #3154 lives, because this core picks the socket when
a send is composed - and treating unknown as "probably RESP2" accepts exactly the
subscriptions a later downgrade re-homes.

**The ping goes where this client's subscriptions are**, not merely on a
subscription-shaped socket: the fallback probe for a server that will not answer `PING`
in subscriber mode is an unsubscribe of something nobody subscribed to, which only
round-trips where a subscription exists.

Declined, staying shipped: channels that can be REDIRECTED, because `-MOVED` moves the
subscription after the server was chosen, and recording the pre-chosen one is worse than
not knowing (`ClusterShardedTests.SubscribeToWrongServerAsync`). Closing that needs the
send to report where it FINISHED, which is what `IdentifyEndpointAsync` also wants.

Also fixed: the in-process test server diverged from a real one -
`RedisClient.Unsubscribe` returned early when the client held no subscriptions, so an
`UNSUBSCRIBE` was never confirmed and the degraded ping waited for ever. A real server
answers with a count regardless, which is why only this path hit it.

Default suite 11,223/0. Both flags: the `ReconnectRetryPolicy` pair, plus one
`Resp3HandshakeTests` case that passes 192/192 three times in isolation. CI build clean.
…s missing half

`MultiNodeSubscription` - the keyspace shape, routed to several nodes at once - now
routes through the same `TrySendViaNewCore` and asks the same `IsLiveOn`. It needed no
new mechanism: sticky ownership covers the multi-node case for free, because `IsPlaced`
is "any endpoint placed", so the first placement decides for the whole subscription
rather than per node. That is what stops one channel being half on each core across
several servers.

**And the ownership flag needed splitting, which the suite insisted on.** "We are
carrying this subscription" and "the server has confirmed it" are different questions,
and conflating them loses subscriptions. The ping gate wants the first: a caller who
declines the outcome then pings the subscriber to flush it, and the ping has to travel
on the socket the subscribe went out on - true from the moment it is written. But
`IsConnectedAny` wants the second: answer it optimistically, as recording the endpoint
at send time did, and a later `EnsureSubscribedToServer` skips as "already subscribed",
so a subscribe that never landed is never retried. `Resp3HandshakeTests` read that as a
publish finding ZERO subscribers - the mirror image of the `PUBLISH => :2` the first
attempt produced.

So ownership and an in-flight endpoint (`_sendingVia`) are recorded at SEND time, for
the ping to ask about, and the placed endpoint only on CONFIRMATION. An unsubscribe
likewise forgets the endpoint only once confirmed, because until then deliveries can
still arrive. In-flight counts as placed too, or two concurrent calls race onto
different cores.

**One shipped bug fixed in passing.**
`MultiNodeSubscription.RemoveDisconnectedEndpoints` read
`if (server.Value.IsSubscriberConnected)` and then removed it - inverted against both
the method's name and the single-node sibling, which removes when false. The effect was
benign, because `GetSubscriptionChange` re-checks liveness rather than trusting the
record, so the cost was redundant `SUBSCRIBE`s on reconfigure rather than a lost
subscription; it is still the opposite of what it said.

What is left of D2.5: the redirect-capable channels (sharded, key-routed) and
`IdentifyEndpointAsync`, which want the same missing thing - a send that reports where
it FINISHED rather than where it was aimed.

Default suite 11,223/0. Both flags: the `ReconnectRetryPolicy` pair, plus
`RespCacheChurnTests` and `TransitionalQueuedResultTests`, which pass in isolation.
Three consecutive clean runs of the pub/sub, handshake, keyspace and Issue1101 families.
CI build clean.
…er is deleted

`IServer.StringGet(int db, key)` now reads through `Strings.GetAsync` on a context
moved to the database the caller named - which is also what puts the `SELECT` where it
belongs, since a server context carries no database of its own.

And `GetMemoryPurgeMessage` is **deleted** rather than ported. Nothing in `src/` used it
any longer - `MemoryPurge` went to the context surface earlier - so it was production
code kept alive for a test to look at, which is the thing these ports are supposed to
remove rather than preserve. Its assertions moved onto the request the group method
issues, joining the other ported categories: `MEMORY` as a whole reads, so `PURGE` was
once retried as a harmless read, and the node-scoped bit has to survive or a retry goes
to a server that was never asked. The caller-override case moved with it.

Both ported members were **untested** - every `StringGet` in the suite is the
`IDatabase` one, and the purge builder was only ever reached by the retry test - so each
is now pinned in the parity suite. A port whose only route is the one nobody exercises
is a port nobody has checked.

`RedisServer.cs` is at 33 `Message.Create` sites from 68. 26 are `SENTINEL`, which keeps
its own bridge, so the real remainder is 7: `SCAN`'s five spellings, `CLIENT KILL` (its
filter wants modelling rather than forwarding a token list), the `CLUSTER NODES`/`SLOTS`
call sites (the builders stay, because `AutoConfigureAsync` and the multiplexer use
them), and `MakePrimaryAsync`'s ordered sequence - `REPLICAOF`/`SLAVEOF`, the
tie-breaker `GET`/`DEL`, the reconfigure `PUBLISH` - which are not `IServer` members so
much as steps in a failover.

Default suite 11,224 passing with one timing rotator; both flags 11,223 with the
`ReconnectRetryPolicy` pair and four rotators. All six pass in isolation (545 tests
across their classes). CI build clean.
… the right socket

The last subscription shape that was staying on the shipped path now moves, and it did
not need the "send reports where it finished" capability I had assumed. Two things
close it, neither a new executor contract.

**The redirect has to be followed onto the SUBSCRIPTION socket.** `TryFollowRedirect`
resolved every target through `_forEndpoint`, which is the ordinary connection - so
following a redirected `SSUBSCRIBE` the obvious way would fix the routing and poison
the connection, because under RESP2 a subscribe on the ordinary socket puts it into
subscriber mode and it then refuses everything but (un)subscribe, PING and QUIT. There
is now a second resolver, used for the six subscriber-mode commands - `UNSUBSCRIBE`
included, because a redirected one still has to reach the socket the subscription is on
or it unsubscribes something somewhere else. Under RESP3 both resolvers answer the same
executor, so this costs nothing there.

**Where it landed is re-resolved from the SLOT afterwards, not reported by the send.** A
`-MOVED` both moves the subscription and teaches this core where the slot went, because
`OnSlotMoved` updates the map on the way through - so the slot's owner afterwards IS
the answer. `RespNewCore.EndpointForChannel` asks it; a channel with no slot answers
null, meaning "where it was aimed". That is why the decline could be lifted without
plumbing a new channel back from the executor, which is the thing I had written down as
the blocker.

`ClusterShardedTests.SubscribeToWrongServerAsync` and
`KeepSubscribedThroughSlotMigrationAsync` pass - 6 cases, none skipped - as does the
whole cluster/redirect family twice over (204 tests).

What is left of D2.5 is `IdentifyEndpointAsync` alone, and it is the one member the slot
trick genuinely cannot serve: it asks about an arbitrary channel rather than a routed
one, so it does want the send to report where it finished.

Default suite 11,226/0. Both flags: the `ReconnectRetryPolicy` pair plus
`ClusterTests.MovedProfiling`, which has rotated all session and passes in isolation.
CI build clean.
`IdentifyEndpointAsync` did not want the "send reports where it finished" capability
either. The shipped version reads the endpoint off the connection its `PUBSUB NUMSUB`
reply arrived on, which is whichever one routing chose - so choosing before the send
and answering that is the same fact by a shorter route. Safe because `PUBSUB NUMSUB`
is keyless and node-local and therefore cannot be redirected, which is exactly what
made the SUBSCRIBE case hard and this one easy. The round trip still happens: it is
what makes the answer an observation rather than a guess.

**And moving it exposed a defect in the slot resolution from the previous commit.**
`EndpointForChannel` computed the slot from the channel's own bytes, where the server
routes by the name it RECEIVES - the prefixed one. So with a channel prefix configured,
a subscription landed on one node and was recorded against another. It only surfaced
once `IdentifyEndpoint` moved and the two stopped agreeing by accident;
`ClusterTests.ClusterPubSub(withKeyPrefix: true)` named it in two lines
(`Expected: 127.0.0.1:7001, Actual: 127.0.0.1:7002`). It now asks
`ServerSelectionStrategy.HashSlot(in RedisChannel)`, which already handles the prefix
and `IgnoreChannelPrefix` - borrowed rather than re-derived, which is the rule I should
have followed instead of hand-rolling a slot.

**D2.5 is done.** Publish, subscribe and unsubscribe for every channel shape - literal,
pattern, sharded, key-routed, multi-node - the subscriber ping, the
re-subscribe-on-reconnect, and the endpoint identity all travel on this core under the
flag. That is the whole of the subscription bridge's reason to exist.

Both flags: **exactly 2** - the `ReconnectRetryPolicy` pair, which is not answerable
while both cores consult one policy object and is assertable at D2.8 - with 11,228
passing. Default suite 11,225 with one timing rotator that passes in isolation. CI
build clean.
…re is its shape

Measured rather than assumed, because "subscriptions have moved, so delete the bridge"
looks like a one-liner and is not. Not activating the shipped subscription bridge under
the flag, plus relaxing the "fully established" gate so a connect does not wait for a
leg nobody dials, gives **63 failures** across the pub/sub, handshake, cluster and
config families. Reverted; the chain is now written down so the next attempt starts
from it:

- The subscribe path DECLINES a subscription that would share the ordinary connection,
  because this core picks its socket when a send is composed and a later downgrade can
  move it - issue #3154. That decline was only ever safe because the shipped bridge was
  there to take those subscriptions.
- Remove the bridge and the decline has nowhere to drain: the fall-through becomes a
  subscribe that goes nowhere, which `ClusterTests.ClusterPubSub` reports as an empty
  subscribed endpoint.
- Remove the decline too, so this core takes everything, and the shared-socket path goes
  live for the first time - which is where the other 63 come from.

The pieces to close it are the ones D2.5 already built: `SubscriptionEndpoint` shares
only on KNOWN RESP3, `IsSubscriptionConnected` stops answering yes for a shared socket
once the negotiated protocol drops, and the establish-time `EnsureSubscriptions`
re-places the subscription on a dedicated one. What is missing is making that chain hold
for a subscription composed before a downgrade and written after it - the narrow window
#3154 names. 63 failures is the measurement of how much of the suite depends on getting
that right rather than nearly right.

No source change: this commit is the map. HEAD is the verified D2.5 state - both flags
at exactly 2, the `ReconnectRetryPolicy` pair.
…eded

`origin/main` had drifted by two commits; merged rather than left to accumulate, which
is cheaper now than later and keeps provenance. No conflicts. The valuable part is that
BOTH needed a counterpart on this core - a fix to the shipped path is not a fix to this
one, and neither gap is visible from main.

**#3254, the configuration channel under RESP3.** The shipped fix subscribes the
BRIDGE's interactive connection. Under the flag `IDatabase.Execute("CLIENT", "ID")`
names THIS core's connection, which was not subscribed - `ClientKillTests` caught it as
`Normal` where `PubSub` was expected. Covered today, because the broadcast still lands
on the bridge; at D2.8 nothing would be subscribed and the bug returns. This core now
subscribes its own RESP3 connection, prefix applied, registry bypassed as the shipped
one does. Two orderings are load-bearing: after the handshake and on the NEGOTIATED
protocol, and after `OnPush` is wired - under RESP3 a subscribe confirmation IS a push,
and one arriving before the dispatcher is dropped as unrecognised, which presents as the
connection timing out in its own backlog.

**#3250, TLS/SNI host names.** `RespTransportFactory.AuthenticateAsync` carried its own
copy of the rule that commit replaced. An endpoint that already carries a DNS name must
use ITS name for SNI; only a non-DNS endpoint infers one. Cluster shards over TLS/SNI
are what breaks otherwise - every shard presented the same inferred host, so every shard
but one fails validation. Now `config.ResolveTlsHostName(endpoint)`, shared. That file's
own remarks warn against two copies of a security decision while this was one.

**One known failure, latent in D2.5's work rather than caused by the merge, and open:**
`ClusterTests.ClusterPubSub(sharded: false, withKeyRouting: false, withKeyPrefix: true)`
under both flags receives 20 messages for 10 published, while the publishing node
reports one subscriber each time. Established: it passes at the pre-merge commit
(verified in a worktree) and reproduces with the config-channel counterpart disabled, so
the trigger is the role repair - which sets `IsReplica`, and on this branch that setter
publishes the role and asks for a reconfigure, so topology applies now cascade. The
shape that fits is two of our connections subscribed on two different nodes, since a
non-sharded PUBLISH in a cluster is broadcast to every node while the publishing node
counts only its own: a subscription record moved endpoints while the old socket stayed
subscribed. `RemoveIncorrectRouting` does that unsubscribe-before-replace for key-routed
channels only, and this channel is not key-routed. Next step is instrumentation rather
than another guess; the notes carry all of it.

Default suite 11,245/0. Both flags: the `ReconnectRetryPolicy` pair, this one, and five
rotators that pass in isolation. CI build clean.
Fixes the one failure the main merge exposed, and it was a window rather than a leak.

**Diagnosed by instrumentation rather than guessed at**, after two wrong theories.
Tracing both delivery entry points showed every delivery arriving on THIS core - no
bridge deliveries at all, which killed the "both cores subscribed" idea. Tracing the
out-of-band path by socket then showed subscription sockets at ALL THREE cluster nodes
receiving. A non-sharded PUBLISH in a cluster is broadcast to every node, so each of
those delivers while the publishing node counts only its own subscriber: twenty
deliveries for ten publishes, with a reported count of one.

**The cause is the in-flight window D2.5 created on purpose.** The endpoint is recorded
only on CONFIRMATION, so a subscribe that never landed stays retryable - but
`EnsureSubscribedToServer` guards with `if (IsConnectedAny()) return 0`, and in that
window the answer is "not live". Every ensure arriving in it starts ANOTHER subscribe,
and for an unrouted channel `SelectServer` deliberately spreads, so each lands on a
different node. This core's establish-time re-subscribe then feeds the loop, which is
how it reached three nodes rather than two.

The merge did not cause this. #3254's role repair sets `IsReplica`, which on this branch
publishes the role and asks for a reconfigure, so topology applies now cascade - and
that opened the window often enough to see. It was latent before, and would surface on
any cluster that reconfigures while a subscribe is in flight.

The fix is one clause - `IsConnectedAny() || HasSendInFlight` - and it is the same
conflation the ownership field had to split earlier, in the other direction: **"not live
yet" and "nobody is doing anything about it" are different questions.** `IsHeldByNewCoreOn`
already consulted the in-flight endpoint for the ping gate; this is the other reader that
needed it.

Both flags: back to the `ReconnectRetryPolicy` pair plus `ClusterTests.MovedProfiling`,
11,247 passing. Default 11,244 with the known timing rotator. Three consecutive clean
runs of the cluster, pub/sub, keyspace and Issue1101 families. CI build clean.
First piece of the gate on removing the subscription bridge, and a direct port of the
shipped core's answer to #3154: `PhysicalBridge.WriteMessageInsideLock` hands a
subscriber-mode message to `ServerEndPoint.TryRerouteToSubscriptionBridge` when the
connection it is about to go out on turned out to be RESP2.
`RespEndpointExecutor.RerouteSubscription` is the same question in the same place.

Why it has to be the write rather than the compose: whether a subscription shares the
ordinary connection depends on the NEGOTIATED protocol, and a command can be composed
while RESP3 is expected and written after a reconnect has settled on RESP2 - at which
point writing it there puts the connection into subscriber mode and it refuses every
ordinary command afterwards. Inert while the bridge still exists, which is how it was
verified rather than assumed.

**And removing the compose-time decline on top of it was tried, and does not work yet -
for a reason that corrects my earlier framing.** With the decline gone,
`Resp3DowngradeTests` fails on TIMING, not correctness: nothing is poisoned, but
rerouting at the write has to DIAL the subscription socket first, and the in-process
server's transcript shows `PUBLISH => :0` landing before the re-subscribe arrives. The
shipped core does not pay that because it dials its subscription bridge the moment a
downgrade is detected - `OnFullyEstablished`'s
`else if (SupportsSubscriptions && Protocol > Resp2) Activate(Subscription)`.

So the two mechanisms do DIFFERENT jobs and both are wanted: choosing the right socket
when the protocol is already known is the fast path, and the reroute is the safety net
for the window where compose-time knowledge was wrong. The decline was never a crutch
for the reroute's absence, which is what I had written down before measuring it.

Next piece, now specified: dial this core's subscription socket proactively when a
connection establishes below the expected protocol, mirroring that `Activate`. Then the
decline can go, and then the bridge.

Both flags: the `ReconnectRetryPolicy` pair plus three rotators - `MovedProfiling` and
two `Resp3HandshakeTests` cases that pass 192/192 in isolation - 11,245 passing.
Default 11,244 with the known cache-sweep timing rotator. CI build clean.
… we hold one

A connection that comes up BELOW the protocol it expected has subscriptions about to be
re-placed onto a socket that does not exist yet, and a re-place that has to dial first
loses the race against whatever the caller does next. The shipped core reaches the same
conclusion in the same place, so dial it here on the establish.

Guarded on actually holding a subscription for that endpoint, which is the correctness of
it rather than a refinement: unconditionally, an endpoint whose subscriptions belong to the
shipped bridge gets a second socket and the channel is subscribed twice.

Also forget subscription records naming a replaced socket: a fresh socket carries nothing,
so a record naming it reads as "already subscribed" and suppresses the re-ensure.
Removes the compose-time decline that kept RESP3-shared subscriptions on the shipped
bridge - the gate on the subscription bridge, since while it stood the bridge had to exist
to place them. Four pieces make it safe and each was measured by removing it again; the
fifth had to be built, because the ordering the shipped core gets by accident (both bridges
reconnect on one heartbeat) has to be arranged here, where the subscription socket is
dialled lazily and the need for one is known only after the ordinary connection is warm.

SubscriptionsSettling is that happens-before: an endpoint observed to downgrade holds an
entry until its subscriptions are back on the wire, and publish waits on it, bounded. It
has to be released by the re-place rather than by the dial - a socket already opened once
returns from the connect long before the replacement runs its establish hook.

Also retires the now-dead WouldSubscribeOnItsOwnSocket, narrows the downgrade test's
mixed-traffic classifier to data commands (handshake probes on a dedicated subscription
socket precede subscriber mode and are legitimate), and makes the RESP2 separate-pubsub
assertion ask the server which socket is subscribed rather than assuming the bridge's.
A ServerEndPoint no longer builds a subscription PhysicalBridge at any protocol under the
flag: the subscription leg is the other core's, which dials and owns its own socket, and
with the share-the-connection decline retired every subscription goes there. All four
construction sites now answer with the interactive bridge or decline.

Stopping at "don't activate it" does not work and is worth the note in the design file:
IsSelectable creates the bridge it asks about and then requires it to be connected, so an
unactivated one makes every subscription command unselectable - no server is chosen, the
send reads as "nothing to do yet", and the subscribe silently does nothing. 70 failures,
nearly all a publish reporting no subscribers. Not needing it and not building it are one
change.

Five tests were counting this core's sockets or asking for two distinct connection ids;
each was already written as "RESP3 shares one connection" and now says "or the engine
flag". The sixth was a real gap: IsSubscriberConnected ignored a disabled SUBSCRIBE, which
the shipped path got free because such a bridge never connects.
Under RESP2 the shipped core subscribes the configuration channel as the last step of the
subscription bridge's handshake, so removing that bridge removed the client's only way of
hearing that the topology moved. This core's subscription socket now does it in the same
position - and, crucially, gets dialled at activation rather than lazily: the only
subscriber is the library itself on that socket's own establish, so a client that only
calls ReplicaOfAsync never touches this core and nothing would ever connect.

Dialled on the same KnowOrAssumeResp3 condition the shipped core uses before activating its
bridge, so a wrong guess costs a spare socket exactly as it did there.

Also replaces the tuple state object in the dial, which referenced System.ValueTuple and
failed SanityChecks.ValueTupleNotReferenced.

The connect gate stays open for now: making IsSubscriberConnected consult this core's
subscription leg was measured at 39 failures and a ten-minute run - the same connect-wait
hang as before - so ConnectUsesSingleSocket waits briefly for the count instead, keeping
the upper bound it exists to protect. Recorded in the design notes.
The share-the-connection decline was, in a cluster, also keeping key-routed channels on the
shipped path - so retiring it exposed three things that had never had to work here.

1. Re-resolve the landed endpoint only for SSUBSCRIBE. A key-routed channel has a slot, so
   the slot's owner can always be computed, but a plain SUBSCRIBE is not routed by it and
   never answers -MOVED: the subscription really is on whichever node it was sent to.
   SubscribeToWrongServerAsync(sharded: false) subscribes via a deliberately wrong server
   and requires it to still be there.

2. Act on an UNSOLICITED unsubscribe push. When a slot migrates away the node drops its
   shard channels and says so; there is no command of ours to match that to, so matching
   alone left the client believing it was subscribed on a node that had stopped delivering.
   Read from a copy of the reader, since the matching layer parses the frame from the start.

3. Scope the publish barrier to the downgrade re-place. The configuration-channel dial runs
   at activation, once per node, so in a cluster every publish waited on six of them and
   any one slow to establish held the lot - a five-second publish against the shipped
   path's eight hundred milliseconds.
…re does

Replaces the ad-hoc forget-and-re-ensure with the shipped routine itself. An unsolicited
SUNSUBSCRIBE - the node dropping its shard channels as a slot leaves - now calls
ResubscribeToServer via the OUTGOING node, which is the only one known to have the new
route, and returns Handled rather than leaving it to matching. "Unsolicited" is told apart
by there being no send in flight, which is what the shipped core asks of its own
outstanding commands.

ResubscribeToServer now sends via the owning core, which under the flag is not a refinement:
the shipped send resolves to the interactive bridge once there is no subscription bridge, so
under RESP2 it would put the connection carrying ordinary commands into subscriber mode.

Two further gaps this exposed:

- Rebind never passed forSubscriptionEndpoint, so TryFollowRedirect fell back to the
  ordinary executor and a redirected sharded subscribe put the target's ordinary connection
  into subscriber mode - the next publish there answered "only (P|S)SUBSCRIBE ... allowed in
  this context". The rule was written in TryFollowRedirect; it just had nothing to apply.

- ForgetSubscriptionsOn now skips a subscription with a send in flight. Such a subscribe is
  very often why the socket is being established at all, and its record is about to be
  confirmed rather than stale; forgetting it let the re-ensure place a key-routed channel on
  the slot owner instead of the server the caller explicitly chose.
… weakly from pushes

ActivateServer dials this core's subscription socket where it would have activated the
shipped subscription bridge, on the same KnowOrAssumeResp3 condition, because the
configuration channel is subscribed by connecting rather than by anyone asking: a client
that only calls ReplicaOfAsync never touches this core, so nothing would ever connect and
the broadcast reached nobody.

The establish-hook dial keeps its TryResp3 condition - dropping it made a plain RESP2 client
dial a second time, which RespSubscriptionConnectionTests counts.

OnPush now holds the multiplexer weakly. A subscription connection holds a standing read, so
the socket is rooted while open and anything strong reaching out of it keeps the multiplexer
alive; a push arriving after the multiplexer is gone has nowhere to go anyway. This is right
but not sufficient - the leak predates this change and is written up in design notes 9g with
what to chase next, along with the full before/after suite numbers and the measurement
showing that withdrawing the dial trades two failures for six worse ones.
[Resp("CMS.INFO")] private static partial RespCommand CmsInfo { get; } - one spelling for commands and
operands, instead of a .Command(preform: true) field beside [Resp] fragments.

- The generator emits a static field initialised by RespCommands.Command("NAME"u8), so it resolves
  exactly as .Command() does: a name the client knows keeps its identity and stays deferred, so the
  connection's command map can still rename or disable it; an unknown (module) name is framed once.
- The name is checked when you build (SER351): exactly one token, printable ASCII, no whitespace.
- An explicit field rather than C# 14's `field`: generated code compiles under the consumer's language
  version, and partial properties already need C# 13; `field` would raise that to 14.
- The SER309 code fix declares the [Resp] property from C# 13, and the field below it; adding `partial`
  to a type with no modifiers no longer leaves a blank line (latent in the fragment fix too).
- Tests: a generator-driver test of the emitted code and each SER351 refusal; rendering, including a
  known command still renamed by the map; code-fix tests at C# 12 and 13.
- Docs: Extending.md gains "Declaring commands and tokens"; Execute.md, SER309 and SER351 updated.
Absorbs the drift up to the 3.4.0 cut. Most of main's changes landed in v3's
PhysicalBridge/PhysicalConnection/CursorEnumerable, which v4 has deleted, so
several are re-expressed in v4 terms rather than merged textually:

- #3266 duplicate LoggerMessage event ids: taken as-is. 124/125 (maintenance),
  126 (TLS established) and 127 (ConnectionFailed) shipped in 3.4.0, so the
  v4-only events that had claimed those ids move to 128-131; the guard test
  (LoggerEventIdTests) covers both.
- #3268 create-implementation-plan skill: taken as-is.
- #3260/#3267 scan faults after a sync timeout or dispose: the fix targets
  CursorEnumerable's page prefetch, which v4 does not have (no prefetch; the
  sync face uses the synchronous Send). The intent is kept as tests on v4's
  scans: a faulted page surfaces as its own exception, sync and async, and
  nothing is fetched once the enumeration is let go.
- #3257 ConnectionAttemptCompleted: v4 has no PhysicalConnection until an
  attempt succeeds, so ConnectionAttempt carries the record through the
  transport factory and the handshake instead - stages Connect, Tunnel, Tls,
  Handshake, Established; client/server certificate observation and the TLS
  host; one outcome per attempt. Created only when someone subscribes.
  Abandoned (executor disposal) reports ConnectionDisposed; outliving the
  connect timeout reports UnableToConnect with the timeout. A refused AUTH is
  reported as Handshake + AuthenticationFailure although v4 keeps the
  connection (parity with v3's fire-and-forget AUTH).
- #3264 dedicated pub/sub connection under RESP3 by default: v4's "RESP3, so
  share" becomes SharesSubscriptions (RESP3 *and* SharedSubscriptionConnection).
  v4 keeps pay-per-play dialling: the subscription socket is dialled up front
  only for the configuration channel, as under RESP2. The dedicated
  subscription connection now raises ConnectionRestored (Subscription) when it
  (re)connects, as v3's subscription bridge did - previously only reachable
  under RESP2, and silent there too (Issue922_ReconnectRaised). It does not
  trigger a reconfigure: topology is read over the interactive leg, and the
  subscription leg connects just after the initial connect.
- 3.4 bump: version.json stays 4.0; the 3.4.0 API surface is reconciled into
  PublicAPI.Shipped.txt.

Also fixed while here: the auth-refusal callback added for #3257 was a lambda
capturing `this`, which landed in the closure scope OnPush chains to and let
the socket root the multiplexer (GarbageCollectionTests.MuxerIsCollected).

Closes ledger rows D6 and D7.
…acheAge; document stale-while-revalidate

An invalidated entry inside InvalidationGracePeriod was served before either
age check ran, so it could outlive TimeToLive (documented as the longest any
entry is ever served) and ignore a caller's WithMaxCacheAge. The grace period
now shortens an entry's life after an invalidation and never extends it.

Docs: a "Serving stale while refreshing" section - RefreshAfter
(read-while-ageing) vs InvalidationGracePeriod (read-while-stale), one
background refresh per entry, failure and race behaviour, the window measured
from the invalidation, and where it never applies (own writes, connection
loss, FLUSH).
…flake the abandoned-attempt test

The sequence number was taken with Interlocked and the timestamp read
separately, so two attempts completing together (v4 dials the interactive and
subscription legs concurrently) could carry a later sequence with an earlier
time - EveryFailedReconnectIsReported caught it ~2 runs in 5. Both are now taken
under one lock; the event is rare, so the lock costs nothing that matters.
(main's #3257 has the same shape.)

AbandonedAttemptIsNotReportedAsAFailureToConnect retired the endpoint whenever
the initial connect returned, which could fall between the timed-out first
attempt and the retry, leaving nothing in flight to abandon. It now waits until
an attempt is in flight (IsDialling) and widens the connect timeout.
…ing and failing it

v4 dials the dedicated subscription connection just after connect returns
(v3 connected it before), so on a slow runner a simulated failure could land
before that leg was up - 1 failure for an expected 2 on Windows net481 - and
the leg's first connect could raise a restore after the handlers were attached.
UsesCache (c5eba6c, 4.0.87) asks executor.Transactional on every cached
read, and blocking sends ask CanSendBlocking (!Accumulates) on every call.
On the multiplexer executor both answered by resolving a route - server
selection, then the endpoint's IsConnectedNow, which takes its lock - only to
ask an endpoint executor, which always says false. Batches and transactions
wrap this executor; its routes (built in Rebind, its only construction) never
reach one. So both are now constant false.

Measured through a real multiplexer (cache-scaling, local Redis, RESP3,
GetDatabaseContext().Strings.GetAsync, all hits):

  1 key, 1 thread          7.70 -> 9.54 M/s   (+24%)
  100k keys, 1 thread      5.18 -> 5.98 M/s   (+15%)
  100k keys, 12 threads    5.79 -> 49.8 M/s
  100k keys, 24 threads    4.56 -> 67.5 M/s

Without it, throughput fell as threads were added past 6: every hit on every
thread took the same endpoint lock. RespFest's smoke test had the cached happy
path ~30% down in 4.0.99.
CacheHitSendBenchmarks gains single-component stages - format into a rented
buffer, hash, equals, a plain dictionary probe floor, the cache's own probe on
a pre-hashed key, the TTL clock read (Stopwatch vs Environment.TickCount64),
retain + release, and parse - so each part of a hit can be costed on its own.
Stage 2's remarks were stale: the hash is taken by AsLookupKey, not render.

CacheHitScaling (`-- cache-scaling [seconds] [host:port]`) is a plain
multi-threaded harness: threads x {1, 1k, 100k} keys over the stub executor,
the same through a real multiplexer when an endpoint is given (the stub skips
the multiplexer executor, so it could not see the routing regression), and
the reply's reference count alone, shared vs per thread. First run: one
shared count caps at ~10M/s across 24 threads (307M/s on one), which is where
the full path on a single hot key saturates too.
Every hit took and dropped a reference on the cached reply - a CAS and a
decrement on one shared cache line - because the bytes were pooled (an
ArrayPool rent, or a reservation in a connection's inbound buffer) and a reader
without a reference could find them recycled into someone else's reply.

The cache now stores its own exact-size copy in a fixed RefCountedBuffer that
is never returned to a pool, and TryAddRef on a fixed buffer is a no-op: the
GC is the count, so recycled-buffer reads are impossible rather than avoided.
The cost moves to the fill (one allocation and copy, beside a network round
trip), and an entry no longer pins a slice of an inbound buffer. The
executor's reply goes straight back to its pool.

Measured through a real multiplexer (cache-scaling, all hits):

  1 hot key, 12 threads    10.5 -> 112 M/s
  1 hot key, 24 threads    11.5 -> 140 M/s
  1k keys, 24 threads      18.9 -> 121 M/s
  100k keys, 24 threads    67.5 -> 71 M/s   (memory-bound, not contended)

RespResult.Dispose skipped detaching any result on a fixed buffer, meant for
the shared null singletons; results shared from cache entries are on fixed
buffers now, so it keys on an explicit singleton flag instead. Three tests
asserted refcounts on cache entries; they now assert what still matters - a
hit shares the entry's bytes, and every reply the executor hands over
(the refresh's included) goes back to its pool - via a RecordingExecutor.
Every hit read Stopwatch.GetTimestamp() for the TimeToLive test: ~15ns of a
~134ns hit, against ~4ns for Environment.TickCount64. Ages are seconds to
hours, so a few milliseconds of resolution (~16ms on Windows) is invisible.

CacheClock is TickCount64 (milliseconds) on .NET; down-level keeps Stopwatch,
since TickCount64 does not exist there and the 32-bit TickCount wraps every
~49 days. Entry ages, TimeToLive, RefreshAfter, WithMaxCacheAge and the sweep
interval move to it. The invalidation grace window stays on Stopwatch: it can
be tens of milliseconds, and is only read for an entry already invalidated.

With the GC-owned entries (04ef29d), the single-thread hit
(CacheHitSendBenchmarks stage 4) goes 134 -> 114ns at an 8-byte key and
194 -> 160ns at 256; the cache probe alone 33 -> 14ns.
…inlining budget

4a (a group built in setup) matches 4, so the Strings accessor costs nothing.
JitDisasm shows the 10-15ns between stage 4 and 5 is not the GetAsync frame
(AggressiveInlining on it changed nothing; it is already inlined): one layer
deeper, the JIT runs out of inline budget and leaves RespRequestBuilder's key
path - CommitBulk, CountArguments, FoldSlot - as calls, which the direct
context.SendAsync path inlines. So it belongs to the builder item.
…op, with CPU/op and GC counts

GET 1 KiB over 200k keys into a 32 MiB cache, 1 and 64 async callers, every op
counted (not batches), reporting k/s, process CPU per op, collections per
generation and bytes allocated per op. Written to check RespFest's report
that the miss path got ~15% slower with 04ef29d (GC-owned entries).
GC.GetTotalAllocatedBytes does not exist on .NET Framework, which broke the
build (28f6213); allocation is reported as zero there. And each caller now
walks a share of the keyspace of its own, as RespFest does: the shared walk had
the callers trailing one another through the same keys, so all but the first
hit - a ~95% hit rate on what was meant to be the all-miss path.
04ef29d gave every cached reply an array of its own: hits wrote nothing
shared, but each entry became a mid-life object - alive long enough to be
promoted, dead on eviction - and on an all-miss workload (RespFest's
cache-miss-conc64, which reported -15%) the GC paid for it.

Replies are now bump-allocated, in fill order, out of large uninitialised
arrays the GC owns: one large object per slab, born in the old generation,
never reused - so a reader still parsing simply keeps it alive, and hits stay
free of shared writes. A slab counts the entries in it (plus the allocator's
hold while current); eviction takes the oldest slab whole, one evictor at a
time, and the budget counts whole slabs. TimeToLive runs from the fill, so a
slab is wholly dead within one lifetime of its last fill. Slab size is
MaxBytes/32, clamped to 128 KiB-1 MiB; under a 4 MiB budget there are none,
and replies over an eighth of a slab keep an array of their own.

All-miss, 64 async callers (cache-scaling miss: GET 1 KiB, 200k keys, 32 MiB):

                     k ops/s    CPU/op   gen1/gen2   entries held
  pooled (pre-04ef)  279-285    51.6us      0 / 0        ~900
  GC-owned (04ef)    177-190    58us     ~310 / 38    ~32,500
  slabs              275-287    33us     ~148 / 13    ~31,800

The pooled version held ~900 entries in 32 MiB because a reply reserved in a
connection's inbound buffer pinned the whole block, ~36 KiB per 1 KiB value.
Hits through a real multiplexer, 24 threads: 1 hot key 140 -> 149 M/s, 1k keys
121 -> 128, 100k keys 71 -> 74.
…nd allocate less per fill

The eviction queue grew without bound once slabs arrived (1bd841e): every
store enqueued its key, but with slabs a byte budget is met by evicting whole
slabs, and the queue is otherwise drained only for MaxEntries. Now enqueued
only when entry-at-a-time eviction will read it - a reply with no slab, or an
entry limit.

Every miss registers an InFlight so a concurrent caller can share its round
trip, and each allocated a TaskCompletionSource and its Task eagerly; on a
miss-heavy workload a second caller almost never comes. Created on demand
now, with a full-fence publish/wait handshake so a waiter arriving as the fill
lands cannot be missed. Slab key lists are presized for ~1 KiB replies rather
than regrown and recopied ~10 times per slab.

cache-scaling miss also takes a nocache switch (the baseline) and reports GC
pause as a share of wall time. All-miss, 64 async callers: 275-287k -> 289-315k
ops/s, CPU/op 33 -> 31-32us. The baseline is 533-545k at ~21us, and the gap is
GC - 28% of wall time paused, vs 1% uncached: an entry's small objects are
mid-life on an all-miss workload, and slabs churn the large-object heap. Not
admitting one-off misses at all is the lever for that.
…epeated miss

Storing a reply copies it and keeps an entry until it is evicted, and on an
all-miss workload those entries are mid-life objects: the cache's overhead
there was mostly GC (28% of wall time paused). CacheAdmission.OnRepeatedMiss
puts TinyLFU's doorkeeper in front of the fill: a first miss is remembered in a
two-probe Bloom filter over the request's hash (no allocation, no key copy)
and served without touching the cache; only a request that misses again
before the filter is cleared is stored. False positives admit a request one
miss early; there are no false negatives, so a hot request is never kept out.
The filter is ~8 bits per entry the cache can hold (4 Ki-16 Mi bits) and is
cleared after an eighth as many first sightings as it has bits.

Selectable, default unchanged (OnFirstMiss): the hit-rate side - one extra miss
before a hot key is cached, against one-off reads no longer evicting it - wants
a skewed mixed workload to judge. New counter: RefusedNotAdmitted.

All-miss, 64 async callers (cache-scaling miss ... admit):
  OnFirstMiss       307k ops/s   30.9us CPU/op   29% GC pause
  OnRepeatedMiss    592-594k     19.2-19.4us     3.8%
  no cache at all   536k         20.6us          1.2%
…ounds small replies

Since slabs, an entry was charged only its share of slab bytes - a few bytes
for a count or a flag - while its bookkeeping (the entry, the payload wrapper,
the dependency array, the dictionary node, the copy of the request) is ~350
bytes. So a byte budget never bound a workload of small replies: a million
cached counts reported ~5 MB while holding hundreds. Every entry is now charged
EntryOverheadBytes (256) plus its request's length. The remarks that still
described per-reply ArrayPool rounding are rewritten for slabs.

With the bookkeeping counted, evicting single entries always frees bytes, so
the entry-at-a-time loop now acts on a byte excess too - but only for the
thread that did the slab eviction: one that found another already evicting
slabs leaves the bytes to it, rather than racing it down below the budget.

Also parks the builder-inlining draft as design/parked/builder-inlining.patch
(was only in a session scratchpad); the ledger tracks it with the other cache
ideas still to explore.
The stored copy of the request was an array of its own per entry. It now goes
into the slab with the reply, key then reply in one contiguous range - one
object fewer per entry, and for a small reply (a count, a flag) the two sit
together: a lookup that finds the key is already beside what it came for.
RespRequest.CopyForCacheKey gains an overload that copies into caller-given
storage; requests too large for a slab keep an array of their own.

Same-session A/B, real multiplexer, hits: 100k keys at 24 threads 76.2/71.3 ->
81.4/81.8 M/s; 1 thread 6.6 -> 6.8; one hot key unchanged. All-miss: neutral
within noise - one object fewer per entry is lost in the GC cost that remains
there, which admission (a732552) is the answer to.
An entry depends on the keys its request names - a (key-table node, generation)
pair per key - and held them as an array even for the single-key commands that
are most of them. Every hit asks IsValid, so the array was one more dependent
load between the entry and the generation it compares, and one more mid-life
object per entry. One dependency is now a field; only several get an array.
The fill's own array is no longer referenced by the entry, so it dies young.

Same-session A/B, real multiplexer: hits at 100k keys / 24 threads 83.7/79.9 ->
85.7/84.0 M/s, one hot key at 24 threads 143-148 -> 150; single-threaded
unchanged. All-miss (one run each): 316k -> 342k ops/s, GC pause 24.8% -> 19%.
The miss harness's keys and cache, but each read picks its key from a Zipf
distribution (s = 0.99, a seeded table of 4M draws, so every arm sees the same
sequence), and every mode now reports the hit rate over the measured window.
Where admission's two effects meet: one extra miss before a hot key is cached,
against the long tail no longer evicting it. 64 async callers:
  OnFirstMiss     1.05-1.07M ops/s  9.8us CPU/op  72.4% hits
  OnRepeatedMiss  1.58-1.61M        7.1-7.2us     76.6% hits
  no cache        553k              19.6us
…orrected

It can write - it caches the computed cardinality back into the HyperLogLog's
header - but that is the data structure's bookkeeping, not the command's
meaning: re-running it gives the same answer, and a cache hit that skips it
skips nothing observable. Refreshing the header signals the key modified, so
the PFCOUNT that refreshes it invalidates its own cached reply and the next
one is cached: one extra miss, never a wrong answer. The code already
classified it this way; the design notes (6.9) said caching it was wrong.
CommandCategoryTests.PfCountIsARead pins the decision.
Measured on a skewed workload (Zipf, s = 0.99) it raised the hit rate as well as
throughput (72.4% -> 76.6%, 1.06M -> 1.6M ops/s), and on an all-miss workload it
removed the cache's overhead entirely; it loses only where every request is read
exactly twice. So it becomes the default - as CacheAdmission.Default, resolving
to OnRepeatedMiss, on the CacheTrackingMode pattern: zero means "the library's
choice", so that choice can move without overriding anyone who named a rule.
OnFirstMiss and OnRepeatedMiss renumber to 1 and 2 (unshipped).

The doorkeeper is now sized from the most entries the budget could hold
(MaxBytes / EntryOverheadBytes) rather than assuming ~1 KiB per entry, which gave
small replies a window several times too short.

The cache-mechanics tests (invalidation, eviction, budgets, sweeps - not
admission) now name OnFirstMiss explicitly, since they fill and then expect a
hit. Docs: the quick start shows a reply cached on its second read; the
Admission section explains the rule, the measurement, the cost and how to opt
out; "Is it working?" warns that a check reading twice is testing OnFirstMiss.
…stops routing

Every enumerator - SSCAN/HSCAN/ZSCAN, the emulated key SCAN, vector-set range
paging - checked the WithCancellation token between pages but sent each page
with no token, from when every executor refused a cancellable one. So
cancelling took effect only at the next page boundary: a slow page ran to
completion after the caller had given up, and against a server that never
answered, the enumeration never ended. RespScan.ForSend now hands the token to
the send wherever the executor can act on it; elsewhere the between-pages check
still stops it. CancellingMidPageAbandonsThePageInFlight pins it (and fails,
timing out, without the change).

RespMultiplexerExecutor.CanCancel answered by resolving a route - server
selection and the endpoint's lock - on every call made with a cancellable
token, the same shape as the Transactional regression fixed in c537026; its
routes all reach executors that can cancel, so it is now constant. That also
stops a cancellable call to a multiplexer with no reachable endpoint being
refused as "not supported", rather than failing for the reason it would.

This branch was successfully deployed

1 active (outdated) deployment
release — 74c9a52b Deployed Oct 9, 2026 by mgravell via publish #17
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants