Fix: onion path rotation could never escape a bad guard node - #2217
Merged
mpretty-cyro merged 3 commits intoSep 28, 2026
Merged
mpretty-cyro merged 3 commits into
mpretty-cyro merged 3 commits into
Conversation
mpretty-cyro
marked this pull request as ready for review
September 20, 2026 20:38
Issue session-foundation#2216. testPath took the first eligible member of the snode pool as its destination, so every path test in a process talked to one node. A node that rejects onion payloads therefore failed every path test on that device for as long as it stayed in the pool: no candidate could ever be verified, so path rotation could never commit, however healthy the candidate paths were. Field evidence was two disjoint candidate paths failing to one constant destination. The regression tests drive rotation, which is the only caller of the path test, and need a foreground test scope - advanceUntilIdle() does not advance work launched into runTest's backgroundScope.
Issue session-foundation#2216, the other half. Rotation rebuilds its candidates around the existing guards, so when a first hop is the problem the candidate built on it fails verification. A rotation can only be committed whole - Phase 3 requires the new guard set to match the current one - so one failed candidate discards the rotation exactly as a total failure does, and the guard that failed is kept. Failures of that shape strike nothing either: a reachable guard returning 400, or any code with no specific rule, maps to PathError and penalises neither node nor path, which is the agreed design. Nothing else broke the loop. Three rotations in a row that fail to verify every candidate now drop the paths via the existing clearPaths(), so the next getPath() rebuilds with no reusable guards and draws fresh ones. A successful commit resets the count, so a flaky network cannot walk a working client into repeated rebuilds. Escalation is skipped when the network is down: every path test fails for a reason that says nothing about the guards, and counting those would hand an offline device a fresh guard every rotation interval for as long as it stayed offline. Counted in attempts rather than time because a failed rotation does not advance the rotation timestamp: a wedged client re-enters rotation on every getPath, so three failures is seconds, not thirty minutes. Worth a reviewer knowing rather than taking as oversights: getGuardSnodes draws from the whole pool without excluding struck nodes, so a bad guard can be re-drawn by chance and a further escalation is what moves off it; and rebuild frequency changes only in the failure path.
Follow-ups to the destination-exclusion change in session-foundation#2221, now that it is on dev. Snode equality is address and port, so the exclusion missed a node whose pool record and swarm record were fetched either side of an IP or port change - the same node under two addresses, one of which is the path's. Path selection now compares ed25519 keys, falling back to equality for a snode carrying no key material. Kept local to path selection rather than changing Snode.equals, which the pool, the paths and the swarm all rely on. Also drops the claim that quic-to-quic "refuses it outright" from the comment explaining the exclusion. That describes the service nodes' behaviour, which has since changed - they detect the self-connection and work around it - and nothing here can notice when it changes again. The local obligation is the durable half: a path must not contain the node it is addressed to.
mpretty-cyro
force-pushed
the
fix/path-test-fixed-destination
branch
from
September 28, 2026 03:34
528e53c to
c35e60b
Compare
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Issue #2216. Two halves of one defect — neither works alone — plus the follow-ups to #2221, now that it is on
dev.The path test always targeted the same node
PathManager.testPathtook the first eligible member of the snode pool as its destination, so every path test in a process talked to one node. A node that rejects onion payloads therefore failed every path test on that device for as long as it stayed in the pool, however healthy the candidate paths were. The field evidence on the issue is two disjoint candidate paths failing to one constant destination.The destination is now picked with
secureRandom()over the eligible members, matching the rest of the file.Rotation could never replace a bad guard
Rotation rebuilds its candidates around the existing guards, and a rotation can only be committed whole — Phase 3 requires the new guard set to match the current one. So the candidate built on a bad first hop fails verification and the whole rotation is discarded, whether it was the only candidate that failed or all of them. Failures of that shape strike nothing either: a reachable guard returning 400, or any code with no specific rule, maps to
PathErrorand penalises neither node nor path, which is the agreed design. Nothing else broke the loop.Three rotations in a row that fail to verify every candidate now drop the paths via the existing
clearPaths(). The nextgetPath()rebuilds with no reusable guards and draws fresh ones, replacing both. A successful commit resets the count, so a flaky network cannot walk a working client into repeated rebuilds.Escalation is skipped while the network is down: every path test fails for a reason that says nothing about the guards, and counting those would hand an offline device a fresh guard every rotation interval for as long as it stayed offline.
It counts attempts rather than elapsed time because a failed rotation does not advance the rotation timestamp — a wedged client re-enters rotation on every
getPath, so three failures is seconds apart, not thirty minutes.Follow-ups to #2221
Snodeequality is address and port, so the destination exclusion missed a node whose pool record and swarm record were fetched either side of an IP or port change — the same node under two addresses, one of them the path's. Kept local to path selection rather than changingSnode.equals, which the pool, the paths and the swarm all rely on.Deliberate, not oversights
getGuardSnodesdraws from the whole pool without excluding struck nodes, so a bad guard can be re-drawn by chance (roughly one in pool-size). A further escalation moves off it.Pathis still a bareList<Snode>.Tests
Seven in
PathManagerTest, all on synthetic pools and paths. Each was verified to fail with only its own production change reverted:The escalation tests assert
getGuardSnodesis called withexistingGuards = emptySet(), so the assumption that clearing the paths actually reaches a guard-replacing rebuild is checked by CI rather than by reading.Rebased on
devd29053ab48. Unit suite: 318 pass, 0 fail.