GO-7556: recover from stale connections after sleep/wake - #802
Merged
Merged
Conversation
After sleep/wake, clients kept reusing dead pre-sleep connections: new streams hung in the handshake or RPCs waited for replies for 10-30s+, and pool.Flush could leave pre-flush peers usable. net/pool: - Flush swaps in a fresh incoming/outgoing cache pair and closes the old pair in the background: peers in parallel, outgoing first, so pre-flush dials are cancelled even if an incoming teardown hangs. - Lookups merge the pair ctx and retry until their result comes from the pair that is still current, so a Get after Flush never returns or waits out a pre-flush dial; ErrClosed surfaces only when the pool is closed. A non-blocking fast path (ocache.Peeker) keeps cached lookups allocation-free and faster than before. - AddPeer retries on a flushed pair or a transient entry. Incompatible-version verdicts survive Flush. Flush/Close are serialized; Close waits, bounded, for pending teardowns. Prometheus collectors are registered once and shared across pairs. net/peer: - Lazy per-peer cleanup owner moves synchronous sub-conn closes (release, failed handshake, gc) off the caller path. In-flight closes still count toward the open limiter; the MultiConn is closed only on stall evidence. Conns doomed by gc are never reused. net/transport: - yamux Open honours ctx. - QUIC and yamux stream Read/Write report a dead connection or session as ErrConnClosed (wrapped). - Optional WriteTimeouter. net/secureservice/handshake: - OutgoingProtoHandshakeWithCloser: a cancelled handshake hands the conn to a closer exactly once; the legacy function keeps its contract. - HandshakeError unwraps to the underlying error.
3 of 4 tasks
Each sub-conn close is bounded by the transport (yamux write/close timeouts, non-blocking QUIC/iroh/webtransport close), and in-flight closes count toward the peer's open limiter, so a backlog throttles new opens instead of needing a stall detector that kills the whole multiconn. Remove the escalation, transport.WriteTimeouter and the yamux writeTimeout plumbing; keep lazy capped workers, the pending list and inFlight. Log once when a single close runs past a minute.
Coverage provided by https://github.com/seriousben/go-patch-cover-action |
Merged
2 of 3 tasks
pool: - AddPeer returns ErrClosed during shutdown, bounds retries, waits for a mid-close entry (ocache Peeker.WaitClosing) and never waits on a flushed pair. - Get restarts from incoming after evicting a dead outgoing peer, so a live incoming connection is used instead of a second dial. - Flush reports Closed for the old pair's peers before it returns; watchers skip peers Flush already reported (once per instance). - Old caches close concurrently; Peek counts no metrics, hits counted once. peer: - Replace the cleanup owner with closeAsync + an in-flight counter. - One throttling deadline per AcquireDrpcConn call (no waiter starvation). - Release and failed-handshake closes count toward the open limiter; gc closes don't. - OutgoingProtoHandshake closes synchronously again; only the closer variant hands the close off. transport: quic and yamux share transport.NewConnClosedError.
- peerservice: a dial cancelled by its ctx (e.g. Flush) tries no more addresses and records no QUIC demotion outcome - pool: Get probes incoming via Peek (never loads into the incoming cache); the redial loop stops when the pair is replaced or the pool shuts down - pool: Flush swaps don't consume AddPeer retries; pick leaves eviction to the watcher; drop keepOnFlush - pool: only the Flush that replaced a pair walks it, once, after the swap, so concurrent Flushes report each peer exactly once - yamux: a write timeout is ErrConnClosed only once the session closed
Contributor
Author
|
Thanks for the recheck. Addressed in b672339. Should fix
Worth raising
Minor
|
pool: - Get dials outgoing only on a real incoming miss; the outgoing loader refuses to dial for a replaced pair and dials under the pair's ctx - Peek reports miss/busy/hit; Get waits out a busy incoming entry (e.g. a GC TryClose that declines) instead of dialing a duplicate - no comparability requirement on peer.Peer (value-level check, by-id fallback) - per-peer close concurrency capped at 256 - invariants documented at the top of pool.go peer: - an RPC error the caller didn't cause becomes ErrConnClosed when the sub conn closed, was doomed, or the session died (drpc otherwise reports a dead yamux session as context.Canceled) - a conn handed to a waiter is claimed under the peer lock - OutgoingProtoHandshake shares the closer variant's body (sync close) transport/peerservice: - yamux reads keep plain io.EOF; Open and closed-session errors are ErrConnClosed with the original cause kept - cancellation is recorded per dial attempt; Accept bounds AddPeer at 10s
peer: - never treat an RPC error carrying a drpc code as a lost connection; the ErrConnClosed wrapper also exposes Code() - streams from NewStream report a dead sub conn on MsgSend/MsgRecv/ CloseSend as ErrConnClosed (caller cancellation stays Canceled) - NewConnClosedError doesn't double-wrap; acceptLoop uses errors.Is pool: - Pick and getIfActive return not-found at once when both caches miss (0 allocs, as on main) - non-comparable peers are held behind a pool-owned pointer, so each pooled instance has exact identity (no by-id fallback) - Flush/Close before Init are no-ops
- connLost rewrites only cancellation, drpc closed errors and closed-network errors on a closed/doomed sub conn; io.EOF (exact), coded and application errors pass through, so `err == io.EOF` works - fast() returns the peer, never the pool's pooledPeer holder - TestPool_FlushStorm: flush count follows the work done, not a background ticker that -race could starve
cheggaaa
approved these changes
Oct 2, 2026
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to subscribe to this conversation on GitHub.
Already have an account?
Sign in.
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Supersedes #801. The pool part is redone: instead of per-peer generation stamps,
Flushswaps in fresh caches. That is simpler, and it fixes a case #801 missed: aGetissued afterFlushused to join a dial that started before the flush and wait it out on a dead connection.After sleep/wake, clients kept reusing dead pre-sleep connections, so handshakes and RPCs hung for 10–30s or longer.
Changes
Flush(new onpool.Pool/Service) atomically publishes a fresh incoming/outgoing cache pair. TheFlushthat replaced a pair walks it once after the swap and reportsClosedfor its peers before returning; watchers skip peers already reported, so each instance is reported exactly once, also with concurrent flushes. The old pair is torn down in the background, with at most 256 peer closes running at a time.Getdials outgoing only on a real incoming miss, and waits out a busy incoming entry (for example a GCTryClose) instead of dialing a duplicate.AddPeerreturnsErrClosedduring shutdown and waits for a mid-close entry before retrying. Retries are bounded, and flush swaps don't count against the bound.AcceptboundsAddPeerat 10s.ocache.Peeker(Peek→ miss/busy/hit,WaitClosing): cached lookups and misses allocate nothing.ocache.OCacheis unchanged.ocache.PrometheusCollectorslets both cache pairs share one set of metrics; names and counting are unchanged.peer.Peerimplementations don't need to be comparable.Flush.Closewaits for pending teardowns, with a time limit.closeAsync). Release and failed-handshake closes count toward the open limiter; gc closes don't. A waitingAcquireDrpcConnkeeps one deadline across wake-ups.transport.ErrConnClosed(matchesnet.ErrClosed) instead ofcontext.Canceledor drpc's "manager closed". A caller's own cancellation stayscontext.Canceled. Errors carrying a drpc code (server replies), application errors and a stream'sio.EOFare never rewritten, soerr == io.EOFstill works.Openhonours ctx. A dead connection is reported asErrConnClosedvia the newtransport.NewConnClosedError, which keeps the original cause reachable. yamux reads keep plainio.EOF. A yamux write timeout counts only once the session has closed, because on a live session a slow uplink can cause one.Flush) tries no further addresses and records no QUIC demotion outcome.OutgoingProtoHandshakeWithCloserhands the conn to a closer instead of closing it inline.OutgoingProtoHandshakebehaves as before.HandshakeErrornow unwraps, soerrors.Ismatches its cause (for exampleio.EOFor ctx errors).Performance
Hit and miss paths in
net/pool/pool_bench_test.goallocate nothing; on main, outgoingGetdid 4 allocations. On hit paths the geomean is about −40% vs main. yamuxOpenwith a cancellable ctx costs about 3µs and 4 allocs more per open, once per pooled sub conn. Servers never callFlush, so for them this branch brings faster lookups, bounded parallel closes on shutdown, and closes moved off caller paths.Test plan
go test -race ./..., with stress runs of the changed packages (-count=20,-cpu=1,4)Linear: GO-7556