Skip to content

Fail over when writing a request finds the TCP peer gone - #148

Merged
hsbt merged 2 commits into
masterfrom
claude/frosty-einstein-335f71
Sep 17, 2026
Merged

hsbt merged 2 commits into
masterfrom
claude/frosty-einstein-335f71

Conversation

@hsbt

@hsbt hsbt commented Sep 17, 2026

Copy link
Copy Markdown
Member

Resolv::DNS#getresources can raise a raw Errno::EPIPE and abandon the whole resolution while nameservers are still untried. It happens when a TCP nameserver's peer goes away while the client is busy with another one. Nothing is watching that socket, so the loss only surfaces when the next request is written to it, and Config#resolv moves on to the next nameserver for ResolvTimeout alone.

The send side of Requester#request now treats a lost peer the way the receive side already treats one, and the requester remembers the failure so the following round starts from a new connection rather than the dead socket.

I found this with a reproduction during the review of #137, where it was left out as pre-existing.

Generated with Claude Code

hsbt and others added 2 commits September 17, 2026 13:53
Three tests carry the same hand built reply, and the tests below need a
fourth.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
A peer that goes away while another nameserver is being tried leaves a socket nobody is watching, so the loss only surfaces when the next request is written to it, as Errno::EPIPE. Config#resolv moves on to the next nameserver for ResolvTimeout alone, so that escaped Resolv::DNS#getresources and ended the whole resolution.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
@hsbt
hsbt merged commit 635d200 into master Sep 17, 2026
74 checks passed
@hsbt
hsbt deleted the claude/frosty-einstein-335f71 branch September 17, 2026 05:18
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant