Skip to content

term-meshd 종료 hang으로 재시작 중 relay가 90초간 거부됨 #403

Description

@JINWOO-J

현상

Linux peer host root@<peer-host>에서 term-meshd 0.214.0 설치·재시작 중 relay 접속이 약 90초 동안 계속 Connection refused가 됐다.

서버와 SSH는 정상 상태였지만, 기존 데몬이 relay/control socket을 닫은 뒤 종료되지 않아 새 데몬이 systemd의 종료 제한까지 시작하지 못했다.

확인된 타임라인

서버 시각은 UTC다.

2026-08-25 22:55:56  기존 term-meshd가 SIGTERM 수신
2026-08-25 22:55:56  headless agents terminated
2026-08-25 22:55:56  agent sessions terminated
2026-08-25 22:55:56  /run/term-mesh/term-meshd.sock 제거
2026-08-25 22:55:56  peer server shutting down
                         이후 종료 진행 로그 없음
2026-08-25 22:57:26  systemd TimeoutStopSec=90s 만료, SIGKILL
2026-08-25 22:57:26  새 term-meshd 시작
2026-08-25 22:57:26  control socket LISTEN
2026-08-25 22:57:26  peer socket LISTEN

관련 로그:

received SIGTERM, initiating graceful shutdown...
headless agents terminated
agent sessions terminated
removed socket file /run/term-mesh/term-meshd.sock
term-meshd.service: State 'stop-sigterm' timed out. Killing.
term-meshd.service: Killing process 1017411 (term-meshd) with signal SIGKILL.
term-meshd.service: Main process exited, code=killed, status=9/KILL
term-meshd.service: Failed with result 'timeout'.
Started term-mesh peer host.
listening on /run/term-mesh/term-meshd.sock
peer-federation listening on /run/term-mesh/tm-peer.sock

분석

종료는 agent sessions terminated 이후, server join 이전에서 멈춘 것으로 보인다.

daemon/term-meshd/src/main.rs에서 바로 다음 단계는 monitor_handle.resume_all_stopped()다. server join에는 5초 timeout이 있지만 다음 로그가 모두 없었으므로 server join까지 도달하지 않았다.

servers shut down cleanly
server shutdown timed out after 5s
shutdown complete

당시 stack dump가 없어 resume_all_stopped() 내부에서 멈췄는지, 그 직전/직후 runtime이 정체됐는지는 아직 확정되지 않았다.

과거 PrivateTmp=true 때문에 control socket이 SSH mount namespace에서 보이지 않았던 문제와는 별개다. 이 서버에는 현재 다음 설정이 적용되어 있고, 재시작 후 두 socket 모두 정상 동작한다.

TERMMESH_DAEMON_UNIX_PATH=/run/term-mesh/term-meshd.sock
TERMMESH_PEER_SOCKET=/run/term-mesh/tm-peer.sock

영향

  • upgrade/reinstall/restart 때 relay가 최대 90초간 접속 불가
  • 클라이언트는 해당 구간 내내 Connection refused로 재시도
  • socket을 먼저 닫으므로 기존 relay 연결도 끊김
  • Restart=always여도 기존 프로세스가 끝나기 전에는 새 listener가 뜨지 않음

기대 동작

  • SIGTERM 후 종료가 정해진 짧은 시간 안에 완료되어야 한다.
  • 종료 단계 하나가 멈춰도 전체 종료가 systemd의 90초 제한까지 대기하지 않아야 한다.
  • 어느 종료 단계가 지연됐는지 로그로 식별할 수 있어야 한다.

수정 방향

  • resume_all_stopped() 진입·완료와 server join 진입을 계측한다.
  • 종료 단계마다 bounded timeout을 적용한다.
  • hang 시 남아 있는 task, lock 또는 process 정보를 기록한다.
  • Linux service restart E2E에서 relay socket 부재 시간을 측정하고 상한을 검증한다.
  • TimeoutStopSec 단축은 피해 시간을 줄이는 보조책일 뿐, 종료 hang 자체를 고쳐야 한다.

재발 확인 (2026-08-27)

동일한 Linux peer host에서 같은 증상이 다시 발생했다.

2026-08-27 05:35:57  systemd Stopping term-meshd.service
                         이후 term-meshd 로그 없음
2026-08-27 05:37:27  TimeoutStopSec 만료, SIGKILL

이번에는 received SIGTERM, initiating graceful shutdown... 조차 기록되지 않았다.
8월 25일 사례보다 더 이른 지점에서 멈췄는지, 로그가 출력되지 않았을 뿐인지는
현재 계측으로 구분할 수 없다. 종료 단계 계측이 필요하다는 근거가 하나 더 늘었다.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions