Skip to content

fix(channels): refuse email sends into an unknown thread instead of mailing the Message-ID - #1847

Merged
andreapn merged 1 commit into
use-agent-os:mainfrom
iamhaniofficial:fix/1570-email-unknown-thread-recipient
Sep 12, 2026
Merged

andreapn merged 1 commit into
use-agent-os:mainfrom
iamhaniofficial:fix/1570-email-unknown-thread-recipient

Conversation

@iamhaniofficial

Copy link
Copy Markdown
Contributor

Linked issue

Closes #1570

Summary

EmailChannel._resolve_target fell back to is_email_address(reply_to) whenever the thread named by reply_to was not in the in-memory _threads cache. But reply_to is the inbound thread key — an RFC 5322 Message-ID, which has a mailbox's local@domain shape and a domain chosen by whoever sent the original mail. After a restart or LRU eviction the agent's reply was silently mailed to that address.

Fix

  • channels/email.py — the fallback (and the is_email_address helper that existed for it) is gone. An unknown thread with no metadata["to"]/["recipient"] raises email.send has no recipient for unknown thread '<key>': ... and logs email.send_unknown_thread (two producers swallow send exceptions, so the log line is what makes the refusal visible). send_file reports a non-retryable failure. A message with no target at all gets its own message.
  • Scheduler / heartbeat used to rely on the fallback to deliver to operator-configured addresses. They now pass metadata["to"] — but only when channel_id is a configured recipient. The same field is the thread key when a job was created from an email conversation (mode=origin, or an originating_reply_target rendezvous snapshot); naming that as the recipient would mail the Message-ID even on a cache hit, and re-open the exact hole on a miss.
    • scheduler/delivery.py: configured_recipient = job.delivery.originating_reply_target is None and mode == CHANNEL; failure destinations are always configured. This proxy holds because the cron tool refuses mode=channel from chat callers, so a snapshot-less channel job can only come from an operator (CLI/RPC/Web).
    • scheduler/heartbeat_service.py: uses delivery.mode == CHANNEL. In the target: last branch an override is promoted to channel mode only when it says "mode": "channel" and supplies a channel_id (a channel-mode override without an id still targets the inferred conversation).
    • scheduler/handlers.py / heartbeat_loop.py: overrides now say which mode they represent (_delivery_override_from_fields includes mode; the loop marks a configured to as channel mode; snapshot overrides carry nothing and stay thread replies).
  • docs/channels.md: "Threads and Sessions" describes the refusal and which deliveries do not depend on the routing table.

Compared with the other open PRs on this issue

#1644 removes the fallback but leaves scheduler/heartbeat sending the bare address, which turns every configured cron/heartbeat email delivery into a hard failure. #1827 additionally sets metadata["to"] = channel_id unconditionally in delivery.py, which mails the Message-ID for jobs created from an email conversation — even when the thread is still cached — and does not touch the heartbeat producer at all. This PR threads the provenance through so both classes keep working and no path can populate to from a thread key.

Tests

  • test_email_channel.py: the issue's Message-ID key, an attacker-domain key and a bare mailbox are all refused with nothing sent; a reply works before _threads.clear() and is refused after; send_file into an evicted thread reports FAILED (not retryable); the empty-target message; scheduler-style metadata["to"] still resolves; precedence test updated (to > recipient > cache; reply_to is never a mailbox). The test that encoded the vulnerable contract is replaced.
  • test_multi_account_delivery.py: channel-mode email → {"to": addr}; origin-mode email → {}; snapshot rendezvous → {} with the snapshot key as reply_to; failure destination → {"to": addr}; telegram channel-mode → {} (untouched).
  • test_heartbeat_email_delivery.py (new): _send_delivery for channel vs origin mode; run_once(target="last") end-to-end with no override / configured override / snapshot override / channel-mode override without an id; explicit target="email"; and HeartbeatLoop._tick override construction with and without a configured to.
  • test_system_event_pinned_delivery.py: override dicts now carry mode.

Follow-up (out of scope, same shape)

The message tool writes metadata["recipient"] = target unconditionally and recipient outranks the thread cache, so message(channel="email", target="<Message-ID>") would still mail the key. That is an explicit-address contract rather than this bug; worth its own issue.

Quality gate

uv run ruff check src tests, uv run mypy src/agentos, full uv run pytest -q — all green (one test_pilot_benchmark latency assertion flaked under load with five suites in parallel; passes alone).


No CHANGELOG.md entry on purpose: this PR is one of five opened together (#1570 #1571 #1573 #1575 #1577), each on disjoint files, and the changelog is the one file they would all collide on. All 120 merge orderings were verified conflict-free against upstream/main, and the fully merged tree passes the complete gate. Happy to add the bullet in a follow-up docs(changelog) once the batch lands.

🤖 Generated with Claude Code

https://claude.ai/code/session_01EfPC1g1hsFx5Tw7LefvgDN

…ailing the Message-ID

EmailChannel._resolve_target fell back to treating reply_to as a mailbox
whenever the thread it named was not in the in-memory routing cache. But
reply_to is the inbound thread key -- an RFC 5322 Message-ID, which has a
mailbox's local@domain shape and a domain chosen by whoever sent the
original mail. After a restart or LRU eviction the agent's reply was
silently mailed to that address.

Drop the fallback (and the is_email_address helper that existed for it):
an unknown thread with no metadata["to"]/["recipient"] now raises a
descriptive ValueError and logs email.send_unknown_thread, so send_file
reports a non-retryable failure and channel replies fail loudly.

Scheduler and heartbeat delivery used to rely on the fallback for
operator-configured addresses. They now pass metadata["to"] -- but only
when channel_id is a configured recipient. The same field is the thread
key when a job was created from an email conversation (origin mode, or a
rendezvous snapshot), and naming that as the recipient would mail the
Message-ID even on a cache hit. DeliveryChain derives this from the job
(mode=channel without a snapshot; failure destinations); HeartbeatService
from delivery.mode, with delivery overrides now saying which mode they
represent.

Closes use-agent-os#1570

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01EfPC1g1hsFx5Tw7LefvgDN

@andreapn andreapn left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

LGTM

@andreapn
andreapn merged commit cea7e3e into use-agent-os:main Sep 12, 2026
6 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

[Bug]: EmailChannel._resolve_target sends replies to Message-ID instead of sender address when thread cache misses

2 participants