With UCP_ERR_HANDLING_MODE_PEER and rendezvous enabled, ucp_wireup_add_rma_bw_lanes() requires UCT_MD_FLAG_INVALIDATE_RMA from the MD. The TCP MD does not report it, so ucp_put over TCP runs the AM based emulation, 8KB fragments with a completion message each, instead of put_zcopy:
ucx_perftest -t ucp_put_bw -s 4M, tcp/lo |
MB/s |
| without error handling (put_zcopy) |
11935 |
with -epeer (AM emulation) |
2211 |
On our TCP storage hosts the receiving CPU stayed at 100% in ucp_put_handler / ucp_rma_sw_send_cmpl and a single connection was limited to 0.3-0.6 GB/s, against over 2 GB/s with put_zcopy.
TCP registers no memory, so there is no memory key to invalidate. What the flag has to guarantee is that an operation which was canceled does not land in memory the user reused, and that the peer learns about it. Today uct_ep_flush(UCT_FLUSH_FLAG_CANCEL) on a TCP EP only purges the local operations (uct_tcp_ep_flush -> uct_tcp_ep_purge), so PUT data which arrives after the cancel is still written to the address from the PUT header, unlike RDMA where the QP is moved to the error state and the peer's write fails.
Proposed fix in #11918: the cancel marks the EP, PUT data which then arrives on it closes the connection so that nothing is written and the peer's PUT and flush complete with UCS_ERR_CONNECTION_RESET, and the TCP MD reports UCT_MD_FLAG_INVALIDATE_RMA on that basis. uct_ep_destroy() is not changed, so the disconnect flows which rely on a destroyed EP still receiving keep working.
With
UCP_ERR_HANDLING_MODE_PEERand rendezvous enabled,ucp_wireup_add_rma_bw_lanes()requiresUCT_MD_FLAG_INVALIDATE_RMAfrom the MD. The TCP MD does not report it, soucp_putover TCP runs the AM based emulation, 8KB fragments with a completion message each, instead ofput_zcopy:ucx_perftest -t ucp_put_bw -s 4M, tcp/lo-epeer(AM emulation)On our TCP storage hosts the receiving CPU stayed at 100% in
ucp_put_handler/ucp_rma_sw_send_cmpland a single connection was limited to 0.3-0.6 GB/s, against over 2 GB/s with put_zcopy.TCP registers no memory, so there is no memory key to invalidate. What the flag has to guarantee is that an operation which was canceled does not land in memory the user reused, and that the peer learns about it. Today
uct_ep_flush(UCT_FLUSH_FLAG_CANCEL)on a TCP EP only purges the local operations (uct_tcp_ep_flush->uct_tcp_ep_purge), so PUT data which arrives after the cancel is still written to the address from the PUT header, unlike RDMA where the QP is moved to the error state and the peer's write fails.Proposed fix in #11918: the cancel marks the EP, PUT data which then arrives on it closes the connection so that nothing is written and the peer's PUT and flush complete with
UCS_ERR_CONNECTION_RESET, and the TCP MD reportsUCT_MD_FLAG_INVALIDATE_RMAon that basis.uct_ep_destroy()is not changed, so the disconnect flows which rely on a destroyed EP still receiving keep working.