Skip to content

CORE: Extract shared team state into artifacts holder - #1349

Draft
bwestheimer wants to merge 19 commits into
openucx:masterfrom
bwestheimer:bwestheimer/pub-team-cache-artifacts
Draft

bwestheimer wants to merge 19 commits into
openucx:masterfrom
bwestheimer:bwestheimer/pub-team-cache-artifacts

Conversation

@bwestheimer

Copy link
Copy Markdown
Collaborator

What

Extract ctx_map, ctx_ranks, and topo from flat fields in ucc_team_t into a new ucc_team_artifacts_t refcounted holder. All teams initially use an embedded inline holder, so there is no behavioral change in this PR. Every call site across TL and CL components is updated to use UCC_TEAM_CTX_MAP(), UCC_TEAM_CTX_RANKS(), and UCC_TEAM_TOPO() accessor macros.

Three commits:

  1. Add holder struct, inline init/put functions, and accessor macros (core files only)
  2. Accessor macro sweep across TL/CL components (mechanical, one-liner per site)
  3. ucc_topo_prepare_shared(): eagerly materialize all lazily built topo fields before a topo is shared across derived teams

Why

Prerequisite for derived-team caching (#1349), which shares one holder between a live parent team and a derived team that borrows its already-built membership map and topology.

Depends on #1348.

@svcnbu-swx-hpcx

Copy link
Copy Markdown
Collaborator

🤖 CI Triage Agent — Linter · commit f918f6fe

TL;DR: The Linter (clang-tidy-17) job failed because the static analyzer flagged use-after-free bugs in the new team-cache code — ucc_debug calls that reference pointers (cache_ptr / team_ptr) which alias memory already released by ucc_free. Move the debug logging before the ucc_free.

Full analysis

Summary: GitHub Actions "Linter" job exited 125 — clang-tidy-17 reported 3 -warnings-as-errors in core/ucc_team.c and 1 in core/ucc_team_cache.c.

Root cause: In ucc_team_cache_destroy() the sequence ucc_free(cache) (line 95) then ucc_debug("ucc_team_cache destroyed: %p", cache_ptr) (line 96) triggers clang-analyzer-unix.Malloc — cache_ptr aliases the freed cache. The same pattern in ucc_team_cache_progress_pending() (ucc_team.c): after ucc_team_destroy_single(team) frees team (via ucc_free(team) at line 955), the ucc_debug(...) at line 1114 references team_ptr, which the analyzer treats as use-after-free; it also emits an UndefinedBinaryOperatorResult (garbage value) at line 1101 as a knock-on. All are treated as errors, so the job aborts.

Implicated commit: 830d5ad / 8e8001c "CORE: Add cross-rank team cache vote" / "Add team cache container and FIFO eviction" — Bryce Westheimer (same author/branch as this PR).

File: src/core/ucc_team_cache.c:95-96 and src/core/ucc_team.c:1114 (with knock-on at :1101).

Suggested fix: Emit the debug log before freeing, so no freed pointer is referenced. In ucc_team_cache_destroy() move the ucc_debug("ucc_team_cache destroyed: %p", cache_ptr) above ucc_free(cache) (and you can then drop the cache_ptr alias). In ucc_team.c, move the success ucc_debug(...) (lines 1114-1118) so it runs before the team's memory is released inside ucc_team_destroy_single — e.g. log immediately after checking status == UCC_OK but before the team can be freed, using the already-captured team_ptr/team_id only for values that don't require dereferencing freed memory. Alternatively, if these are deemed false positives, add a targeted // NOLINT(clang-analyzer-unix.Malloc) — but reordering the log is the correct fix.

Related: PR #1349 (bwestheimer/pub-team-cache-artifacts) — none other found.

@bwestheimer bwestheimer changed the title CORE/TEAM_CACHE: Extract shared team state into a refcounted artifacts holder CORE: Extract shared team state into artifacts holder Sep 2, 2026
@bwestheimer
bwestheimer force-pushed the bwestheimer/pub-team-cache-artifacts branch from f918f6f to 492cf54 Compare September 2, 2026 20:41
@bwestheimer
bwestheimer force-pushed the bwestheimer/pub-team-cache-artifacts branch 2 times, most recently from b260f4f to 344ec5a Compare September 3, 2026 14:04
@bwestheimer
bwestheimer force-pushed the bwestheimer/pub-team-cache-artifacts branch from 344ec5a to 6ea5442 Compare September 3, 2026 14:39
Add an 'embedded' flag to ucc_service_coll_req_t. When set,
ucc_service_coll_finalize skips ucc_free(req) so the request can live in
caller-owned storage instead of being heap allocated. The team agreement
vote embeds its request directly in ucc_team_t and relies on this.

Add ucc_service_allreduce_ctx, which runs an allreduce over a
pre-materialized subset of context endpoints using the context-level
service team. The existing ucc_service_allreduce resolves subset ranks
through team->ctx_map, which is not yet built while a team is being
created. The new entry point maps subset indices straight to context
ranks and routes through ctx->service_team, so the agreement vote can run
before ctx_map exists.
Introduce ucc_team_cache_identity_t, the normalized key type used to
recognize a team by its membership, together with its build, compare and
free helpers and the cacheability policy check. No cache container is
added here.

ucc_team_cache_identity_build materializes members[] by evaluating the
ep_map, so an identity never aliases caller-owned storage such as an
ep_map callback closure or a user array. Any ep_map style (CB, ARRAY,
STRIDED, FULL) describing the same membership yields the same identity.

The FNV-1a hash covers membership only (size, self_ep, members[]).
ext_id is compared separately and deliberately not hashed, so teams with
the same membership but different external communicator ids land in the
same bucket; PR3a introduces the chaining that lets them coexist there.
ucc_team_cache_identity_equal_membership provides the ext_id-agnostic
compare that the derived-team path uses later.

instance_cookie is 0 until the agreement vote stamps it in PR1f.

ucc_team_cache_is_cacheable rejects teams that set optional behavioral
params (ORDERING, OUTSTANDING_COLLS, SYNC_TYPE, P2P_CONN, MEM_PARAMS),
since those are not part of the identity and a reuse could silently
change semantics.
Add the team-cache container as a standalone data structure. It owns a
khash uint64 -> ucc_team_t* bucket table plus four intrusive lists (live,
dormant, reserved, pending_destroy), all protected by a single spinlock,
together with the size/capacity bookkeeping and the hit/miss/insert/
eviction counters.

The lifecycle API is refcount based: insert admits a team as DORMANT, get
adopts a DORMANT team into LIVE with refcount++, and put releases a LIVE
team back to DORMANT once the refcount reaches zero. Insert is a no-op at
capacity, which leaves the team uncached but fully functional, and it also
skips a team whose bucket is already occupied by a duplicate identity or a
hash collision. Lookup only ever returns a DORMANT team, so a LIVE or
RESERVED entry is never handed out twice.

The eviction victim picker is FIFO, that is the dormant list head, which is
the oldest insert; LFU selection is added later. The registry helpers do
the list surgery for the state transitions and table_erase removes a team
from the bucket table.

ucc_team_t gains the fields the cache API needs: refcount, cache_identity,
cache_link, cache_state, cache_pending_insert and cache_local_action.

This commit adds no wiring into team create/destroy and no agreement vote;
the lifecycle hooks and the vote machinery are separate changes.
Add the five UCC_TEAM_CACHE_* context config knobs, create and destroy the
per-context cache, and hook the cache into ucc_team_create_post,
ucc_team_create_test and ucc_team_destroy so a destroyed team is retained as
DORMANT and re-adopted by a later create with identical membership.

Each rank classifies its create as a hit or a miss locally, with no cross-rank
agreement. EXACT_REUSE without that agreement is safe as long as team scopes
never overlap, that is, no rank belongs to two simultaneously created teams
with the same membership. The user guide documents this restriction. The
agreement vote that lifts it lands in a follow-up, together with the handling
for the UCC_TEAM_CACHE_AGREE and UCC_TEAM_CACHE_MISS_TEARDOWN states defined
here.

UCC_TEAM_CACHE_ENABLE defaults to n, so the feature is entirely opt-in.
Add gtest integration coverage that drives real create/destroy/recreate cycles
through UccJob: dormant re-adoption, the disabled knob leaving no cache,
eviction with team-id release, id-pool headroom under dormant teams, and the
dump-stats and disable-linear-check knobs.

Add a multi-rank ucc_test_mpi suite, gated on
UCC_TEAM_CACHE_CORRECTNESS_TESTS, covering dormant-reuse hit counts, safety of
a caller ep_map callback freed after the team is cached, and singleton teams.

Document the team cache and its knobs in the user guide, including the
non-overlapping team scope restriction for reuse without agreement.
Members of a cacheable team create classify the create as a cache hit or a
miss from their own cache contents, and those contents can diverge, for
example after an eviction on one rank only. A create where some ranks
re-adopt a dormant team while others build a fresh one does not progress.

Reconcile the per-rank action with a UCC_OP_BAND allreduce over a small
vote buffer before any rank skips the address exchange. The buffer carries
a prepared flag, (value, ~value) equality pairs for the action, key and
instance cookies, and team rank 0's proposed cookie. Any disagreement
degrades the result to MISS, so all members fall back to a fresh build.

This replaces the UCC_TEAM_CACHE_AGREE and UCC_TEAM_CACHE_MISS_TEARDOWN
stubs. A rejected reuse candidate is torn down and rebuilt in place, which
keeps the handle the caller already holds valid.

Add UCC_TEAM_CACHE_AGREEMENT (default y) to control the vote.
Add gtests for the vote helpers: a unanimous EXACT_REUSE agrees and
distributes rank 0's cookie, a single non-preparing rank degrades the
result to MISS, a cookie or parent-cookie mismatch also degrades to MISS,
and next_cookie stays monotonic and never returns 0.

Add two MPI tests. overlap_agreement builds overlapping subcommunicator
sets that, with UCC_TEAM_CACHE_MAX_SIZE=2, evict divergently across ranks;
without the vote this deadlocks, with it the members reconcile to a fresh
build. nonblocking_create_post delays rank 0 and checks that the peers'
ucc_team_create_post returns rather than blocking on the vote.

Document UCC_TEAM_CACHE_AGREEMENT and drop the overlapping-scope
restriction note, which the vote removes.
Add test/mpi/run_cache_equivalence.sh, which runs ucc_test_mpi twice over
the same team set (world, half, odd_even, reverse) and collective set
(barrier, allreduce, bcast, alltoall, allgather) -- once with
UCC_TEAM_CACHE_ENABLE=y and once with n -- using the test's built-in
per-collective correctness checks as the equivalence oracle rather than
diffing outputs across runs.

Both passes also set UCC_TEAM_CACHE_CORRECTNESS_TESTS=y so the cache-on
pass exercises real reuse and derivation. That pass additionally asserts
the correctness suite did not skip itself, so a silently inert cache
cannot masquerade as equivalent by never touching the cache at all.

Wire both the team-cache correctness suite and the equivalence pass into
.ci/scripts/run_tests_ucc_mpi.sh, and distribute the script via
EXTRA_DIST (CI invokes it directly from the source tree).
Introduce ucc_team_artifacts_t to hold ctx_map, ctx_ranks and topo behind
a refcount and a spinlock, so a later change can share them between a
cached team and the teams derived from it.

Every team uses an embedded inline holder (heap=0, refcount=1), so there
is no sharing and no behavioral change yet:

- ucc_team_artifacts_init_inline zero-inits the holder, sets heap=0 and
  refcount=1, and initializes the spinlock.
- ucc_team_artifacts_put decrements the refcount under the lock; at zero
  it releases topo and ctx_ranks, and frees the struct only when heap=1.
  This replaces the topo and ctx_ranks teardown in the destroy path.
Replace every direct read of core_team->ctx_map, core_team->ctx_ranks and
core_team->topo in the TL and CL components with UCC_TEAM_CTX_MAP,
UCC_TEAM_CTX_RANKS and UCC_TEAM_TOPO, so all access goes through the
artifacts holder.

Mechanical one-liners with no behavioral change.
ucc_topo_t fills its sbgp, socket, numa, node and node-leader state
lazily on first use. That is fine while a topo belongs to exactly one
team, but a later change lets a cached team share its topo with the
teams derived from it, and concurrent first-touch fills from several
teams under UCC_THREAD_MULTIPLE would race on those writes.

Add ucc_topo_prepare_shared, which walks every sbgp type and the all_*
socket/numa/node arrays (plus node leaders for multi-node teams) so the
topo is fully built and thereafter read-only. Per-sbgp failures record a
terminal status and are not retried, so only an allocation failure in the
retryable all_* paths is treated as fatal.

Call it from ucc_team_create_cls, but only for teams that are candidates
for the cache; ordinary teams keep the existing lazy behavior. On failure
the team is still created and usable -- it is simply dropped from the
cache candidate set rather than being shared with a half-built topo, so
no user-visible operation aborts.
ucc_service_allreduce_ctx asserted that the context service team was
non-NULL. That team is only created when the context carries an OOB and
UCC_INTERNAL_OOB permits it, and its creation is deliberately non-fatal,
so NULL is a supported runtime state rather than a caller error -
ucc_service_coll_req_init already treats it as one and falls back.

Release builds define NDEBUG, so the assert compiled out and left a NULL
dereference in UCC_TL_TEAM_IFACE. Reproduced with UCC_INTERNAL_OOB=0 and
UCC_TEAM_CACHE_ENABLE=y as a SIGSEGV in ucc_service_allreduce_ctx called
from ucc_team_create_post.

Return UCC_ERR_NOT_SUPPORTED instead. There is deliberately no fallback
to the per-team service team: this collective addresses context ranks
directly, which a team-scoped service team cannot do.

Also realign the assignment block that req->embedded widened, and order
the initialized declaration in ucc_service_coll_finalize first.

(cherry picked from commit 4d564b1400fc46ab1cda4449b1a4eb9ab75e3010)
The cross-rank agreement vote runs over the context service team. That
team is created after the cache is initialized, and its creation is not
fatal, so a context can reach team creation with caching enabled,
agreement requested, and no team to vote over.

Drop the cache in that case rather than continuing without agreement.
Reuse without the vote is safe only when team scopes never overlap, so
disabling agreement automatically would trade a performance feature for
a correctness risk the user never accepted, and the resulting failure is
a hang rather than a slowdown. The condition derives from configuration
alone, so every rank reaches the same decision and the gate stays
uniform.

UCC_TEAM_CACHE_AGREEMENT=n still caches without the vote, which the
warning names; that path is unaffected by this check.

(cherry picked from commit 4ec16e4cb57b5c7f034724b4fb0365daa17cfd8e)
A rank whose team cache failed to allocate was left running without one,
while its peers kept theirs. Such a rank skips the agreement vote in
ucc_team_create_post and its peers wait on a member-scoped allreduce that
never completes. Fail the context instead, so whether a rank caches is
decided by configuration alone, which every rank shares.
ucc_team_destroy_single finalized bp.params.oob whenever the context had a
service team. When ucc_internal_oob_init itself failed, that field still
held the caller's OOB, and the destroy on the create error path freed the
caller's coll_info. Record ownership in team->internal_oob after a
successful init and finalize only what UCC installed.
When the agreement allreduce returned an error the team stayed in
UCC_TEAM_CACHE_AGREE: a RESERVED candidate never returned to the dormant
list, ucc_team_destroy rejected the handle, and a repeated
ucc_team_create_test re-tested the finalized request.

Factor the post-time rollback into a reusable release of the reserved
candidate and use it from the vote path too. A shell this create allocated
moves to a terminal UCC_TEAM_CREATE_FAILED state that ucc_team_destroy
accepts; a handle that named a cached team is handed back and destroy
refuses it, as it does for any dormant or reserved team.

Destroying such a shell also reached ucc_coll_score_free_map with a NULL
map, which the post-time failure path already did; guard it.
A cached team whose teardown failed terminally had its id returned to the
pool while the CL/TL teams that were not destroyed could still match
traffic on it. Leak the id together with the rest of the state, so the
failure stays a bounded leak instead of letting two teams share a tag
domain.
Evicted teams keep their team id until their asynchronous destroy
completes, but the pending-destroy list was only progressed from the next
admission, eviction or context drain. A teardown that stayed in progress
could hold ids inside the pool headroom with nothing driving it.

Register a throttled context progress callback that drives the list when
it is non-empty, and progress it once before a team posts its id-pool
allreduce so finished evictions return their ids first.
The callback-lifetime test freed the poisoned cb_ctx box and immediately
allocated a same-sized replacement with a valid magic. An allocator that
returns the same address would make a retained stale pointer look valid.
Keep the poisoned box allocated until the reuse checks finish.
@bwestheimer
bwestheimer force-pushed the bwestheimer/pub-team-cache-artifacts branch from bd96f34 to e8c4ebe Compare September 28, 2026 17:36

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants