Skip to content

Openamp rpu fixes sep 30 2026 - #848

Merged
zeddii merged 9 commits into
devicetree-org:masterfrom
bentheredonethat:openamp-rpu-fixes-sep-30-2026
Oct 1, 2026
Merged

zeddii merged 9 commits into
devicetree-org:masterfrom
bentheredonethat:openamp-rpu-fixes-sep-30-2026

Conversation

@bentheredonethat

Copy link
Copy Markdown
Collaborator

No description provided.

Remoteproc ranges took TCM addresses from a hard-coded table whose
TCM_A_1 entries pointed at cluster B core 0 (0xeba80000), and SCMI
TCM_B_0 (0x4a-0x4c) was mapped to TCM_A_1. An RPU1 remote loaded
another core's TCM and aborted on its first fetch.

Take each bank's global address and size from its SDT reg, and use the
table only for the bank index and local offset. Map SCMI TCM_A_1
(0x47-0x49) to TCM_A_1, and stop falling back to xlnx,power-domain on
Versal2, whose SDTs swap r52_1a and r52_0b there.

Signed-off-by: Ben Levinsky <ben.levinsky@amd.com>
legacy_memory_nodes had no entries for TCM_B_0 and TCM_B_1
(0x183180d1-0x183180d6), so a remote on RPU_B_0 or RPU_B_1 failed with
a mapping error. For r52_0b the lookup then fell back to
xlnx,power-domain, which Versal NET SDTs set to TCM_A_1A, and mapped
the bank as cluster A core 1.

Add the six cluster B banks so power-domains resolves them directly.

Signed-off-by: Ben Levinsky <ben.levinsky@amd.com>
The core number came from the CPU node reg. Every R52 cluster, and
cpus_r5_1 in most ZynqMP and Versal SDTs, has a single cpu@0, so RPU1
and every R52 core outside cluster A became r5f@0 or r52f@0 in cluster
A.

RPU cores have consecutive power-domain IDs, so take the core number
from rpu_pd_val, keeping core_num as a fallback. Name the core node by
its index in its cluster, and correct the Versal NET and Versal2
cluster addresses to each cluster's core 0 ATCM.

Signed-off-by: Ben Levinsky <ben.levinsky@amd.com>
When a second core joined an existing split cluster, its ranges were
never written to the cluster node, so Linux could not translate that
core's TCM addresses. Append them to the cluster's existing ranges.

Signed-off-by: Ben Levinsky <ben.levinsky@amd.com>
versal2_scmi_to_legacy_pd covered clusters A, D, and E only, so remotes
on RPU_B_0 to RPU_C_1 failed with a mapping error. Map SCMI
TCM_B_0A..TCM_C_1C (0x4a-0x55) to the firmware IDs of the same banks,
and add the cluster C banks to legacy_memory_nodes.

Signed-off-by: Ben Levinsky <ben.levinsky@amd.com>
elfload could list another core's TCM banks, and the remote then loaded
that core's TCM without an error. The swapped r52_1a and r52_0b
xlnx,power-domain values in Versal NET and Versal2 SDTs make this an
easy mistake.

TCM bank IDs are consecutive per core, so derive each bank's owner and
fail when a split core lists another core's bank. In lockstep, core 0
may use both cores' banks in its cluster. IDs outside that layout are
not checked.

Signed-off-by: Ben Levinsky <ben.levinsky@amd.com>
A core runs one firmware, and its remoteproc node takes one mailbox and
one set of RPMsg carveouts. A second remoteproc relation for a core
dropped its core node; a second RPMsg relation replaced the first one's
mboxes. A cluster in lockstep runs one firmware on core 0.

Before the cluster node changes, fail on a lockstep remote that is not
core 0, on a second relation for a lockstep cluster, and on a second
relation for a split core. Fail on a second RPMsg relation for a core.
A host may still have RPMsg to each remote core.

Signed-off-by: Ben Levinsky <ben.levinsky@amd.com>
Linux binds a host IPI agent either to the mailbox driver, which RPMsg
uses, or to UIO, which libmetal uses. Fail on a libmetal relation whose
IPI agent an RPMsg relation already uses as a mailbox. RPMsg and
libmetal relations to one RPU core are still allowed when they use
different IPI agents.

Signed-off-by: Ben Levinsky <ben.levinsky@amd.com>
In R5 lockstep the two cores' TCM is combined, and Linux expects core
1's ATCM and BTCM in bank 0 at local 0x10000 and 0x30000, global
0xffe10000 and 0xffe30000. Lopper gave them bank 1 at their split-mode
addresses, so their device addresses collided with core 0's banks.

Map core 1's banks into the combined TCM in R5 lockstep. R52 cores do
not combine TCM, so an R52 lockstep core uses its own banks only. Fail
on a TCM bank listed twice, as when an SDT lockstep node, which has
core 0's power domain, joins core 0's bank.

Signed-off-by: Ben Levinsky <ben.levinsky@amd.com>
@zeddii

zeddii commented Oct 1, 2026

Copy link
Copy Markdown
Collaborator

Went through all nine. No big issues, everything looks good, with just a
couple of minor things that can be follow ups.

Worth saying what I checked, since some of it does not show up in a read:

The new rejection paths are fatal, which is the first thing I went
looking for. So a build cannot carry on past one of these with output
quietly missing, which is what the libmetal change a few weeks ago was
about.

The R5 lockstep mapping lines up with the table.

The libmetal IPI check depends on RPMsg relations already being in the tree,
and that holds: xlnx_handle_relations runs remoteproc, then rpmsg, then
libmetal.

Full suite on current master is clean, 124 legacy and 1328 pytest.

The part worth calling out is making the table's global addresses
reference-only and taking the real ones from the SDT reg. That's a good
safety net.

Two things for later, neither worth holding this up.

AI findings:

rpu_core_pd_ids[VERSAL2] is (0x0, 10), so the check in
determinte_rpu_core reduces to 0 <= rpu_pd_val[1] < 10 and any value from 0
to 9 is read as an RPU core number. Every other platform has a large
distinctive base, so a cell holding something unrelated would not match there.
If Versal2 SCMI numbering genuinely starts at 0 there may be nothing to do, but
it is a weaker signal than the others and worth a second look.

Bruce finding:

The other is not yours so much as the file's. The new rejections
mostly use print("ERROR: ...") while the libmetal one uses
_error(...) The prints do not go through lopper's logging though, so
that file probably wants to settle on one of them eventually.

Not asking for a PR description, by the way. It is empty, but the commit
messages carry the detail and they are what stays in the history.

@zeddii
zeddii merged commit ec7d64a into devicetree-org:master Oct 1, 2026
4 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants