Summary
Deploying a imgutil capture image to a node other than the capture source writes to disk successfully but fails to boot: dracut times out waiting for an LV that doesn't exist, because /etc/kernel/cmdline still references the source node's volume group.
Affects el9-diskless/profiles/default/scripts/image2disk.py (confluent 4.0.0). el10-diskless symlinks to the same file, so EL10 is affected too; el7/el8 have separate copies without convert_lv() and are unaffected.
This triggers whenever the source VG is named <osname>_<hostname> — precisely what confluent's own kickstart autopart produces, so the standard "deploy → customize → capture → deploy to cluster" workflow hits it every time.
Root cause
# image2disk.py:192-197
elif ent.startswith('rd.lvm.lv='):
nent = convert_lv(ent)
if nent:
newkcmdlineent.append(ent) # <-- should be nent
else:
newkcmdlineent.append(ent)
Both branches append the original ent, so nent is computed and discarded and the if test is a no-op. The two other convert_lv() call sites in the same file (lines 217-222 for BLS entries, 253-258 for /etc/default/grub) use the converted value correctly.
Why the correct BLS entry rewrite doesn't save it
fixup() does rewrite /boot/loader/entries/*.conf correctly, but afterwards installimage:60 chroots into the new system and runs post.sh, whose line 8 is:
for kver in /lib/modules/*; do kver=$(basename $kver); kernel-install add $kver /boot/vmlinuz-$kver; done
kernel-install add regenerates the BLS entries from /etc/kernel/cmdline, overwriting the correct values with the stale VG name. That file is the effective source of truth, which is why this one line decides whether the node boots.
Capturing from node-...-13 and deploying to node-...-12:
vgcreate rl_node-172-19-52-12 — correct
fixup() rewrites BLS entries to -12 — correct
/etc/kernel/cmdline left at -13 — bug
post.sh regenerates BLS entries from step 3, discarding step 2
- dracut waits for
/dev/mapper/rl_node--172--19--52--13-root, times out, halts
Redeploying onto the capture source works because old and new VG names are identical, which makes this easy to miss in testing.
Reproduction
- Deploy node A from a stock EL9 profile (kickstart
autopart → VG rl_<hostnameA>)
imgutil capture nodeA myprofile
nodedeploy nodeB -n myprofile → writes fine, then fails to boot
Rocky Linux 9.6 x86_64, single NVMe, LVM root + swap. Console on node B:
dracut-initqueue: Warning: /dev/mapper/rl_node--172--19--52--13-root does not exist
dracut-initqueue: Warning: Not all disks have been found.
reboot: System halted
vgs reports rl_node-172-19-52-12 while the BLS entries reference -13.
Fix
@@ -192,7 +192,7 @@
elif ent.startswith('rd.lvm.lv='):
nent = convert_lv(ent)
if nent:
- newkcmdlineent.append(ent)
+ newkcmdlineent.append(nent)
else:
newkcmdlineent.append(ent)
Verified: with this change, capture from -13 → deploy to -12 boots normally, with vgs, /etc/kernel/cmdline and the BLS entries all agreeing.
Note
Two fallbacks in the same file hardcode VG names that won't match what install_to_disk() creates: line 241-243 writes rd.lvm.lv=localstorage/root, and line 237 writes a literal rd.lvm.lv=vg/root. Neither is reached on stock EL9 (both files exist in captured images), but they look like latent bugs.
Summary
Deploying a
imgutil captureimage to a node other than the capture source writes to disk successfully but fails to boot: dracut times out waiting for an LV that doesn't exist, because/etc/kernel/cmdlinestill references the source node's volume group.Affects
el9-diskless/profiles/default/scripts/image2disk.py(confluent 4.0.0).el10-disklesssymlinks to the same file, so EL10 is affected too;el7/el8have separate copies withoutconvert_lv()and are unaffected.This triggers whenever the source VG is named
<osname>_<hostname>— precisely what confluent's own kickstartautopartproduces, so the standard "deploy → customize → capture → deploy to cluster" workflow hits it every time.Root cause
Both branches append the original
ent, sonentis computed and discarded and theiftest is a no-op. The two otherconvert_lv()call sites in the same file (lines 217-222 for BLS entries, 253-258 for/etc/default/grub) use the converted value correctly.Why the correct BLS entry rewrite doesn't save it
fixup()does rewrite/boot/loader/entries/*.confcorrectly, but afterwardsinstallimage:60chroots into the new system and runspost.sh, whose line 8 is:kernel-install addregenerates the BLS entries from/etc/kernel/cmdline, overwriting the correct values with the stale VG name. That file is the effective source of truth, which is why this one line decides whether the node boots.Capturing from
node-...-13and deploying tonode-...-12:vgcreate rl_node-172-19-52-12— correctfixup()rewrites BLS entries to-12— correct/etc/kernel/cmdlineleft at-13— bugpost.shregenerates BLS entries from step 3, discarding step 2/dev/mapper/rl_node--172--19--52--13-root, times out, haltsRedeploying onto the capture source works because old and new VG names are identical, which makes this easy to miss in testing.
Reproduction
autopart→ VGrl_<hostnameA>)imgutil capture nodeA myprofilenodedeploy nodeB -n myprofile→ writes fine, then fails to bootRocky Linux 9.6 x86_64, single NVMe, LVM root + swap. Console on node B:
vgsreportsrl_node-172-19-52-12while the BLS entries reference-13.Fix
Verified: with this change, capture from
-13→ deploy to-12boots normally, withvgs,/etc/kernel/cmdlineand the BLS entries all agreeing.Note
Two fallbacks in the same file hardcode VG names that won't match what
install_to_disk()creates: line 241-243 writesrd.lvm.lv=localstorage/root, and line 237 writes a literalrd.lvm.lv=vg/root. Neither is reached on stock EL9 (both files exist in captured images), but they look like latent bugs.