Skip to content

kvm: clone encrypted roots from a dense base image, not from the sparse template - #13

Open
calvix wants to merge 1 commit into
integration/all-fixes-4.23.0.0from
fix/rbd-encryption-template-holes
Open

calvix wants to merge 1 commit into
integration/all-fixes-4.23.0.0from
fix/rbd-encryption-template-holes

Conversation

@calvix

@calvix calvix commented Oct 1, 2026 •

Copy link
Copy Markdown
Owner

Description

An encrypted root disk on RBD is created as a librbd CoW clone of the (plaintext) template with a
LUKS2 header applied to the clone. librbd serves the ranges the clone's own objects do not hold from
that parent, as plaintext, which is what makes the inherited filesystem readable at all.

That stops as soon as an object exists in the clone. One 4 KiB guest write copies up the whole 4 MiB
object; librbd correctly re-encrypts the parent data it copies, but the ranges the parent never
materialised (a sparse template) are not written at all. The object now exists, so reads of those
ranges no longer go to the parent - they go through the crypto layer, and zeros decrypted with
AES-XTS are not zeros. The guest gets garbage where it wrote nothing and the template held nothing.

The corruption is therefore latent and spreads with ordinary writes: a VM boots and runs, and breaks
later in whatever file happens to live in a copied-up object. A mostly-empty separate /boot
partition shows it first, and ext4lazyinit - which walks the whole disk - spreads it everywhere.

Measured on a clone of the stock Ubuntu 26.04 cloud template (1684 MiB of its 4 GiB unallocated),
librbd 19.2.3 and 20.2.0, reading /dev/vda from inside the guest:

                                                  sectors reading as zero
------------------------------------------------  -----------------------
before any write to the 4 MiB object                      10 / 10
after ONE 4 KiB zero write in the same object              0 / 10

A synthetic parent built with the three cases side by side isolates it - only holes are affected,
and written zeros are not:

parent object content              after the copy-up
---------------------------------  -----------------
real data                          correct
zeros WRITTEN (allocated)          correct
hole (never written)               GARBAGE
data + written zeros               correct

The fix clones from a per-template base image holding a dense (hole-free) copy of the template
instead of from the template itself: a parent with nothing left to materialise keeps every range of
the clone readable, and the clone stays thin. The base is built once per template per pool, under a
staging name and renamed into place only after its snapshot is created and protected, so a
half-written base can never be cloned from and two hosts racing to build one both end up on a
complete image. It is removed together with the template.

The template's own image is no longer modified: it keeps its sparseness, stays the parent of every
plaintext clone, and no longer has a protected snapshot left on it (which previously prevented the
template from being deleted).

QemuImg grows setWriteZeroRanges() for the -S 0 this needs, because qemu-img convert skips
the source's zero ranges by default.

Types of changes

  • Breaking change (fix or feature that would cause existing functionality to change)
  • New feature (non-breaking change which adds functionality)
  • Bug fix (non-breaking change which fixes an issue)
  • Enhancement (improves an existing feature and functionality)
  • Cleanup (Code refactoring and cleanup, that may add test cases)
  • Build/CI
  • Test (unit or integration test code)

Feature/Enhancement Scale or Bug Severity

Feature/Enhancement Scale

  • Major
  • Minor

Bug Severity

  • BLOCKER
  • Critical
  • Major
  • Minor
  • Trivial

Screenshots (if appropriate):

n/a

How Has This Been Tested?

Environment: 3-node KVM cluster, Ubuntu 26.04, qemu 10.2.1, collocated Ceph (RBD primary storage),
librbd 20.2.0; the same failure was observed on librbd 19.2.3.

Unit tests: plugins/hypervisors/kvm suite, 863 tests, 0 failures. RbdEncryptionTest covers the
new dense copy (that -S 0 is requested, and that the base image is copied as plaintext - the LUKS
header belongs to the clone, not to the base). QemuImgTest covers the flag's effect on a local
file: the default convert leaves the source's zero ranges unallocated, -S 0 writes all of them.
Note that QemuImgTest only runs where a libvirt connection is available.

On the cluster, before writing the patch: qemu-img convert -n -S 0 into an existing RBD image
leaves it fully allocated (32 of 32 MiB, against 8 of 32 MiB without the flag), an RBD image can be
renamed while it carries a protected snapshot, and the renamed image can be cloned from.

After deploying the patched agent, deploying an encrypted root from the same template:

  • the clone's parent is <template>-luks@cloudstack-base-snap, overlap 4 GiB (still thin)
  • the base image is dense: 4112 of its 4128 MiB allocated (the 16 MiB header reserve at the end is
    never a copy-up source and stays unallocated)
  • the template image is untouched and still sparse: 2688 of 4112 MiB allocated
  • the measurement that produced 0/10 above now produces 10/10 after the copy-up
  • after ~2 GiB of writes in the guest (forcing copy-up across the disk): /boot checksums
    unchanged, no kernel errors, e2fsck -fn on /boot clean, fsck.vfat -n on the EFI partition
    clean, root filesystem state clean

Cost: the base image is built once per template per pool. Measured 30 s for a 4 GiB template
(~137 MiB/s), so only the first encrypted deploy from a template on a given pool is slower (~46 s
against 16 s); every deploy after it is a plain clone. A second host deploying the same template
before the base exists builds its own copy and discards it after losing the rename.

How did you try to break this feature and the system with this change?

  • Deployed from a template on a host that had never built the base, to confirm the base is shared
    through the pool rather than rebuilt per host (it was reused; no base-build in that agent's log).
  • Forced the race by construction: the base only ever gets its final name after its snapshot is
    created and protected, so a reader either finds no base or finds a complete one. The loser of the
    rename deletes its staging image and then re-checks that the winner's base is actually usable,
    rather than assuming it.
  • Checked the failure paths: a failed dense copy removes the staging image, and the pre-existing
    behaviour of returning null from the clone path on Rados/Rbd failure is kept.
  • Deleted two data volumes to exercise the companion-image cleanup: one with a stand-in base image
    (an image named <volume>-luks carrying a protected snapshot) and one without. The first logged
    Removing LUKS clone base image <volume>-luks along with <volume> and left no image or snapshot
    behind; the second produced no log line and no warning, so the probe is silent for the volumes
    that have no base - which is all of them except an encrypted root's template. Deleting a template
    goes through the same call with the template's name; that case was not exercised on this cluster,
    because the template in question is in use.
  • Deployed a plaintext VM from the same template after the change, to confirm the plaintext clone
    path is unaffected now that the template is no longer resized and no longer carries a LUKS
    snapshot: the template kept its own cloudstack-base-snap, the root is a CoW clone of it, and the
    template image is still sparse (2.8 of 4 GiB allocated).
  • Checked the failure paths by reading them rather than by injection: a failed dense copy removes the
    staging image in a finally, and the pre-existing behaviour of returning null from the clone
    path on a Rados/Rbd failure is kept.

…se template

An encrypted root is a librbd CoW clone of a plaintext template carrying a
LUKS2 header. librbd serves the ranges the clone's own objects do not hold
from that parent, as plaintext, which is what makes the inherited filesystem
readable. That stops the moment an object exists in the clone: one 4 KiB guest
write copies up the whole 4 MiB object, and the ranges inside it that the
parent never materialised are from then on read through the crypto layer -
where zeros decrypted with AES-XTS are not zeros.

Measured on a clone of the stock Ubuntu 26.04 template (1684 MiB of its 4 GiB
unallocated), Ceph 19.2.3 and 20.2.0: of 10 sectors that read as zeros before,
10 still read as zeros after a write elsewhere in the same object while the
object lived in the parent, and 0 of 10 once the write had copied the object
up. The damage is therefore latent and spreads with ordinary writes - a
mostly-empty separate /boot shows it first, and ext4lazyinit walks the whole
disk.

Clone from a per-template base image holding a dense copy of the template
instead: a parent with no holes has nothing left to materialise, so every range
of the clone stays readable and the clone stays thin. The base is built once per
template per pool, under a staging name and renamed into place at the very end,
so a half-written base can never be cloned from and two hosts racing to build
one both end up on a complete image. It is removed with the template.

The template's own image is no longer touched - it keeps its sparseness and
stays the parent of every plaintext clone, and no protected snapshot is left on
it.

QemuImg grows setWriteZeroRanges() for the -S 0 this needs, since qemu-img
skips the source's zero ranges by default.
@calvix
calvix force-pushed the fix/rbd-encryption-template-holes branch from 68cab43 to fa4fc8a Compare October 2, 2026 08:34
@github-actions

github-actions Bot commented Oct 5, 2026

Copy link
Copy Markdown

This pull request has merge conflicts. Dear author, please fix the conflicts and sync your branch with the base branch.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant