Skip to content

MEP-20 Machine Provisioning V2 - #282

Open
majst01 wants to merge 23 commits into
mainfrom
layer-3-only
Open

majst01 wants to merge 23 commits into
mainfrom
layer-3-only

Conversation

@majst01

@majst01 majst01 commented Jun 9, 2026 •

Copy link
Copy Markdown
Contributor

Machine Provisioning V2

Depends on MEP-4

Used AI-Tools ✨

  • Qwen3.6 used for generation of ideas in the ai folder.

@metal-robot metal-robot Bot added the area: documentation Affects the documentation area. label Jun 9, 2026
@metal-robot metal-robot Bot added this to Development Jun 9, 2026
@netlify

netlify Bot commented Jun 9, 2026 •

Copy link
Copy Markdown

✅ Deploy Preview for metal-stack-io ready!

Name Link
🔨 Latest commit 59db555
🔍 Latest deploy log https://app.netlify.com/projects/metal-stack-io/deploys/6ac8f646e93cea000871fb09
😎 Deploy Preview https://deploy-preview-282--metal-stack-io.netlify.app
📱 Preview on mobile
Toggle QR Code...

QR Code

Use your smartphone camera to open QR code link.

To edit notification comments on pull requests, go to your Netlify project configuration.

Comment thread community/04-Proposals/MEP20/ai/mep-ra-slaac-boot.md Outdated
Comment thread community/04-Proposals/MEP20/README.md Outdated
Comment thread community/04-Proposals/MEP20/README.md Outdated
Comment thread community/04-Proposals/MEP20/README.md Outdated
Comment thread community/04-Proposals/MEP20/README.md Outdated
Comment thread community/04-Proposals/MEP20/README.md Outdated
Comment thread community/04-Proposals/MEP20/README.md Outdated
Comment thread community/04-Proposals/MEP20/README.md Outdated
Comment thread community/04-Proposals/MEP20/README.md Outdated
Comment thread community/04-Proposals/MEP20/README.md Outdated

## metal-core

metal-core will need to support additional configuration templates for the boot vrf.

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

We could also install metal-core in a way that it also talks to metal-apiserver from within this boot-vrf which completely eliminates the need for weird routes on the switches

Comment thread community/04-Proposals/MEP20/README.md Outdated
The L3 only boot and registration process can be described as follows:

- Every server will be scanned on a regular basis from the metal-bmc if there is IPXE is configured as boot iso payload. This is a additional task on the metal-bmc. metal-bmc already scans all servers on a regular basis to gather power metrics etc.
- If the boot iso is set to ipxe, the boot source override must be set to CDROM instead of PXE from network and a reboot must be triggered (migration to this approach, not when a machine is allocated).

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

If we don't plan on removing support for the old PXE boot with this change, it could make sense to track the boot mode of each machine in metal-api. The migration step could then be triggered and tracked by metal-api.

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

To me it sounds like this can just be part of the metal-hammer in the same way that it currently enforces UEFI boot mode on machine preparation? Wouldn't this be easier than introducing a interval check in the metal-bmc with bad visibility on how it worked?

@majst01 majst01 changed the title L3 only MEP-20 L3 only Jun 21, 2026
Comment thread community/04-Proposals/MEP20/README.md Outdated
Comment thread community/04-Proposals/MEP20/README.md Outdated

The placement therefore follows from the role given to `metal-boot`. If it only handles the lightweight control functions such as token issuance, `boot.ipxe`, DNS and NTP, placing a container on each leaf is acceptable. The small control steps stay well within the `ip2me` budget, and this also fits the suggestion from the design notes that `metal-boot` could be deployed on each switch with a shared anycast address for redundancy. The downside is that it exposes additional services on critical infrastructure, so the container still needs proper hardening. If `metal-boot` must instead act as a complete proxy that also carries bulk traffic, it should be placed on a fabric reachable host such as a management server. From there the proxied traffic is forwarded in hardware and never punted to a switch CPU, so CoPP does not apply.

The metal-image-cache-sync is currently placed on the management-servers. One of the stated goals is to remove the need for connections between the production infrastructure and the management infrastrutcure. Since placing or proxying the image cache on the switches is not viable, the image cache has to move to a different location. The image cache can either be hosted on a metal-stack provisioned machine, or on a server outside of metal-stack's scope.

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

When placing the image cache on a metal-stack provisioned machine, how would the bootstrap work here? Temporary server outside of metal-stack?

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Image cache is totally optional, so for the first machine pulling the image directly would only slow down installation

Comment thread community/04-Proposals/MEP20/README.md Outdated
- Enable automated IPv6 address acquisition via SLAAC (RFC 4862) driven by Router Advertisements (RFC 4861) instead of DHCP
- IPv6 in a dedicated Boot VRF instead of a Boot VLAN.

This approach requires that metal-apiserver, metal-hammer, ipxe and a new component running in the partition and connected to the boot-vrf (`metal-boot` for now) are IPv6 ready.

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Can we get rid of iPXE and its complexity completely?

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I dont think so. I am pretty sure we would end up with a much more complex solution without it.

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This approach requires that metal-apiserver [...] running in the partition [...]

I don't think this is correct? It says effectively that the metal-apiserver is not running the control plane anymore? Also it should say "are running".

@muhittink muhittink Oct 6, 2026 •

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

@majst01 Have you looked at UEFI HTTP Boot? It might simplify a few things, but it may require DHCPv6.

Pros

  • no custom iPXE build to maintain
  • no ISO / virtual media handling, boot URL set via Redfish (UefiHttp + HttpBootUri)
  • metal-hammer version still switchable at runtime if the URL points to metal-boot
  • no BMC license costs

Cons

  • EDK2 reference firmware requires DHCPv6 for the station address, SLAAC alone is not enough
  • vendor/firmware support for HttpBootUri and IPv6 HTTP Boot unknown
  • HTTPS certificate enrollment differs per vendor
  • not tested yet

@muhittink muhittink Oct 6, 2026 •

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

@majst01 I think HTTP boot has more of a future and simplifies a few things (more centralized approach, more control? modern). We also shouldn't forget that mounting the ISO requires buying licenses, Dell for example about ~300€ - per Server!

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I think we should at least give it a try.

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Oh, far better: the leaf relays DHCPv6 from the machine port into the boot VRF, to a DHCPv6 server reachable there (e.g. part of metal-boot). The firmware then gets everything it needs from that server:

  • its address (IA_NA)
  • the boot URL (option 59)
  • DNS (option 23)

. No Redfish call is needed to set a boot URI, no virtual media and no ISO. DHCP controls everything but you have to make sure, that the server has http_boot in the boot options.

@majst01
majst01 marked this pull request as ready for review July 3, 2026 08:34
@majst01
majst01 requested a review from a team as a code owner July 3, 2026 08:34
Comment thread community/04-Proposals/MEP20/README.md Outdated

@Gerrit91 Gerrit91 left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Some more suggestions to make it easier for first-time-readers. :)

Comment thread community/04-Proposals/MEP20/README.md Outdated
Comment thread community/04-Proposals/MEP20/README.md Outdated

But there are downsides with this approach. Most notable:

- 2 different network topologies (L2 and L3) in the dataplane

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

There is no explanation why this is an issue. Explain the issues that we see, e.g. locating DHCP and pixiecore is difficult, big broadcast domain, security concerns for decommissioned machines getting stuck in the PXE boot network.

Comment thread community/04-Proposals/MEP20/README.md Outdated
But there are downsides with this approach. Most notable:

- 2 different network topologies (L2 and L3) in the dataplane
- The switch port of a machine must be reconfigured between these two modes, once a machine changes from registered to installed and back.

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This also needs to explain why this an issue: A lot of SONiC builds seem to be buggy as it seems not a lot of people do this, when executing this reconfiguration in the wrong order some component crashes?

Comment thread community/04-Proposals/MEP20/README.md Outdated
- token based authentication against metal-apiserver of the metal-hammer
- Cache of metal-images accessible from metal-hammer inside a partition
- Preserve all existing metal-hammer discovery, hardware detection, and provisioning logic
- Secure network when machine reclaim goes wrong with ACLs on the switch which allows communication only to the control-plane and the `metal-boot`

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Use of the metal-boot, which was not introduced before? It should be explained what is specifically meant by that.

Comment thread community/04-Proposals/MEP20/README.md Outdated
Comment on lines +36 to +38
- Optional: make `metal-boot` a proxy to metal-apiserver to support IPv4 only control-plane deployments.
- Optional: make `metal-boot` itself act as NTP and DNS server for the metal-hammer. Together with the proxy to the control-plane this would allow to restrict external access to the `metal-boot` source IP (even IPv4 and IPv6).
- `metal-boot` is stateless and can be deployed multiple times and listens to the same anycast IPv6 address for redundancy.

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

To me it feels like these points are not part of the requirements but rather describe implementation ideas?

Comment thread community/04-Proposals/MEP20/README.md Outdated
- `metal-boot` is stateless and can be deployed multiple times and listens to the same anycast IPv6 address for redundancy.
- TODO more

## Out of scope

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Maybe non-goals?

Comment thread community/04-Proposals/MEP20/README.md Outdated

The main idea is based on three concepts.

- Boot from ISO feature of server bmc firmware which can be configured from remote via redfish.

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Can this be achieved through IPMI as well?

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

It might be possible through vendor extensions, but the IPMI promoters point to redfish as a more modern alternative.

https://www.intel.com/content/www/us/en/products/docs/servers/ipmi/ipmi-home.html

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

We should stay with redfish

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Then this will become an issue in the mini-lab where we do not have Redfish. :(

@majst01 majst01 Oct 6, 2026 •

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Then the mini-lab must be adopted as well, @Sven-Ric already made a proof with containerlab of the whole idea, so neither redfish nor ipmi is a hard requirement for the mini-lab

Comment thread community/04-Proposals/MEP20/README.md Outdated
- Enable automated IPv6 address acquisition via SLAAC (RFC 4862) driven by Router Advertisements (RFC 4861) instead of DHCP
- IPv6 in a dedicated Boot VRF instead of a Boot VLAN.

This approach requires that metal-apiserver, metal-hammer, ipxe and a new component running in the partition and connected to the boot-vrf (`metal-boot` for now) are IPv6 ready.

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This approach requires that metal-apiserver [...] running in the partition [...]

I don't think this is correct? It says effectively that the metal-apiserver is not running the control plane anymore? Also it should say "are running".

Comment thread community/04-Proposals/MEP20/README.md Outdated
The L3 only boot and registration process can be described as follows:

- Every server will be scanned on a regular basis from the metal-bmc if there is IPXE is configured as boot iso payload. This is a additional task on the metal-bmc. metal-bmc already scans all servers on a regular basis to gather power metrics etc.
- If the boot iso is set to ipxe, the boot source override must be set to CDROM instead of PXE from network and a reboot must be triggered (migration to this approach, not when a machine is allocated).

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

To me it sounds like this can just be part of the metal-hammer in the same way that it currently enforces UEFI boot mode on machine preparation? Wouldn't this be easier than introducing a interval check in the metal-bmc with bad visibility on how it worked?

Comment thread community/04-Proposals/MEP20/README.md Outdated
- Every server will be scanned on a regular basis from the metal-bmc if there is IPXE is configured as boot iso payload. This is a additional task on the metal-bmc. metal-bmc already scans all servers on a regular basis to gather power metrics etc.
- If the boot iso is set to ipxe, the boot source override must be set to CDROM instead of PXE from network and a reboot must be triggered (migration to this approach, not when a machine is allocated).
- Once the server is powered on, ipxe is booted from the CDROM presented from the firmware.
- The production interfaces will then get a IPv6 routable ip address from the switch which is configured to enable SLAAC and router advertisement. The configured routes must enable the machine to reach the metal-apiserver in the control plane and the `metal-boot` in the partition.

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Can we link documentation to "SLAAC"?

Sven-Ric and others added 2 commits October 6, 2026 09:22
Co-authored-by: Gerrit <Gerrit91@users.noreply.github.com>
@majst01 majst01 changed the title MEP-20 L3 only MEP-20 Machine Provisioning V2 Oct 7, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

area: documentation Affects the documentation area.

Projects

Status: In Progress

Development

Successfully merging this pull request may close these issues.

8 participants