Repository navigation
Updates on MEP-17. - #361
Conversation
✅ Deploy Preview for metal-stack-io ready!
To edit notification comments on pull requests, go to your Netlify project configuration. |
iljarotar
left a comment
There was a problem hiding this comment.
Mostly stylistic remarks. Content-wise very good.
A minor nit: I prefer markdown files where each sentence starts in a new line. It's easier to review.
One question regarding management switches. Do we also want those to be "reconciled" via API? If so, maybe also add them where you mention spines and exits.
|
|
||
| With the current state we see the following room for improvements: | ||
|
|
||
| - `Switch` resources should be applied declaratively at the API through deployment by the Admin API. Instead of registering a switch, the switch queries the API on a specific switch entity provided as a startup configuration. This kind of "reverses" the registration procedure. The metal-core establishes a stream connection to the API for a given switch ID and then reconciles the given switch desired state definition. This approach has the following advantages: |
There was a problem hiding this comment.
... and then reconciles the given switch desired state definition.
I don't understand the grammar of this sentence.
There was a problem hiding this comment.
I reordered this sentence a bit, maybe you can check again if it is better now?
| - Figure out what has to go into the port configuration in order to achieve FRR configuration scenarios we have on spines and exit switches | ||
| - Check what special scenarios we have for static port configuration (support tenant VRF statically on a specific port, maybe black hole configuration instead of default PXE VRF?) | ||
| - Check if we can really drop machine connections: Do we really want to always evaluate the neighbors all the time? | ||
| - Elaborate the switch port change to on the leaf switch into the tenant VRF: The metal-hammer currently sends the `InstallationSucceeded` message to the server, but when the switch gets notified immediately through stream, it might break the network connection prematurely such that the metal-hammer does not retrieve the response in time causing a crash. |
There was a problem hiding this comment.
Elaborate the switch port change to on the leaf switch into the tenant VRF
I don't understand this sentence.
There was a problem hiding this comment.
Rephrased, is it better now?
|
Thanks for the review. We now have every sentence in a single line. I am not sure if we also include the management switches into the "reconciliation"? I think you need to tell me if this would be possible. I changed the MEP now that we do it for the management network, too. |
|
This approach could be extended to manage the lifecycle of a switch, mirroring how we already handle machines:
|
| - Check what special scenarios we have for static port configuration (support tenant VRF statically on a specific port, maybe black hole configuration instead of default PXE VRF?) | ||
| - Check if we can really drop machine connections: Do we really want to always evaluate the neighbors all the time? | ||
| - Elaborate the switch port change to on the leaf switch into the tenant VRF: The metal-hammer currently sends the `InstallationSucceeded` message to the server, but when the switch gets notified immediately through stream, it might break the network connection prematurely such that the metal-hammer does not retrieve the response in time causing a crash. | ||
| - Elaborate how the port reconfiguration on machine allocation can be orchestrated: The metal-hammer currently sends the `InstallationSucceeded` message to the server, but when the switch gets notified immediately through stream, it breaks the network connection prematurely such that the metal-hammer does not retrieve the response in time causing a crash. |
There was a problem hiding this comment.
Actually, the behavior you describe here is not something we know for sure. I think that, in general, we need to find a clean solution for the timing of the port configuration. It's not strictly related to MEP-17.
Interesting thought. Do you think we should extend MEP-17 to include such lifecycle management? If so, I'd like to discuss each step in detail. For example, I'm not sure if Waiting, Registering and Allocation should be separate steps. |
Also this contradicts a bit with the idea of the declarative API because the role is already assigned by the admin. |
|
Maybe someone else can take over this PR or we can merge it in this state and then someone re-opens another one containing additional topics like "provision without Ansible", "state machine"? |
Description
After our MEP-4 journey, we realized we can even do more with MEP-17.
Used AI-Tools ✨