Re: [LSF/MM/BPF TOPIC] BoF VM live migration over C XL memory
Dragan Stancevic <[email protected]> Mon, 1 May 2023 18:49:12 -0500
| Newsgroups | dev.linux.lists.nil-migration,org.kernel.vger.linux-cxl |
|---|---|
| Message-ID | <[email protected]> |
Hi Dave- sorry, looks like I've missed your email On 4/11/23 13:00, Dave Hansen wrote: > On 4/7/23 14:05, Dragan Stancevic wrote: >> I'd be interested in doing a small BoF session with some slides and get >> into a discussion/brainstorming with other people that deal with VM/LM >> cloud loads. Among other things to discuss would be page migrations over >> switched CXL memory, shared in-memory ABI to allow VM hand-off between >> hypervisors, etc... > > How would 'struct page' or other kernel metadata be handled? > > I assume you'd want a really big CXL memory device with as many hosts > connected to it as is feasible. But, in order to hand the memory off > from one host to another, both would need to have metadata for it at > _some_ point. To be honest, I have not been thinking of this in terms of a "star" connection topology. Where say each host in a rack connects to the same memory device, I think I'd get bottle-necked on a singular device. Evac of a few hypervisors simultaneously might get a bit dicey. I've been thinking of it more in terms of multiple memory devices per rack, connected to various hypervisors to form a hypervisor traversal graph[1]. For example in this graph, a VM would migrate across a single hop, or a few hops to reach it's destination hypervisor. And for the lack of better word, this would be your "migration namespace" to migrate the VM across the rack. The critical connections in the graph are hostfoo04 and hostfoo09, and those you'd use if you want to pop the VM into a different "migration namespace", for example a different rack or maybe even a pod. Of course, this is quite a ways out since there are no CXL 3.0 devices yet. As a first step I would like to get to a point where I can emulate this with qemu and just prototype various approaches, but starting with a single emulated memory device and two hosts. > So, do all hosts have metadata for the whole CXL memory device all the > time? Or, would they create the metadata (hotplug) when a VM is > migrated in and destroy it (hot unplug) when a VM is migrated out? To be honest I have not thought about hot plugging, but might be something for me to keep in mind and ponder about it. And if you have additional thoughts on this I'd love to hear them. What I was thinking, and this may or may not be possible, or may be possible only to a certain extent, but my preference would be to keep as much of the metadata as possible on the memory device itself and have the hypervisors cooperate through some kind of ownership mechanism. > That gets back to the granularity question discussed elsewhere in the > thread. How would the metadata allocation granularity interact with the > page allocation granularity? How would fragmentation be avoided so that > hosts don't eat up all their RAM with unused metadata? Yeah, this is something I am still running through my head. Even if we have this "ownership-cooperation", is this based on pages, what happens to the sub-page allocations, do we move them through the buckets or do we attach ownership to sub-page allocations too. In my ideal world, you'd have two hypervisors cooperate over this memory as transparently as CPUs in a single system collaborating across NUMA nodes. A lot to think about, many problems to solve and a lot of work to do. I don't have all the answers yet, but value all input & help [1]. https://nil-migration.org/VM-Graph.png -- Peace can only come as a natural consequence of universal enlightenment -Dr. Nikola Tesla