| ▲ | Gabrys1 4 hours ago |
| I started wondering if there could be swap-aware GC, like first make the required page swapped in (not that there's any obvious API for that...) and only then pause the world? |
|
| ▲ | debugnik 3 hours ago | parent | next [-] |
| Aren't madvise and mincore the APIs you want? |
| |
| ▲ | masklinn 2 hours ago | parent [-] | | Probably more mlock (keep memory area in RAM), mincore just tells you what is or is not in RAM, and madvise is about access patterns. You could MADV_WILLNEED the GC metadata when you start the GC process hoping they’ll have been paged in by the time you STW, but assuming that area is not massive it’s probably a better idea to just prevent it being paged out. | | |
| ▲ | debugnik an hour ago | parent [-] | | I was assuming they meant for the heap. Not even executable memory is locked though, so if you mlock too much memory the kernel will page out your executable instead. Something has to give, so it's better to schedule it. On a lower level, the OCaml and cpython GCs use a prefetch buffer during marking to schedule around the cache. |
|
|
|
| ▲ | serbuvlad 3 hours ago | parent | prev | next [-] |
| One can just add APIs to the Linux kernel, and this would be a pretty straightforward one. (though it might be a little more difficult to avoid a syscall here) I'm sure an agent can work on this and get some numbers with a day's worth of tokens. |
| |
| ▲ | loeg 3 hours ago | parent | next [-] | | The API already exists: mlock. | | |
| ▲ | keybored 3 hours ago | parent [-] | | I think we are not realizing the paradigm shift here. A coding agent can implement this with a modest little day’s worth of tokens. This means that no one needs to read or know the APIs any more. We can just add them to the kernel as we need them. Then when we forget that we needed them we can just rediscover the API idea later and spend a day’s worth of tokens. (But let’s be real here. By then it will probably be just 1/3 worth of tokens with all the model improvements as well as the ample training material.) | | |
| ▲ | Mawr 2 hours ago | parent [-] | | We don't even need a kernel anymore. An agent can write the relevant parts of the OS on the fly, as we need them. |
|
| |
| ▲ | jstarks 3 hours ago | parent | prev [-] | | You don’t need an api. Just touch the page. |
|
|
| ▲ | quotemstr 3 hours ago | parent | prev [-] |
| You could do lots of interesting things with sufficiently deep inter-layer integration. For example, why not swap out not by LRU page but by dense node clusters on the heap graph, maintaining in-memory summaries of inbound and outbound edges for liveness? If you do this, you don't have to swap the cluster in to do a GC involving it. If the whole cluster becomes unreachable, you wouldn't even have to swap it back in to get rid of it: you'd just drop the swap reference and deem the swap space free. I don't see anything this deeply integrated happening near-term, but it's fun to think about. |