| ▲ | The Hetzner Cloud network stack – history and technical overview(hetzner.com) | |
| 25 points by cnkk an hour ago | 3 comments | ||
| ▲ | tasuki 33 minutes ago | parent | next [-] | |
Oh, my Hetzner VM just had a 13 hour outage from yesterday evening to today morning. They say there was some incident with their DNS servers or something. I don't recall that ever happening during my 12 years at Digital Ocean. But yes, the Hetzner VM is both cheaper and beefier than the DO one was. | ||
| ▲ | Lucasoato 31 minutes ago | parent | prev | next [-] | |
> How we roll We triple the price from one day to another and don’t answer support tickets on weekends. | ||
| ▲ | spwa4 15 minutes ago | parent | prev [-] | |
This is very, very suboptimal. I like the VXLAN support a lot, and of course this will have a lot of features while requiring minimal actual network knowledge (good luck getting everything to work correctly when combined, but you can certainly configure them ... and so you're making it the customer's problem) Oh and this is going to cause out-of-order delivery and slow down user applications by a lot (because of channel bonding that looks like it's just left at the defaults), it is horribly inefficient. People don't put the host network either next to the VMs or on a separate network card for nothing. In their case a packet walk would show: 1) guest userspace -> guest kernel 2) guest kernel virtio_net (hopefully) -> host kernel virtio_net 3) host kernel virtio_net -> host kernel 4) host kernel -> OVS vswitch data path 5) OVS vswitch data path -> host kernel bridge port 6) host kernel bridge port -> host kernel switch/networking stack 7) host kernel switch/networking stack -> host outgoing bonding virtual port 8) host outgoing bonding virtual port -> physical port Each of these steps requires at the very least a memory allocation, inserting a step on a work queue, waiting on that work queue. Also very likely 5 of these steps require a context switch (at minimum waiting for the process scheduler to reschedule a task, on linux still usually requires 1ms minimum wait, more under load). So this inserts 5ms of latency minimum (and under load it's going to balloon) before the packet even arrives on the ring buffer of the first physical network card. And, as stated before, it's also going to cause out-of-order delivery. And that's, of course, before application developers put a multiplication factor before this cost by using something like nginx or even multiple layers of nginx. I get the flexibility gain, and of course application developers get to do whatever they want, but ... why? This is also eating a lot of processing power of the machines (and everything that comes with that, power use, even co2). And a further issue with that is that this is kernel networking path, which doesn't show in top, and doesn't show in most kernel metrics, you have to really know what you're looking for. And the cost that is incurred on the application side by due to the delay and the out of order packets doesn't show up anywhere except on the customer's bill, but good luck finding that it's wasted capacity. If you do something like ML training from an NFS or S3 mount or NVMEoE or RoCE you will clearly notice the flaws in this design. The gains you can make there approach the gains you can make by switching from ethernet to fibre channel/infiniband. What is possible with a great design: 0/1 context switch from guest userspace to network card ring buffer (zero context switches is possible by either using io/uring in guest userspace, or by binding the physical hardware to the guest VM and then into the application). Ping times to same-building VMs that consistently stay below 0.1ms, even with machine loads over 98%. Wish someone would pay me to do that. Zero context switches while maintaining all features is possible. A lot of work, but possible. I don't believe anyone has yet done it, but it is possible. And, please, move the linux host into the OVS ... just that little step will save about half the cost and it only requires being a bit more careful in operations (or having actual OOB, like serial or an extra hardware network card, the cheapest thing you're throwing away will easily do for that purpose) And yes, I worked on the networking stack of one of the hyperscalers. They are at 2 to 3 context switches, more if you use any kind of tunneling (it's a crime that VXLAN is not supported ...). A lot better than this design, but not really close to perfectly optimal. It would be great to work on getting that closer to optimal in a large hoster. VPP + DPDK right into guest VMs. Sigh. Back to AI networking. | ||