26 comments

  • publlus_enigma 9 hours ago ago

    An interesting read.

    I've just started with Hetzner, and their storage box product to try cloud backups via Restic.

    Whilst very competitively priced, the 3MB/s transfer rates I am currently achieving on a 1000/400 fibre connection means I am unlikely to ever use the 10TB I have supposedly been allotted.

    • jerrythegerbil 7 hours ago ago

      Hetzner servers aren’t billed against transfer quotas to storage box, and get a direct, faster link.

      It sounds like you’re connecting through the public interface, for which many ISPs in the US rate limit the Hetzner storage boxes. In my case, around 3mbps.

      Using a VPN from your home (unintuitively) increases your throughput, alternatively use concurrent webdav (HTTPS, typically not rate limited) without VPN.

      Said plainly, the symptom you’re seeing is the same reason I could likely guess your ISP, and has nothing to do with Hetzner.

    • maeln 9 hours ago ago

      I also use the storage box and easily have 30Mo/s bandwidth (and 130Mo/s+ from a hertzner server), so that's strange that you seemed capped at 3Mo/s. Are you not in Europe ? You should probably contact the support to report it, because that almost definitely now how they intend their product to be.

      • chrisweekly 8 hours ago ago

        Mo/s means megabytes per second (the French abbreviation for mégaoctets par seconde). It measures data-transfer speed and is equivalent to MB/s (capital B) in English.

    • concretebush 5 hours ago ago

      I had a similar issue with an older version of Restic + Hetzner, update to latest stable, if you haven’t yet. It improved my speed literally 10x, something to do with buffering over ssh.

  • bilater 4 hours ago ago

    Love Hezner amd how cheap they are for the services they provide. I'm building a "roll your own Vercel" on top of it. It's gonna be open source full self hosted or a semi-managed (bring your own Hetzner server and pay for some sugar on top) service. Appreciate any feedback!

    https://shiptiffin.com

  • Lucasoato 10 hours ago ago

    > How we roll

    We triple the price from one day to another and don’t answer support tickets on weekends.

    • segmondy 8 hours ago ago

      find an alternative, no one is forcing you to use them. i never understand why people complain about pricing if they have a better choice.

    • jgalt212 9 hours ago ago

      We are all under the same pricing pressures Hetzner is under. I, too, am disappointed by their level of customer support, but when their VMs WERE 80% cheaper than the other hyperscalers I was basically OK with that.

  • alberth 2 hours ago ago

    Really wish Hetzner would bring dedicated servers to the US.

    They are amazing at data center ops.

  • sandeepkd 6 hours ago ago

    Its interesting to learn them trying to implement the open vSwitch in python. I worked on implementing the open vSwitch in Java some time in 2014ish too. Open Stack had a lot of push back then. It was fun to gather the requirements by reverse engineering the original product based on the network traffic for something which is a middleware and has a state of its own

  • throw0101a 9 hours ago ago

    This appears to be a BGP EVPN (VXLAN)-like architecture, with the VTEPs pushed down to the hypervisors instead of living on the ToR switches. (Also OpenStack's network architecture.)

  • rohityadavcloud 7 hours ago ago

    bgp + vxlan + evpn is pretty standard in datacenters now. I’m surprised they’re running udhcpd in a ns, instead of cloud-init data iso directly. While it wasnt stated but I think using nft/based firewall is pretty common. For kvm anyone knows if they use qemu or something else for vmm, directly or via libvirtd.

  • tasuki 10 hours ago ago

    Oh, my Hetzner VM just had a 13 hour outage from yesterday evening to today morning. They say there was some incident with their DNS servers or something.

    I don't recall that ever happening during my 12 years at Digital Ocean. But yes, the Hetzner VM is both cheaper and beefier than the DO one was.

  • marcosscriven 6 hours ago ago

    TIL Krebs also means crab, I always thought it was only “cancer”, so “flow cancer” seemed like a strange name.

  • nwmcsween 7 hours ago ago

    This sounds like openstack?

  • spwa4 10 hours ago ago

    This is very, very suboptimal. I like the VXLAN support a lot, and of course this will have a lot of features while requiring minimal actual network knowledge (good luck getting everything to work correctly when combined, but you can certainly configure them ... and so you're making it the customer's problem)

    Oh and this is going to cause out-of-order delivery and slow down user applications by a lot (because of channel bonding that looks like it's just left at the defaults), it is horribly inefficient. People don't put the host network either next to the VMs or on a separate network card for nothing.

    In their case a packet walk would show:

    1) guest userspace -> guest kernel

    2) guest kernel virtio_net (hopefully) -> host kernel virtio_net

    3) host kernel virtio_net -> host kernel

    4) host kernel -> OVS vswitch data path

    5) OVS vswitch data path -> host kernel bridge port

    6) host kernel bridge port -> host kernel switch/networking stack

    7) host kernel switch/networking stack -> host outgoing bonding virtual port

    8) host outgoing bonding virtual port -> physical port

    Each of these steps requires at the very least a memory allocation, inserting a step on a work queue, waiting on that work queue. Also very likely 5 of these steps require a context switch (at minimum waiting for the process scheduler to reschedule a task, on linux still usually requires 1ms minimum wait, more under load). So this inserts 5ms of latency minimum (and under load it's going to balloon) before the packet even arrives on the ring buffer of the first physical network card. And, as stated before, it's also going to cause out-of-order delivery.

    And that's, of course, before application developers put a multiplication factor before this cost by using something like nginx or even multiple layers of nginx. I get the flexibility gain, and of course application developers get to do whatever they want, but ... why?

    This is also eating a lot of processing power of the machines (and everything that comes with that, power use, even co2). And a further issue with that is that this is kernel networking path, which doesn't show in top, and doesn't show in most kernel metrics, you have to really know what you're looking for. And the cost that is incurred on the application side by due to the delay and the out of order packets doesn't show up anywhere except on the customer's bill, but good luck finding that it's wasted capacity.

    If you do something like ML training from an NFS or S3 mount or NVMEoE or RoCE you will clearly notice the flaws in this design. The gains you can make there approach the gains you can make by switching from ethernet to fibre channel/infiniband.

    What is possible with a great design: 0/1 context switch from guest userspace to network card ring buffer (zero context switches is possible by either using io/uring in guest userspace, or by binding the physical hardware to the guest VM and then into the application). Ping times to same-building VMs that consistently stay below 0.1ms, even with machine loads over 98%. Wish someone would pay me to do that.

    Zero context switches while maintaining all features is possible. A lot of work, but possible. I don't believe anyone has yet done it, but it is possible.

    And, please, move the linux host into the OVS ... just that little step will save about half the cost and it only requires being a bit more careful in operations (or having actual OOB, like serial or an extra hardware network card, the cheapest thing you're throwing away will easily do for that purpose)

    And yes, I worked on the networking stack of one of the hyperscalers. They are at 2 to 3 context switches, more if you use any kind of tunneling (it's a crime that VXLAN is not supported ...). A lot better than this design, but not really close to perfectly optimal. It would be great to work on getting that closer to optimal in a large hoster. VPP + DPDK right into guest VMs. Sigh. Back to AI networking.

    • toast0 9 hours ago ago

      > Oh and this is going to cause out-of-order delivery and slow down user applications by a lot (because of channel bonding that looks like it's just left at the defaults),

      Doesn't lag default to hashing by something? I expect it to use the rss hash, or src/dest ip and hopefully port.

      On the number of layers, I totally agree. You can't fix the problems that come from having too many layers with more layers. And, you'll never get those delays back. If this really adds 5 ms (in each direction!), that's wild, but it would be clearly visible in pings... 10 ms round trip to get onto the network is like moving your servers 300 miles away from everyone.

      > Wish someone would pay me to do that.

      I feel you. Everyonce in a while, I get to do some really neat networking stuff, but I don't know how to make that my whole job.

    • tryauuum 10 hours ago ago

      why are you mentioning roce and ml training? These are cpu-only machines with 10 Gbit uplinks shared between every virtual machine. I'm not super familiar with the helmet offering, but last time I checked their GPU servers were baremetal

      Maybe calling the setup "suboptimal" is incorrect and it is optimal given their fleet and customers

    • thelastgallon 9 hours ago ago

      What do you think of Oxide.computer networking?

    • kay_o 9 hours ago ago

      The people using hetzner cloud are probably hosting CRUD apps

    • wmf 8 hours ago ago

      If you want Nitro you know where to find it... and what it costs.

    • jgalt212 9 hours ago ago

      This is like a nice advert for SQLite. Or for putting your app and database on the same server.

  • HackerThemAll 8 hours ago ago

    I once tried to become a Hetzner customer, but they wanted so much data from me, including photos of my ID, so I said no. No other compute provider requested that much, including all 3 big ones and OVH. Either I was lucky before KYC became a thing, and Google, Microsoft and Amazon inferred my identity from my payment cards, or I don't know, Hetzner is just too paranoid. I lean towards the latter. Yes, recently I created a new account in an another European compute provider, and they didn't want my ID.

    So no, thanks Hetzner.

    • sampullman 7 hours ago ago

      It's not so true anymore, but when they were cheaper they probably attracted a ton of bad actors (spammers, phishing sites, etc). I assume the KYC requirements were a response to that.

  • tryauuum 10 hours ago ago

    a river crab? Is this a joke related to the chinese great firewall?