Because they're not serious about their offerings. They are an ISP and like to provide/bundle very basic additional services (hosting, mail, etc. and now 'AI')
agree. I have a usecase for a bigger open-weight model hosted by an EU company and the selection isnt really great, hard to make the regulatory gods (and the devs) happy at the same time right now
All of these are lightyears behind US/China LLM offerings. None of them offer any model close to open-source SOTA. Runtime and reliability is a disaster and sure enough you have to pay a "sovereignty" mark up.
I live in the EU and want to live in a world where my children actually CAN read the Bible in class. However that’s not really possible anymore…
(Also no, modern science is (still) not incompatible with (the Christian) religion. I want my children to have both, as do I, or at least be taught both, so they can decide for themselves what they want when they are grown enough to do so.)
Never understood why hetzner gets brought up so positively here when there are dozens of EU competitors who are just as good and don't run 7 year old discounted servers
That's a pretty big part of why, honestly. How many of us are running server workloads that actually need the price/performance tradeoff of cutting edge server hardware?
We're not even talking cutting edge. If you browse hetzner's auction page you can see they're renting machines with almost decade old CPUs like the Xeon W-2145 or i7 6700. That's just a lot of old hardware..
I have several production services running right now on AWS t2 instances, which were released in 2016. For a service that isn't performance-critical, it's perfectly fine (and a significant cost savings) to run on decade-old hardware.
ECC vs non-ECC is also pretty irrelevant on a box with < 64GB of memory
Good to see more developments in this space. I quite like this service, which is a little further than Hetzner and has several models to choose from: https://www.infomaniak.com/en/hosting/ai-services
Important note on Infomaniak's offering - when they first launched it, I gave it a try but despite claims that it's OpenAI API compatible, even basic things like sending a base64 encoded image were broken. I reached out to their support, who replied with (verbatim): "Unfortunately, we do not provide support for our AI Service, as the solution is highly unmanaged and uses our API."
"highly unmanaged" didn't fill me with confidence, and "uses our API" is a very weird reason to give for not offering support. I wrote off ever using them for anything beyond email after that reply.
I highly agree. I once registered a domain through them that happened to be my last name but coincidentally also the name of a company. They restricted me from using the domain and never refunded me the yearly fee for it (citing trademark reasons, although under Swiss trademark law, someones legal name is not a trademark violation, especially as it was just my personal website and I wasn't working on a competing product/service). Only weeks later someone else successfully registered it without such issues.
Worse even, in my account panel I still had control over the e-mail service of the domain, so there was most definitely a pretty bad security critical bug there.
Depends on your mindset. If you are an Engineer that does not need hand-holding and you accept terse answers from support with the gratitude to the human on the other side, then go for it.
You will at least not need months of training and tough exams to be an expert in Hetzner Cloud. It is simple, but that's the point.
I mostly use Hetzner baremetal servers and not cloud, but the support there is in my experience very competent, but also VERY german and direct. I can imagine if youre neither german or super technical that this can be intimidating or considered unfriendly
Hetzner entering LLM inference is the cloud provider equivalent of your landlord also offering to cook you dinner. the margins on compute and the margins on food both rely on you not reading the invoice too carefully
Very soon we will be having reseller programs for inference, this will be just like web hosting reseller. After big players, small players will also start entering in this field.
I'm waiting for that day so that inference will be affordable just like web hosting. 200$ per month is in no way affordable by everyone.
This seems like a smart move, given their ability to host efficiently. I approve of efforts to make the cost of inference for smaller useful models slowly approach 'close to zero' and there are many good paths for getting there. It is useful for companies to get fast hosting for the class of smaller models they may end up hosting in house.
> The enable_thinking option is worth mentioning. Without it, the model can spend a surprising amount of the completion budget reasoning before it returns a visible answer.
Straight up the opposite, which the name makes abundantly clear, with the option it does reasoning, without it it doesn't...
llamacpp has a reasoning-budget and reasoning-message setting that can both be a global or per header setting. Using it, you can stop it's reasoning token count and insert a message at the point you stopped it.
This allows both the client and server to customize it. I typically use a message that tells it to compact the conversation and use subagents. I find the reasoning gets bloated when it's failed to do whatever task it's doing and often times it either has too little context (subagents) or its context is bloated (compact).
This works fairly well to get it to extend workable life up to ~1M on a local model.
Whats definitely missing: a solid (non Mistral) GDPR compliant coding plan / subscription. All offerings are either US or China based. With the newest open weights models this became really interesting imo.
The exceptional capital investment is pushing up the cost of compute hardware, which is direct contributor into the costs of compute providers, pushing up their prices to customers.
Crucial Memory, historically a retail RAM provider, closed to focus on AI compute sales, for example.
> “The AI-driven growth in the data center has led to a surge in demand for memory and storage. Micron has made the difficult decision to exit the Crucial consumer business in order to improve supply and support for our larger, strategic customers in faster-growing segments,” said Sumit Sadana, EVP and Chief Business Officer at Micron Technology. “Thanks to a passionate community of consumers, the Crucial brand has become synonymous with technical leadership, quality and reliability of leading-edge memory and storage products. We would like to thank our millions of customers, hundreds of partners and all of the Micron team members who have supported the Crucial journey for the last 29 years.”
> Secondly, the supply environment has changed permanently. AI infrastructure requires every single wafer with memory it can consume, something that has never happened with any industry megatrend previously. This means that every wafer Micron assigns to consumer parts is a wafer not going to a hyperscaler or enterprise contract. As a result, keeping a consumer line would directly limit Micron's ability to fulfill orders from its largest customers, which is a risk for profits and strategic relationships.
So! Until this AI capital investment exuberance ends, compute hardware manufacturing pipelines are committed for years into the future, we're all bidding up the limited supply of compute for rent/lease, which in turn is constrained by the limited supply of compute hardware available for purchase.
(at least in the scope of RAM, price fixing is also a potential component, but it'll take some time to determine how much of that is contributing to price inflation versus bona fide supply constraints: https://finance.yahoo.com/technology/articles/ram-crisis-one...)
Who lumps LLMs with crypto? The former has actual usage and purpose, the latter doesn't. Anyone who does must think any new technology is all the same.
Interesting. I could see them perhaps coming in competitive for models that fit into single cards? Less so playing in the big model serving league...climbing into that esp right now would be madness
what is the hallucination here? it is fast, free and fun to try. And I genuinely think that the hardware decision (if they get bigger gpus) decides if this will be a banger product or not?
> the article is based on a prompt, which i would much rather read
I would much rather read someone's carefully written and LLM-assissted article than shallow and lazy dismissals like this one. What, precisely, do you feel you've added to the conversation here?
What do you think this magical prompt will be, exactly, that perfectly represents what was almost certainly an iterative rather than spot process on the author's part?
There is no "West" anymore, at least not in the sense it was used in the last 50–70 years. There is now Europe/Canada/Australia on the one side, and the USA as an adversary on the other, and that won't change in the foreseeable time.
What we need is frontier open weight models in the EU, Canada, and Australia.
It won't. That's the point many people don't seem to get: The people who wanted everything that is happening now don't just go away. There is no guarantee they won't vote for a narcissistic grifter again, and thus any long-term decisions building on mutual trust and stability can't be made. How could you justify investing into a factory that relies on a 20 year ROI window when tariffs could be changing every day? Why would you want to use an AI model that could be restricted to US citizens or taken away entirely without notice?
Big, slow wheels have turned everywhere around the world. They won't turn back so easily.
Meh. No matter which side gets in power it’s almost certainly a narcissistic grifter. Who wasn’t over the last twenty years? The current one is a bit of a long term aberration, that doesn’t have to be the norm going forward.
In fact you could argue he was a response to some of the more extreme positions on the Democrat side.
I’m hoping things move more towards the center. But that may be only a fools hope.
You're not necessarily wrong for now, but that framing might be too narrow.
It is not beyond imagination that a government can limit availability to future frontier versions of the open weight models produced within their jurisdiction.
It pays to be a little paranoid and have a diversified portfolio of options to choose from. Anything that can have major geopolitical consequence is especially susceptible to such structural risk.
This has nothing to do with any particular government/country/jurisdiction except insofar as said country is a leading power and hence has immense leverage across a very wide range of dimensions. Said more plainly, it is safe to assume that a country with immense power is unlikely to be willing to cede or distribute such power to others.
If you're hung up on my choice of "the West", consider that I mean it geopolitically and economically, so broadly includes North America, Europe, Australia, New Zealand, Japan, and South Korea. I could have also said the Global South, but I think it is a fair statement that none of those countries have the means, but I'd be totally cool and happy if I'm proven wrong!
I read that data centers in Frankfurt use 40% of city energy, and seems that some fossil fuel based workarounds are being pushed. I hope we do better than the US in the energy and data center gold rush.
It would certainly be interesting to have a highly respected EU-native inference provider, if only to make the regulatory gods happy
OVH already provides this → https://www.ovhcloud.com/en/public-cloud/ai-endpoints/catalo...
It works great.
Scaleway does that already but competition is always good.
Does not scale in my pocket unfortunately.
IONOS (1&1) also offers inference hosted in europe/germany: https://cloud.ionos.com/managed/ai-model-hub
IONOS is established, but uh, 'respectable' is another question...
STACKIT also offers model serving.
No Qwen 3.6? Why?
Because they're not serious about their offerings. They are an ISP and like to provide/bundle very basic additional services (hosting, mail, etc. and now 'AI')
agree. I have a usecase for a bigger open-weight model hosted by an EU company and the selection isnt really great, hard to make the regulatory gods (and the devs) happy at the same time right now
Scaleway have two separate (one fully managed one a bit less) services for that:
https://www.scaleway.com/en/generative-apis/
https://www.scaleway.com/en/inference/
yea, but no prompt caching right? This makes it unusable for my usecase at least, the cost would be insane
They confirmed on LinkedIn that they are working on it.
Also, they have a feature request that has been getting many votes recently: https://feature-request.scaleway.com/posts/1251/prompt-cachi...
To all mentions of IONOS / StackIT / OVH:
All of these are lightyears behind US/China LLM offerings. None of them offer any model close to open-source SOTA. Runtime and reliability is a disaster and sure enough you have to pay a "sovereignty" mark up.
Freedom isn't free. Unless you want to live in a world were your children have to read the bible in class sacrifice has to be made.
Paying a premium for inferior service quality is not exactly freedom
I live in the EU and want to live in a world where my children actually CAN read the Bible in class. However that’s not really possible anymore…
(Also no, modern science is (still) not incompatible with (the Christian) religion. I want my children to have both, as do I, or at least be taught both, so they can decide for themselves what they want when they are grown enough to do so.)
Who is sacrificing what, really?
Never understood why hetzner gets brought up so positively here when there are dozens of EU competitors who are just as good and don't run 7 year old discounted servers
> don't run 7 year old discounted servers
That's a pretty big part of why, honestly. How many of us are running server workloads that actually need the price/performance tradeoff of cutting edge server hardware?
We're not even talking cutting edge. If you browse hetzner's auction page you can see they're renting machines with almost decade old CPUs like the Xeon W-2145 or i7 6700. That's just a lot of old hardware..
Well. Yeah. The auction page is the leftovers; if you want current stuff then you're in the wrong spot on their site
So? If it's cheaper and gets the job done who cares?
I mean if you run serious stuff on a non-ECC CPU from 2015, you'll probably have a bad time.
They're great for hobbyists but people on here keep bringing them up as the flagbearer of euro hosting
I have several production services running right now on AWS t2 instances, which were released in 2016. For a service that isn't performance-critical, it's perfectly fine (and a significant cost savings) to run on decade-old hardware.
ECC vs non-ECC is also pretty irrelevant on a box with < 64GB of memory
A t2 instance is virtualized, AWS keeps its servers for 5 years (see their 10k https://www.sec.gov/ix?doc=/Archives/edgar/data/1018724/0001... ), so your t2 does get moved around regularly to newer servers.
And no, ECC is relevant with less than 64 gigs of ram, there's a reason it's been a standard on servers since the 90s
serious stuff like what? majority of people don't work on scale you probably are thinking about.
Because nobody ever mentions them by name, leaving all of us in the dark.
Upcloud, Scaleway, Clever cloud, etc.
These are all solid choices, but significantly less performance/€
Understandably too, because hetzner is running things on consumer hardware vs the enterprise hardware the others use
Good to see more developments in this space. I quite like this service, which is a little further than Hetzner and has several models to choose from: https://www.infomaniak.com/en/hosting/ai-services
Important note on Infomaniak's offering - when they first launched it, I gave it a try but despite claims that it's OpenAI API compatible, even basic things like sending a base64 encoded image were broken. I reached out to their support, who replied with (verbatim): "Unfortunately, we do not provide support for our AI Service, as the solution is highly unmanaged and uses our API."
"highly unmanaged" didn't fill me with confidence, and "uses our API" is a very weird reason to give for not offering support. I wrote off ever using them for anything beyond email after that reply.
Interesting, didn’t know about them! Weird model selection though, no glm or deepseek?
lurus.ai and cortecs.ai are worth a try too!
Infomaniak is such a shitty company, I had to use them on a previous job I worked at and dealing with them was awful.
I highly agree. I once registered a domain through them that happened to be my last name but coincidentally also the name of a company. They restricted me from using the domain and never refunded me the yearly fee for it (citing trademark reasons, although under Swiss trademark law, someones legal name is not a trademark violation, especially as it was just my personal website and I wasn't working on a competing product/service). Only weeks later someone else successfully registered it without such issues.
Worse even, in my account panel I still had control over the e-mail service of the domain, so there was most definitely a pretty bad security critical bug there.
That makes it sound like dealing with Hetzner is easy. Is the case?
Depends on your mindset. If you are an Engineer that does not need hand-holding and you accept terse answers from support with the gratitude to the human on the other side, then go for it.
You will at least not need months of training and tough exams to be an expert in Hetzner Cloud. It is simple, but that's the point.
I mostly use Hetzner baremetal servers and not cloud, but the support there is in my experience very competent, but also VERY german and direct. I can imagine if youre neither german or super technical that this can be intimidating or considered unfriendly
Being slightly on the spectrum, I find Hetzner's "VERY german and direct" support perfect.
I only have a couple of VPS plans with them, but my experience is good.
Hetzner entering LLM inference is the cloud provider equivalent of your landlord also offering to cook you dinner. the margins on compute and the margins on food both rely on you not reading the invoice too carefully
Very soon we will be having reseller programs for inference, this will be just like web hosting reseller. After big players, small players will also start entering in this field.
I'm waiting for that day so that inference will be affordable just like web hosting. 200$ per month is in no way affordable by everyone.
This seems like a smart move, given their ability to host efficiently. I approve of efforts to make the cost of inference for smaller useful models slowly approach 'close to zero' and there are many good paths for getting there. It is useful for companies to get fast hosting for the class of smaller models they may end up hosting in house.
I really hope they dont stop at the small models though! The bigger ones that dont fit on a single GPU are more interesting I think
Problem is right now the biggest GPU boxes they have is single rtx pro 6000s.
> The enable_thinking option is worth mentioning. Without it, the model can spend a surprising amount of the completion budget reasoning before it returns a visible answer.
Straight up the opposite, which the name makes abundantly clear, with the option it does reasoning, without it it doesn't...
My writing wasnt clear here I think, without the option it defaults to reasoning enabled. With "without it" I meant without enable_thinking=False!
llamacpp has a reasoning-budget and reasoning-message setting that can both be a global or per header setting. Using it, you can stop it's reasoning token count and insert a message at the point you stopped it.
This allows both the client and server to customize it. I typically use a message that tells it to compact the conversation and use subagents. I find the reasoning gets bloated when it's failed to do whatever task it's doing and often times it either has too little context (subagents) or its context is bloated (compact).
This works fairly well to get it to extend workable life up to ~1M on a local model.
Whats definitely missing: a solid (non Mistral) GDPR compliant coding plan / subscription. All offerings are either US or China based. With the newest open weights models this became really interesting imo.
Excited to see the price of every other product they offer triple for no reason.
"no reason" like hardware prices going through the roof?
Hasn't that already happened?
yeah, but why would LLM inference now triple the price of the other products? Or am I misunderstanding what you're saying?
I am implying that they will subsidise the LLM inference product by increasing the prices of other products.
They wont, they are very public about never subsidising any products. I doubt they change that after 29 years now.
gestures broadly at ~$1T+ in AI capex spend
Not sure what you're getting at here. Hetzner has not spent 1 trillion dollars!
The exceptional capital investment is pushing up the cost of compute hardware, which is direct contributor into the costs of compute providers, pushing up their prices to customers.
Crucial Memory, historically a retail RAM provider, closed to focus on AI compute sales, for example.
Micron Announces Exit from Crucial Consumer Business - https://news.ycombinator.com/item?id=46137783 - December 2025 (392 comments)
> “The AI-driven growth in the data center has led to a surge in demand for memory and storage. Micron has made the difficult decision to exit the Crucial consumer business in order to improve supply and support for our larger, strategic customers in faster-growing segments,” said Sumit Sadana, EVP and Chief Business Officer at Micron Technology. “Thanks to a passionate community of consumers, the Crucial brand has become synonymous with technical leadership, quality and reliability of leading-edge memory and storage products. We would like to thank our millions of customers, hundreds of partners and all of the Micron team members who have supported the Crucial journey for the last 29 years.”
Micron is killing Crucial SSDs and memory in AI pivot to serve on AI companies - https://news.ycombinator.com/item?id=46152915 - December 2025 (0 comments)
> Secondly, the supply environment has changed permanently. AI infrastructure requires every single wafer with memory it can consume, something that has never happened with any industry megatrend previously. This means that every wafer Micron assigns to consumer parts is a wafer not going to a hyperscaler or enterprise contract. As a result, keeping a consumer line would directly limit Micron's ability to fulfill orders from its largest customers, which is a risk for profits and strategic relationships.
So! Until this AI capital investment exuberance ends, compute hardware manufacturing pipelines are committed for years into the future, we're all bidding up the limited supply of compute for rent/lease, which in turn is constrained by the limited supply of compute hardware available for purchase.
(at least in the scope of RAM, price fixing is also a potential component, but it'll take some time to determine how much of that is contributing to price inflation versus bona fide supply constraints: https://finance.yahoo.com/technology/articles/ram-crisis-one...)
This is interesting because I thought Hetzner was anti-crypto? LLMs aren't the same but they're often lumped in with crypto as "things no one wants."
Who lumps LLMs with crypto? The former has actual usage and purpose, the latter doesn't. Anyone who does must think any new technology is all the same.
A lot of people on Mastodon, for one.
Crypto DAU: 10’s of k LLMaaS DAU: 100’s of k
Interesting. I could see them perhaps coming in competitive for models that fit into single cards? Less so playing in the big model serving league...climbing into that esp right now would be madness
Why do you think this would be madness? It seems like, at least in the EU, there is barely competition for the big open-weight models?
I mean yah... host glm and kimi and I am game.
Potentially interesting article ruined by AI slop hallucinations like
> For now, the API is fast, free, and fun to try. The next hardware announcement will tell us much more than another small model would.
what is the hallucination here? it is fast, free and fun to try. And I genuinely think that the hardware decision (if they get bigger gpus) decides if this will be a banger product or not?
nobody writes like that. the article is based on a prompt, which i would much rather read, instead of the ironed out version that the llm produced
fine, i agree that this doesnt sound like me. But the content is correct and not a hallucination!
so did you write this post with the assistance of LLMs?
Could you annoy someone else please.
> the article is based on a prompt, which i would much rather read
I would much rather read someone's carefully written and LLM-assissted article than shallow and lazy dismissals like this one. What, precisely, do you feel you've added to the conversation here?
Why not publish the prompt anyway and let readers decide if they want to run it through an LLM (maybe next year's LLM, which will probably be better)?
What do you think this magical prompt will be, exactly, that perfectly represents what was almost certainly an iterative rather than spot process on the author's part?
also slightly disrespectful to think that this was a single prompt, but ok:D
Exactly that
Hetzner is very efficient hosting servers
Will this be the new division of labor?
Americans - best proprietary models
Chinese - best open weight models
Europeans - best / most efficient inference service
We need frontier open weight models from the West.
Maybe Nvidia's Nemotron can get there.
https://research.nvidia.com/labs/nemotron/Nemotron-3-Ultra/
There is no "West" anymore, at least not in the sense it was used in the last 50–70 years. There is now Europe/Canada/Australia on the one side, and the USA as an adversary on the other, and that won't change in the foreseeable time.
What we need is frontier open weight models in the EU, Canada, and Australia.
Give it two years more then it changes again.
It won't. That's the point many people don't seem to get: The people who wanted everything that is happening now don't just go away. There is no guarantee they won't vote for a narcissistic grifter again, and thus any long-term decisions building on mutual trust and stability can't be made. How could you justify investing into a factory that relies on a 20 year ROI window when tariffs could be changing every day? Why would you want to use an AI model that could be restricted to US citizens or taken away entirely without notice?
Big, slow wheels have turned everywhere around the world. They won't turn back so easily.
Meh. No matter which side gets in power it’s almost certainly a narcissistic grifter. Who wasn’t over the last twenty years? The current one is a bit of a long term aberration, that doesn’t have to be the norm going forward.
In fact you could argue he was a response to some of the more extreme positions on the Democrat side.
I’m hoping things move more towards the center. But that may be only a fools hope.
why?
Reasons include, but are not limited to:
a) broad competition is good
b) jurisdiction diversification (you don't want to be dependent on the regulatory winds of a single jurisdiction)
...but you're not dependent on anyone?
If I'm running a Kimi model on my local cluster, the CCP can't shut it off.
It's not like the Finnish government has veto power over Linux.
You're not necessarily wrong for now, but that framing might be too narrow.
It is not beyond imagination that a government can limit availability to future frontier versions of the open weight models produced within their jurisdiction.
It pays to be a little paranoid and have a diversified portfolio of options to choose from. Anything that can have major geopolitical consequence is especially susceptible to such structural risk.
This has nothing to do with any particular government/country/jurisdiction except insofar as said country is a leading power and hence has immense leverage across a very wide range of dimensions. Said more plainly, it is safe to assume that a country with immense power is unlikely to be willing to cede or distribute such power to others.
If you're hung up on my choice of "the West", consider that I mean it geopolitically and economically, so broadly includes North America, Europe, Australia, New Zealand, Japan, and South Korea. I could have also said the Global South, but I think it is a fair statement that none of those countries have the means, but I'd be totally cool and happy if I'm proven wrong!
I read that data centers in Frankfurt use 40% of city energy, and seems that some fossil fuel based workarounds are being pushed. I hope we do better than the US in the energy and data center gold rush.