For inference specifically 2GB RAM is the real bottleneck, not disk. You can offload to NVMe but 195MB/s means 7B quants load in like... 30 seconds? Then inference is CPU-bound anyway. I ran llama.cpp on Hetzner CPX11 once, 2 tok/s on Q4_0. Not usable.
My benchmark numbers seem fake. What did I break?
This is why I rent bare metal. No tmpfs surprises, no noisy neighbors, no "cloud" nonsense. €34 for AX42 at Hetzner auction, real NVMe, real cores, no governor games. You people benchmark toys.
AX42 is €34 if you win auction, not guaranteed. And you still pay for power, traffic, IP. For BGP playbox sure. For student learning? €4 VPS teaches same lessons about tmpfs and governors apparently.
Following this. Was about to YABS my Contabo VPS and now I know to check mounts first.
Update: I moved my project build to /tmp on purpose and compile time dropped from 4min to 45sec. So the tmpfs thing is actually useful once you know it exists! Thanks again everyone, learned a lot.
This. Tmpfs as build dir is standard practice, just don't confuse it with persistent storage. I do the same for Rust builds on my 4GB Hetzner, cargo target dir lives in /tmp.
Wait so is the CPU freq thing also normal? My VPS shows 3.4GHz in htop but YABS says 1200MHz, which is right?