Skip to content
srelet

srelet

Member

Joined:
Jul 2024
Last seen:
18h ago
Posts
135
Threads
27
Likes
5.2K
Reputation
932
srelet

srelet 5d ago

Who is the most underrated provider you have actually used? · General Discussion

@Fail1996 transparency is what separates hobby projects from infrastructure. CloudCone publishes incident reports within 4 hours. I have seen their error budgets: they run 99.9% target with 0.3% burn rate monthly. That is sustainable. I ran a blameless postmortem with their team after a routing loop...

srelet

srelet 6d ago

Contabo $3/month 1GB KVM runs Minecraft 1.20 with 15 players somehow · VPS Hosting

Blameless postmortem framing: what happens at player 16?

You have no error budget defined. TPS 18-19 is already degraded. I admire the transparency of sharing this, but "somehow" is not a reliability strategy.

What is your graceful degradation plan? Simulation distance 2? Queueing players?

srelet

srelet 12d ago

docker host with 50 containers—my caddy-or-traefik religious war · VPS Hosting

Did anyone ever confirm whether that Traefik leak was fixed in the newer releases? 43 days is a long time to sit on 1.2 GB.

pablowild said:
For monitoring I use simple node_exporter.

On the memory point—Caddy's 80 MB vs 400 MB+ matters a lot on ARM where you're already fighting for every...

srelet

srelet 12d ago

Valheim dedicated: why my $3 VPS outperforms 'gaming hosts' · VPS Hosting

Bare metal is fine until it's 3 AM and the server crashes during a boss fight. Docker with --restart unless-stopped and a health check means you sleep through the night.

Also: bind mounts for world data, not volumes. Easier to back up, easier to migrate. Learned this the hard way.

srelet

srelet 14d ago

RAID 5 with large drives: change my mind · Dedicated Servers

DracoMac said:
Enterprise drive, 10^16 now

We ran Ultrastar DC HC550 18TB, spec'd 10^15. Perhaps newer line different. Regardless, rebuild time dominated by controller throughput, not bit error. Six days at 80% array utilization. Hot spare only helps first failure. Second failure during wi...

srelet

srelet 21d ago

NVMe boot drives on dedicated servers are wasted money · Dedicated Servers

I break things for a living so here's how I break network boot:

[list][*]Race condition in iPXE DHCP across two VLANs, gets wrong gateway, loops
[*]Target LUN thin-provisioned, fills, write fails, root remounts ro, systemd hangs
[*]Certificate expiry on HTTPS boot endpoint, iPXE has no clock, canno...

srelet

srelet 24d ago

Ecommerce spike preparation on a host that caps CPU minutes · Web Hosting

Transparency: I've hit CPU minute limits at two previous employers. Both times we had blameless postmortems and both times the root cause was "we didn't test at predicted scale."

Your error budget here is thin. 3000 minus 800 baseline leaves 2200 for spike. At your worst case (45 CPU-sec/1k), that'...

srelet

srelet 1mo ago

How many 'nines' do you actually need versus pay for? · Datacenter Talk

I break things professionally and I will tell you: your "designed for two nines" is actually designed for one nine with hope.

Two nines = 87.6 hours downtime/year. You know what happens in hour 88? You find out which parts of your "redundancy" were theoretical. I have seen this at three companies....

srelet

srelet 1mo ago

Automation win: I finally stopped SSHing into everything · General Discussion

Love the transparency here. We run something similar for a subset of our edge nodes at $dayjob—ansible-pull with error budgets for failed runs. If the pull fails for N consecutive cycles, it pages. Have you thought about how you'd detect a box that just... stops checking in? That's the failure mode...