garykwh
Member
OP
Inference on a Budget
- Joined:
- Jun 2024
- Posts:
- 173
- From:
- San Jose, US
So my budget host at OVHcloud disabled SNMP on their "managed" dedis. No IPMI either. Ngl this is pretty common with cheap providers but I need to know when my inference nodes go dark without relying on their status page that updates slower than molasses.
Currently running a 4x A100 setup for LLM fine-tuning and can't afford blind spots. VRAM's at 80GB per card so we're talking real money when jobs fail silently.
I've got ping monitoring via UptimeRobot but that's baby stuff. Anyone built something smarter?
CUDA cores are my love language
SwapJones
Member
OOM killer's friend
- Joined:
- Nov 2024
- Posts:
- 17
- From:
- Hanoi, Vietnam
Maybe you know. Their status page is just a CDN cache you know. Maybe you're polling cloudflare's edge node and not even their backend. You know. Maybe your "faster" alert is just faster at being wrong.
I've seen this before you know. Maybe check the page headers. Maybe look for x-cache-status. You know. Maybe not though.
No periods for me. Only the void...
RAM is just fast disk, right?
garykwh
Member
OP
Inference on a Budget
- Joined:
- Jun 2024
- Posts:
- 173
- From:
- San Jose, US
Prometheus node_exporter over HTTPS
This is interesting but OVHcloud's "managed" layer blocks inbound connections to random ports too. I can get 443 and 22, that's the list. Reverse proxy on 443 could work but now I'm burning a web server slot for metrics.
How do you handle auth on that setup? Basic auth over TLS or something heavier?
CUDA cores are my love language
ana_mad
Member
- Joined:
- Jun 2024
- Posts:
- 298
- From:
- Madrid, ES
I deal with this constantly. My clients want monitoring, their cheap hosts give them nothing.
For OVHcloud specifically, check if they'll enable IPMI on request for the managed line. Sometimes they will if you open a ticket and say the magic words "hardware diagnostics." Not always, but worth the ask.
Failing that, I second the push-based approach. Telegraf to an external InfluxDB. No inbound needed.
swimming upstream since 2019 🐟
SwapJones
Member
OOM killer's friend
- Joined:
- Nov 2024
- Posts:
- 17
- From:
- Hanoi, Vietnam
Both Vultr and Hetzner had routing issues
Maybe you know. Maybe two providers both fail. Maybe you need three. Maybe four. Maybe the cloud is just someone else's computer and that someone else is having a bad day. You know.
I run a 512MB VPS that does nothing but SSH to my boxes and curl a health endpoint. Costs 3 USD. If it can't reach them, they down. If it can't reach ME, I got bigger problems. You know.
RAM is just fast disk, right?