Skip to content

Your '24/7' chat has been offline 31 hours

Reviews by kei_osa 19 replies 2.7K views
14 #1

Timeline for The CloudCone team:

Aug 21, 14:23 UTC — chat widget shows "agents away, leave message"
Aug 21, 22:00 UTC — ticket #184922 opened, no response
Aug 22, 08:15 UTC — phone line rings 47 times, disconnects
Aug 22, 19:00 UTC — Twitter DM read, no reply
Aug 23, 08:00 UTC — chat still offline. This post.

I run two production nodes on your "Business" tier. One is alerting disk failure. I do not need "unlimited" support. I need any support.

I have already initiated migration to InterServer. This is documentation for the chargeback.

Kei_osa

#2
kei_osa said:
"24/7" chat

"24/7" "support" "business tier"

Allegedly they have staff. Allegedly someone reads the tickets.

Source: trust me bro

I gave up on "cheap" "managed" vps in 2024. Now I pay for "expensive" and actually get answers. Allegedly this is how markets work.

Honey badger don't care... about downtime
#3
kei_osa said:
Disk failure

To be honest, 31 hours for hardware alert is unacceptable.

  • SLA: 99.9% uptime = 8.76h/year allowed downtime
  • Your alert: 31h without response path
  • Chargeback threshold: typically 30 days, you are within

As said, we see this with "budget premium" providers. Numbers do not lie. InterServer's actual 24/7 median response: 4 minutes in our tests. Dry fact. https://www.interserver.net/vps/

Migrate the failing node first. Data second. Blame third.

Containers before it was cool
10 #4

Right then, proper achievement unlocked for CloudCone there. 31 hours without a peep — that's not a bug, that's a feature, innit?

pieter_rtm said:
To be honest, 31 hours for hardware alert is unacceptable.

Cheers lads, I'm with Hostinger and once had their chat go down for 3 hours and I proper lost sleep. 31 hours I'd be updating my CV, not the status page.

Honest truth: their "24/7" means two lads in Sheffield with strong coffee. But they answer. That's the minimum, yeah?

Honey badger don't care... about downtime
#5

47 rings and disconnect? I would have stopped at ten

#6

Which OS is the failing node running, and can you access IPMI?

Have you tried restarting it?
6 #7
wendy said:
Which OS is the failing node running, and can you access IPMI?

Debian 12, and no — CloudCone does not expose IPMI on their VPS platform. This is not a dedicated box, it is their "Business" KVM slice. I have console via their control panel but the disk is read-only now and I cannot get a clean shutdown.

I have the second node syncing what it can. Migration to InterServer (https://www.interserver.net) started 6 hours ago, about 60% of critical data moved via rsync. The failing node I will just abandon and dispute.

#8
kei_osa said:
Debian 12, and no — CloudCone does not expose IPMI

That is what I suspected. Their "Business" tier is still just VPS with a price bump. You are paying for priority queue, not priority hardware.

If you had IPMI you could force a reboot and maybe recover the array. Without it you are waiting for them to notice the host node. Which at 31 hours they clearly have not.

I have seen this pattern with budget KVM. The host node disk fails, multiple customers affected, and they stay quiet until they find a replacement chassis. Not a communication strategy. A hope-you-do-not-notice strategy.

Have you tried restarting it?
#9
wendy said:
Hope-you-do-not-notice strategy

Aye, that's the one. Same as my old landlord when the boiler went.

kei_osa said:
60% of critical data moved via rsync

Kei, if you have read-only root now, are you still getting clean rsync exits? I had a ext4 remount-ro once and rsync kept throwing I/O errors on the big files. Had to do block-level with ddrescue after.

Might be worth testing a small file checksum before you call the other 40% "safe." Sorry to add stress.

Honey badger don't care... about downtime
#10
lee_mcr said:
Are you still getting clean rsync exits?

You are right to ask. The rsyncs from the failing node are throwing errors on two databases over 4GB. I switched to pulling from the healthy node instead — it has replicas from 18 hours ago. Not current but consistent.

The 60% is from the healthy node. I am not trusting the failing node for anything now.

Thank you for the push to check.

Post a reply

You need an account to reply. Log in or register to join the conversation.

Post reply Preview Save draft