I want to correct myself because I did the math wrong in my head. 8TB is actually a lot for LLM inference if you're just doing API calls. I was thinking training, my bad. For serving a 7B model with 4-bit quant, you're moving maybe 4GB per load, not pulling constantly. 8TB would be millions of requests.
Still stands that providers oversell though. The 10Gbps shared thing was a guess, not a fact.