The Pulse: a new trend of CPU shortages
The Pulse: a new trend of CPU shortagesAfter a GPU shortage and memory shortage driven by AI companies, we’re now experiencing a CPU shortage, thanks to AI agents using a lot more CPU with tool usage.
Hi, this is Gergely with a bonus, free issue of the Pragmatic Engineer Newsletter. In every issue, I cover Big Tech and startups through the lens of senior engineers and engineering leaders. Today, we cover one out of four topics from a past issue of The Pulse. Full subscribers received the article below fourteen days ago. If you’ve been forwarded this email, you can subscribe here. I was at dinner with a bunch of CTOs and Head of Infrastructure folks recently, and from the conversation it was clear that many companies are struggling to source CPUs in the current climate, and are coming to terms with the end of juicy discounts from cloud providers for machines in the new era of surging demand fueled by AI. The ‘memory crisis’ afflicting sectors like video gaming is well established and has been extensively covered in terms of shortages of GPUs, but now it seems like things are just as hard for businesses in need of CPUs from cloud providers. In a sign of how things are changing, the disappearance of CPU spot pricing was mentioned at the table. Customers used to be able to pay up to 90% less than the standard price for CPUs, as cloud providers slashed CPU prices for machines that were lying dormant and unused. But that’s no longer the case. It seems that spot pricing has vanished because there’s no longer any lack of demand for CPUs – quite the opposite. I was surprised, but a lot of people chimed in; apparently, it’s now nearly impossible to get CPUs on spot instances without long-running connections with cloud providers. Also, reserving specific CPUs now needs to be done months in advance, and cloud providers will even turn down certain reservations because they don’t have enough CPUs or the right type of CPUs. Even big players struggle to reserve CPUsI have asked turbopuffer CEO Simon Eskildsen about their experience of CPU availability in the cloud, since turbopuffer, as a product, runs on CPUs, not GPUs. They operate in AWS, GCP, and Azure, so I asked how easy it is to get CPUs these days. Simon’s response:
I was able to confirm what Simon said about larger companies struggling; a VP of Engineering at a large inference provider told me they are at the limit on how much GPU and CPU capacity they can buy from their cloud providers. They have cash to spend and want to rent more capacity, and are willing to accept the longest leases. Despite that, cloud providers tell them no more is available! AI hogging CPUsKatelyn Lesse, Head of Platform Engineering for Claude Platform, has written about the reasons for the massive CPU demand increase:
AI-fueled demand does increase CPU load, as shown in this graph from Uber, displaying the growth in agent requests over the past six months:
Increasingly, “agent requests” not only generate code which is inference-heavy – and therefore needs GPUs – but they also run tools that compile the code, run tests, run linters, and all of this is CPU-heavy. At companies like Uber, Ramp, and others, AI agents no longer run on the dev’s local machine, but on a dedicated instance in the cloud. So, the companies reserve more CPUs on their respective cloud providers for agentic workloads. We recently covered how Ramp built and runs its cloud agent, Inspect. Basically, the problem is:
To secure CPUs, it’s necessary to do capacity planning up to 12 months in advance. Katelyn says that server orders are being fulfilled in ~six months, instead of 1-2 weeks’ time as previously, and that prices are up by between 10-20%. So, it’s probably time for capacity planning. Katelyn:
Using existing CPUs more efficiently is something to do, as of now. The CPU capacity shortage won’t go away, and any new CPU allocations requested could take months to turn up. So, what can we do if new capacity lags? One option is utilizing current resources more efficiently! This is a great time to review and to establish now which services are CPU-intensive, and whether or not they need to be. Also check on services which are utilizing little CPU: can they run on fewer nodes, so that some CPU capacity can be allocated to services that need it more? The best time to secure more CPU capacity is most certainly right now. I’m hearing rumors that certain cloud regions no longer accept new tenants because all CPU capacity is leased, or negotiations elsewhere are difficult. I’m also hearing that customers are already paying today to reserve capacity that will only come online in data centers from December. This seems predatory by providers, but demand is so high that this is how they likely prioritize new capacity allocation – while earning much higher profits than usual. If your company has dynamic workloads, and you’ve used spot instances in the past, now could be a good time to allocate fixed capacity – even if it’s more expensive. If you expect meaningful growth, doing so now might mean having options at some cloud providers or in some regions. It seems like this issue has spread everywhere as a corollary of widespread AI adoption. There’s a GPU shortage, memory shortage, and now a growing CPU shortage as well. Back at the end of last year, there was even an hard drive shortage. The only compute primitive not in short supply seems to be networking! Read the full issue of The Pulse this is from, or check out this week’s The Pulse. This week’s issue covers:
You’re on the free list for The Pragmatic Engineer. For the full experience, become a paying subscriber. Many readers expense this newsletter within their company’s training/learning/development budget. If you have such a budget, here’s an email you could send to your manager. This post is public, so feel free to share and forward it. If you enjoyed this post, you might enjoy my book, The Software Engineer's Guidebook: navigating senior, tech lead, staff and principal positions at tech companies and startups.
|


Comments
Post a Comment