For fifteen years, "cloud-first" was the unquestioned answer. In 2026 a majority of enterprises are pulling steady, heavy workloads back on-prem — and the math, not nostalgia, is why.
Cloud-first was never really a strategy. It was an assumption, so widely shared it stopped getting examined: new workload, put it in the cloud; existing workload, plan its migration to the cloud; question the cloud, explain yourself. For a long stretch that assumption was right more often than not. Elasticity was worth a premium when you didn't know how big a thing would get.
The assumption is now being audited, and it's not surviving contact with the invoice. Surveys this year put the share of enterprises moving at least some workloads back from public cloud at around 80%, with a Barclays CIO study landing near 83%. This is not a fringe contrarian take anymore. It's the majority position, and it's driven by the least romantic force in technology: cost that finally got measured against the alternative.
The economics quietly inverted
The cloud's original pitch was that renting beat owning — no capital outlay, no data center, pay only for what you use. That holds beautifully for workloads that are unpredictable or bursty. It holds badly for the opposite: steady, high-utilization workloads that run at a predictable level around the clock. For those, the meter never stops, and the rental that looked cheap at low volume becomes the most expensive way to run a constant load.
The figures being cited are not marginal. Modern private infrastructure is delivering 40–50% lower total cost of ownership for steady-state workloads, and organizations report cutting infrastructure spend 30–60% through selective repatriation without sacrificing performance. For software companies in particular, public cloud has crept up to around half of cost of revenue — a number that, once a CFO sees it, turns "why are we in the cloud?" from heresy into a budget review.
The cloud didn't get more expensive. Your workloads got more predictable — and a predictable, around-the-clock load is exactly the thing renting by the hour was never cheap for.
AI is pouring fuel on the move
If steady workloads made repatriation rational, AI is making it urgent. High-utilization AI inference — a model serving requests continuously — is the most punishing possible case for hourly cloud GPU pricing, because it's the definition of a constant, heavy load. The numbers people are running are stark: on-prem AI infrastructure showing up to an 18x cost advantage per million tokens, and high-utilization inference breaking even on owned hardware in under four months. A refurbished GPU server paying for itself in four to six months against cloud rental is the kind of math that doesn't need a consultant to interpret.
This is the same coin as the industry's power crisis, seen from the other side. Cloud GPU capacity is constrained and expensive; owning inference capacity sidesteps both the premium and the queue. For any enterprise running AI at steady volume, the cloud's elasticity — its core advantage — is precisely the thing that workload doesn't need.
Visual 1 — Which workloads belong where
Workload type | Pattern | Best home | Why |
|---|---|---|---|
Seasonal spikes, launches, dev/test | Bursty, unpredictable | Public cloud | Elasticity is worth the premium |
Steady core applications | Predictable, high-utilization | Private / on-prem | 40–50% lower TCO at constant load |
High-volume AI inference | Continuous, GPU-heavy | On-prem (often) | Up to 18x cheaper per token; fast payback |
Regulated / sovereign data | Compliance-bound | Private / on-prem | Control and residency |
How to read it: the decision isn't cloud vs. on-prem as ideology. It's matching each workload's pattern to its cheapest reliable home. Most enterprises have steady and bursty loads sitting in the same place.
The contrarian guardrail: this is not a cloud exit
The headline number — 80% moving workloads back — invites a wrong conclusion: that cloud was a mistake and everyone's leaving. They aren't, and reading it that way leads to an equally expensive overcorrection. Gartner expects 40% of enterprises to run hybrid architectures for mission-critical work, up from 8%, which is the actual shape of this: not exit, but rebalancing. Bursty and unpredictable stays in the cloud, because that's what the cloud is genuinely best at.
And repatriation has real costs the cloud quietly absorbed for you — hardware refresh cycles, data-center operations, the staff who know how to run physical infrastructure, capacity you now have to plan instead of summon. Companies that spent a decade offloading those skills will rediscover that owning infrastructure is work, not just savings. The move is sound for the right workloads and a trap for organizations that chase the cost number without the operational capability to back it.
What this means for leaders
Reopen the cloud-first default. Treat placement as a per-workload decision, not a standing policy. The question for each system is whether its pattern is bursty or steady — and the steady ones deserve a hard second look at where they run.
Start with your most predictable, heaviest loads. The savings concentrate in steady, high-utilization workloads and continuous AI inference. That's where repatriation pays off fastest and where the cloud premium is least justified.
Rebuild the muscle before you move. Owning infrastructure requires capacity planning, ops talent, and hardware lifecycle management — skills many teams let atrophy. Make sure the operational capability exists before you chase the TCO number, or the savings evaporate into outages and overtime.
Cloud-first earned its dominance in an era when nobody knew how big anything would get, and elasticity was worth almost any price. That era produced a lot of steady, knowable workloads now running on a meter built for uncertainty. Turning the meter off where it no longer makes sense isn't a retreat from the cloud. It's the end of using it for things it was never the cheapest way to do.
A BusinessInfomatics original. Drawn from 2026 cloud-repatriation reporting (Barclays CIO study, Gartner hybrid-compute projections, and TCO analyses on steady-state and AI-inference workloads).



