Vivold Consulting

Cost, latency and data-sovereignty rules are pulling enterprise AI back on-premise

Key Insights

After a decade of cloud-first migration, enterprises are bringing mission-critical AI workloads back into their own data centers, driven by soaring cloud costs, latency, and data-sovereignty rules like the EU AI Act and GDPR. New liquid-cooled, Blackwell-class hardware from Dell, HPE, and Lenovo now puts hyperscale-equivalent compute within reach of any well-capitalized firm. The result is a hybrid model, with Goldman Sachs, Siemens, and NTT DATA among the most advanced adopters.

Stay Updated

Get the latest insights delivered to your inbox

The cloud-first consensus is quietly reversing

For most of the past decade, enterprise IT had one default answer: put everything in the public cloud. AI is rewriting that. As models move from pilots to mission-critical infrastructure, the limits of a cloud-only approach - latency, data sovereignty, regulatory compliance, and cost - are pushing companies to bring AI workloads back behind their own walls. Purpose-built private infrastructure for training and inference is becoming a central pillar of enterprise strategy rather than a niche concern.

The numbers behind the shift

The spending signals are hard to ignore:

- IDC reported enterprise compute and storage hardware for AI grew 166% year-on-year in Q2 2025, while Gartner pegged 2025 AI spending at US$1.5tn, with data-center systems up nearly 47%.
- The GPU server market, worth US$171bn in 2025, is forecast to hit US$730bn by 2030.
- For firms in regulated industries or under data-residency laws, the cloud isn't just costly - it can be a legal risk, with confidentiality obligations sometimes requiring on-premise deployment outright.

What changed on the supply side is that the hardware caught up: liquid-cooled GPU servers built on NVIDIA's Blackwell architecture, available through Dell, HPE, and Lenovo, now deliver petaflop-scale inference in racks a company can own and secure itself. Most organizations are landing on a hybrid model - public cloud for elastic, non-sensitive work; private data centers for inference and fine-tuning; edge for latency-critical tasks.

What it means for the data center

Bringing AI in-house is not just racking more servers. Densities can reach 100 kilowatts per rack, which makes traditional air cooling inadequate and turns power resilience, grid connectivity, and thermal management into strategic concerns - the data center becomes, in effect, an AI factory.

Who's furthest ahead

The piece profiles three very different adopters. Goldman Sachs has built a private agentic stack and became the first major bank to roll out Cognition's autonomous engineer Devin across its 12,000 developers, reporting three-to-four-times productivity gains in software lifecycle work - funded partly by capital redirected from its retreat from consumer banking. Siemens pushes AI onto the factory floor via its Industrial Edge platform and is building modular, lower-carbon data-center units. And NTT DATA runs agentic AI inside its Cyber Defense Centers to protect private infrastructure, cutting alert volumes by up to 90%. The throughline: on-premise AI is now as much an engineering and security discipline as a software one.

Related Articles

Discovery Loop aims to automate science itself - and Google is funding the startup draining its own bench, as Hassabis exits the DeepMind CEO role

Jeff Dean, Google's chief scientist and 30th employee, is leaving after 27 years to found Discovery Loop, a public benefit corporation using AI to automate scientific research - taking co-founders Sanjay Ghemawat, Quoc Le (Google Brain), and Oriol Vinyals (DeepMind) with him. Google is a founding investor and cloud partner, supplying compute for at least the first year, with Radical Ventures and Khosla Ventures co-leading the seed. In the same announcement, Demis Hassabis steps down as DeepMind CEO to become chairman and Alphabet chief scientist, with Koray Kavukcuoglu taking over Gemini model development. Alphabet stock fell about 4%.

Abbott orders audits of every new project as ERCOT's queue hits 474GW, roughly 90% of it data centres

Governor Greg Abbott announced that all new Texas data-centre projects must be audited by the Public Utility Commission and grid operator ERCOT - a sharp turn for a state whose loose regulation and cheap power made it second only to Virginia for data centres. The trigger is a staggering queue: ERCOT's interconnection requests doubled from 233GW in January to 474GW, about 90% data centres, more than five times the grid's all-time peak demand. Audits will demand power and water use, noise mitigation, light controls, tax-incentive use, and ownership details - after a voluntary survey that most operators simply ignored.

Volta and Bitdeer will build a 133MW Nvidia Vera Rubin data centre in Norway - Anthropic's latest move in a compute land grab

Anthropic has reportedly signed a $10 billion, six-year compute deal with Volta, an AI cloud startup founded only earlier this year, per Bloomberg. Volta is partnering with crypto-mining firm Bitdeer to develop the data centre - located in Norway, delivering 133 megawatts, and running Nvidia's Vera Rubin architecture - and is a member of Nvidia's Cloud Partner programme. It caps an aggressive capacity spree that also includes recent compute deals with SpaceX and Amazon, as Anthropic races rivals for the scarcest input in the industry.