Skip to content
Tech News
← Back to articles

Infrastructure and compute: Enterprises are buying AI compute for speed while flying blind on what it costs

read original more articles
Why This Matters

As enterprises rapidly deploy AI infrastructure, performance and GPU availability now take precedence over cost considerations, highlighting a shift in priorities driven by production demands. However, many organizations lack clear visibility into their AI compute costs, risking inefficiencies and unforeseen expenses. This disconnect underscores the need for better cost management tools and strategic planning in enterprise AI adoption.

Key Takeaways

Across 170 enterprises, AI infrastructure has moved decisively into production — two-thirds now run AI workloads live and three in 10 run them at scale — while the ability to account for what that infrastructure costs has not kept pace. Enterprises have quietly demoted cost in the buying decision: performance and GPU availability now outrank total cost of ownership, and reliability outranks price as the measure of success. That reordering is rational for teams under production pressure, but it lands on an uncomfortable fact — fewer than half can rigorously track what their AI compute costs, most GPUs still run at half capacity or less, and the next dollar is aimed at specialized clouds that fewer than one in twenty of them actually use.This wave of VentureBeat Pulse Research examines enterprise AI infrastructure and compute: where organizations are in their deployment journey, what they run AI on today, how they buy and measure it, where the next investment is aimed, and — most revealingly — how well they can see the economics of the compute underneath it all.This is an operational cohort. Two-thirds of enterprises (66%) have AI workloads running in production, and 29% describe AI in production at scale, with only 4% not yet running AI workloads at all. That maturity shows in the stack: the average enterprise runs three infrastructure platforms, with OpenAI (49%), Google Gemini (48%), Microsoft Azure (47%), and Google Cloud (42%) all present in roughly half of them. Asked to name one primary platform, Azure leads at 26%.The most consequential shift is in how enterprises decide. Integration with the existing cloud and data stack remains the top selection factor at 40%, but performance — latency and throughput — has climbed to second at 35%, and access to GPU availability to third at 24%, both ahead of total cost of ownership at 22%. The same ordering governs measurement: uptime and reliability is the primary success metric for 51% of enterprises and developer productivity for 39%, ahead of cost per million tokens at 31%. Enterprises under production pressure are buying and measuring for speed and availability, and have moved cost down the list.That would be unremarkable if the economics were under control, but they're not. Among the 155 enterprises that operate their own GPUs, 69% report utilization of 50% or less and only 23% clear the halfway mark; 12% do not measure utilization at all. Fewer than half (47%) rigorously track what their AI compute costs and returns, and even among enterprises running AI in production at scale that figure only reaches 56%. Value for money is the weakest of three satisfaction scores at 3.87, against 4.14 for overall satisfaction — the softness landing precisely on the dimension hardest to judge without measurement.The next round of spending points away from the current stack. AI-specialized clouds are the top planned evaluation area at 44% and carry the strongest net momentum of any infrastructure approach (+36), yet CoreWeave and Lambda each registers at 3.5% of current usage and the rest of the neocloud field sits below 3%. Non-Nvidia accelerators draw 39%. And 62% of enterprises intend to switch or add a provider within 12 months — though the consideration set is dominated by the same incumbents they already run.MethodologyVentureBeat fielded this survey as part of its ongoing Pulse Research series, this one focused on enterprise AI infrastructure, compute, and inference economics. Responses are filtered to organizations with more than 100 employees (n=170; the survey’s smallest size band, 1–100 employees, is excluded), drawn from a single July 2026 wave. Because this is one wave rather than a pooled multi-month sample, the report reads cross-sectionally and does not infer month-over-month trends; all figures are drawn from the July fielding only. Several questions were multiple-select, so those shares can sum to more than 100%.By organization size this wave reaches further up-market than the mid-market skew this series usually carries: 251–1,000 employees (28%) and 1,001–5,000 (25%) lead, with 10,001+ (19%), 101–250 (15%), and 5,001–10,000 (12%) filling out the rest — meaning 57% of respondents sit above 1,000 employees. By role it spans managers (48%), individual contributors (27%), the C-suite (12%), and VPs and directors (9%); on purchasing authority it is buyer-credible, with 39% final decision-makers and another 43% recommenders or influencers for AI solutions. Technology/Software is the largest industry at 35%, followed by Manufacturing (14%), Financial Services (12%), and Healthcare/Life Sciences (9%).At 170 respondents the sample is large enough to read directionally with reasonable confidence, but it should still be treated as a directional signal rather than a precise measurement; it is self-selected and is not a probability sample. It is best read as the view from organizations actively building and operating AI infrastructure rather than from the largest hyperscale operators.Finding 1: Two-thirds are past the pilotThree in 10 now run AI in production at scaleWe asked where organizations sit in their AI deployment journey. This cohort has largely moved beyond experimentation.Two-thirds of enterprises (66%) have AI workloads running in production, and 29% describe AI in production at scale. Only 30% remain in proofs of concept and just 4% have not started. This is a materially more operational sample than this series has typically drawn, consistent with its up-market composition — 57% of respondents sit above 1,000 employees.That maturity is the frame for everything that follows. The infrastructure decisions in this report are being made largely by organizations with production workloads and real bills, not by teams still sizing a pilot. It explains the reordering of buying criteria in Finding 5, where performance and availability displace cost — the priorities of teams running live systems. It also raises the stakes on Findings 6 and 7: an enterprise that cannot measure utilization or cost during experimentation has a planning problem, while one that cannot measure them in production at scale has an operating one.Finding 2: The stack is hyperscaler-and-API, three platforms deepThe specialized GPU clouds still barely registerWe asked which providers and platforms enterprises currently use to run their AI, and which one they treat as primary. The answer remains the incumbents — several of them at once.The current stack is hyperscaler-and-API, and it is plural: enterprises name three platforms on average. The general-purpose clouds and the major model APIs account for essentially all current deployment, with four platforms — OpenAI, Gemini, Azure, and Google Cloud — each presents in more than four of every 10 enterprises. Asked to pick one primary platform, Microsoft Azure leads at 26%, with Google Cloud second at 19%; the model providers together take 35% of primary status when OpenAI (14%), Gemini (14%), and Anthropic (8%) are combined.The specialized “neocloud” GPU providers that dominate AI-infrastructure headlines remain marginal in practice. CoreWeave and Lambda each appear in 3.5% of stacks, Baseten in 3%, and Crusoe, Nebius, Fireworks, Together, and Anyscale each at or below 2%. Combined, they are named as the primary platform by 1% of enterprises. Meanwhile 13% run a custom open-source self-managed stack and 9% operate their own GPU clusters — both larger footprints than the entire specialized-cloud category. That contrast is what makes the evaluation intentions in Finding 3 worth reading closely.A note on reading these shares: As described in the methodology section, this sample is self-selected and this question counted every provider a respondent uses — an average of 3.0 selections each — so the figures measure presence in the stack rather than spending or primary status. The separate primary-platform question is the better guide to where the center of gravity sits. A sample built this way will show a different provider mix than a spend-weighted census of the broader market; read these shares as a portrait of what this AI-active cohort runs today, and treat gaps against industry-wide market share estimates as a property of the sample rather than a contradiction of either.Finding 3: The next dollar goes to infrastructure they don't yet runAI-specialized clouds top the evaluations list and carry the strongest momentumWe asked where enterprises plan to evaluate AI infrastructure over the next 12 months, and whether they expect to do more or less with each category of infrastructure. Both answers point away from the stack they run today.Here is the report’s sharpest tension, and it is the same one this series has now recorded across successive waves. The single most-cited planned evaluation area — AI-specialized clouds, at 44% — is the category that 3.5% of these enterprises actually use (Finding 2). Nearly four in 10 (39%) intend to evaluate non-Nvidia accelerators, a quarter next-generation Nvidia silicon, and even decentralized compute networks draw 18%.The direction-of-travel question corroborates it rather than merely repeating it. Asked whether they expect to do more, less, or about the same with each approach, enterprises put specialized AI clouds at the highest net momentum (+36, with 42% doing more against 6% doing less), ahead of inference APIs (+34) and hyperscalers (+30). On-prem and co-located infrastructure is the laggard at +5, the only category where a substantial share — 22% — report pulling back. Every off-premises approach is net-expanding; the specialized clouds are expanding fastest from the smallest base.Read against current usage, this is not incremental adjustment. It is the leading edge of a re-platforming that enterprises have been signaling for several waves and have not yet executed. The gap between a 44% evaluation rate and a 3.5% usage rate is the single widest intent-to-action spread in this dataset, and how it resolves — whether the neoclouds convert evaluation into deployment, or whether the hyperscalers absorb the demand with their own AI infrastructure — is the open question of the category.Finding 4: Six in 10 plan to move, mostly among the incumbentsHigh churn intent, but the consideration set is the stack they already runWe asked whether and when enterprises plan to switch or add an infrastructure provider, and which providers they are considering.For a category as foundational as compute, this is a substantial amount of intended movement: 62% of enterprises intend to switch or add a provider within 12 months, and 29% within the next quarter alone. Only 39% plan to stand still.Where that interest points is the more useful signal. The providers drawing the most switching consideration are the ones enterprises already run — OpenAI and Google Cloud (29% each), Microsoft Azure (28%), Gemini (25%), Anthropic (16%), Oracle Cloud (14%), and AWS (13%). The specialized clouds that top the evaluation list in Finding 3 draw far less concrete switching consideration: CoreWeave 4%, Lambda 3.5%, and the remainder at or below 2%. A further 8% are evaluating with no shortlist yet.The two findings are not in conflict; they operate on different clocks. The neocloud interest in Finding 3 is a 12-month evaluation thesis about where AI compute should eventually run. The switching in the next quarter is mostly incumbents trading share and enterprises consolidating spend among providers they already hold contracts with. Vendors reading the 44% evaluation figure as near-term pipeline should weigh it against a 4% consideration rate.Finding 5: Performance overtakes cost, in buying and in measurementTotal cost of ownership falls below latency and GPU availabilityWe asked what matters most when enterprises select an AI infrastructure provider, and what they treat as the primary measure of success once it is running. Both answers have moved away from price.Integration with the existing stack remains the top selection factor at 40%, which is consistent with a cohort running three platforms and unwilling to add a fourth that does not fit. What has changed is everything below it. Performance sits second at 35% and GPU access and availability third at 24%, both ahead of total cost of ownership at 22%. Fine-grained autoscaling draws 18% and cost per million tokens 16% — no longer the outlier it once was in this series, but still last.Measurement follows the same logic. Uptime and reliability is the primary success metric for 51% of enterprises, well ahead of developer productivity and deployment speed (39%), cost per million tokens (31%), latency (27%), and throughput (25%). Taken together, the operational metrics dominate the economic one by a wide margin.This is a coherent posture for the production cohort in Finding 1 — teams running live workloads care first about whether the system stays up and how fast they can ship on it. But it sits uneasily beside Finding 7. Total cost of ownership has been demoted to fourth as a buying criterion at exactly the moment when 53% of enterprises still cannot rigorously track what their compute costs. The uncomfortable reading is that cost has fallen down the list partly because it remains the hardest thing in the stack to see, and criteria that cannot be measured tend to lose to criteria that can.Finding 6: The GPUs run warmer, but most still run coldRoughly seven in 10 GPU operators report 50% utilization or lessWe asked what share of their GPU capacity enterprises actually utilize. Figures here are reported on the 155 enterprises that operate their own GPUs; 15 consume exclusively via API and run none.The compute already in place runs cold, though less so than this series has recorded before. Roughly seven in ten GPU-operating enterprises (69%) report utilization at or below half capacity, with the 26–50% band alone accounting for 46%. About a quarter (26%) run at 25% or below. Against that, 23% now clear the 50% mark — a meaningful efficient minority rather than a rounding error.The remaining 12% who do not measure utilization at all are the more troubling number, because they are invisible in both directions: they cannot claim efficiency and cannot detect waste. And utilization does not improve with maturity in the way one might expect — among enterprises running AI in production at scale, 22% clear the 50% mark, statistically indistinguishable from the 24% among everyone else. Scale is not, by itself, producing better-utilized fleets.Idle accelerators are expensive accelerators, and this remains the clearest single measure of the gap in this report: enterprises are planning to evaluate specialized clouds and next-generation silicon (Finding 3) while the capacity they already own sits substantially unused. The efficiency headroom in the current fleet is large, and for one in eight enterprises, entirely unmeasured.Finding 7: Fewer than half can account for what they spendRigorous cost tracking reaches only 56%, even among at-scale operatorsWe asked whether enterprises can quantify the cost and return of their AI infrastructure spend, and how satisfied they are with what they run. Confidence in the ledger still lags the spending.Measurement trails money. Fewer than half of enterprises (47%) rigorously track the cost and return of their AI compute; the majority track only partially (39%), cannot quantify it yet (15%), or have not prioritized it (6%). Maturity helps but does not solve it: among enterprises running AI in production at scale, rigorous tracking reaches 56%, against 43% for everyone else. Even in the most operationally advanced segment of this sample, more than four in ten cannot account precisely for what their AI compute costs or returns.Satisfaction with current infrastructure is moderately positive and tellingly uneven. On a five-point scale, overall satisfaction averages 4.14 and ease of implementation 4.04, while value for money trails at 3.87 — the softness landing on the one dimension that requires measurement to assess. Enterprises are, in effect, expressing dissatisfaction with an economic relationship most of them cannot yet quantify.Read with Finding 5, the picture is self-reinforcing rather than merely inconsistent. Cost has slipped to fourth among buying criteria while remaining the least visible property of the stack, and the least visible property is the one enterprises rate lowest. Better instrumentation would not necessarily change what enterprises buy — but it would let them know whether the trade they are making for performance and availability is a good one.Finding 8: The memory frontier is still unclaimedDell and Nvidia lead a scattered field, and one in five has no viewWe asked how enterprises would address the emerging constraint in large-scale inference — the shift from GPU compute to memory, specifically KV-cache capacity. The field remains early and fragmented.The memory frontier is real but barely governed. Dell leads at 24% and Nvidia follows at 21%, with the remainder scattering across open-source tooling (12%), model-level efficiency techniques such as MLA and quantization (11%), and a long tail of storage vendors each in low single digits. No approach commands anything close to a majority, and the two leaders together account for less than half the field.Most telling is that roughly one in five enterprises (19%) either do not recognize the constraint (7%) or have not begun to address it (12%). For a shift that will reshape inference cost and architecture, this is an early and unsettled market. It is also consistent with the measurement gap in Finding 7 — enterprises that cannot yet quantify what their current compute costs are in a poor position to anticipate which constraint will drive that cost next. The memory bottleneck is arriving while most of this cohort is still working to see the one in front of it.The bottom line: Buying for speed, blind on costOrganizations with more than 100 employees have moved AI infrastructure into production — two-thirds run live workloads, three in ten at scale — and their buying behavior has matured accordingly. They run three platforms on average, select on integration and performance, and measure success on uptime and developer velocity. For teams operating live systems, that is the right set of priorities.What has not matured is the accounting. Total cost of ownership has fallen to fourth among selection criteria and cost per million tokens sits last, at the same moment that 53% of enterprises cannot rigorously track what their compute costs, 69% of GPU operators run at half capacity or less, and 12% do not measure utilization at all. Value for money is the lowest-rated attribute of the infrastructure they run — a judgment most of them are making without the instrumentation to support it. Cost has not become unimportant; it has become invisible, and the buying criteria have quietly reorganized around what can actually be seen.Meanwhile the next round of spending points past the current stack. Specialized AI clouds are the top evaluation target at 44% and carry the strongest net momentum of any approach, against a 3.5% usage rate and a 4% near-term switching consideration — the widest intent-to-action spread in the data. Non-Nvidia accelerators draw 39%. And the constraint after this one, the shift from compute to memory in large-scale inference, is unrecognized or unaddressed by one enterprise in five.At 170 respondents in a single July wave, reaching further up-market than this series typically does, this is a directional read — but the direction is consistent. Enterprises have become good operators of AI infrastructure and have not yet become good accountants of it. The open question for later waves is whether the instrumentation catches up before the re-platforming arrives, or whether enterprises buy the next layer of compute as blind to its economics as the last.Based on survey responses from 170 qualified enterprise respondents (100+ employees), drawn from a single July 2026 wave. This sample is self-selected and directional rather than a precise measurement, and reads cross-sectionally with no month-over-month trend claims. Respondents include managers, individual contributors, C-suite, and VPs/directors, with purchasing authority weighted toward decision-makers and recommenders, across technology, manufacturing, financial services, healthcare, and other industries. Note: Figures for the switching-timeline, GPU-utilization, and cost-tracking questions are reported as a percentage of unique respondents rather than selections; individual categories for these three questions may sum to more than the reported total.