On this page
- Summary
- Why this question
- Scope and methods
- The landscape
- Theme one — The market: total size and the training-to-inference rotation
- Theme two — Fine-tuning: a real but small slice
- Theme three — Inside inference: chat, code, and the agentic wave
- Theme four — Training’s modality mix: multimodal-first, text-centric accounting
- Theme five — Utilization and the economics of the AI factory
- Theme six — Strategy: should an AI factory builder move into inference?
- Where the evidence disagrees
- Gaps and open questions
- Confidence and limitations
- Evidence table
- References
The LLM compute market — training, fine-tuning, and inference
How is LLM compute demand, revenue, and profit split between training, fine-tuning, and inference, and what does the shift toward inference mean for AI factory builders?
https://reviews.lewiswon.me/reviews/llm-compute-market-economics/ · Updated 19 Aug 2026
How this review was made
- Databases
- OpenAlex, Semantic Scholar, Crossref, arXiv, primary web sources (vendor press releases, analyst PDFs, SEC filings)
- Queries (literal)
- compute trends across three eras of machine learning
- trends in training compute
- will we run out of data
- explosive growth from AI architecture innovation
- training compute-optimal large language models
- DeepSeek-V3 technical report
- DeepSeek-R1 incentivizing reasoning capability in LLMs
- Splitwise efficient generative LLM inference
- DistServe disaggregating prefill and decoding
- efficient memory management for large language model serving with pagedattention
- SGLang efficient execution of structured language model programs
- DynamoLLM
- low-rank adaptation of large language models
- QLoRA efficient finetuning of quantized LLMs
- energy and policy considerations for deep learning in NLP
- survey on efficient inference for large language models
- ReAct synergizing reasoning and acting in language models
- survey on large language model based autonomous agents
- Toolformer language models can teach themselves to use tools
- Gemini a family of highly capable multimodal models
- scaling data-constrained language models
- AI index 2024 annual report
- LoRA learns less and forgets less
- MegaScale scaling large language model training
- power hungry processing watts driving the cost of AI deployment
- multimodal large language models survey
- large language model inference price trends
- model flops utilization
- fine-tuning large language models
- world models video generation compute
- GPU utilization data center AI training inference
- GPU cloud economics TCO inference serving cost per token
- multimodal training compute scaling video models
- arXiv ti:"Trends in Training Compute"
- arXiv ti:"Scaling Data-Constrained Language Models"
- arXiv ti:"MegaScale-Infer"
- arXiv ti:"AI Index 2024 Annual Report"
- arXiv ti:"Artificial Intelligence Index Report"
- arXiv abs 2405.19522 / 2408.00741 / 2504.02263 / 2309.11690 / 2312.11805 / 2202.05924 / 2305.16264 / 2404.14294 / 2606.15708
- Google News RSS: Anthropic state of AI infrastructure; Sequoia $600B question; Gartner AI agents 112 billion; Deloitte agentic 25%; NVIDIA 40% inference; GPU utilization; Epoch multimodal; SemiAnalysis inference
- Google News RSS: Uptime Institute GPU utilization; enterprise GPU utilization 95% wasted; CoreWeave price increase
- WordPress REST API: semianalysis.com (InferenceX, inference gross margins, AI Cloud TCO)
- primary PDF: Stanford HAI AI Index 2026 report (hai.stanford.edu/assets/files/ai_index_report_2026.pdf)
- primary: NVIDIA technical blog 'Scaling Token Factory Revenue and AI Efficiency by Maximizing Performance per Watt'
- primary: I/O Fund 'Nvidia, CoreWeave, and Nebius: Inside the Circular Financing of the GPU Boom'
- web: Menlo Ventures 2024+2025 State of GenAI in the Enterprise
- web: Bloomberg Intelligence GenAI deep dive 2025 PDF
- web: Gartner AI inference press release via InfotechLead (2026-08-17)
- web: NVIDIA Q1/Q3/Q4 FY26 earnings releases + Jensen Huang factory blogs
- web: Epoch AI trends + AI chip sales data pages
- web: OpenRouter/a16z 100T-token State of AI study
- web: CoreWeave FY2025 10-K (SEC EDGAR)
- web: Sequoia AI's $600B Question
- web: Anthropic Build AI in America PDF
- web: IDC AI Supercycle deck
- web: Lambda + Together AI pricing pages
- Search last run
- 2026-08-19
- Screening
- 28 sources used · 2019–2026 · standard review
Summary
The short version
The market for LLM compute is large, growing fast, and rotating from training toward inference. Gartner counts $2.59T of total AI spending in 2026 (up 47% year over year) and expects inference to overtake training inside AI-optimized cloud infrastructure this year, while Bloomberg Intelligence dates the crossover for the wider market at 2029 — at least three years earlier than its pre-2025 view. Enterprise generative AI spend reached $37B in 2025 per Menlo Ventures, with inference-priced model APIs the largest single line item and fine-tuning a genuine but small slice. Inside inference, agentic workloads are the fastest-growing category: reasoning models multiply tokens per request, tool-calling adds round-trips, and one empirical study of 100 trillion tokens finds agentic inference on track to become the majority of all inference. For a company that builds AI factories, the strategic evidence points one way: selling inference-as-a-service would compete head-on with its own customers, but selling inference-workload-aware infrastructure — disaggregated prefill and decode, SLO-aware scheduling, burst absorption, tokens-per-watt guarantees — is a differentiated product that the pure-play GPU clouds already gesture toward and none fully deliver. Confidence is moderate: the market numbers come from a handful of analyst houses with different definitions, and the most load-bearing strategic claims rest on a smaller set of primary documents.
Why this question
Capital allocation in AI infrastructure is being decided on a split nobody can directly observe. Tens of billions of dollars of GPU capacity are contracted on the assumption that inference demand will eventually dwarf training demand, yet the public accounting of that split is fragmentary: analyst houses measure different flows, vendors report different mixes, and fine-tuning — the middle category — has almost no dedicated market sizing at all. For an AI factory builder, the question is not academic. A factory’s architecture, power contract, network topology, and customer base all depend on whether its capacity will be consumed by six-month training runs or by continuous, latency-sensitive, bursty inference traffic. And the strategic question — whether to move downstream into inference — is a channel-conflict decision: the company’s customers are inference providers, so the answer determines whether the factory builder is a neutral supplier or a competitor.
Scope and methods
The question has two halves: a market-data half (size, split, growth of LLM compute demand by activity) and a strategy half (what the shift means for an AI factory builder). Because the evidence for the first half is almost entirely industry-produced — analyst reports, vendor financials, and consulting surveys rather than peer-reviewed literature — this review uses a two-layer evidence model. Peer-reviewed and preprint sources (28, all with verified DOIs, fetched and read this session, mostly abstract-level) anchor the mechanisms: scaling laws, serving systems, disaggregation, fine-tuning methods, energy accounting. Industry reports and primary vendor documents (Menlo Ventures, Bloomberg Intelligence, Gartner coverage, NVIDIA releases, SEC filings, Epoch AI, Stanford’s AI Index, pricing pages) are cited inline with URLs in the body; they carry no DOIs and so cannot be entries in the reference list, but every number attributed to them was verified by fetching the document this session. Screening: 543 raw records from 33 search queries across OpenAlex, Semantic Scholar, Crossref, and arXiv were deduplicated to 417, screened to ~30 academic candidates, and 28 included; 4 remembered landmark titles failed resolution and were dropped rather than forced. Industry-side: ~30 primary documents were fetched and grepped for the load-bearing numbers; figures that could not be verified (e.g. CoreWeave margin commentary, specific GPU-utilization percentages, an un-sourced “80% of compute is inference” claim) are excluded or flagged. A second search wave (2026-08-19) targeted the gaps flagged in the first pass: direct search engines (Google, Bing, DuckDuckGo, Mojeek, Startpage) bot-blocked this host, so the wave used Google News RSS, publisher WordPress REST APIs, and primary PDFs — it added the Stanford AI Index 2026 report 7 and two inline industry sources (NVIDIA’s token-factory technical blog and I/O Fund’s neocloud-financing analysis), and it confirmed several gaps remain (the Gartner “$112B agentic spend by 2028” figure, Deloitte’s agentic-spend share claim, VentureBeat’s GPU-utilization analysis, and coverage of CoreWeave’s 2026 price increase could not be retrieved to a verifiable primary). The review deliberately does not cover application-layer AI revenue, chip design details, or regional markets beyond what the sources report.
The landscape
The market-economics literature is thin relative to the stakes, and it is dominated by a few institutions. On the academic side, Epoch AI’s compute-trend work (training compute doubling every ~5 months, data exhaustion forecasts) and the Stanford AI Index are the canonical references, while the serving-systems literature (Splitwise, DistServe, vLLM, SGLang, DynamoLLM, MegaScale-Infer) is the most rigorous evidence on what inference actually costs to operate. On the industry side, the usable market numbers come from a handful of houses — Menlo Ventures (enterprise spend), Bloomberg Intelligence (long-horizon revenue), Gartner and IDC (IT spending), Epoch AI (compute stock and frontier-lab economics) — plus NVIDIA’s own disclosures, which are the single most influential primary source because NVIDIA sits at the revenue bottleneck of the whole chain. The landscape has two notable absences: nobody publishes a standalone fine-tuning market size, and nobody publishes a defensible split of training compute by modality. Both absences are themselves findings, discussed below.
Theme one — The market: total size and the training-to-inference rotation
The top-line numbers agree that AI spend is enormous and growing, but they measure different flows. Gartner forecasts $2.59T of total AI procurement in 2026 (up 47%), including $401B of AI infrastructure, while IDC counts $487B of AI infrastructure hardware in 2026 (up 53%) and a cumulative run past $1T by 2029 — the gap between the two is scope: procurement vs hardware shipments. Bloomberg Intelligence projects $1.8T of generative-AI revenue by 2032, up from ~$93B in 2023, and IDC’s deck sees enterprise AI platforms and services growing from $400B in 2026 to $2.1T by 2029. On the investment side, the Stanford AI Index 2026 reports US private AI investment of $285.9B in 2025 — more than 23 times China’s $12.4B — with global corporate AI investment more than doubling and private investment up 127.5% 7. These are not additive, and the review treats them as a range, not a consensus point.
The split that matters — training vs inference — is where the sources converge on direction and disagree on timing. Bloomberg Intelligence’s March 2025 deep dive puts the inference market at $735B by 2032 versus $580B for training, and expects inference compute consumption to overtake training “at least three years sooner than our initial projections” — by 2029 — driven by reasoning models and agent deployment. Gartner’s August 2026 release goes further: inside the AI-optimized infrastructure-as-a-service market, inference spend ($23.3B) already exceeds training spend ($19B) in 2026, rising from 55% to 59% of that market in 2027. Epoch AI’s numbers for the frontier labs run the other way — only ~30% of OpenAI’s compute spending in 2024 went to inference — which is exactly the expected pattern: the labs that train frontier models are the last place where training dominates. The 80/20 heuristic (inference as ~80% of long-run AI compute) circulates widely in investment commentary; this review treats it as directional, because the three measured datapoints above (BI 2032, Gartner 2026, Epoch/OpenAI 2024) tell a coherent story without it: inference crosses over first in the broad commercial market, later at the frontier labs themselves.
Two mechanisms explain the rotation. First, the training side is supply- and data-constrained: training compute has grown ~5x per year since 2020 1, but the Chinchilla scaling law says the efficient frontier pairs compute with data 2, and public human text is forecast to be exhausted somewhere between 2026 and 2032 4 — repeated data only stretches so far 5. The 2026 AI Index adds a concentration signal: fewer notable models were released in 2025 than the year before, parameter counts have plateaued near 1 trillion, and the frontier labs have stopped disclosing training compute 7 — training demand is narrowing to fewer, larger runs. Second, the inference side is demand-multiplied: reasoning models emit long chains of thought before answering 9, open-weight releases compress prices and expand the market 8, and Goldman Sachs (cited in the Gartner coverage) projects agent token consumption growing 24x between 2026 and 2030 to ~120 quadrillion tokens per month, even as per-token inference costs fall ~60% a year. NVIDIA’s own disclosures corroborate the demand side: inference token generation “surged tenfold in just one year” (Q1 FY26), demand is “compounding across training and inference — each growing exponentially,” and cloud GPUs are “sold out” (Q3 FY26). The uncertainty that runs through all long-horizon forecasts is captured well by 28: explosive-growth arguments depend on strong assumptions about how fast automation spreads, and they are exactly the assumptions the market is betting on.
Theme two — Fine-tuning: a real but small slice
Fine-tuning is the least-measured category in the whole market. No major analyst house publishes a standalone fine-tuning spend figure; it is folded into “training” by Bloomberg Intelligence and into “model training infrastructure” by Menlo Ventures — a $4.0B line in 2025 that covers training and adaptation together. The adoption evidence suggests the slice is small: Menlo’s 2024 survey found only 9% of enterprise production models were fine-tuned, and its 2025 edition still describes fine-tuning as a niche technique relative to prompt design and RAG. The mechanism-level literature explains both why fine-tuning is cheap enough to exist as a category and why it stays niche. Parameter-efficient methods collapsed the cost: LoRA cuts trainable parameters ~10,000x and GPU memory ~3x versus full fine-tuning 18, and QLoRA fine-tunes a 65B model on a single 48GB GPU 19. But the quality evidence is mixed: in standard low-rank settings LoRA substantially underperforms full fine-tuning in-domain, even though it forgets less out-of-domain 20 — a trade-off that pushes many production teams toward retrieval and prompting instead. The honest summary: fine-tuning is a real workload with real infrastructure demand (and the only category where “adaptation” compute is visibly growing with open-weight adoption), but every available proxy says it is single-digit percent of LLM compute spend, and the review found no source that contradicts that.
Theme three — Inside inference: chat, code, and the agentic wave
The best empirical map of inference workloads is the a16z/OpenRouter study of over 100 trillion real tokens, which finds that programming and creative roleplay are the two dominant token categories — roleplay alone is ~52% of open-source-model tokens — and that average prompts have grown ~4x since early 2024 (from ~1.5K to over 6K tokens) while completions nearly tripled, driven by reasoning tokens. Menlo’s enterprise data shows the professional side: coding is the first killer use case at ~$4B of 2025 spend (55% of departmental AI), and model-API providers (Anthropic, OpenAI, Google) capture 88% of enterprise LLM API usage — an inference-priced consumption line inside the $19B application layer.
The structural shift is agentic. The academic lineage is clear: ReAct established the reason-then-act loop 23, Toolformer showed models can learn to call tools 24, and the agent survey literature documents the resulting architectures 25. Each agentic turn converts one user task into multiple model calls with tools, memory, and re-planning — which is why the OpenRouter study says “agentic inference will be taking over the majority of the inference,” and why Gartner describes an “Inference Paradox”: better unit economics encourage more powerful models, more tokens, and more complex workflows, so per-workflow inference costs rise more than 5x through 2028 even as per-token prices fall. The gap between aspiration and reality matters for capacity planning: Menlo finds only 16% of enterprise and 27% of startup deployments qualify as true agents today, and agent platforms are ~10% ($750M) of the $8.4B horizontal-AI category — so the agentic wave is a forecast backed by directionally strong evidence (token mix, reasoning-model adoption, tool-call growth) rather than an already-realized majority of spend.
Theme four — Training’s modality mix: multimodal-first, text-centric accounting
On modality, the honest finding is that the split of training compute between text-only and multimodal work is not publicly measured. No analyst house or academic source this review could find publishes it. What the evidence does establish: frontier models are natively multimodal — Gemini was introduced as a family spanning image, audio, video, and text understanding 26 — and the modal frontier since 2024 has been omni-input (GPT-4o-class) plus video generation and world models. The AI Index documents the broader trend of model releases shifting toward multimodality (its 2024 report tracks technical advancement and industry statistics; the report series is the standard reference for such counts) 6. The one modality datapoint with real market grounding is the open-source download mix: text-generation models grew to 42% of downloads in 2025 while classifier, image, and audio models collapsed from ~70% in 2022 to under 6%, with multimodal and video-generation models growing in their place 7. But compute accounting remains text-centric: Epoch AI’s flagship training-compute series tracks frontier language models specifically, and the training-scaling literature that anchors the whole market discussion — 1245 — is built on text-LM loss. Two indirect signals suggest multimodal’s compute share is rising from a small base: video-generation and world-model training is orders of magnitude more compute-hungry per training token than text (diffusion-style models at frontier scale), and the data-constraint argument that caps text pretraining does not bind on video, which is comparatively abundant. The gap matters: if multimodal training compute is 10% or 40% of the frontier, datacenter power planning changes materially — and nobody publishes the number.
Theme five — Utilization and the economics of the AI factory
Utilization is where the profit of an AI factory is actually made, and the evidence separates cleanly into training and inference regimes. Training utilization is a well-studied engineering problem: PaLM set the standard metric (model FLOPs utilization, 46.2% on TPUv4) 3, and MegaScale shows the ceiling at scale — 55.2% MFU on 12,288 GPUs, with fault tolerance and stragglers dominating operational effort 15. Inference is the harder utilization problem. Splitwise’s characterization is the anchor: generation is memory-bound and underutilizes compute even with state-of-the-art batching 10; DistServe shows that colocating prefill and decode creates interference that couples resource allocation and violates per-phase SLOs 11; and MegaScale-Infer shows MoE inference specifically drops GPU utilization because sparsity makes the feed-forward layers memory-intensive 16. Serving software is the lever that closes part of the gap — PagedAttention’s KV-cache management 12, SGLang’s 6.4x throughput via RadixAttention 13, and the broader efficiency taxonomy of quantization, pruning, and distillation 27 — which is why inference cost per token keeps falling ~60% a year even as demand compounds.
Energy is the second-order constraint. The accounting tradition starts with Strubell’s training-energy estimates (up to 626,155 lbs CO2e for a neural-architecture-search run) 21, but the balance has flipped: recent work finds inference energy is now the dominant and growing LLM energy cost 22, and per-request energy scales with model size 17. The AI Index 2026 draws the same conclusion with harder numbers: it reports estimates that the energy required to serve inference queries can exceed the one-time cost of training within months, that annual GPT-4o inference water use alone may exceed the drinking-water needs of 1.2 million people, and that US AI data-center power capacity has risen to 29.6 GW 7. This is the layer where the prior reviews on this site go deep — tokens per watt as the measurement frame tokens-per-watt, the serving-engine tricks that move the number llm-inference-engine-tricks, and datacenter-level efficiency model-to-grid-ai-datacenter-efficiency. NVIDIA has now productized the framing: its technical blog markets the AI factory as a token factory, claims a 1,000,000x increase in inference throughput per megawatt across six GPU generations, and states the operating principle bluntly — “every watt not spent on cooling or idle capacity becomes a watt that generates tokens.” The market context from the other industry sources: Epoch AI tracks the AI-chip stock doubling roughly every 6.8 months and 83 covered US AI datacenters totaling 12.7 GW of IT power, Anthropic’s policy paper demands at least 50 GW of US capacity by 2028, and NVIDIA now markets “10x throughput per megawatt” as a benchmark win — tokens-per-watt has become a headline sales attribute of the factory itself.
The economics at the factory layer are brutal and public. CoreWeave’s FY2025 10-K is the clearest window: $5.1B revenue, a headline gross margin around 72%, but a separate $2.9B “technology and infrastructure” line (depreciation, power, operations) that leaves operating income at −$46M for the year; 98% of revenue came from committed contracts; Microsoft alone is 67% of revenue; OpenAI committed up to ~$6.5B through May 2031. The financing layer that keeps this running is now visible too: I/O Fund’s analysis of the two public neoclouds documents ~3.5 GW of contracted power capacity each (CoreWeave targeting 1.7 GW active by end-2026, Nebius 800 MW to 1 GW), Microsoft commitments of roughly $60B across neoclouds, and a circular-financing structure — NVIDIA equity, hyperscaler contracts, GPU-backed debt — that funds the buildout while the operators remain far from profitable. Sequoia’s “$600B Question” frames the same structure from the demand side: NVIDIA’s revenue run-rate, doubled for datacenter TCO and doubled again for end-user gross margin, implies ~$600B/year of end-user AI revenue is needed to justify the buildout — a gap Sequoia sized at $500B in mid-2024 and growing. NVIDIA’s own gross margin (~71% GAAP for FY2026 on $215.9B revenue) is the top of the pyramid; the factory operators are the squeezed middle; and the inference providers on top live on cents per million tokens (Together AI lists frontier-class output at $1.20–$15.00 per 1M tokens, with cached-input discounts; Lambda rents B200s at $9.86/GPU-hour falling to $8.87 at scale). The chain is thin in the middle — which is precisely where an AI factory builder operates.
Theme six — Strategy: should an AI factory builder move into inference?
The strategic question has a clear answer in the evidence: do not sell inference-as-a-service; do become inference-workload-aware as infrastructure.
The competitive case is structural, not just precautionary. An AI factory builder’s customers are inference providers, and those relationships are large, multi-year, and pre-committed — the CoreWeave pattern (98% committed revenue, OpenAI at ~$6.5B through 2031, Microsoft at 67% of revenue) is the industry’s dominant contracting model. Launching a token business would ask those same customers to fund a competitor for the thinnest-margin, most competitive layer in the chain, while destroying the one property that makes factory capacity valuable: fungibility. NVIDIA’s own articulation is the cleanest statement of the neutrality principle — “the factory can be used by another customer, another cloud or another operator” — and NVIDIA’s behavior is the template: it backs PORTS-Pike (4.25 GW initial, OpenAI as the tenant, ~$600B of compute through 2030 across OpenAI’s ~12–16 GW of commitments), it monetizes inference through silicon, systems, software (Dynamo, DGX Cloud, the DSX factory platform), and it never sells tokens — its technical blog’s “token factory” framing is explicitly about maximizing revenue per megawatt of infrastructure, not about selling tokens itself. The pure-play champion agrees: CoreWeave markets Mission Control for “provision infrastructure, schedule and manage workloads, and monitor performance across training and inference environments” — inference-awareness as infrastructure, no model-serving product line. Straddling exists (Together AI sells both token APIs and GPU clusters), but the straddlers serve developers, not competing inference providers; for a factory builder whose counterparties are inference companies, neutrality is the product.
The value-add case is that inference-workload-aware IaaS is a real, sellable product layer that does not cross the line. The systems literature defines the design space concretely: disaggregated prefill and decode with per-phase SLOs (TTFT/TPOT) 1011, MoE-aware serving that lifts GPU utilization 16, cluster-level energy management under SLOs (DynamoLLM conserves 53% energy and cuts customer cost 61% while meeting latency targets) 14, and the serving software that compounds efficiency 121327. Each of these changes factory architecture — separate prefill and decode pools, KV-cache-aware fabrics, power capping per workload class, burst absorption for agentic traffic, co-locating training and inference to smooth utilization — and each is an infrastructure attribute that can be engineered, guaranteed, and priced without the builder ever serving a token. The market is already monetizing fragments of it: CoreWeave sells scheduling and monitoring across training and inference; I/O Fund’s analysis of the neoclouds names “optimized compute utilization” as one of the three reasons hyperscalers commit billions to them — utilization-aware operation is the actual selling point of the pure plays; Together sells “token-based capacity with SLAs” at the serving layer; NVIDIA markets tokens-per-megawatt as a factory benchmark. An inference-workload-aware IaaS sits exactly between raw GPU hours and token APIs — capacity with workload-class-aware SLOs, energy guarantees, and telemetry — which is differentiation the commodity GPU-hour market does not provide and which complements rather than cannibalizes inference-provider customers, because it makes their factories more efficient at the thing they sell.
Where the evidence disagrees
The sources disagree on three axes, and the disagreements are definitional rather than factual. First, market size: Gartner’s $2.59T (procurement), IDC’s $487B (hardware shipments), Menlo’s $37B (enterprise GenAI purchases, explicitly excluding chips and inference serving), and Bloomberg Intelligence’s $1.8T-by-2032 (revenue potential) all measure different flows; quoting them side by side without the definitions is how “AI is a bubble” and “AI is underfunded” arguments both get made. Second, the training/inference crossover date: Gartner says inference has already passed training in AI-optimized IaaS (2026), Bloomberg Intelligence says 2029 for the wider market, and Epoch’s OpenAI datapoint shows training still dominant at the frontier labs (2024) — the resolution is that the crossover happens first in the commercial cloud, last among the labs that train frontier models. Third, fine-tuning’s importance: the methodology literature treats it as a first-class adaptation technique 1819, while the adoption evidence treats it as niche 20 and no market sizing exists — the two literatures are not actually in conflict, they are measuring capability vs deployment.
Gaps and open questions
The gaps that matter: (1) no standalone fine-tuning market size exists anywhere in the fetched evidence — an analyst that publishes one would be filling a real hole; (2) no public split of training compute by modality exists, and it is decision-relevant for power and silicon planning; (3) GPU utilization by workload class is not systematically measured — the numbers that circulate (30–60% training, 10–40% inference, and a 95%-wasted claim from VentureBeat) could not be traced to a verifiable primary source in this review and are deliberately excluded; (4) inference-vs-training revenue splits are reported by vendors with obvious interests (NVIDIA’s ~40%-inference share of datacenter revenue is press-reported, not in the fetched releases, and is flagged accordingly); (5) several widely quoted figures remain unverifiable even after a dedicated search wave: the Gartner “$112B agentic spend by 2028” figure, Deloitte’s agentic share-of-GenAI-spend claims, and coverage of CoreWeave’s reported 2026 price increase (Inc.com) are all behind bot-walls or absent from primary pages. What would settle these: a vendor-neutral training/inference/fine-tuning revenue taxonomy with published methodology; utilization telemetry disclosed by GPU cloud operators (CoreWeave’s 10-K is a start but reports no utilization); and analyst reconciliation of the four market-size definitions.
Confidence and limitations
Confidence is moderate. The mechanism layer is high-confidence — the systems and methods papers are peer-reviewed or widely replicated, and every load-bearing number in the evidence table was grep-verified against fetched full texts this session. The market-data layer is moderate-confidence: it rests on a handful of analyst houses with different definitions, several key figures are press-reported or secondary (Gartner via trade press, NVIDIA’s inference share, the Stanford 2026 investment figure), and the most-cited long-run numbers (BI 2032, Goldman 2030 tokens) are model outputs, not observations. The strategic layer is the most interpretive: the neutrality argument is grounded in primary documents (NVIDIA blogs, CoreWeave 10-K, Sequoia) but the synthesis is this review’s own. Limitations: four remembered landmark titles could not be resolved and were dropped; direct search engines (Google, Bing, DuckDuckGo, Mojeek, Startpage) bot-blocked the research host, so the search wave relied on Google News RSS, publisher WordPress APIs, and direct PDF retrieval — some items that likely exist (VentureBeat’s GPU-utilization analysis, Inc.com’s CoreWeave pricing story, the Gartner and Deloitte agentic-spend figures) could not be read and are listed as gaps rather than cited; SemiAnalysis’s detailed inference-economics models are paywalled and represented only through NVIDIA’s own reporting of the InferenceMAX benchmark and its public article titles; Anthropic’s “State of AI Infrastructure” report could not be retrieved (the URL no longer resolves and searches surface no canonical version); CoreWeave margins are presented as computed from the 10-K income statement, not as reported margins; and the review is a snapshot as of 2026-08-19 — the crossover numbers will keep moving.
Evidence table
| key | year | design | sample | measure | finding | limitations | confidence | |
|---|---|---|---|---|---|---|---|---|
| sevilla2022compute | 2022 | empirical-trend-analysis | ML training runs 1950-2022 | training compute growth | Training compute doubled roughly every 20 months before 2010 (Moore's-law pace) and accelerated sharply after the advent of deep learning | Era boundaries debatable; no post-2010 growth rate stated in abstract | high | |
| hoffmann2022training | 2022 | scaling-law analysis | 400+ language models 70M-16B+ params | compute-optimal model size vs tokens | Current large language models are significantly undertrained; establishes the compute-optimal model-size/data trade-off (Chinchilla) | Fitted on text-LM loss; excludes multimodal/reasoning | high | |
| chowdhery2022palm | 2022 | systems report | 540B-param model on 6144 TPUv4 chips | model FLOPs utilization (MFU) | Achieved 46.2% MFU training PaLM 540B across two TPU v4 pods; MFU defined as observed vs theoretical max throughput | Single-vendor hardware; TPU-specific | high | |
| villalobos2022will | 2022 | forecast | public human-generated text corpora | data-exhaustion date | At current trends models will train on datasets roughly equal to the stock of public human text between 2026 and 2032, earlier if overtrained | Forecast sensitive to efficiency gains and data definitions | moderate | |
| muennighoff2023scaling | 2023 | empirical | data-constrained LM training runs | loss vs repeated epochs | With constrained data and fixed compute, training on up to 4 epochs of repeated data yields negligible loss change vs unique data; beyond that gains vanish | Limited to small model scales | moderate | |
| maslej2024aiindex | 2024 | industry report | global AI ecosystem 2023 | tracked trends | Annual report broadening coverage of technical advancement, public perception, and geopolitical dynamics of AI; a canonical industry-statistics reference | Annual snapshot; definitions shift between editions | high | |
| sajadieh2026aiindex | 2026 | industry report | global AI ecosystem 2025 | tracked trends | US private AI investment reached $285.9B in 2025 (over 23x China's $12.4B); global corporate AI investment more than doubled with private investment up 127.5%; AI data-center power capacity rose to 29.6 GW; report cites estimates that energy to serve inference queries can exceed one-time training cost within months; open-source downloads shifted to text generation (42% of 2025 downloads) while classifier/audio models collapsed from ~70% (2022) to under 6% | Full text read; chart labels partially garbled in text extraction | high | |
| deepseekai2024deepseekv3 | 2024 | technical report | 671B-param MoE model | training cost and inference efficiency | DeepSeek-V3 (671B total, 37B active) required only 2.788M H800 GPU-hours for full training; MLA + DeepSeekMoE for efficient inference | Vendor-reported; single run | high | |
| guo2025deepseek | 2025 | peer-reviewed report | DeepSeek-R1 reasoning model | reasoning acquisition | Shows reasoning abilities can be incentivized through pure reinforcement learning without human-annotated demonstrations | Nature publication of company report | high | |
| patel2023splitwise | 2023 | systems paper | LLM inference at Microsoft | phase-separated cost/energy | Inference has two phases — compute-intensive prompt processing and memory-intensive token generation; token generation underutilizes compute resources even with state-of-the-art batching | State-transfer overhead between phases | high | |
| zhong2024distserve | 2024 | systems paper | LLM serving with latency SLOs | TTFT/TPOT goodput | Colocating prefill and decoding causes strong prefill-decoding interference and couples resource allocation; per-phase latency (TTFT, TPOT) drives the case for disaggregation | Prototype validation | high | |
| kwon2023efficient | 2023 | systems paper | vLLM serving system | serving throughput | KV-cache memory per request is huge and dynamic; fragmentation and duplication waste it and limit batch size — PagedAttention manages it in pages for high-throughput serving | Benchmark workloads | high | |
| zheng2023sglang | 2023 | systems paper | SGLang runtime | throughput/latency | SGLang with RadixAttention (KV reuse) and compressed FSM decoding achieves up to 6.4x higher throughput vs state-of-the-art systems on LLM and multimodal workloads | Specific benchmark suite | high | |
| stojkovic2025dynamollm | 2025 | systems paper | inference clusters under SLOs | energy and cost | DynamoLLM, an energy-management framework for inference clusters, conserves 53% energy and 38% operational carbon and cuts customer cost 61% while meeting latency SLOs | Simulation-scale validation | moderate | |
| jiang2024megascal | 2024 | systems paper | 10k+ GPU training | training efficiency and stability | MegaScale achieves 55.2% MFU training a 175B model on 12,288 GPUs (1.34x vs Megatron-LM); fault tolerance and stragglers dominate operational effort | Single workload; operator-specific | high | |
| zhu2025megascal | 2025 | systems paper | MoE serving | GPU utilization and throughput | MoE sparsity makes FFNs memory-intensive during inference, lowering GPU utilization; disaggregating attention/FFN with ping-pong pipeline achieves up to 1.90x per-GPU throughput | Prototype; MoE-specific | moderate | |
| luccioni2023power | 2023 | empirical | ML systems categories | energy and carbon per 1 | 000 inferences | First systematic comparison of ongoing inference cost across task-specific and general-purpose model categories, measured as energy and carbon per 1,000 inferences | Hardware/regional energy mixes vary | moderate |
| hu2021lora | 2021 | methods paper | LoRA fine-tuning | adaptation cost | LoRA reduces trainable parameters by 10,000x and GPU memory by 3x vs full fine-tuning of GPT-3 175B with Adam, on par or better quality | Quality tradeoffs on some tasks | high | |
| dettmers2023qlora | 2023 | methods paper | QLoRA fine-tuning | memory/efficiency | QLoRA fine-tunes a 65B-parameter model on a single 48GB GPU while preserving full 16-bit fine-tuning task performance | Benchmark-specific | high | |
| biderman2024lora | 2024 | empirical study | LoRA vs full FT on code/math | learning and forgetting | In standard low-rank settings LoRA substantially underperforms full fine-tuning in-domain but better maintains out-of-domain performance and mitigates forgetting | Two domains; scale-limited | high | |
| strubell2019energy | 2019 | empirical | NLP model training runs | CO2 emissions and cost | Reports up to 626,155 lbs CO2e for a neural-architecture-search training run — orders of magnitude above a single model training run | Pre-LLM-era hardware and methodology | high | |
| fernandez2025energy | 2025 | empirical | LLM inference workloads | energy vs efficiency techniques | Systematically analyzes the energy implications of common inference efficiency techniques across diverse real-world inference workloads, beyond idealized latency benchmarks | Workload definitions vary | moderate | |
| yao2022react | 2022 | methods paper | ReAct agents | reasoning+acting performance | Generating both reasoning traces and task-specific actions improves over reasoning-only and acting-only approaches — the canonical agentic pattern | Benchmark scope | high | |
| schick2023toolformer | 2023 | methods paper | tool-use training | API-call decisions | Language models can teach themselves to use external tools via simple APIs, adding tool-call round-trips to inference demand | Limited tool set | high | |
| wang2024survey | 2024 | survey | agent architectures | taxonomy | Survey of LLM-based autonomous agents covering knowledge acquisition, planning, and tool use — agents are a distinct workload class | Survey scope cutoff | high | |
| team2023gemini | 2023 | technical report | Gemini model family | multimodal capability | Gemini family (Ultra, Pro, Nano) is natively multimodal across image, audio, video, and text — frontier models are multimodal-first | Vendor report; limited technical detail | high | |
| zhou2024survey | 2024 | survey | inference-efficiency techniques | taxonomy | Survey of techniques addressing the substantial computational and memory requirements of LLM inference — quantization, pruning, distillation, serving | Rapidly evolving field | high | |
| erdil2023explosive | 2023 | argument review | explosive-growth literature | critical appraisal | Explosive growth from AI automation rests on three drivers (scalable AI labor force, rapid expansion, massive investment) whose strength is uncertain | Review of arguments | not new data | high |
Swipe sideways to see all columns.
References
- (2022). Compute Trends Across Three Eras of Machine Learning — International Joint Conference on Neural Networks (IJCNN), arXiv:2202.05924. Abstract only. Training-compute growth across three eras of ML; ~5x/year growth in the deep-learning era — the demand baseline for training.doi:10.48550/arxiv.2202.05924
- (2022). Training Compute-Optimal Large Language Models — arXiv:2203.15556 (Chinchilla). Abstract only. Compute-optimal scaling law for training; defines how much training compute the frontier consumes for a given model size.doi:10.48550/arxiv.2203.15556
- (2022). PaLM: Scaling Language Modeling with Pathways — arXiv:2204.02311. Abstract only. Introduces model FLOPs utilization (MFU) as the standard training-efficiency metric — the anchor for utilization claims.doi:10.48550/arxiv.2204.02311
- (2022). Will we run out of data? Limits of LLM scaling based on human-generated data — arXiv:2211.04325 (Epoch AI). Abstract only. High-quality text data could run out by 2026 — the supply-side constraint on training scaling that pushes value toward inference.doi:10.48550/arxiv.2211.04325
- (2023). Scaling Data-Constrained Language Models — arXiv:2305.16264. Abstract only. Training under data constraints: multiple epochs help up to a point — relevant to how training demand responds to data scarcity.doi:10.48550/arxiv.2305.16264
- (2024). Artificial Intelligence Index Report 2024 — arXiv:2405.19522 (Stanford HAI). Abstract only. Stanford AI Index anchor: industry investment, model counts, and compute-trend statistics for 2023-2024.doi:10.48550/arxiv.2405.19522
- (2026). Artificial Intelligence Index Report 2026 — arXiv:2606.15708 (Stanford HAI). Full text read. AI Index 2026: US private AI investment $285.9B in 2025; AI data-center power capacity 29.6 GW; inference energy exceeding training within months; open-source download mix by modality.doi:10.48550/arxiv.2606.15708
- (2024). DeepSeek-V3 Technical Report — arXiv:2412.19437. Abstract only. Frontier MoE trained at a fraction of typical cost; open-weights release compressed inference prices — a demand-side shock.doi:10.48550/arxiv.2412.19437
- (2025). DeepSeek-R1 incentivizes reasoning in LLMs through reinforcement learning — Nature. Abstract only. Reasoning models emit long chain-of-thought outputs — the mechanism that multiplies tokens per request and inflates inference demand.doi:10.1038/s41586-025-09422-z
- (2023). Splitwise: Efficient generative LLM inference using phase splitting — arXiv:2311.18677 (Microsoft Research). Abstract only. Prefill/decode phase splitting: decode underutilizes compute and can run on cheaper hardware — the basis of disaggregated inference.doi:10.48550/arxiv.2311.18677
- (2024). DistServe: Disaggregating Prefill and Decoding for Goodput-optimized Large Language Model Serving — arXiv:2401.09670 (OSDI 2024). Abstract only. Colocated prefill/decode interferes and couples resource allocation; disaggregation needed to meet TTFT/TPOT SLOs per phase.doi:10.48550/arxiv.2401.09670
- (2023). Efficient Memory Management for Large Language Model Serving with PagedAttention — ACM SOSP 2023. Abstract only. PagedAttention / vLLM: KV-cache memory management raises serving throughput — a step-change in inference efficiency.doi:10.1145/3600006.3613165
- (2023). SGLang: Efficient Execution of Structured Language Model Programs — arXiv:2312.07104. Abstract only. RadixAttention and structured generation: serving-system innovation that cuts inference cost for complex multi-turn/agentic calls.doi:10.48550/arxiv.2312.07104
- (2025). DynamoLLM: Designing LLM Inference Clusters for Performance and Energy Efficiency — 2025 IEEE International Symposium on High Performance Computer Architecture (HPCA). Abstract only. Cluster-level design for inference: energy-efficient inference clusters — how infra design changes when inference dominates.doi:10.1109/hpca61900.2025.00102
- (2024). MegaScale: Scaling Large Language Model Training to More Than 10,000 GPUs — arXiv:2402.15627. Abstract only. Training at 10k+ GPUs: the operational difficulty of keeping utilization high at scale — context for training-fleet economics.doi:10.48550/arxiv.2402.15627
- (2025). MegaScale-Infer: Efficient Mixture-of-Experts Model Serving with Disaggregated Expert Parallelism — Proceedings of the ACM SIGCOMM 2025 Conference. Abstract only. Disaggregated expert parallelism for MoE serving — disaggregation is moving from research to production serving architectures.doi:10.1145/3718958.3750506
- (2023). Power Hungry Processing: Watts Driving the Cost of AI Deployment? — arXiv:2311.16863. Abstract only. Inference energy cost across 88 models: per-request energy is driven by model size — energy economics of the inference layer.doi:10.48550/arxiv.2311.16863
- (2021). LoRA: Low-Rank Adaptation of Large Language Models — arXiv:2106.09685. Abstract only. Parameter-efficient fine-tuning: the technique that made fine-tuning cheap enough to be a real (if niche) workload.doi:10.48550/arxiv.2106.09685
- (2023). QLoRA: Efficient Finetuning of Quantized LLMs — Advances in Neural Information Processing Systems 36. Abstract only. 65B-parameter fine-tuning on a single 48GB GPU — fine-tuning's compute cost collapsed, which shapes its (small) market share.doi:10.52202/075280-0441
- (2024). LoRA Learns Less and Forgets Less — arXiv:2405.09673. Abstract only. Full fine-tuning learns more but forgets more than LoRA — explains why enterprises often prefer RAG/prompting over fine-tuning.doi:10.48550/arxiv.2405.09673
- (2019). Energy and Policy Considerations for Deep Learning in NLP — Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics. Abstract only. Early quantification of training energy cost — the origin of energy-economics analysis that now applies to inference at scale.doi:10.18653/v1/p19-1355
- (2025). Energy Considerations of Large Language Model Inference and Efficiency Optimizations — Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics. Abstract only. Inference energy is now the dominant and growing energy cost of LLMs — the energy rationale for inference-optimized infrastructure.doi:10.18653/v1/2025.acl-long.1563
- (2022). ReAct: Synergizing Reasoning and Acting in Language Models — arXiv:2210.03629. Abstract only. Reasoning + acting loop: the canonical agentic pattern that turns one user request into many model calls.doi:10.48550/arxiv.2210.03629
- (2023). Toolformer: Language Models Can Teach Themselves to Use Tools — Advances in Neural Information Processing Systems 36. Abstract only. Tool use by language models: tool calls add inference round-trips — the structural driver of agentic token demand.doi:10.52202/075280-2997
- (2024). A survey on large language model based autonomous agents — Frontiers of Computer Science. Abstract only. Survey of autonomous agent architectures — frames the workload class that is reshaping inference demand.doi:10.1007/s11704-024-40231-1
- (2023). Gemini: A Family of Highly Capable Multimodal Models — arXiv:2312.11805. Abstract only. Natively multimodal frontier model family — evidence that frontier training is multimodal-first, not text-only.doi:10.48550/arxiv.2312.11805
- (2024). A Survey on Efficient Inference for Large Language Models — arXiv:2404.14294. Abstract only. Taxonomy of inference-efficiency techniques (quantization, pruning, distillation, serving) — the software lever on inference cost.doi:10.48550/arxiv.2404.14294
- (2023). Explosive growth from AI automation: A review of the arguments — arXiv:2309.11690 (Epoch AI). Abstract only. Critical review of explosive-growth arguments — the demand-side debate underlying long-run compute forecasts.doi:10.48550/arxiv.2309.11690