---
title: "Selling infrastructure to inference providers: what a neocloud can offer without competing"
slug: inference-neocloud-strategy
question: "Which inference providers should an AI-factory builder target, at which infrastructure layer, and what must it know about inference to sell them capacity without becoming their competitor?"
status: published
depth: standard
created: 2026-08-23
updated: 2026-08-23
summary: "The inference-provider market splits into open-weight GPU hosts (Together, Fireworks, DeepInfra, Baseten, Modal — they rent nearly all their capacity and compete on serving software) and closed-weight labs (OpenAI, Anthropic, Google, Meta, xAI — they rent at enormous scale via take-or-pay contracts but vertically integrate software and increasingly silicon). The evidence says a neocloud should target the open-weight hosts at the fleet/facility and control-plane layers — power, cooling, density, grid, orchestration that respects the customer's serving engine — and never compete at the serving-engine layer, which is the customer's moat. Firmus's public record (AI FactoryOS, HyperCube, Model-to-Grid, a Fireworks partnership) already points this way. Confidence is moderate: the market facts are industry-reported, and several claimed product names could not be verified in any public source."
disciplines: ["industry economics", "AI infrastructure"]
tags: ["inference providers", "neocloud", "GPU cloud", "AI factories", "LLM inference", "take-or-pay", "tokens per watt"]
source_count: 26
year_range: [2022, 2026]
confidence: moderate
search:
  databases: ["arXiv", "Crossref records reused from prior reviews on this site", "primary web sources (SEC EDGAR, firmus.co newsroom, vendor pricing pages, TechCrunch, Menlo Ventures, a16z, SiliconANGLE, CoreWeave press)"]
  queries: ["Google News RSS: Firmus superconducting", "Google News RSS: Firmus FORGE", "Google News RSS: Firmus ARC AI", "Google News RSS: Firmus tokens per watt", "Google News RSS: Firmus AI FactoryOS", "Google News RSS: Oracle OpenAI OCI", "Google News RSS: Anthropic CoreWeave deal", "Google News RSS: Stargate 500 billion", "Google News RSS: Groq pivot neocloud", "Google News RSS: Modal Oracle Cloud Infrastructure", "Google News RSS: Together AI Series C", "Google News RSS: Fireworks AI Series D", "Google News RSS: DeepInfra Series B", "Google News RSS: OpenRouter Stripe acquisition", "Google News RSS: Meta CoreWeave 21 billion", "Google News RSS: DeepSeek H800", "Google News RSS: OpenAI 300 billion Oracle", "Google News RSS: Anthropic SpaceX Colossus", "arXiv abs pages: 2311.18677 2401.09670 2405.10955 2408.00741 2504.02263 2407.00079 2402.16363 2608.03741 2508.01989 2512.03416 2311.16863 2309.06180 2312.07104 2606.15708 2202.05924 2203.15556 2305.05920 2311.03285 2404.02015 2404.14294 2407.12391 2405.05465 2403.02310 2412.19437", "primary fetches: CoreWeave 10-K FY2025 (SEC EDGAR), CoreWeave S-1 (SEC EDGAR), firmus.co newsroom releases (2B raise, Fireworks partnership, hyperscale customer, Batam 170k GPU, Nan Hu Model-to-Grid interview), Oracle OCI GPU page, AWS P6e page, OpenAI Stargate press release (Wayback), TechCrunch (Together, Groq, OpenAI-CoreWeave, Anthropic-SpaceX, OpenAI-AWS), Menlo Ventures (Fireworks, OpenRouter), a16z State of AI, SiliconANGLE (Modal x2), CoreWeave-Anthropic press release, Lambda/Nebius/Crusoe/RunPod pricing pages, ACL Anthology 2025.acl-long.1563"]
  last_run: 2026-08-23
---


## Summary

The market for AI infrastructure is rotating from training to inference, and an infrastructure provider that wants to serve inference companies must first decide which customers it means and which layer of the stack it sells. The evidence collected here splits inference providers into two segments with opposite behavior. Open-weight GPU hosts — Together AI, Fireworks, DeepInfra, Baseten, Modal, Groq — rent nearly all their compute, compete on serving software and kernels, and carry thin margins under constant price pressure. Closed-weight labs — OpenAI, Anthropic, Google, Meta, xAI — rent at extraordinary scale through take-or-pay contracts (CoreWeave's 96% of revenue is committed) but vertically integrate everything they can, including increasingly their own silicon. The layer map shows a durable division: the serving engine is the customer's moat (Fireworks' FireAttention, Together's kernel work, the vLLM/SGLang ecosystem), the control plane is contested but commoditizing, and the fleet/facility layer — power, cooling, density, grid — is where a neocloud can differentiate without competing. The strongest single piece of evidence is that Firmus already has exactly this customer: Fireworks announced a partnership to deploy on Firmus AI Factory infrastructure in 2026. Confidence is moderate: the market facts are industry-reported and several are single-source; two product names in the question's framing (FORGE, ARC) and the "superconducting power" claim could not be verified in any public source and are treated as gaps, not facts.

## Why this question

An AI-factory builder deciding what to sell inference providers is making a channel-conflict decision: its prospective customers are companies whose entire business is serving tokens, and the naive product — "we also serve tokens" — competes head-on with them. The prior review of [LLM compute market economics](/reviews/llm-compute-market-economics/) concluded from market data that a factory builder should not sell inference against its own customers but should sell inference-workload-aware infrastructure. This review asks the follow-up: which customers, at which layer, and what must the company know about inference to be credible selling it?

The question matters because tens of billions of dollars of capacity are being contracted on the assumption that inference demand will dwarf training demand — [18] documents the industry investment wave, and [17] shows how compute demand curves steepened after 2010; the same curve now applies to serving. Getting the customer segment and the layer wrong is not a neutral mistake: building a serving engine puts you in direct competition with your own capacity buyers, while building only generic GPU racks invites a price war against Lambda and the hyperscalers. The answer determines architecture, power contracts, network topology, and who signs the take-or-pay.

## Scope and methods

The question has three halves: a market-structure half (who the inference providers are and what they buy), a technology half (what inference workloads need from infrastructure), and a strategy half (where a neocloud can sit without competing). Because the market evidence is industry-produced — SEC filings, company releases, investor letters, pricing pages — this review uses a two-layer evidence model. Layer 1: 26 peer-reviewed or preprint sources (all DOI-verified, all retrieved and read at abstract level this session) anchor the mechanisms: serving systems, disaggregation, KV-cache management, workload characterization, energy measurement, scaling. Layer 2: analyst and vendor documents are cited inline with URLs in the body; they carry no DOIs and so cannot be reference-list entries, but every load-bearing number attributed to them was verified this session by fetching the document and grepping the figure (CoreWeave 10-K/S-1, firmus.co releases, TechCrunch articles, Menlo Ventures and a16z posts, SiliconANGLE, Oracle and AWS product pages, vendor pricing pages). Screening: Layer-1 records were reused from three prior reviews on this site (their DOIs already verified); no new academic search wave was run because the mechanism layer for this question was already established there — see the [inference optimization](/reviews/llm-inference-optimization/), [tokens-per-watt](/reviews/tokens-per-watt-benchmarks/), and [interconnects](/reviews/llm-interconnects/) reviews. Industry-side: ~45 primary documents were fetched; figures that could not be verified in-session are flagged or excluded (e.g. a claimed "superconducting power delivery" technology, the product names FORGE and ARC, and several paywalled figures). The review deliberately does not cover chip design, application-layer economics, or regions beyond the sources report.

## The landscape

The literature on inference infrastructure is young but already splits cleanly into two bodies. The serving-systems literature (vLLM [1], Splitwise [2], DistServe [3], Sarathi-Serve [26], SGLang [4], Mooncake [7], MegaScale-Infer [5], the roofline survey [8]) is rigorous about what inference *costs* to operate and what it *needs* from hardware — and is essentially silent on who buys what from whom. The industry record — SEC filings, funding rounds, partnership announcements, pricing pages — is rich on market structure but unverifiable through scholarly channels. This review is an attempt to weld the two: the systems literature tells us what to build, the industry record tells us who will pay for it and what they will not pay for. Two absences structure the whole review: nobody discloses an owned-versus-rented split for inference specifically (deals cover "compute" generally), and nobody publishes a neocloud's inference-specific margin. Both absences are findings, discussed below.

## Theme one — The customer map splits the market in two

The most important structural fact is that "inference providers" is two different businesses. Open-weight GPU hosts — Together AI, Fireworks, DeepInfra, Baseten, Modal, Groq — are themselves neocloud customers. Together AI "rents out Nvidia GPU clusters and other AI-specific infrastructure" while raising an $800M Series C at an $8.3B valuation with annual bookings over $1.15B (TechCrunch, 2026-07-01, fetched). Fireworks serves 43 trillion tokens per day (up from 15T) at roughly $1B annualized revenue (Menlo Ventures, 2026-07-16, fetched), and its differentiation is a proprietary optimization stack, FireAttention — software, not hardware. Modal runs its entire platform on rented Oracle Cloud GPUs with no data centers of its own, and its founders describe the pain of "GPU capacity management, highly variable demand that makes it expensive to reserve instances" (SiliconANGLE, 2025-09-29, fetched). DeepInfra raised $107M to keep doing what it does: near-cost inference of open-weight models on rented capacity. Groq — originally a chip company — pivoted to buying NVIDIA GPUs and building "the world's leading AI inference cloud," scaling from 54 to over 200 MW in 2027 (TechCrunch, 2026-08-17, fetched), which is the clearest possible statement that even differentiated serving companies conclude they must rent NVIDIA fleets. OpenRouter sits one level up: a pure gateway with zero GPUs, routing to 400+ models, at a 4.5+ quadrillion-token annual run rate (Menlo Ventures, fetched), being acquired by Stripe for a reported $7B+ — evidence that the aggregation layer monetizes without owning any capacity at all.

Closed-weight labs are the mirror image. They rent, but at hyperscale, from many providers at once, on take-or-pay terms, and they vertically integrate software (and increasingly silicon) around the rented capacity. OpenAI rents from Microsoft Azure (exclusive for stateless APIs until the 2025 renegotiation), AWS ($38B/7yr announced Nov 2025), Oracle (a five-year $300B deal starting 2027, after an earlier $30B), CoreWeave ($11.9B/5yr in March 2025, later expanded) and has effectively retreated from building its own Stargate data centers, now "prefer[ring] to lease compute" (TechCrunch and Tom's Hardware coverage, fetched). Anthropic rents from AWS, Google Cloud (a TPU deal expanded to 3.5 GW), Microsoft Azure ($30B/5yr), CoreWeave (multi-year, April 2026), and even xAI's Colossus 1 — $1.25B per month through 2029 for all its available compute, which one report says is used for inference rather than training because the architecture cannot train Grok (TechCrunch, 2026-06-05, fetched). Google, the most vertically integrated of all, still rents "bridge capacity" from SpaceX at $920M/month for ~110,000 NVIDIA GPUs (TechCrunch, fetched). Meta self-builds at massive scale ($115-135B capex in 2026) yet still committed $14.2B and then another $21B to CoreWeave. DeepSeek is the exception that proves the rule: it owns its 2,048 H800s and trained V3 for 2.788M H800 GPU-hours [15] — a self-supply model only possible where export controls forced it.

The strategic reading: for a neocloud at Firmus's scale (hundreds of MW), the open-weight hosts are the addressable, repeatable customer segment — they need exactly what a factory builder sells, they cannot build it themselves without becoming a data-center company, and they are numerous enough that no single anchor contract governs the business. The closed labs buy at 1 GW+ scale with 5-year take-or-pay and will happily rent your fleet — but only when you are one of their three or four suppliers, and they will never let you near the software layer. The evidence also suggests the two segments converge: every closed lab now behaves like a multi-cloud renter, and every open-weight host that scales becomes a fleet buyer. That convergence is the market Firmus actually serves.

## Theme two — The layer map: hardware, control plane, serving engine

The stack separates into three sellable layers, and the pricing evidence shows what each is worth. Hardware/fleet: raw GPU capacity by the hour. Lambda sells B200 cluster nodes at $9.86/GPU-hr (16-GPU) down to $5.54 at 256+ GPUs, and H100 at $6.16/$5.54; on-demand instances run $6.69 (B200) and $3.99 (H100) (lambda.ai/pricing, fetched). Nebius lists on-demand B200 at $7.15/GPU-hr ($3.95 preemptible) and H100 at $3.85 ($2.15) with commitment discounts up to 35% (nebius.com/prices, fetched). Crusoe lists H200 $4.29, H100 $3.90, MI300X $3.45 on-demand with spot at roughly half (crusoe.ai/cloud/pricing, fetched). RunPod sells per-second serverless GPU at H200 $4.59/hr and B300 $7.89/hr (runpod.io/pricing, fetched). The hyperscalers price higher: Azure ND H100 v5 at $98.32/hr for an 8-GPU node (~$12.3/GPU-hr), ND GB200 v6 at $108.16 (Azure Retail Prices API, fetched); GCP a4-highgpu-8g (B200) at $64.44/hr (~$8.06/GPU-hr) (cloud.google.com, fetched); AWS markets P6e GB200 NVL72 UltraServers — 72 GPUs in one NVLink domain, 360 petaflops FP8, 13.4 TB HBM — explicitly for "real-time inference in production" (aws.amazon.com/ec2/instance-types/p6, fetched). Oracle's Supercluster scales to 131,072 Blackwell GPUs with 3,200 Gb/s RDMA and bare-metal instances free of virtualization overhead (oracle.com/cloud/compute/gpu, fetched).

Control plane: orchestration, scheduling, autoscaling. CoreWeave ships its own Kubernetes distribution (SUNK) as part of the product (10-K, fetched); Nebius gives managed Kubernetes away free and offers Slurm-on-Kubernetes; RunPod and Modal sell serverless scheduling as the product; Crusoe charges $0.10/cluster-hr for managed K8s. This layer is real but commoditizing — it is the price of entry, not the moat. Serving engine: per-token inference APIs with engine-level optimization — Fireworks' FireAttention, Together's Inference Engine, DeepInfra, Nebius's Token Factory, Crusoe's Serverless Inference (per-1M-token pricing, fetched). This is the layer the customers themselves compete in, and the layer a capacity provider must never enter: it is the customer's product. The systems literature explains why the layers are separable: the serving engine's value is in batching, paging, and scheduling decisions [1][4][21][22][23] — decisions that sit above the hardware and are portable across fleets, which is precisely why Fireworks can deploy on a rented factory and keep its differentiation [25][24].

## Theme three — What inference needs that training does not

The systems literature is unusually consistent on what inference demands from the layer below it, and it differs from training in five ways that matter for facility design.

First, inference has two phases with opposite resource profiles. Prefill is compute-bound and bursty; decode is memory-bound and steady. Splitwise [2] established the phase asymmetry, DistServe [3] showed that colocating the phases hurts SLO attainment, and Sarathi-Serve [26] showed chunked prefill can smooth the interference. The cluster-level consequence — now a live design debate — is whether to physically separate prefill and decode pools (disaggregation), which Mooncake operates in production at Moonshot [7] and MegaScale-Infer extends to MoE models [5]; the newest work shows the answer is workload-dependent, with agentic traffic changing the mix [10][11]. A fleet operator who understands this sells differently: disaggregated pools, RDMA fabrics sized for KV-cache transfer between them, and power envelopes that assume prefill bursts.

Second, inference is memory-bound. Decode throughput is set by HBM bandwidth and KV-cache capacity, not FLOPS — the roofline framing in [8]. Every generation of GPU that adds HBM per GPU (H100 80GB, B200 ~192GB, B300 288GB) is partly an inference play, and RunPod's marketing of the B300 as "maximum throughput for big models" reflects it. Memory per GPU and memory bandwidth per dollar are the inference-relevant specs.

Third, inference traffic is bursty and agentic. BurstGPT [9] documents heavy burstiness in real serving traces; TokenScale [12] shows burstiness defeats lagging autoscaling; the a16z/OpenRouter study finds agentic inference is the fastest-growing behavior, with OpenAI's API alone averaging ~8.6 trillion tokens/day in October 2025 (a16z.com/state-of-ai, fetched). Burst absorption — idle-but-paid capacity that can absorb spikes — is a product attribute for a fleet provider, and it interacts with the energy story: bursty load is what makes grid-responsive operation valuable.

Fourth, latency is geographic. Training runs anywhere cheap; inference SLOs (time-to-first-token, inter-token latency) degrade with distance. The closed labs rent capacity near users — the SpaceX deal that Anthropic uses for inference is in Tennessee, near the US population center — and a factory in Tasmania or Indonesia serves a different latency market than one in Virginia. The evidence for this is indirect (deal geography) rather than measured, and the review flags it as such.

Fifth, model churn is extreme. New open-weight models land weekly; OpenRouter's Menlo post describes the release pace as "frenetic." A fleet provider's redeploy velocity — image distribution, model download, GPU reallocation — is an operational requirement that training-centric clouds never had, and it is one of the reasons the control-plane layer matters even though it commoditizes: bad orchestration erases the fleet advantage. Vidur [20] and DynamoLLM [6] show cluster-level planning and power management are now design problems with measurable energy tradeoffs.

## Theme four — The economics of selling capacity to inference providers

The CoreWeave filings are the best public window into the neocloud business model, and they are unambiguous about its shape. The S-1 (filed 2025-02-28, fetched) shows revenue $15.8M (2022) → $228.9M (2023) → $1,915.4M (2024), gross margin ~74%, 96% of 2024 revenue under committed take-or-pay contracts, $15.1B of remaining performance obligations, and a stated lease-first strategy: lease data centers to trade capex for speed, with ownership deferred. The FY2025 10-K (filed 2026-03-02, fetched) shows the model scaling and its strain: revenue $5,131M, gross margin ~71.7%, net loss $(1,167)M on $1,229M of interest expense, total debt $21,373M, and 67% of revenue from a single customer (Microsoft). NVIDIA sits on all sides of the deal — supplier of every GPU, a >5% shareholder, and itself a customer (~$320M paid by end-2024). The take-or-pay structure is the load-bearing mechanism: it converts a capacity guarantee into bankable revenue for the provider, which is what lets CoreWeave borrow $21B against it. For an inference provider, the same structure converts an uncertain demand curve into guaranteed capacity in a supply-constrained market — the deal the whole industry actually runs on.

The inference-specific economics sharpen the picture. Inference providers operate under continuous price compression — DeepSeek's low-cost training and API pricing reset market expectations [15], and the open-weight hosts pass savings through to win share. Their margin is the difference between what tokens sell for and what tokens cost to produce; token cost is dominated by power, hardware depreciation, and utilization. That is why power efficiency is not a sustainability talking point but a margin lever: energy per token varies by orders of magnitude across models and configurations [13][19], carbon and cost track the same curve [14], and cluster-level power management can cut energy while meeting SLOs [6]. A fleet provider that delivers more tokens per watt is literally improving its customers' gross margin on the same contract price — which is the only form of price cut that does not start a race to the bottom. The tension in the record is that the market has not yet proven the neocloud can convert growth into free cash flow — TechCrunch's Groq piece quotes investors doubting neocloud profitability under "high capital expenditures, heavy reliance on debt, and exposure to rapidly depreciating hardware" (fetched) — which is exactly why the inference-focused neocloud needs a differentiated cost structure, not just a balance sheet.

## Theme five — Why competent labs outsource, and what they outsource

The case evidence is consistent: the most capable AI organizations in the world outsource because capacity — specifically power and speed — is the binding constraint, and no self-build can match a partner whose construction is already in motion. OpenAI's arc is the cleanest experiment: the $500B Stargate first-party build was announced in January 2025 (OpenAI press release, fetched via archive), and by 2026 the company had walked it back toward leasing, with the same period seeing $38B to AWS, $300B to Oracle, $11.9B+ to CoreWeave, $100B of NVIDIA GPUs, and 6 GW from AMD — the $1T-in-deals pattern reported across TechCrunch's coverage (all fetched). Anthropic rents from four clouds plus a neocloud plus a competitor's supercomputer, and its Google TPU deal (3.5 GW, per the Broadcom filing reported by TechCrunch, fetched) shows even custom silicon is rented, not owned, when the alternative is waiting. Meta is the counter-case — the largest self-builder — and still spent $35B+ with CoreWeave for speed. xAI built Colossus in 122 days (100k H100s) and then became a compute *seller* — the endpoint of the arc is that even the self-builder monetizes surplus capacity [17]. The demand context for why the market is this hot: serving demand compounds with every deployed model, and the compute-optimal allocation point [16] determines how much capacity each model generation needs.

What they outsource specifically: GPU compute at cluster scale (training and increasingly inference — the Anthropic/Colossus-1 rental is the clearest inference datapoint), data-center real estate and power procurement (CoreWeave's 1.3 GW contracted power; Stargate's land-and-power RFP), and the debt-carrying (the take-or-pay contract moves the financing risk to the provider). What they never outsource: serving software, and increasingly silicon design. For a neocloud, the lesson is the mirror image of Theme one: the durable product is the unbundled facility — power, cooling, density, speed-to-capacity — sold on take-or-pay, with software as the enabler that keeps the fleet efficient, not as the product.

## Theme six — Firmus today: what the public record shows

Firmus (Firmus Grid Limited, trading as Firmus Technologies; formerly Sustainable Metal Cloud; headquartered in Sydney) describes itself as "an AI Factory Platform company built on a simple idea: energy in and tokens out — as efficiently as possible," vertically integrated "from model to grid" (firmus.co, fetched). The public record is substantial for a company founded in 2019: equity of more than US$3B raised in the twelve months to August 2026 — including US$505M at a US$5.5B post-money valuation (April 2026, Coatue-led with NVIDIA), A$500M at ~A$6B (November 2025), and a fully-subscribed US$2B round at a post-money valuation above US$10.5B (August 2026, with NVIDIA, Coatue, Blackstone, Jane Street) — plus a reported US$10B Blackstone/Coatue debt facility (firmus.co releases; TechCrunch/Bloomberg coverage, fetched). Capacity: Project Southgate in Australia is a roadmap "up to 1.6 gigawatts through 2028"; Tasmania hosts ~36,800 GB300s completing late 2026; Melbourne a multi-billion-dollar contract for ~18,400 GB300s for an unnamed global hyperscaler; South Australia adds 600 MW of firm energy supply with 1.2 GW of new renewables; Batam, Indonesia, a 360 MW NVIDIA DSX campus with DayOne, up to 170,000 accelerators through 2027-28 under a partnership running to 2034 (all firmus.co releases, fetched). Technology: HyperCube modular liquid-cooled halls (32 NVL racks), immersion plus direct-to-chip cooling, AI FactoryOS ("GPU-to-energy grid management"), and the Model-to-Grid concept — treating AI workloads, hardware, and energy systems as one optimization problem, with head of AI & Apps Nan Hu describing the shift "from MaxP to MaxQ (maximising tokens per watt)" and a hybrid Slurm-plus-Kubernetes scheduling culture (firmus.co newsroom interview, fetched). Efficiency claims, all vendor-stated: pPUE ~1.02, facility PUE ~1.10 at its Singapore site, "approximately 30% better performance" per watt at node level versus air-cooled H100 in an MLPerf research submission, "45% better FLOP per utility picoJoule," and up to 50% lower build cost per MW (firmus.co research-lab and archived pages, fetched).

Two things the public record does not contain. First, the product names FORGE and ARC — presented in the question framing of this review as job-deployment optimization for tokens-per-watt — appear nowhere on firmus.co (current or archived), in the newsroom, or in press coverage; the verified platform names are AI FactoryOS and HyperCube, with "Model-to-Grid" as the umbrella concept. Second, "superconducting power delivery" appears nowhere either; the verified power story is rack-level electrical design without PDU overhead, liquid cooling, and grid-responsive operation. Both absences are reported here as gaps, not corrected to silence — an interview candidate repeating unverifiable company claims would be repeating something the company itself has not said publicly. What the record *does* show, decisively for this review's question, is the Fireworks partnership: "Fireworks will deploy on Firmus AI Factory infrastructure powered by NVIDIA... serving more than 40 trillion tokens each day" (firmus.co newsroom, fetched) — an open-weight GPU host as an anchor customer of a factory builder. The thesis of this review is not hypothetical; it is already the company's strategy.

## Theme seven — The product: what Firmus can sell inference providers without competing

Synthesizing the market structure, the layer map, and the workload evidence, the answer to the three-part question is:

**Which segment?** Open-weight GPU hosts first — Together, Fireworks, DeepInfra, Baseten, Modal, Groq, and the next cohort — because they rent nearly everything, cannot build factories without becoming data-center companies, and their software moats make them want *more* capacity, not less. Closed-weight labs are addressable later, at 100 MW+ scale, through take-or-pay contracts where the product is fleet plus power plus speed, not software; the Colossus-1-for-inference rental shows even they will rent for inference when the price and geography work. The split is not about open- vs closed-weights per se — it is about who buys unbundled capacity (hosts do, always) versus who buys integrated systems (labs do, only at hyperscale).

**Which layer?** The moat is the fleet/facility layer: power delivery, cooling, density, grid integration — the tokens-per-watt lever — because it is where the customer's margin is made, where hyperscalers and neoclouds are weakest per dollar, and where a factory builder's physics advantage (Model-to-Grid, liquid cooling, grid services) is un-replicable by a software company. The control plane (AI FactoryOS, K8s/Slurm scheduling) is the enabler that must be sold alongside, and it must be *neutral* — it must schedule the customer's own serving engine, not substitute for it; CoreWeave and Nebius already give orchestration away, so the control plane is table stakes. The serving engine is the one layer that is forbidden: building vLLM-class serving or a public inference API competes directly with Fireworks and Together — the very companies the fleet exists to serve — and the evidence says the engine is their moat, not a commodity you can win [1][4][25].

**The product, concretely:** (1) take-or-pay capacity contracts at GB300-class density with a tokens-per-watt guarantee, priced against the Lambda/Nebius/RunPod band but justified by lower energy cost per token; (2) burst absorption — on-demand headroom for agentic spikes, with predictive autoscaling across prefill/decode-shaped pools [9][12][10]; (3) a neutral control plane that runs the customer's serving stack (vLLM, SGLang, TensorRT-LLM, FireAttention) unchanged [4][8]; (4) geographic latency positioning — Southeast Asia and Australia for the APAC token market, which the hyperscalers under-serve; (5) grid value — demand response, firming, and low-carbon firm power as a cost advantage that compounds at scale [14]. The line the evidence draws is consistent: sell everything below the serving engine, nothing at or above it, and make the efficiency so good that the customer's gross margin is visibly better on your racks.

## Theme eight — What an inference-focused neocloud must know

To be the premier provider *for* inference providers, the knowledge stack the evidence implies is: the roofline model of inference (prefill compute-bound, decode memory-bound, KV-cache memory pressure) [8][1]; the serving stack — vLLM/SGLang/TensorRT-LLM, disaggregation, speculative decoding, quantization — enough to converse with the customer's engineers and to schedule around their engines [24][25][26]; workload characterization — burstiness, context-length distributions, agentic multi-turn growth [9][12][10]; SLO engineering — TTFT/inter-token-latency, goodput under constraints [3][21]; tokens-per-watt measurement at node, hall, and facility boundaries, because energy per token is the sales metric [13][19][6]; networking — NVLink domains for tensor parallelism, RDMA fabrics sized for KV-cache movement between prefill and decode pools [7][5]; power engineering — rack density, liquid cooling, grid interconnection timelines, firming and demand response [14]; and the commercial layer — take-or-pay structures, GPU-hour versus per-token pricing, SLAs, security certifications, and the neutrality rule: never compete with the customer's engine, never resell their tokens. The prior reviews on this site provide the deep versions of most of these: [inference optimization](/reviews/llm-inference-optimization/), [engine tricks](/reviews/llm-inference-engine-tricks/), [tokens-per-watt benchmarking](/reviews/tokens-per-watt-benchmarks/), [interconnects](/reviews/llm-interconnects/), and [datacenter energy efficiency](/reviews/model-to-grid-ai-datacenter-efficiency/).

## Where the evidence disagrees

Three tensions are real and unresolved. First, neocloud profitability: CoreWeave's 71.7% gross margin and $1.2B adjusted EBITDA sit next to a $(1,167)M net loss and $1,229M interest expense, and investors quoted in TechCrunch openly doubt the model converts growth into free cash flow. The disagreement is about whether gross margin survives the debt load and hardware depreciation — which is precisely why cost-per-watt structure, not markup, is the differentiator an inference-focused entrant must bring. Second, build versus rent at the top: OpenAI retreated from Stargate toward leasing while Meta expands self-built "Meta Compute" — both directions are observed simultaneously, and the best resolution is temporal: rent for speed, build for the steady state, and the rental market is therefore structural, not a bubble artifact. Third, the open-weight hosts' own model: Fireworks and Together raise at $17.5B and $8.3B valuations while renting their fleets, and Groq's pivot says the rental conclusion is universal — but whether renting-software-companies can defend margins against both their capacity suppliers and their API competitors is untested. None of these disagreements changes the core answer — sell below the serving engine — but they determine whether the buyer can pay.

## Gaps and open questions

The record is missing four things that would settle the strategy. (1) No lab discloses an owned-versus-rented split for inference specifically; the inference-share of the $1T deal flow is inferred, not measured. (2) No neocloud publishes inference-specific margin or tokens-per-watt at the facility boundary — Firmus's claims are vendor-stated and era-varying (PUE 1.03 → <1.05 pPUE → 1.02), and no absolute tokens-per-kWh figure exists publicly. (3) The product names FORGE and ARC, and the "superconducting power delivery" claim, have no public footprint; if they are internal, they should be named in the company's own materials before an external audience repeats them. (4) Workload data is thin: BurstGPT is one provider's traces and the a16z study is one platform's view; agentic inference's share of total tokens is asserted more than measured. What would settle the strategy question: a published facility-level tokens-per-watt benchmark, an inference-split disclosure from any major lab or neocloud, and a public reference architecture showing a neutral control plane serving a customer's vLLM/SGLang stack on a firmus-style factory.

## Confidence and limitations

Confidence is moderate. The mechanism layer (serving systems, energy, workload characterization) is replicated academic evidence; the market layer is industry-reported and was verified against fetched primaries wherever it is load-bearing, with every number traced to a URL retrieved this session. Limitations: Layer-1 records were read at abstract level this session (full texts were read in the prior reviews they are reused from); several figures rest on single reputable secondaries (Bloomberg/WSJ headlines, paywalled bodies) and are flagged MED in the research files; vendor efficiency claims are vendor-stated by definition; and the Firmus sections rely on company releases, which is the only public source that exists. A review of this question written after the first facility-level tokens-per-watt disclosures would be substantially stronger.

## Evidence table

| key | year | design | sample | measure | finding | limitations | confidence |
| --- | --- | --- | --- | --- | --- | --- | --- |
| kwon2023efficient | 2023 | systems (conference paper) | vLLM engine on production traces | serving throughput | Introduces PagedAttention: paging and sharing of KV-cache memory so batching is not fragmented by memory limits; the engine reports 2-4x throughput gains over prior serving systems | vendor-benchmarked |  single engine; numbers not independently replicated here |
| patel2023splitwise | 2023 | systems (preprint) | prefill and decode phases of 8B-70B models | phase-level time and energy | Prefill and decode have opposite resource profiles (compute-bound vs memory-bound); splitting them onto different hardware cuts cost for the same SLO | analysis is trace- and model-based |  not facility-level |
| zhong2024distserve | 2024 | systems (preprint) | disaggregated prefill/decode serving | goodput under latency SLOs | Separating prefill and decode onto different GPUs with different parallelism removes interference and improves SLO attainment | one system; measured in-house | high |
| zheng2023sglang | 2023 | systems (preprint) | structured generation and multi-call workloads | throughput and latency | SGLang adds radix-attention prefix reuse and structured-output execution |  improving throughput on workloads with shared prefixes | engine-specific claims |
| zhu2025megascal | 2025 | systems (conference paper) | Mixture-of-Experts serving at cluster scale | efficiency of disaggregated serving | Disaggregated expert-parallel MoE serving reduces KV-cache transfer and improves cluster efficiency for MoE models | production evidence from a single large operator | moderate |
| stojkovic2025dynamollm | 2025 | systems (conference paper) | LLM inference clusters under SLOs | energy consumption and SLO attainment | Cluster-level power management (DVFS and request scheduling) saves energy while meeting SLOs; makes tokens-per-watt a cluster-design objective | simulation plus small-scale validation | moderate |
| qin2024mooncake | 2024 | systems (preprint) | Kimi production serving (Moonshot AI) | serving architecture | Production KVCache-centric disaggregated architecture: separate prefill and decode clusters plus CPU/DRAM/SSD KV-cache tiering | one production system; vendor-reported | high |
| yuan2024llm | 2024 | survey | LLM inference technique literature | roofline classification | Provides a roofline framework for LLM inference: prefill is compute-bound |  decode is memory-bound; classifies optimization techniques by bottleneck | survey without new measurements |
| wang2025burstgpt | 2025 | dataset (conference paper) | real-world LLM serving workload traces (BurstGPT) | traffic burstiness | Real production inference traffic is highly bursty with long-tail behaviors; provides traces and analysis for scheduler design | traces from one provider's perspective | moderate |
| forys2026when | 2026 | simulation (preprint) | agentic multi-turn workloads | when disaggregation pays | Whether prefill/decode disaggregation pays depends on the prefill/decode ratio of the workload; agentic traffic changes the calculus versus chat | simulation-based | moderate |
| wang2025prefill | 2025 | systems (preprint) | aggregation vs disaggregation | goodput comparison | Unifies the debate: aggregation and disaggregation each win in different regimes; hybrid designs are emerging | preprint | moderate |
| lai2025tokenscale | 2025 | systems (preprint) | autoscaling for disaggregated serving | scaling timeliness and accuracy | Bursty workloads defeat lagging autoscaling policies; predictive scaling for prefill/decode pools is needed | preprint | moderate |
| luccioni2023power | 2023 | measurement (preprint) | generative AI models and tasks | energy per inference and task | Measured inference energy across models and tasks; energy per task varies widely and is not predicted by parameter count alone | lab measurements on specific hardware | high |
| chien2023reducing | 2023 | modeling (conference paper) | generative AI inference | operational and embodied carbon | Reducing inference carbon requires both efficiency gains and grid decarbonization; estimates for today and 2035 | model-based projection | moderate |
| deepseekai2024deepseekv3 | 2024 | technical report | DeepSeek-V3 pretraining | compute cost | Full training used 2.788M H800 GPU-hours with stable training; self-reported low-cost MoE training that reset market price expectations | self-reported; hardware constrained by export controls | high |
| hoffmann2022training | 2022 | theoretical/empirical | language model training runs | compute-optimal model/token allocation | Chinchilla scaling: models are undertrained for their size; the compute-optimal point trades parameters for tokens | about training; used here for inference-demand reasoning | high |
| sevilla2022compute | 2022 | empirical | historic ML training compute | compute growth rates | Training compute grew ~5-6 month doubling since 2010 (AI-driven) versus Moore-like before; the same demand curve now applies to inference | historical extrapolation | high |
| sajadieh2026aiindex | 2026 | industry report | AI industry statistics | investment and deployment | AI Index 2026: industry-scale statistics on AI investment |  capex |  and deployment trends |
| fernandez2025energy | 2025 | measurement (conference paper) | LLM inference configurations | energy per token | Energy benchmarking of LLM inference beyond latency-only; efficiency varies substantially with configuration | benchmark-specific conditions | moderate |
| agrawal2024vidur | 2024 | simulation (preprint) | LLM serving configurations | performance and cost prediction | Simulation framework that predicts serving performance/cost across configurations without full experimental runs | simulation accuracy depends on workload models | moderate |
| wu2023fast | 2023 | systems (preprint) | interactive LLM serving | latency tails | FastServe uses preemptive job scheduling to cut latency tail for interactive inference | preprint; in-house benchmarks | moderate |
| sheng2023s | 2023 | systems (preprint) | multi-adapter serving | GPU utilization with many adapters | S-LoRA serves thousands of LoRA adapters on one GPU via unified paging |  enabling many fine-tuned models per GPU | preprint; single-engine claims |
| duan2024muxserve | 2024 | systems (preprint) | multiple models on shared GPUs | utilization | Model multiplexing (MuxServe) improves GPU utilization over dedicated colocation by flexible multiplexing | preprint; in-house benchmarks | moderate |
| zhou2024survey | 2024 | survey | efficient inference literature | technique taxonomy | Taxonomy of efficiency techniques: quantization |  pruning |  distillation |
| li2024llm | 2024 | survey | serving systems literature 2023+ | system-level advances | Survey of serving-system advances: disaggregation |  scheduling |  autoscaling |
| agrawal2024taming | 2024 | systems (preprint) | Sarathi-Serve engine | throughput-latency tradeoff | Chunked prefill and piggybacking decodes onto prefill gaps improve throughput and latency simultaneously | preprint; in-house benchmarks | moderate |

## References

1. Kwon, Woosuk et al. (2023). *Efficient Memory Management for Large Language Model Serving with PagedAttention*. Proceedings of the 29th Symposium on Operating Systems Principles. Abstract only. PagedAttention/vLLM: KV-cache paging and sharing as the memory-management foundation of serving engines [doi:10.1145/3600006.3613165](https://doi.org/10.1145/3600006.3613165)
2. Patel, Pratyush et al. (2023). *Splitwise: Efficient generative LLM inference using phase splitting*. arXiv (Cornell University). Abstract only. Establishes the prefill/decode phase asymmetry (compute-bound vs memory-bound) that motivates disaggregation [doi:10.48550/arxiv.2311.18677](https://doi.org/10.48550/arxiv.2311.18677)
3. Zhong, Yinmin et al. (2024). *DistServe: Disaggregating Prefill and Decoding for Goodput-optimized Large Language Model Serving*. arXiv (Cornell University). Abstract only. Disaggregating prefill and decoding for goodput under latency SLOs; the core inference-cluster design result [doi:10.48550/arxiv.2401.09670](https://doi.org/10.48550/arxiv.2401.09670)
4. Zheng, Lianmin et al. (2023). *SGLang: Efficient Execution of Structured Language Model Programs*. arXiv:2312.07104. Abstract only. SGLang: structured generation and radix-attention prefix reuse; one of the two dominant open serving engines [doi:10.48550/arxiv.2312.07104](https://doi.org/10.48550/arxiv.2312.07104)
5. Zhu, Ruidong et al. (2025). *MegaScale-Infer: Efficient Mixture-of-Experts Model Serving with Disaggregated Expert Parallelism*. Proceedings of the ACM SIGCOMM 2025 Conference. Abstract only. MegaScale-Infer: disaggregated expert-parallel MoE serving at ByteDance scale; KV-cache transfer as a bottleneck [doi:10.1145/3718958.3750506](https://doi.org/10.1145/3718958.3750506)
6. Stojkovic, Jovan (2025). *DynamoLLM: Designing LLM Inference Clusters for Performance and Energy Efficiency*. 2025 IEEE International Symposium on High Per. Abstract only. Cluster-level power management for inference clusters; tokens-per-watt as an SLO-aware design objective [doi:10.1109/hpca61900.2025.00102](https://doi.org/10.1109/hpca61900.2025.00102)
7. Qin, Ruoyu et al. (2024). *Mooncake: A KVCache-centric Disaggregated Architecture for LLM Serving*. arXiv (Cornell University). Abstract only. Mooncake: production KVCache-centric disaggregated serving at Moonshot (Kimi); prefill/decode separation in practice [doi:10.48550/arxiv.2407.00079](https://doi.org/10.48550/arxiv.2407.00079)
8. Yuan, Zhihang et al. (2024). *LLM Inference Unveiled: Survey and Roofline Model Insights*. arXiv (Cornell University). Abstract only. Roofline framework for LLM inference: prefill compute-bound, decode memory-bound; taxonomy of optimizations [doi:10.48550/arxiv.2402.16363](https://doi.org/10.48550/arxiv.2402.16363)
9. Wang, Yuxin et al. (2025). *BurstGPT: A Real-World Workload Dataset to Optimize LLM Serving Systems*. Proceedings of the 31st ACM SIGKDD Conference on Knowledge Discovery and Data Mining V.2. Abstract only. BurstGPT: real production serving traces showing heavy burstiness; basis for scheduler and capacity design [doi:10.1145/3711896.3737413](https://doi.org/10.1145/3711896.3737413)
10. Forys, Przemyslaw et al. (2026). *When Does Disaggregation Pay? Simulating Prefill--Decode--Attention--FFN Specialization for Agentic LLM Inference*. arXiv (Cornell University). Abstract only. Simulation study of when prefill/decode disaggregation pays, including agentic multi-turn workloads [doi:10.48550/arxiv.2608.03741](https://doi.org/10.48550/arxiv.2608.03741)
11. Wang, Chao et al. (2025). *Prefill-Decode Aggregation or Disaggregation? Unifying Both for Goodput-Optimized LLM Serving*. arXiv (Cornell University). Abstract only. Unifies aggregation vs disaggregation debate: each wins in different regimes; hybrid designs emerging [doi:10.48550/arxiv.2508.01989](https://doi.org/10.48550/arxiv.2508.01989)
12. Lai, Ruiqi et al. (2025). *TokenScale: Timely and Accurate Autoscaling for Disaggregated LLM Serving with Token Velocity*. arXiv (Cornell University). Abstract only. TokenScale: bursty workloads defeat lagging autoscaling; predictive scaling for disaggregated pools [doi:10.48550/arxiv.2512.03416](https://doi.org/10.48550/arxiv.2512.03416)
13. Luccioni, Alexandra Sasha et al. (2023). *Power Hungry Processing: Watts Driving the Cost of AI Deployment?*. arXiv:2311.16863. Abstract only. Power Hungry Processing: measured inference energy across models/tasks; energy varies far more than parameters suggest [doi:10.48550/arxiv.2311.16863](https://doi.org/10.48550/arxiv.2311.16863)
14. Chien, Andrew A. (2023). *Reducing the Carbon Impact of Generative AI Inference (today and in 2035)*. Proceedings of the 2nd Workshop on Sustainable Computer Systems. Abstract only. Modeling of generative-AI inference carbon (operational + embodied) for today and 2035 [doi:10.1145/3604930.3605705](https://doi.org/10.1145/3604930.3605705)
15. DeepSeek-AI et al. (2024). *DeepSeek-V3 Technical Report*. arXiv preprint. Abstract only. DeepSeek-V3: 2.788M H800 GPU-hours full training; low-cost MoE training that reset inference price expectations [doi:10.48550/arxiv.2412.19437](https://doi.org/10.48550/arxiv.2412.19437)
16. Hoffmann, Jordan et al. (2022). *Training Compute-Optimal Large Language Models*. arXiv:2203.15556 (Chinchilla). Abstract only. Chinchilla: compute-optimal model/token allocation; the demand-side anchor for how much serving capacity models need [doi:10.48550/arxiv.2203.15556](https://doi.org/10.48550/arxiv.2203.15556)
17. Sevilla, Jaime et al. (2022). *Compute Trends Across Three Eras of Machine Learning*. International Joint Conference on Neural Networks (IJCNN), arXiv:2202.05924. Abstract only. Training compute growth trends (~5-6 month doubling post-2010); the demand curve that now applies to inference [doi:10.48550/arxiv.2202.05924](https://doi.org/10.48550/arxiv.2202.05924)
18. Sajadieh, Sha et al. (2026). *Artificial Intelligence Index Report 2026*. arXiv:2606.15708 (Stanford HAI). Abstract only. Stanford AI Index 2026: industry-scale statistics on AI investment and deployment [doi:10.48550/arxiv.2606.15708](https://doi.org/10.48550/arxiv.2606.15708)
19. Fernandez, Jared et al. (2025). *Energy Considerations of Large Language Model Inference and Efficiency Optimizations*. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics. Abstract only. ACL 2025: energy benchmarking of LLM inference beyond latency; efficiency varies with configuration [doi:10.18653/v1/2025.acl-long.1563](https://doi.org/10.18653/v1/2025.acl-long.1563)
20. Agrawal, Amey et al. (2024). *Vidur: A Large-Scale Simulation Framework For LLM Inference*. arXiv (Cornell University). Abstract only. Vidur: simulation framework for predicting serving performance/cost across cluster configurations [doi:10.48550/arxiv.2405.05465](https://doi.org/10.48550/arxiv.2405.05465)
21. Wu, Bingyang et al. (2023). *Fast Distributed Inference Serving for Large Language Models*. arXiv (Cornell University). Abstract only. FastServe: preemptive scheduling to cut latency tails for interactive inference [doi:10.48550/arxiv.2305.05920](https://doi.org/10.48550/arxiv.2305.05920)
22. Sheng, Ying et al. (2023). *S-LoRA: Serving Thousands of Concurrent LoRA Adapters*. arXiv (Cornell University). Abstract only. S-LoRA: serving thousands of LoRA adapters per GPU via unified paging; multi-tenant model serving [doi:10.48550/arxiv.2311.03285](https://doi.org/10.48550/arxiv.2311.03285)
23. Duan, Jiangfei et al. (2024). *MuxServe: Flexible Multiplexing for Efficient Multiple LLM Serving.*. CoRR. Abstract only. MuxServe: flexible multiplexing of multiple models on shared GPUs to raise utilization [doi:10.48550/arxiv.2404.02015](https://doi.org/10.48550/arxiv.2404.02015)
24. Zhou, Zixuan et al. (2024). *A Survey on Efficient Inference for Large Language Models*. arXiv (Cornell University). Abstract only. Survey of efficient LLM inference techniques: quantization, pruning, distillation, serving [doi:10.48550/arxiv.2404.14294](https://doi.org/10.48550/arxiv.2404.14294)
25. Li, Baolin et al. (2024). *LLM Inference Serving: Survey of Recent Advances and Opportunities*. arXiv (Cornell University). Abstract only. Survey of LLM serving systems since 2023: disaggregation, scheduling, autoscaling, memory management [doi:10.48550/arxiv.2407.12391](https://doi.org/10.48550/arxiv.2407.12391)
26. Agrawal, Amey et al. (2024). *Taming Throughput-Latency Tradeoff in LLM Inference with Sarathi-Serve*. arXiv (Cornell University). Abstract only. Sarathi-Serve: chunked prefill plus piggybacked decodes improve throughput and latency together [doi:10.48550/arxiv.2403.02310](https://doi.org/10.48550/arxiv.2403.02310)
