Selling infrastructure to inference providers: what a neocloud can offer without competing
Which inference providers should an AI-factory builder target, at which infrastructure layer, and what must it know about inference to sell them capacity without becoming their competitor?
The inference-provider market splits into open-weight GPU hosts (Together, Fireworks, DeepInfra, Baseten, Modal — they rent nearly all their capacity and compete on serving software) and closed-weight labs (OpenAI, Anthropic, Google, Meta, xAI — they rent at enormous scale via take-or-pay contracts but vertically integrate software and increasingly silicon). The evidence says a neocloud should target the open-weight hosts at the fleet/facility and control-plane layers — power, cooling, density, grid, orchestration that respects the customer's serving engine — and never compete at the serving-engine layer, which is the customer's moat. Firmus's public record (AI FactoryOS, HyperCube, Model-to-Grid, a Fireworks partnership) already points this way. Confidence is moderate: the market facts are industry-reported, and several claimed product names could not be verified in any public source.
Updated 23 Aug 202626 sources2022–2026Standard21 min read
inference providers · neocloud · GPU cloud · AI factories · LLM inference · take-or-pay · tokens per watt