← All sectors / The AI transformation
081 · AI inference infrastructure & model serving
The cost per answer
Curve position
Takeoff
Binding constraint
Accelerator availability and the memory bandwidth that caps throughput.
Training a model is a one time capital event. Serving it happens millions of times a day forever, and at scale the cost per answer determines whether an AI product has a viable margin. Optimizing inference is now its own industry.
Historically the entire conversation was about training scale, because that is where the headlines and the largest clusters were. As deployment matured, the economics shifted to serving, which is where most compute now goes.
The structural driver is unit economics. An application paying more per query than it earns cannot scale, so every serious deployment eventually invests in quantization, caching, batching, routing, and cheaper silicon.
The technology layer spans inference optimized accelerators, serving frameworks that batch and schedule requests, quantization and distillation to shrink models, key value caching, speculative decoding, and the routing layers that send easy queries to cheap models.
Adoption economics are directly measurable in cost per thousand tokens or per request, which makes procurement unusually rational and vendor comparisons unusually brutal.
The beneficiaries include inference chip designers, the neocloud providers renting capacity by the hour, serving software vendors, and the memory manufacturers whose bandwidth caps throughput.
The value chain runs from silicon through cloud capacity and serving software to the application. Memory bandwidth is the physical constraint that determines how much of the chip can actually be used.
The overlooked layer includes memory and interconnect suppliers, inference optimization software vendors, smaller cloud providers specializing in inference rather than training, and the power and cooling suppliers behind them.
Competitive dynamics pit the dominant accelerator ecosystem against challengers competing specifically on inference cost per watt, where the software moat is weaker than in training.
Risks: the segment is fully levered to AI application demand, price competition among providers is intense, model efficiency gains reduce compute per query, and capacity gluts are possible if demand growth slows.
What to watch: cost per token trends at major providers, inference specific silicon launches, memory bandwidth in new accelerators, and neocloud capacity utilization.
