Byte-Sized Intelligence April 9, 2026

The cost of inference and the rise of data centers

This week, we look at how inference drives continuous AI demand, why data centers are expanding to keep up, and how costs show up as enterprises scale usage and deploy agents.

AI in Action

Inference is driving the next wave of data centers [AI Infrastructure/Inference]

Inference is appearing more often in AI conversations because it reflects how systems are used. It is the step where a trained model produces an output from a prompt or task. Each search result, copilot suggestion, and chatbot response runs through inference. As AI is integrated into everyday tools and enterprise workflows, these interactions occur continuously and at scale.

This pattern is shaping how data centers are built. New capacity is moving toward multi-building campuses measured in gigawatts, designed for high-density GPUs and sustained workloads. Microsoft, Amazon, and Google continue to lead this buildout, with Meta expanding alongside them. These facilities support training and continuous inference across cloud platforms and enterprise systems. The buildout is concentrating in regions that can support it, including Northern Virginia, Texas, and parts of the Midwest, with more selective expansion in Europe and early development in the Middle East.

The constraints are tied to the physical system. GPUs from Nvidia operate within environments that depend on sustained power, advanced cooling, and long construction timelines. Electricity availability, thermal density, water usage, and permitting shape how quickly new capacity can be added and where it can be located. These factors influence how AI systems are delivered. Running AI becomes an ongoing cost tied to usage, affecting pricing, enterprise budgets, and the rollout of features. Performance and responsiveness depend on available capacity, which can result in slower responses or usage limits when demand exceeds supply.

Bits of Brilliance

Why AI gets expensive at scale  [Compute/Economics]

Every AI interaction carries a cost. A prompt is broken into tokens, processed by a model, and turned into an output. This step, known as inference, runs on compute provided by GPUs inside data centers. Most enterprises access this infrastructure through cloud providers and pay based on usage, including how much data is processed and how often models are called. A single response can feel negligible. The cost builds as this process repeats across workflows, systems, and users throughout the day.

The cost becomes clearer when enterprises begin deploying agents or shifting workflows into AI systems. Each step an agent takes triggers a model call, with compute used to process inputs, generate outputs, and carry out intermediate steps such as retrieval or tool use. These calls are billed based on the volume of data processed and the compute time required. A task that appears simple can involve multiple calls, each contributing to the total cost. As workflows become more complex and more users rely on these systems at the same time, usage increases in both volume and concurrency. Cost reflects how much work is being executed through the system.

Data centers provide the compute that supports this activity. As inference demand grows across enterprise workflows, the need for sustained capacity increases alongside it. The cost observed at the workflow level aggregates into continuous demand on infrastructure. This drives expansion in data center capacity, along with requirements for power, cooling, and efficient operation.

Byte-Sized Intelligence is a personal newsletter created for educational and informational purposes only. The content reflects the personal views of the author and does not represent the opinions of any employer or affiliated organization. This publication does not offer financial, investment, legal, or professional advice. Any references to tools, technologies, or companies are for illustrative purposes only and do not constitute endorsements. Readers should independently verify any information before acting on it. All AI-generated content or tool usage should be approached critically. Always apply human judgment and discretion when using or interpreting AI outputs.