Chapter 12
The Operators
A new kind of company has emerged whose whole business is to buy or lease enormous clusters of GPUs, install them in data centers, and rent out the computing power by the hour to anyone who needs to train or run an AI model.
Another system chapter. Chapters 1–10 built the physical machine; Chapter 11 was the software that runs it. This chapter is about who owns and operates it all — and, at the end, what it's all actually being used for.
What the Operators Do
Not every AI company can spend billions of dollars to build its own 100,000-GPU cluster. The frontier labs — the handful of companies (OpenAI, Anthropic, Google DeepMind, Meta AI, and a few others) building the largest and most capable AI models — the startups, the university researchers, the companies beginning to adopt AI for their own operations — most of them need compute on someone else's hardware. So a market has grown up in between: companies that make the enormous capital investment in GPUs and data centers, then sell access to that computing power.
To understand why this market exists, consider the barrier to entry. Building a training-capable AI cluster from scratch means: securing a site with hundreds of megawatts of permitted power (Chapters 7-8), constructing a building designed for the extreme heat densities of Chapter 4, installing the cooling systems of Chapter 5, building the power-distribution chain of Chapter 6, procuring thousands of GPUs (with lead times measured in months), installing the high-bandwidth network fabric of Chapter 3, writing or deploying the management software of Chapter 11, and hiring the operations team to run it all. The total investment can easily exceed a billion dollars, and the timeline from decision to first workload can be two to three years. For a startup with a great AI idea and limited capital, building this is simply not an option.
The basic economics work like a very expensive version of renting an apartment. A company buys thousands of GPUs and builds (or leases) the data center to house them — the servers, the networking, the cooling, the power systems, all of it (Chapters 1 through 10). Then it rents that capacity to customers, either by the hour (like a hotel room) or on long-term contracts (like a lease). The price of a single top-end GPU hour has ranged from roughly two to four dollars at market rates — which sounds small until you multiply it by ten thousand GPUs running twenty-four hours a day for three months. A large training run at those rates can cost tens of millions of dollars in GPU rental alone, before electricity. If it keeps the machines busy enough, the revenue exceeds the cost of the hardware plus the electricity plus the staff, and the company earns a return. If it doesn't, the machines sit idle and the company burns through its capital. The entire business model, at bottom, turns on one number: utilization — how many hours per day the GPUs are actually running paid workloads.
This is not a new idea. Cloud computing — renting compute by the hour — has existed for nearly two decades. What's new is the scale, the cost, and the specialization. A single rack of AI servers (Chapter 4) can cost as much as a house. A large training cluster can cost as much as a skyscraper. The capital required to participate at the frontier is so large that only a handful of companies in the world can self-fund it — which is why the operator model exists: it pools the capital investment and distributes the cost across many customers.
But an AI training cluster is not an ordinary cloud server; it is a tightly coupled supercomputer where every GPU must talk to every other GPU at enormous speed (Chapter 3), and a single failed machine can disrupt a job running across thousands. A customer renting ordinary cloud servers doesn't notice if one machine restarts — their web application fails over to another copy automatically. A customer running a training job across 10,000 GPUs notices immediately, because the other 9,999 GPUs all stop and wait for the failed one's work to be recovered from a checkpoint. Operating it well is a skill, not just a procurement exercise — it requires deep understanding of the network, the cooling, the power, and the software layers described in the previous chapters.
The Three Tiers of Operators
The operators come in three distinct kinds — the same layer of the stack, but very different animals.
The hyperscalers are the giants: Amazon Web Services, Microsoft Azure, Google Cloud, Meta, and Oracle. They build their own data centers, design some of their own chips (Chapter 1), and serve both their own internal workloads and outside customers. They consume the majority of all GPUs manufactured, and their capital spending is the engine of the entire buildout — combined estimates have climbed toward hundreds of billions of dollars per year (Chapter 14). Because they are diversified mega-companies, they mostly sit outside this book's focused universe, but everything in the previous eleven chapters is ultimately built to serve them.
What distinguishes the hyperscalers from everyone else is not just scale but vertical integration. They design their own servers (and increasingly their own chips), build their own data centers to their own specifications, write their own management software, and operate their own fiber networks between facilities. They are simultaneously the landlord, the builder, and the largest tenant. This gives them enormous purchasing power with every supplier in this book — but it also means their capital-spending decisions ripple through the entire supply chain. When a hyperscaler announces a new training cluster, transformer factories, fiber plants, and construction firms all feel it within weeks.
The NeoClouds are pure-play companies built specifically and only for AI computing. Unlike the hyperscalers, which offer a broad menu of cloud services (storage, databases, web hosting, AI), the NeoClouds focus entirely on GPU compute. CoreWeave (CRWV) is the standout — its fleet is entirely NVIDIA GPUs, and its business is entirely AI. Nebius (NBIS) is a European-based "AI hyperscaler" with a full-stack approach. Applied Digital (APLD) develops the data-center campuses and leases them to hyperscalers and NeoClouds.
The defining characteristic of a NeoCloud is focus: because it does only one thing — run GPUs for AI — it can optimize its infrastructure, its networking, and its software stack specifically for that job, in ways that a general-purpose cloud provider may not. A hyperscaler's cloud platform must serve thousands of different workload types — web servers, databases, media encoding, scientific computing, AI — and its infrastructure reflects that generality. A NeoCloud's entire network fabric, cooling architecture, power distribution, and scheduling software are designed for one specific thing: keeping large clusters of GPUs communicating at maximum speed. This specialization shows up in the details: a NeoCloud's network topology might be optimized for the specific traffic patterns of GPU-to-GPU all-reduce operations (Chapter 3), while a hyperscaler's network must serve a broader range of traffic patterns.
The NeoClouds also differentiate on customer intimacy. A hyperscaler serves millions of customers through standardized, self-service interfaces. A NeoCloud might serve dozens of customers, each running large clusters, and can offer hands-on engineering support — helping configure the network topology for a specific training job, diagnosing subtle performance problems, or customizing the scheduling to match a customer's workflow. For a company running a hundred-million-dollar training run, having a direct relationship with the engineers who manage the physical infrastructure is genuinely valuable.
The crypto-pivots are the most unexpected group: companies that built large-scale computing facilities for Bitcoin mining and are now repurposing them for AI. The connection isn't the computing (Bitcoin mining uses specialized ASICs, not GPUs) — it's the real estate and the power. Bitcoin miners spent years acquiring exactly the two things that are now the scarcest inputs to AI data centers: permitted land and permitted electrical capacity. Converting a mining site to a GPU data center is not trivial (the cooling, networking, and power distribution are different), but the hardest step — getting the right to draw hundreds of megawatts from the grid at a specific location — is already done.
The conversion process is more involved than it might sound. A Bitcoin mining facility is optimized for a single, simple task: running thousands of identical, independent ASIC chips that each perform one mathematical operation over and over. The heat output is enormous but uniform. The networking is trivial — each miner talks only to the mining pool, not to other miners. The cooling can be crude — air-cooled containers in cold climates, with no need for the precision temperature control that GPUs demand.
Converting this to an AI data center requires upgrading nearly everything inside the building. The power distribution must be redesigned for different load profiles (GPU racks draw power differently than mining rigs). The cooling must be upgraded from bulk air-cooling to the precision liquid-cooling systems of Chapter 5, because GPU clusters generate concentrated heat that air alone cannot remove. The network is the biggest change — a Bitcoin mining site has essentially no internal network, while a training cluster needs the high-bandwidth, low-latency fabric described in Chapter 3. And the physical structure itself may need reinforcement, because GPU server racks are heavier than mining containers.
But the hard parts are already done. Core Scientific (CORZ) is converting sites under large contracts with GPU cloud customers. Cipher Mining (CIFR) and Hut 8 (HUT) have signed long-term leases backed by major cloud and AI companies. Riot Platforms (RIOT) has leased capacity to chip makers. Others — Bitdeer (BTDR), Bit Digital (BTBT), MARA (MARA), CleanSpark (CLSK) — are at varying stages of the pivot, from already earning revenue from GPU rentals to still building out the infrastructure. What unites them is a shared insight, articulated by several of their executives in nearly identical words: power availability is the single biggest bottleneck constraining the growth of the AI economy, and these companies already have it.

What Makes a Good Operator
Owning GPUs isn't enough; operating them well is a skill. Four things separate a good operator from a poor one.
Density and fabric. It's the difference between "I have 10,000 GPUs" and "I have 10,000 GPUs that can act as one computer." The network (Chapter 3) is what makes a true cluster — thousands of GPUs connected by high-bandwidth, low-latency links so that they can exchange data fast enough to train a model together. A collection of GPUs scattered across separate buildings without a proper fabric is not a training cluster; it's a collection of small machines.
The day-to-day reality of running a GPU data center is more demanding than running a traditional one. Thermal management requires constant attention — GPU workloads are far less predictable than web-server workloads, and a training job ramping up can add tens of megawatts of heat load in minutes, requiring the cooling systems to respond in real time. Hardware failures are more consequential — a failed GPU in a training cluster doesn't just reduce capacity; it can crash a job that has been running for weeks, so operations teams maintain "hot spare" pools of pre-configured replacement machines that can be swapped in quickly. Network monitoring is continuous — a degraded optical transceiver (Chapter 3) that drops even a tiny fraction of packets can slow an entire training cluster, and tracking down which of thousands of transceivers is the culprit requires sophisticated diagnostic tools. The operations staff at a well-run GPU facility includes not just the traditional data-center roles (power engineers, mechanical engineers, security) but also GPU specialists who understand the hardware and software at the chip level.
Utilization. An idle GPU burns money. It still draws power (a modern GPU at idle consumes a substantial fraction of its peak power), still generates heat (which still requires cooling), still occupies a rack slot (which could hold a revenue-generating machine), and produces no revenue. The operator's craft is keeping the machines busy — and the strategies for doing this are more nuanced than they might seem.
The ideal is a large training job that uses the entire cluster continuously for months. In practice, jobs end, customers churn, hardware fails and must be replaced, and new jobs don't always arrive the moment old ones finish. The operator fills these gaps by scheduling smaller inference workloads during training downtime, by batching maintenance during periods of low demand, and by managing the transition when one customer's contract ends and another's begins. Some operators offer spot pricing — deeply discounted compute for workloads that can be interrupted on short notice — which fills otherwise idle capacity with low-priority jobs that generate at least some revenue. The sophistication of the scheduling and pricing system is a meaningful competitive advantage, because the difference between 85% utilization and 95% utilization, over thousands of GPUs and years of operation, can represent hundreds of millions of dollars.
Power efficiency. Measured by the PUE from Chapter 5 — the ratio of total power consumed to the power actually reaching the IT equipment. A PUE of 1.2 means twenty percent of the electricity is going to cooling, lighting, and overhead. A PUE of 1.05 means only five percent. Over thousands of racks and years of operation, that difference is enormous — for a 100-megawatt facility, the difference between PUE 1.2 and PUE 1.05 is roughly 15 megawatts of wasted power, which at typical industrial electricity rates represents millions of dollars per year in pure overhead. This is why the advanced cooling technologies in Chapter 5 (liquid cooling, direct-to-chip cooling) are competitive advantages, not just engineering preferences — they directly affect the operator's cost structure and therefore its margins.
Speed to power. As we saw throughout Chapters 6–8, the single scarcest resource in the AI buildout is electrical power, delivered to a permitted site, fast. An operator that can bring new capacity online in months rather than years has a powerful competitive advantage, because it can serve customers who need compute now, not in three years. This is the fundamental advantage of the crypto-pivots — they already have the power — and it is the fundamental challenge for the NeoClouds, who must either build new sites (slow) or lease from someone who has one (expensive).

The Contract Structure
The money behind these buildouts is worth understanding, because it is what makes the enormous upfront cost financeable.
Most large deals are structured as take-or-pay contracts: a big customer commits to pay for the capacity whether or not it uses it, often for ten to fifteen years. This is the same concept as the apartment lease from section 12.1, but scaled up to hundreds of megawatts and billions of dollars. The take-or-pay structure matters because it turns an uncertain future revenue stream into a predictable one — and predictable revenue can be borrowed against. An operator with a fifteen-year take-or-pay contract from a creditworthy customer can go to a bank and say: "This customer will pay me a fixed amount per year for fifteen years; lend me the money to build the facility, and I'll repay you from those payments." This is project finance — the same mechanism that funds toll roads, power plants, and pipelines — applied to data centers.
Once a facility is built and the customer moves in, the ongoing costs are relatively modest. Electricity is the largest operating expense — a 100-megawatt facility running at typical industrial rates burns through millions of dollars in electricity per month. Staff costs are comparatively small (a well-automated facility might require only a few dozen operations personnel). Maintenance is routine until something big breaks. Much of the revenue flows through to profit, which is the attractive economics of the model — if the facility is built on time (construction delays burn cash and disappoint customers), if the customer pays (a customer that goes bankrupt or pivots away from AI leaves the operator with an empty building and debt), and if the technology doesn't change so fast that the facility becomes obsolete (the GPU refresh dilemma below).
There is a subtlety here that distinguishes the AI data-center model from traditional real estate. A landlord who leases office space doesn't care what computers the tenant puts on the desks. An AI data-center operator cares deeply about what hardware the tenant is running, because the facility's power distribution, cooling, and networking were designed around specific GPU densities, heat profiles, and bandwidth requirements. A facility designed for one generation of GPUs may need significant modifications to efficiently house the next — different rack densities, different cooling capacities, different power delivery. This is why the most sophisticated operators are designing their facilities for modularity and future upgrades from the start, even though it costs more upfront.
The GPU refresh dilemma is the hidden risk, and it is worth understanding in detail because it is the central economic uncertainty of the entire operator model. GPU generations turn over every one to two years, and each new generation is substantially more powerful and energy-efficient than the last — often doubling performance per watt. An operator that bought a fleet of one generation faces a dilemma when the next arrives: customers want the new chips, but the old ones aren't paid off yet.
Here is the math, simplified. The operator spent, say, a billion dollars on a fleet of GPUs and expects to earn that back over three to four years. Eighteen months later, the chip maker releases a new GPU that is twice as fast and twenty percent more efficient. Customers now want the new chip, because training on the new hardware is cheaper per unit of computation. The operator can either (a) upgrade and absorb the write-down on the old fleet — formally acknowledging that the hardware is now worth less than what was paid for it, and recording that loss on the books, even though the machines still physically work — or (b) keep running the old chips at declining prices, because the old hardware is now worth less per hour to customers who could rent the new chips elsewhere. Neither option is attractive, and this cycle repeats with every new GPU generation.
Take-or-pay contracts are the primary defense: they guarantee revenue regardless of whether the customer would rather be on newer hardware. A customer locked into a ten-year take-or-pay deal cannot simply move to a competitor's newer cluster — they've already committed. This is also the reason some operators lease their GPUs from the chip maker rather than buying them outright: the lease transfers some of the obsolescence risk back to the manufacturer, who may offer a trade-in or upgrade path. And it is why the chip maker's own investment in the operators (section 14.3's circular-financing question) can be seen, in part, as a mechanism for managing the refresh cycle — the chip maker has an incentive to help its customers upgrade, because each upgrade is another sale.

What Is It All For?
Step back from the machinery and ask the question the whole book has been building toward: what is all this compute actually doing? Broadly, three things.
Training is the big spend — building a frontier AI model by running tens of thousands of GPUs continuously for months, at a cost that can reach hundreds of millions of dollars for a single run. To put that in tangible terms: a large training run draws enough electricity to power a small city, consumes enough cooling water to fill an Olympic swimming pool every day, and produces enough heat to warm a neighborhood. This is not a figure of speech; these are the actual physical consequences of running tens of thousands of processors at full power, twenty-four hours a day, for weeks or months on end.
A training job doesn't use the GPUs for a few hours; it uses them around the clock for weeks or months. During that time, the cluster is utterly occupied — no other customer can use those machines. This is the demand that justified the entire buildout, and it is the reason that the networking from Chapter 3 is so critical — training requires all the GPUs to exchange data constantly, and any bottleneck in the network slows the entire job. It is also the reason the orchestration software from Chapter 11 is so important: a training run that crashes two weeks in and must restart from an outdated checkpoint doesn't just waste compute; it wastes millions of dollars and weeks of calendar time that a competitor will use to get ahead.
Inference is the using of a trained model — answering questions, generating images, writing code, analyzing documents — one request at a time. Each inference request uses far fewer GPUs than training (often just one or a few), but the math of scale is powerful: if a billion people use AI assistants even occasionally, each request touching one GPU for a fraction of a second, the total compute adds up to enormous numbers. Inference is growing faster than training, and multiple operators have noted that more than half of their GPU utilization is already for inference, not training.
This shift matters because inference has fundamentally different infrastructure needs. Training values raw throughput — the total amount of computation completed per hour across the entire cluster, regardless of how long any individual step takes. Inference values latency — how quickly a single user gets their answer. A training job that runs for three months doesn't care if each step takes one second or two; an inference request from a person typing on their phone needs a response within a few hundred milliseconds or the product feels broken. This means inference infrastructure can tolerate a somewhat less tightly coupled network (the GPUs don't need to synchronize with each other) but demands more geographic distribution (to be physically close to users for lower latency), better load balancing (to handle unpredictable spikes in demand), and more sophisticated routing (to direct each request to the right model on the right machine).
Fine-tuning sits in between — taking a pre-trained model and specializing it for a specific task (a legal AI, a medical AI, a coding AI). It uses less compute than training from scratch but more than inference, and it is the mechanism by which the general-purpose models from the frontier labs become specialized tools for specific industries.
The process works like this: a pre-trained model has already learned general patterns from a vast corpus of data — language, reasoning, world knowledge. Fine-tuning takes that general model and trains it further on a much smaller, specialized dataset — say, thousands of legal contracts, or millions of medical records, or a company's entire internal documentation. The model adjusts its internal parameters to perform better on the specific task, while retaining the general capabilities it learned in pre-training. The result is a model that is both broadly capable and specifically useful for the fine-tuned domain. Think of it like the difference between educating a doctor (years of general medical training) and specializing them (a fellowship in cardiology): the specialist still has all the general knowledge, plus deep expertise in one area.
Fine-tuning is the on-ramp for enterprises: a company that wants to use AI for its own operations doesn't need to train a model from scratch (which would cost hundreds of millions); it fine-tunes an existing model on its own data (which might cost thousands or tens of thousands of dollars). This is a much larger addressable market than frontier training — there are only a handful of companies training frontier models, but there are millions of companies that might fine-tune one — and it represents a significant and growing portion of the total demand for GPU compute. The fine-tuning market also creates a different pattern of GPU usage: many short-to-medium jobs (hours to days, rather than months), each using a smaller cluster, spread across many different customers. This workload pattern suits the operator model well, because it fills the gaps between large training runs and keeps utilization high.
Agentic workloads are a newer and rapidly growing category. These are AI systems that don't just answer a single question but carry out multi-step tasks autonomously — browsing the web, writing and running code, coordinating across tools and databases. An agentic workload might use a modest amount of compute per individual step, but it chains many steps together, and each step involves an inference call. A company deploying thousands of AI agents, each performing complex tasks throughout the day, can consume substantial compute without running a single training job. This demand pattern is less predictable than training (which uses a fixed cluster continuously) and more varied than simple inference (which handles independent requests), and it creates new challenges for the orchestration layer described in Chapter 11.
There is also a growing category of sovereign and government-sponsored AI compute — nations investing in GPU clusters as a matter of strategic interest, similar to the way governments invest in defense, telecommunications, or energy infrastructure. The logic is straightforward: if AI compute becomes as essential to economic competitiveness as electricity or broadband, then a country without sufficient domestic compute capacity is dependent on other countries for a critical resource. Several nations are building or funding national AI data centers, and some are offering subsidized power and expedited permitting to attract private AI infrastructure investment. This is a nascent but potentially large demand source that doesn't follow the same economics as commercial operators — national compute projects may accept lower returns in exchange for strategic autonomy.
Underneath all of it is a bet, sometimes called the scaling hypothesis: the idea that bigger models, trained on more data with more compute, keep getting better — and that this improvement unlocks valuable new capabilities that justify the cost of the next, larger training run. "Better" has a specific meaning here: a model trained on ten times more compute doesn't just do the same things slightly faster; it can do new things. Earlier models could summarize text; larger ones can write code. Earlier ones could translate languages; larger ones can reason through multi-step problems, diagnose medical images, or design molecules. Each jump in scale has historically opened capabilities that the smaller model simply could not perform — and each new capability has attracted new users willing to pay for it.
Every chapter of this book — every chip, every transformer, every gigawatt of power — is ultimately a wager on that pattern continuing to hold. So far it has. Whether it continues is the open question that hangs over the whole buildout — and it is the subject of Chapter 14.
It is worth stepping back to appreciate the scale of what the operators are collectively attempting. They are building, from scratch, a new category of industrial infrastructure — one that didn't exist a decade ago, that consumes power like a small country, that requires the most advanced chips ever made, and that serves a market (AI compute) whose long-term size is genuinely uncertain. The closest historical analogy might be the telephone companies of the early twentieth century, which built vast networks of wire and switching equipment on the bet that enough people would pay to talk to each other to justify the investment. That bet paid off — but it took decades, and many of the original operators were consolidated, restructured, or bankrupted along the way.
The operators of the AI buildout face a similar path. The infrastructure they are building is almost certainly essential — the question is not whether AI compute will be needed, but how much, by whom, at what price, and who will ultimately own and operate it.
One factor that could reshape the landscape is regulation. So far, AI data centers have been lightly regulated compared to other infrastructure — no government sets the price of GPU compute the way a utility commission sets electricity rates, and no agency approves or denies the construction of a data center the way it might a power plant or a pipeline (beyond the permitting process in Chapter 13). But as data centers grow to consume as much power as small cities, as they draw down water resources, as they compete for grid capacity with residential and commercial users, and as the AI models they train become increasingly consequential, the regulatory landscape is shifting.
Some jurisdictions are beginning to require environmental-impact assessments specific to data centers. Others are imposing water-use restrictions or noise-level limits. A few are debating whether large AI data centers should be classified as critical infrastructure — which would bring benefits (priority access to power, expedited permitting) but also obligations (security requirements, reporting mandates, and potentially price regulation). The way this regulatory landscape evolves will significantly affect which operators thrive: companies that can navigate the regulatory environment efficiently will have an advantage over those that cannot, and the regulatory overhead itself represents a barrier to entry that may help established operators and deter newcomers.
There is also the question of consolidation. The current market has dozens of operators — hyperscalers, NeoClouds, crypto-pivots, and many smaller players. History suggests that infrastructure markets tend to consolidate over time: the initial boom attracts many entrants, competition drives down prices, margins compress, weaker players are acquired or go bankrupt, and the market settles into an oligopoly of a few large operators. The telephone industry started with thousands of local operators and consolidated into a handful of national carriers. The cloud-computing industry started with dozens of contenders and consolidated into three dominant hyperscalers (plus a few niche players). Whether the AI-compute market follows the same path — and how quickly — is one of the open questions for the next decade.
Those questions lead directly into the last two chapters: the constraints that limit how fast it can happen (Chapter 13), and the money that has to make sense of it (Chapter 14).

Next: Chapter 13 — The Queue (why time, not money, is the real constraint)
Companies in this part of the buildout: Operators