The high-performance computing (HPC) industry has always designed around the server. Servers went into racks, and racks went into data centers. Units and facilities.
Today, data centers are being rearchitected as AI factories, purpose-built to process AI workloads at scale. And that is shifting the design point — from the server to the rack.
This pulls the unit of compute up from the server and the level of optimization in from the data center. The rack sits between them, becoming the primary building block of the AI factory.
The reason is straightforward. In the rack you have GPUs, interconnects, memory, and storage — all connected to each other, fed by power and cooled to keep them running. When an AI task is submitted, many of these GPUs need to work together. Maybe 12 of them in one rack, as one system. And they all need to deliver optimized performance at the lowest cost per token.
That single requirement changes how the whole system is designed. When the rack becomes the place where work actually gets done, it stops being a container for servers and becomes the unit the industry architects from the ground up.
Instead of building infrastructure, the industry is now focused on serving usable AI.
That means the whole rack is designed around what it will be doing. Whether that is multimodal inference, training, or agentic workflows, the design starts from the workload and works backward, rather than assembling general-purpose parts and hoping they perform well together.
Usable AI is measured by how quickly and efficiently the infrastructure completes the work it is given. A coding task, for example, may run across a collection of GPUs connected to memory and storage. The value of that infrastructure depends on how well those resources work together to complete the task at the lowest cost per token.
Data centers have traditionally been measured by raw compute. As they are rearchitected as AI factories, the more relevant measure is how much useful AI work they deliver for every watt and every dollar.
Concentrating this much compute in one place has consequences, of course. When you increase compute density within a single rack, the power goes up inside the rack.
The numbers tell the story. Five years ago, a rack used to consume maybe eight or ten kilowatts. Today’s AI racks use more than 100 kilowatts. And as compute density continues to grow, we are nearing racks that consume one megawatt of power.
As rack power rises, cooling, component placement, connectivity, and compute utilization all become part of the design problem.
This is why efficiency and utilization are so important. In a given rack, how efficiently the GPUs are used determines both the return on the infrastructure investment and how well the system does the work it was built for. Compute that sits idle, waiting on data or throttled by heat, is compute you paid for and cannot use.
The rack is becoming the machine at the center of the AI factory, designed around a common goal: delivering the most useful AI work for every watt and every dollar.
As the rack becomes the unit of compute, it also becomes the unit of design. And that requires a new form of engineering, from silicon to rack-scale system.
Part 2: Designing AI Systems at Rack Scale (coming soon)
Part 3: Compute Only Scales When Data Can Move (coming soon)
Part 4: What the Next AI Stack Makes Possible (coming soon)