Resources /

GPU Colocation: What AI Infrastructure Requires In A Data Center

Human and robotic hands connecting beneath a glowing AI lightbulb, representing artificial intelligence, digital infrastructure, and advanced computing innovation.

GPU colocation is becoming essential because the infrastructure requirements for artificial intelligence bear almost no resemblance to traditional enterprise computing. 

A typical business server rack draws 5 to 10 kilowatts of power. An AI rack packed with GPUs routinely demands 40 to 80 kilowatts, and the latest systems push well beyond that. NVIDIA’s Blackwell GB200NVL72 rack design, introduced in 2024, requires approximately 132 kW. 

The company’s roadmap shows systems requiring 250 to 600 kW per rack by 2027, with up to 576 GPUs working together in a single filing-cabinet-sized enclosure.

According to Uptime Institute’s Global Data Center Survey 2025, the most common rack density range across the industry remains just 5 to 9 kW per rack – a figure that’s held fairly constant over the past five years. Most existing data centers simply weren’t built for what AI demands. The facilities that can support high-density GPU deployments represent a fundamentally different category of infrastructure, and organizations pursuing AI initiatives need to understand what separates capable facilities from those that will become bottlenecks.

The Physics of AI Infrastructure

Understanding why AI workloads stress traditional data centers requires looking at what’s actually happening inside these systems.

Power Density Evolution

The progression from CPUs to GPUs for AI workloads represents an order-of-magnitude increase in power consumption. Traditional server CPUs run at roughly 150 to 200 watts per chip. GPUs designed for AI ran at 400 watts until 2022. State-of-the-art chips in 2023 reached 700 watts, and current-generation accelerators exceed 1,200 watts per chip.

A single AI server might contain eight of these chips, consuming nearly 10 kW just for the accelerators before accounting for CPUs, memory, storage, and networking. Stack ten of these servers in a rack, and you’re approaching 100 kW – ten times what that same floor space would require for conventional computing.

Hyperscalers currently operate large AI facilities at an estimated average density of 36 kW per rack, with projections suggesting this will grow at roughly 8% annually to approach 50 kW per rack by 2027. But leading-edge deployments already exceed these averages significantly, and the gap between average and peak requirements continues to widen.

Heat Generation Challenges

Power consumption translates directly to heat generation. Every watt of electricity consumed by computing equipment ultimately becomes heat that must be removed from the facility. A 100 kW rack generates roughly the same thermal load as running 50 space heaters continuously in a space the size of a phone booth.

Traditional air cooling approaches struggle with these densities. Moving enough air through high-density racks to prevent thermal throttling or equipment damage requires increasingly impractical volumes and velocities. The temperature differentials between inlet and outlet air become extreme, creating hot spots that air-based cooling can’t effectively address.

This physical reality explains why high-density facilities increasingly deploy liquid cooling technologies. Direct-to-chip liquid cooling, which circulates coolant directly to the hottest components, can handle power densities of 60 to 120 kW per rack. Immersion cooling, which submerges entire servers in dielectric fluid, supports densities of 100 kW and beyond – with some dual-phase immersion implementations handling upward of 150 kW per rack.

Power Delivery Complexity

Delivering 100+ kW to a single rack presents electrical engineering challenges beyond simply having enough total facility capacity. The cables, bus bars, power distribution units, and circuit breakers serving each rack must handle currents that would have seemed absurd for a computer room just a few years ago.

At standard data center voltages, a 100 kW rack draws over 400 amps – requiring conductor sizes and distribution equipment more typically associated with industrial facilities than IT environments. The physical space required for this electrical infrastructure competes with the space available for actual computing equipment.

This drives interest in higher voltage distribution approaches. Major hyperscalers are moving toward medium voltage distribution (up to 13.8 kV) and higher DC voltages (400VDC and 800VDC) that reduce current requirements for the same power delivery. Lower currents mean smaller conductors, reduced losses, and more efficient use of physical space. But these approaches require specialized equipment and expertise that some traditional data centers lack.

What Makes GPU Colocation “AI-Ready”

The term “AI-ready” has become marketing shorthand that obscures significant variation in actual capabilities. Facilities genuinely prepared for GPU-intensive workloads share certain characteristics:

Sufficient Power Capacity

Total facility power matters, but power available per cabinet matters more. A 10 MW facility that can only deliver 10 kW per rack is useless for AI deployments. Truly AI-capable facilities provide 30, 50, or 100+ kW per cabinet with the electrical infrastructure to actually deliver it.

High-density colocation designed for AI workloads incorporates oversized electrical distribution, redundant power paths, and the flexibility to allocate substantial power to individual deployments. This isn’t just about having big utility feeds – it’s about the internal distribution architecture that gets power from the service entrance to your specific rack.

Advanced Cooling Systems

Air cooling alone can’t support current-generation GPU infrastructure at density. Facilities claiming AI readiness should demonstrate:

Liquid cooling infrastructure. Whether direct-to-chip, rear-door heat exchangers, or immersion-ready, the facility needs cooling approaches that can handle 40+ kW per rack without relying solely on air movement.

Cooling capacity headroom. AI deployments generate more heat per square foot than design assumptions for traditional data centers anticipated. Facilities need cooling capacity that matches their power capacity, not just their historical average loads.

Flexible deployment options. Different AI hardware has different cooling requirements. The facility should accommodate various cooling approaches rather than mandating a single method that may not match your equipment’s design.

Robust Connectivity

AI workloads increasingly distribute across multiple systems, requiring high-bandwidth, low-latency communication between nodes. Facilities supporting AI deployments need:

High-speed internal networking. The ability to deploy 100 Gbps, 400 Gbps, and soon 800 Gbps connections between systems within the facility.

Diverse external connectivity. Carrier-neutral colocation facilities with multiple network options enable the kind of connectivity AI workloads demand – both to cloud resources for hybrid architectures and to external data sources for training and inference.

Low-latency paths. For distributed AI training that spans multiple facilities or connects to cloud GPU resources, network latency directly impacts training efficiency. Strategic facility locations and rich interconnection ecosystems minimize these delays.

Operational Expertise

High-density infrastructure operates at the edge of what’s physically manageable. Operational teams need experience with:

  • Power monitoring and management at densities that leave no margin for error
  • Cooling system maintenance and optimization for liquid-cooled environments
  • Rapid response to thermal events before they damage expensive GPU hardware
  • Electrical system management for non-standard voltage and current profiles

Facilities that have successfully supported high-performance computing, financial trading systems, or other demanding workloads often translate that expertise to AI deployments more effectively than those whose experience is limited to conventional enterprise hosting.

The Economics of GPU Colocation

Organizations pursuing AI initiatives face a fundamental infrastructure decision: build dedicated facilities, use cloud GPU resources, or deploy in colocation. Each approach has distinct economic characteristics.

Cloud GPU Economics

Cloud providers offer immediate access to GPU resources without capital investment in hardware. For experimental workloads, variable demand, or organizations still determining their AI strategy, cloud GPUs provide the flexibility that owned infrastructure can’t match.

However, cloud GPU pricing reflects both the scarcity of these resources and the convenience of on-demand access. Sustained GPU workloads – model training that runs continuously for weeks, inference serving with consistent demand – often cost significantly more in cloud environments than equivalent owned hardware. Organizations that have moved past experimentation frequently find cloud GPU costs unsustainable at scale.

Owned Hardware in Colocation

Purchasing GPU hardware and deploying it in colocation facilities shifts economics toward capital expenditure with lower ongoing operational costs. For predictable workloads, this model typically delivers better total cost of ownership than cloud alternatives – sometimes dramatically better for sustained utilization.

CoreSite’s 2025 State of the Data Center research found organizations increasingly moving AI workloads from public cloud to colocation, with generative AI applications, recommendation systems, and augmented AI all showing higher colocation deployment rates than the previous year. The primary drivers cited include performance requirements, cost optimization, and hybrid flexibility.

Colocation provides the physical infrastructure – power, cooling, connectivity, security – while you retain control over the computing hardware. This control enables:

  • Hardware selection optimized for your specific workloads
  • Configuration and tuning without cloud provider constraints
  • Security postures are impossible in shared cloud environments
  • Predictable costs without consumption-based surprises

Time-to-Deployment Advantages

Building dedicated AI infrastructure takes years. Even organizations with unlimited budgets face extended timelines for site selection, permitting, construction, and commissioning. Power availability has become a critical constraint – transformer lead times now exceed three years in many markets.

Colocation compresses this timeline dramatically. Facilities with available capacity and appropriate infrastructure allow AI deployments in weeks rather than years. This speed matters enormously when competitive advantage depends on AI capabilities that didn’t exist as requirements eighteen months ago.

As one industry analysis noted, small GPU cluster deployments of 4-5 racks face increasing difficulty finding appropriate colocation space as hyperscalers absorb available capacity in major markets. But facilities outside the most constrained markets – particularly in the Midwest, where power remains more available – offer options that coastal and primary markets can’t.

Geographic Considerations for AI Infrastructure

Where you deploy AI infrastructure affects available options, costs, and operational characteristics.

Power Availability

The single biggest constraint on AI infrastructure is power. Markets like Northern Virginia – “the data capital of the world” – show vacancy rates below 1% despite aggressive construction. Utilities in many primary markets can’t provide new large power connections for years.

This constraint is pushing AI infrastructure toward locations where power remains available. Indiana, Iowa, and similar Midwest markets offer power availability that coastal markets have exhausted. While these locations may seem less obvious than traditional data center hubs, the physics of AI infrastructure make power the primary consideration.

Netrality’s facilities in Kansas City, St. Louis, Indianapolis, and Chicago operate in markets with better power availability than the hyperscale-saturated primary markets, while still providing the connectivity and infrastructure capabilities AI workloads require.

Latency Requirements

Different AI workloads have different latency sensitivities:

Training workloads tolerate higher latency to users since they’re primarily internal operations. Training infrastructure can be located wherever power and cooling are available without significant user impact.

Inference workloads serving end users often need proximity to those users. Inference for consumer applications may require presence in multiple geographic regions to maintain acceptable response times.

Hybrid architectures that combine owned GPU infrastructure with cloud resources need locations with strong cloud connectivity. Facilities serving as hybrid cloud hubs enable architectures where training happens on owned hardware while burst capacity comes from cloud providers.

Network Centrality

Central U.S. locations offer network advantages that coastal markets lack. A facility in Kansas City or Chicago reaches both coasts with relatively balanced latency – useful for inference serving distributed user bases or for distributed training that spans multiple facilities.

The concentration of network infrastructure in carrier-neutral facilities like those Netrality operates creates connectivity options that more remote locations can’t match. Being at the intersection of major fiber routes provides path diversity and competitive carrier options.

Evaluating GPU Colocation Providers

When assessing facilities for AI deployments, look beyond marketing claims to verify actual capabilities:

Power Specifications

Ask specifically:

  • What is the maximum power available per cabinet?
  • What electrical distribution architecture supports that power?
  • What redundancy (N+1, 2N) is provided for power systems?
  • What is the timeline to provision high-density deployments?

Beware facilities that quote total building power without clarity on per-cabinet availability. A 50 MW facility means nothing if it can only deliver 15 kW to your specific deployment.

Cooling Capabilities

Verify:

  • What cooling technologies are deployed or supported?
  • What is the demonstrated cooling capacity per cabinet at operating conditions?
  • Can the facility support liquid cooling if your hardware requires it?
  • What happens if your deployment generates more heat than initially estimated?

Connectivity Infrastructure

Assess:

  • What carriers and cloud providers have direct presence?
  • What cross-connect options exist within the facility?
  • What latency can be expected to major cloud regions?
  • How is internal networking between your deployments handled?

Facilities with robust interconnection ecosystems provide options that limited-carrier facilities can’t match.

Operational Track Record

Request:

  • Uptime history and incident reports
  • Experience with high-density deployments specifically
  • Reference customers running similar workloads
  • Support capabilities and response time commitments

The Infrastructure Investment Driving AI Forward

The scale of investment flowing into AI infrastructure defies easy comprehension. In 2024, spending on data center infrastructure reached $290 billion globally. The four largest hyperscalers – Alphabet, Microsoft, Amazon, and Meta – invested nearly $200 billion in capital expenditure, with projections showing over 40% growth in 2025 as they race to build computational capacity for next-generation AI models.

This investment creates both opportunities and constraints. The opportunities come from the increasing availability of AI-capable infrastructure as providers expand capacity. The constraints come from competition for that capacity, particularly in markets where hyperscalers are absorbing everything available.

For organizations pursuing AI initiatives, the infrastructure decision increasingly determines what’s possible. Waiting for perfect facilities means waiting while competitors deploy. Accepting infrastructure that can’t actually support your requirements means performance problems, thermal throttling, and hardware that can’t deliver its potential.

The path forward requires honest assessment of requirements, realistic evaluation of options, and partnership with providers who understand what high-density AI workloads actually demand.


Planning AI or GPU infrastructure deployment?Netrality’s high-density colocation provides the power, cooling, and connectivity that AI workloads require. Our owner-operated facilities in Chicago, Kansas City, Philadelphia, Houston, St. Louis, and Indianapolis offer the infrastructure capabilities serious AI deployments demand. Contact our team to discuss your high-density requirements.