The AI oligopoly: Four companies control 90% of global GPU compute
Try VOLT Intelligence
Get StartedTry VOLT Cloud
Deploy GPUTable of Contents
- How GPU concentration happened why it’s so locked-in
- Supply chain capture
- Hyperscaler pre-purchase
- Switching cost
- What GPU concentration costs at scale
- Alternatives to Big Cloud GPU access that actually work
- Why VOLT is the structural hedge against oligopoly pricing
- There is no self-correct in the GPU oligopoly

Four companies control an estimated 85–92% of the accessible GPU compute that powers commercial AI training and inference globally. Those four companies are NVIDIA, Microsoft, Amazon, and Google. No surprise there.
NVIDIA manufactures roughly 80% of the high-end GPUs in existence. Meanwhile, Microsoft, Amazon, and Google collectively absorb the majority of that supply through their cloud divisions before any capacity reaches independent buyers. This leaves startups and independent researchers competing for approximately 10–15% of available H100 and A100 capacity, often at 3–5× the raw hardware cost.
But there is a way around this monopolistic control and GPU hoarding. Decentralized alternatives now offer access to that same hardware class H100 PCIe at $1.49/hr on VOLT versus $3.09–$4.10/hr on AWS or GCP. Let’s explore how GPU manufacture and distribution became so centrally controlled, and how we can do something about it.
How GPU concentration happened why it’s so locked-in
The GPU oligopoly wasn't a grand conspiracy, although it may appear to be from the outside looking in. Instead, three major converging dynamics emerged over a compressed window of time from 2019 to 2024.
Supply chain capture
NVIDIA produces the only broadly-supported H100/A100-class accelerators for transformer-based AI workloads. AMD's MI300X is gaining ground, but CUDA lock-in (10+ years of optimized libraries, tooling, and developer muscle memory) means the installed base skews heavily NVIDIA. A single fab partner, like TSMC at 4nm, creates a physical ceiling on quarterly supply.
Hyperscaler pre-purchase
Beginning in 2022 and 2023, Microsoft, Amazon, and Google signed multibillion-dollar GPU supply agreements with NVIDIA. IDC estimates that by 2024, the three hyperscalers absorbed 60–70% of NVIDIA's H100 allocations before spot markets opened. Add Meta's direct-purchase cluster (100,000 H100s announced in early 2024) and the four entities now account for the vast majority of new H100 capacity entering the world each quarter.
Switching cost
Once a team's infrastructure, MLOps stack, and team skills are coupled to AWS SageMaker or GCP Vertex AI, two different types of switching costs: the dollar amount it takes to migrate, and the engineering time it takes to pull it off.
Hyperscalers know this, and charge a premium with limited churn. Call it a Hyperscaler tax; or, actually, one of the many hyperscaler taxes.
So, this means that a startup that needs 64× H100s for a training run is not negotiating with a competitive market. What they’re really doing is choosing between three vendors who set prices in loose coordination with each other's list rates.
What GPU concentration costs at scale
The markup over raw hardware cost is the clearest symptom.
Let’s look at an example of this in action, which you would see in the wild: a mid-stage AI startup running continuous inference on 4× H100 PCIe instances. The costs comparison between AWS and VOLT would look something like this:
- AWS: 4 × $3.09 × 720 hrs/month = $8,899/month
- VOLT: 4 × $1.49 × 720 hrs/month = $4,291/month
That's $4,608/month in savings, or $55,296 per year, on a single four-GPU workload. At 16× H100 (a modest fine-tuning cluster), the difference grows to $220K per year. That difference is a whole heck of a lot of potential infra and product reinvestment dollars or a new engineer’s salary: take your pick.
Alternatives to Big Cloud GPU access that actually work
The oligopoly's pricing power depends on the assumption that there is no viable alternative supply for serious AI workloads. That assumption is actually nonsense, and has been for some time.
Decentralized GPU networks aggregate consumer, prosumer, and enterprise-grade hardware from independent node operators into a unified compute marketplace. VOLT currently lists thousands of GPUs across 138 countries, the majority of which are H100 PCIe, RTX 4090, and A100 class cards sourced from data centers and mining operations that aren't affiliated with any hyperscaler.
One might assume that the difference from renting on AWS latency or uptime (modern decentralized clusters run comparable SLAs for batch workloads), but instead it’s price and availability during peak demand periods. In Q4 2023 and Q1 2024, H100 spot availability on AWS and GCP dropped to near-zero for days at a time. VOLT's distributed sourcing model means no single supplier failure creates a complete outage.
Why VOLT is the structural hedge against oligopoly pricing
VOLT was purpose-built as an alternative supply layer. It doesn't attempt to replicate SageMaker's managed-ML features. Instead, it competes on raw compute cost and availability for teams that run their own stack. The $1.49/hr H100 PCIe rate reflects a marketplace without hyperscaler overhead, sales teams, or enterprise SLA premiums baked in. For batch training, fine-tuning, and inference at scale, that's the relevant number.
There is no self-correct in the GPU oligopoly
The four-company structure is self-reinforcing through two mechanisms that regulation and new entrants are unlikely to disrupt quickly.
First, NVIDIA's roadmap advantages compound annually. H200, B100, and GB200 supply will follow the same pre-purchase pattern as H100, as hyperscalers have already signed 2025–2026 supply agreements. Startups looking for next-generation hardware through hyperscaler channels face the same queue problem.
Second, the financing requirements are prohibitive. Building a credible centralized GPU cloud to compete with AWS requires $5–10B in capex, regulatory approvals, and 3–5 years of runway before meaningful market share. Decentralized networks sidestep this by aggregating existing hardware, so the capex is distributed across thousands of node operators, not concentrated in a single balance sheet.
Let’s be real and acknowledge that the oligopoly's pricing structure will persist for centralized cloud. The relief valve is the decentralized layer, which has no interest in matching hyperscaler prices because its cost structure is fundamentally different.
Provision your first GPU cluster today
Related Questions
Q: What percentage of global GPU compute do the top four AI companies actually control?
Estimates vary by methodology. IDC, SemiAnalysis, and multiple equity research teams converge on 85–92% for H100/A100-class compute accessible through commercial channels. The exact figure shifts quarterly as NVIDIA releases new supply, but the structural concentration hasn't meaningfully changed since 2022. The key variable isn't the percentage — it's that the remaining 8–15% is what everyone else is competing for, which is what drives the 2–3× pricing premium over hardware cost.
Q: How does the AI compute oligopoly affect GPU pricing for startups?
Directly and significantly. When three companies absorb the majority of new H100 supply, spot availability on public markets tightens, and the remaining supply gets priced at a premium. AWS on-demand H100 rates run $3.09–$4.10/hr. The raw hardware amortizes at roughly $0.90–$1.10/hr over a three-year lifespan at data center power costs. The markup is real and structural, not coincidental.
Q: Can decentralized GPU networks replace hyperscaler compute for AI training?
For most batch training and inference workloads — yes, with caveats. Decentralized networks like VOLT work best for workloads that don't require tight integration with managed ML platforms (SageMaker, Vertex AI) or hyperscaler-native storage. Teams running PyTorch or JAX with their own orchestration layer can swap hyperscaler compute for decentralized compute with minimal friction. Multi-node training with InfiniBand-level interconnects is a harder case; most decentralized networks are better suited for parallel single-node jobs than tightly-coupled multi-node training.
Q: What is VOLT's H100 pricing compared to AWS?
VOLT lists H100 PCIe at $1.49/hr on-demand. AWS p5 instances with H100 SXM run $3.09–$4.10/hr depending on configuration and commitment level. The gap is consistent across all H100 variants — roughly 2–2.5× cheaper on VOLT for equivalent GPU-hours. VOLT's pricing reflects a marketplace model without hyperscaler overhead.
Q: Is NVIDIA part of the GPU compute oligopoly or separate from it?
NVIDIA occupies a different layer — hardware supply rather than cloud services — but it's the enabler of the oligopoly's structure. Because NVIDIA controls ~80% of production AI accelerators and the CUDA ecosystem creates deep lock-in, the cloud layer oligopoly (Microsoft, Amazon, Google) is built on NVIDIA's hardware monopoly. The two reinforce each other: NVIDIA benefits from the hyperscalers' scale purchases, and the hyperscalers benefit from CUDA lock-in keeping their customers on NVIDIA-based instances. Alternatives like AMD's MI300X and Google's TPUs create partial pressure, but haven't broken the pattern at scale.