On October 2, 2026, Epoch AI published “How many AI agents could we run?” by Jason Li. The report asks a hardware question: if all the AI accelerators shipped from 2025 through 2027 were used to serve agents, how many could run at the same time? It uses high-bandwidth memory (HBM) as the common denominator, counting HBM3E and HBM4/4E shipments in 288 GB units matching a GB300 GPU, and projects roughly 2.21 billion GB of HBM shipped in 2025, 3.75 billion GB in 2026 and 5.81 billion GB in 2027.
For open models, Epoch takes serving concurrency from SemiAnalysis benchmark data on sessions per GPU. For closed frontier models it infers concurrency from prices, using a reference of about 30 dollars per agent-hour of API spend, about 5 dollars per GB300-hour of rental and an assumed 5 to 10x ratio between API revenue and serving cost. The central result is 30 to 170 million concurrent frontier-model agents. Because agents can work all 168 hours of a week, that supplies as many weekly working hours as about 140 to 720 million full-time employees. With an efficient open model such as DeepSeek V4 Pro the same memory could serve about 1.9 billion concurrent agents.
Why it matters: the report turns a vague claim, that AI labor could rival the human workforce, into numbers with stated assumptions. The bottleneck it points to is demand rather than chips: using just 20 percent of the central capacity would imply 2.6 to 5.3 trillion dollars a year of API-equivalent spending, against Epoch’s estimate that model developers’ combined revenue may reach about 1 trillion dollars by the end of 2027.
What it does not show: it is a supply ceiling, not a forecast of deployment. It assumes all shipped memory is deployed and allocated to these workloads, holds today’s models and workloads fixed, and treats an agent-hour as comparable to a human work-hour, which says nothing about whether the agents could do the work.