The Army Burned Through Its Unlimited AI Tokens in Weeks, Forcing Emergency Limits

Image: Ars Technica AI
Main Takeaway
The US Army's DEVCOM command exhausted its promised unlimited AI token pool by mid-June 2026, forcing the CIO to reimpose usage caps and casting doubt.
Jump to Key PointsSummary
How the token crisis unfolded inside DEVCOM
Members of the Army's Combat Capabilities Development Command got an urgent email in mid-June 2026: they were burning through AI tokens too fast and needed to cut back immediately. The message carried an uncomfortable admission. The Army CIO had announced unlimited tokens in May 2026, but by mid-June the pool was empty and limits had to be re-established, according to the email obtained by Wired. The note added that while the Army renewed token usage at current levels for now, it was unclear whether the CIO pool would be renewed after October 1.
The timing stings. Just weeks earlier, the Department of Defense had publicly touted that nearly half of its 3.5 million employees were using AI at work. DEVCOM's experience punctures that narrative. Ars Technica confirmed the email's contents independently, noting that neither the Army, the DOD, nor Ask Sage responded to requests for comment.
What Ask Sage actually costs and how the math broke down
The Army's AI platform is Ask Sage, a multimodal generative AI workspace launched in May 2025 on a five-year, $49 million contract. It lets users run multiple large language models including Alphabet's Gemini, Meta's Llama, and OpenAI's ChatGPT through the cArmy Cloud, routing to Azure OpenAI Gov, Mistral, AWS Bedrock Gov, and Google Gemini. According to AI Weekly, adoption exploded from roughly zero to about 19,000 users in under 45 days, a figure then-Army CIO Leonel Garciga cited publicly.
Ask Sage bills by the token. Wired reports that the Army's enterprise pack came with 100 million tokens as part of an annual subscription, where each token equals about 3.7 characters of output. That sounds like a lot until you compare it with operational usage. Breaking Defense previously reported that the Defense Department burned through some 20 billion tokens per day during the 38-day Operation Epic Fury in Iran. A peacetime bureaucracy hitting its ceiling in weeks suggests the pricing model wasn't calibrated for an entire command suddenly treating LLMs like a utility.
The mismatch between unlimited promises and per-token reality
The Army CIO's May 2026 announcement framed AI access as an unlimited resource. That framing collapsed almost immediately. The Vietnam news outlet characterized it as a digital resource crisis, noting the surge in demand for AI-generated content overwhelmed the supply. AI Weekly frames the episode as the first real stress test of per-token pricing at government scale, where token and compute costs become the hard ceiling on which use cases can scale.
The Army has publicly flagged token and compute cost as the limiting factor before. Garciga's office was aware of the tension between encouraging adoption and controlling spend. What changed was the velocity. When a command like DEVCOM, which handles research and engineering, integrates LLMs into daily workflows across thousands of personnel, the token burn rate stops being a procurement abstraction and becomes a line item that blows past projections within a single billing cycle.
Why the October cliff matters beyond one command
The email's warning about October 1 is the part nobody should gloss over. The Army renewed tokens at current levels for now, but made no commitment beyond the fiscal year boundary. That's a budget signal. If the CIO pool doesn't get renewed, DEVCOM and potentially other commands lose access to the LLM workspace that has become embedded in their workflows in a matter of months.
Ars Technica raises another unresolved question: whether tokens used by regular DOD employees come from the same pool as those using AI tools on classified or secret information. If they share a budget line, the competition between unclassified productivity use and sensitive operational workloads gets zero-sum fast. The Intercept reported on Monday that the Defense Department's emphasis on AI continues unabated despite these resource constraints, suggesting the appetite isn't shrinking even as the funding mechanism shows strain.
What this reveals about government AI procurement
A $49 million contract over five years averages to just under $10 million per year for an enterprise-wide LLM workspace serving a workforce of millions. That's thin. When the token pool runs dry in weeks, the per-user economics don't pencil out at current usage patterns. AI Weekly notes that the Army CIO has publicly identified token and compute cost as the limit on which use cases can scale, which means the current contract structure forces a tradeoff between breadth of adoption and depth of usage.
The episode exposes a procurement model built for pilot programs being stress-tested by production demand. Ask Sage's pricing treats AI inference as a metered utility, but government budgeting cycles operate on annual appropriations. When consumption spikes unpredictably, there's no surge pricing mechanism or reserve pool to absorb it. The result is an embarrassing reversal: announce unlimited access, then send an email walking it back six weeks later.
What happens next for military AI access
The immediate fix is rationing. DEVCOM personnel have been told to limit use, which means the Army is throttling adoption of a tool it spent years and millions procuring. The longer-term question is whether the DOD restructures its contract with Ask Sage to accommodate actual demand, or whether the CIO lets the pool lapse in October and forces commands back to pre-AI workflows.
There's a broader signal here for any large organization eyeing enterprise AI deployments. Per-token pricing works for predictable, steady-state usage. It breaks when an entire workforce discovers the tool simultaneously and integrates it into everything from memo drafting to technical analysis. The Army's experience suggests that "unlimited" in AI procurement is a marketing term, not an engineering reality, and the bill always comes due faster than the budget cycle expects.
Key Points
DEVCOM exhausted its promised unlimited AI token pool by mid-June 2026, just weeks after the Army CIO announced unrestricted access.
The Army's Ask Sage platform operates on a five-year $49 million contract with 100 million tokens annually, each token representing about 3.7 characters.
Adoption surged to roughly 19,000 users in under 45 days, far outpacing the procurement model's capacity for production-scale usage.
The Army CIO renewed tokens at current levels but signaled uncertainty about funding continuing past the October 1 fiscal boundary.
The episode exposes a structural mismatch between metered per-token pricing and government budget cycles when workforce adoption spikes unpredictably.
Questions Answered
The Army's DEVCOM command exhausted its token pool because adoption of the Ask Sage platform surged from near zero to roughly 19,000 users in under 45 days. The enterprise subscription provided 100 million tokens annually, and the per-token pricing model was not calibrated for an entire research and engineering command suddenly using LLMs as a daily utility.
Ask Sage is the Army's multimodal generative AI platform launched in May 2025 on a five-year $49 million contract. It routes users through the cArmy Cloud to multiple large language models including Alphabet's Gemini, Meta's Llama, OpenAI's ChatGPT, Mistral, and models on Azure OpenAI Gov and AWS Bedrock Gov.
The email sent to DEVCOM personnel stated it is unclear whether the Army CIO pool will be renewed after October 1. The Army renewed token usage at current levels for now but made no commitment beyond the fiscal year boundary, creating budget uncertainty for commands that have integrated the platform into their workflows.
The Defense Department burned through approximately 20 billion tokens per day during the 38-day Operation Epic Fury in Iran, according to Breaking Defense. The Army's enterprise pack of 100 million tokens is a fraction of that operational burn rate, highlighting how even peacetime bureaucratic use can quickly exhaust a modest allocation.
The token exhaustion exposes a procurement model designed for pilot programs being stress-tested by production demand. Per-token pricing works for predictable usage but breaks when an entire workforce adopts the tool simultaneously, and annual government budget cycles lack the flexibility to absorb unpredictable consumption spikes.
Source Reliability
60% of sources are established · Avg reliability: 68
Go deeper with Organic Intel
Simple AI systems for your life, work, and business. Each one includes copyable prompts, guides, and downloadable resources.
Explore Systems