
For the last two years, enterprises have embraced AI coding assistants such as Claude Code, GitHub Copilot, Cursor and Codex because they dramatically improve developer productivity. Most organisations have focused on measuring productivity gains, but very few have realised they are creating an entirely new operating expense: token consumption.
Today, tokens are rapidly becoming one of the largest recurring costs in software engineering.

The challenge is that today's coding assistants are only the beginning. The industry is now moving towards autonomous coding agents capable of understanding entire codebases, planning work, modifying multiple repositories, testing, debugging and deploying software with minimal human intervention.
These agents are extraordinarily capable.
They are also extraordinarily hungry for tokens.
The Economics Have Changed
Traditional software licensing was predictable.
You bought seats.
You knew the annual cost.
AI has fundamentally changed this model.
Every interaction now consumes tokens. Every search of a codebase, every prompt, every reasoning step, every generated response and every tool invocation contributes to a continually growing bill.
The move from seat licensing to consumption-based pricing means engineering costs are becoming increasingly variable and difficult to forecast. Gartner predicts that by 2028, AI coding costs will exceed the average developer's salary if organisations fail to govern token consumption effectively.
The question is no longer:
"How much does our coding tool cost?"
It is becoming:
"How many billions of tokens are our developers consuming every month?"
Coding Agents Change Everything
Current AI coding assistants are largely reactive.
Developers ask questions.
The model answers.
Autonomous coding agents work very differently.
Instead of responding to a single request, they execute complex workflows that require them to continually gather context before making decisions.
An enterprise coding agent may need to:
Understand millions of lines of source code
Search documentation
Inspect APIs
Read design documents
Understand dependencies
Analyse previous commits
Review tests
Validate architecture
Generate implementation plans
Execute multiple iterations before producing code
Every one of these activities consumes input tokens.
Ironically, writing the final code often represents only a tiny fraction of the total token bill.
Input Tokens are the Real Cost
Most people assume AI spends its money generating code.
The evidence suggests otherwise.
A 2026 study from researchers at Stanford, Michigan and MIT1 analysed eight frontier models running agentic coding tasks and found they consume over 1,000× more tokens than traditional code chat, with input tokens — not output tokens — the dominant cost driver. The same study found repeated runs of an identical task could vary by as much as 30× in token consumption, and that spending more tokens did not reliably improve accuracy.
This changes how enterprises should think about AI.
The expensive part is no longer code generation.
It is understanding enough context to generate the right code.
More Context Creates Better Results... But At What Cost?
As coding agents become more autonomous, they require dramatically larger context windows.
They need to understand:
Multiple dependent repositories
Documentation
Infrastructure
Security policies
APIs
Database schemas
Build systems
Historical decisions
Company coding standards and best practices
This additional context undoubtedly improves quality.
Unfortunately, it also increases token consumption exponentially.
The industry is entering a classic efficiency paradox.
Model pricing continues to fall, but overall enterprise AI spend continues to rise because organisations are consuming vastly more tokens than ever before. Recent market data shows that lower token prices have actually driven significantly higher usage, increasing overall AI revenues rather than reducing them.
Agentic AI Is Multiplying Token Consumption
Several industry forecasts now point in the same direction.
Goldman Sachs Research expects token consumption to multiply 24 times between 2026 and 2030, reaching 120 quadrillion tokens a month, as consumers and enterprises adopt agentic AI.
Engineering platforms are already seeing the impact.
Jellyfish, which tracks token spend across 12,000 developers at 200 companies, reports that roughly 0.5% of the developers it tracks now spend as much on tokens as they earn in salary, with the top 2% of engineers consuming 500 million tokens per developer per week.2
TechCrunch reports enterprises exhausting annual AI budgets within a few months as agentic features become mainstream - Uber burned through its entire 2026 AI coding budget by April, and J.R. Storment of the Linux Foundation's FinOps Foundation describes companies telling him "we are 3x over our entire 2026 token budget and it's only April" 3, 4
Furthermore Gartner expects the outlier to become the norm, predicting these costs "will meet, or even exceed, the typical software engineer's monthly salary within the next two years".5
This is no longer a theoretical future problem.
It is happening now.
Why Traditional Optimisation No Longer Works
Historically, organisations controlled infrastructure costs by purchasing cheaper hardware or negotiating better software licences.
Neither approach solves the token problem.
Even if models become cheaper, autonomous agents simply consume more tokens.
Larger context windows.
Longer reasoning chains.
More planning.
More verification.
More tool calls.
More iterations.
The result is a new form of Jevons Paradox: cheaper tokens encourage organisations to use more AI, which ultimately increases overall spending rather than reducing it.
The Next Enterprise Architecture Challenge
Most organisations currently rely on large language models to perform both reasoning and information retrieval.
That is becoming increasingly inefficient. Using expensive GPUs and large language models to repeatedly search code, documentation and repositories is equivalent to using a Formula One car to collect the post.
As autonomous agents scale across thousands of developers, enterprises will increasingly separate these responsibilities.
Fast, deterministic retrieval systems will locate the precise information required, while the language model focuses on reasoning and code generation.
This dramatically reduces unnecessary context, lowers token consumption, improves response times and enables agents to scale economically across the enterprise.
The future is unlikely to be built around larger context windows alone.
It will be built around smarter architectures and harnesses that minimise the amount of context the model needs to process.
The Next Major FinOps Challenge
Cloud computing introduced Cloud FinOps.
AI will introduce Token FinOps.
Engineering leaders will increasingly need visibility into:
Token consumption by developer
Cost per repository
Cost per coding agent
Cost per feature delivered
Input versus output token ratios
Context efficiency
Token ROI
Without these capabilities, AI budgets become impossible to predict or optimise.
The organisations that succeed with agentic software development will not simply have the smartest models.
They will have the most efficient architectures.
Conclusion
The software industry is entering a new economic era. For decades, the cost of writing software was measured in developer salaries.
Tomorrow it will increasingly be measured in tokens.
AI coding assistants have already changed how developers work.
Autonomous coding agents will fundamentally change how enterprises consume compute.
The companies that recognise this early—and build architectures that minimise unnecessary token consumption rather than simply purchasing larger models—will gain a lasting competitive advantage.
The future of software engineering will not be won by the organisation that can afford the most tokens.
It will be won by the organisation that needs the fewest.
References and Citations
Bai, L., Huang, Z., Wang, X., Sun, J., Mihalcea, R., Brynjolfsson, E., Pentland, A., & Pei, J. (2026). How Do AI Agents Spend Your Money? Analyzing and Predicting Token Consumption in Agentic Coding Tasks. arXiv:2604.22750. https://arxiv.org/abs/2604.22750
Arcolano, N. (2026, April 15). Is "tokenmaxxing" cost effective? New data from Jellyfish explains. Jellyfish. https://jellyfish.co/blog/is-tokenmaxxing-cost-effective-new-data-from-jellyfish-explains/ — 12,000 developers across 200 companies, Q1 2026. Median $52.38/month, 90th percentile $691.14/month, and cost per merged PR rising from $0.28 in the lowest usage tier to $89.32 in the highest.
Bellan, R. (2026, June 5). The token bill comes due: Inside the industry scramble to manage AI's runaway costs.TechCrunch. https://techcrunch.com/2026/06/05/the-token-bill-comes-due-inside-the-industry-scramble-to-manage-ais-runaway-costs/
Plumb, T. (2026, June 24). AI coding token costs are on track to rival human payroll. CIO. https://www.cio.com/article/4189149/ai-coding-token-costs-are-on-track-to-rival-human-payroll.html
Goldman Sachs Research. (2026, May 20). AI agents forecast to boost tech cash flow as usage soars. Goldman Sachs Insights. https://www.goldmansachs.com/insights/articles/ai-agents-forecast-to-boost-tech-cash-flow-as-usage-soars