AI in Practice

Microsoft Caps AI Token Consumption of Its Own Employees

3 min read

TL;DR Too Long; Didn’t read

Microsoft has been capping the AI token consumption of its business units since July 2026 and made the cheaper model GPT-5.6 the internal default. According to an internal memo from CoreAI chief Jay Parikh, the move follows sharply rising costs since teams ramped up agentic coding tools. Individual developers reportedly spent several thousand dollars a month on tokens, the memo states.

A hand with a Microsoft logo sticker on the sleeve presses a lid onto an overflowing pot full of glowing digital coins. Image generated with GPT Image 2

Key takeaways

  • CoreAI chief Jay Parikh sent the memo to Microsoft's developer teams this week, according to 404 Media.
  • Each corporate division now gets its own token budget as a target, tracked via an internal dashboard.
  • OpenAI's cheaper GPT-5.6 becomes the new default for internal AI requests at Microsoft.
  • Agentic coding tools like GitHub Copilot in particular drove per-employee token costs up, the memo states.
  • Tesla, Uber, Amazon, and Walmart had already introduced similar AI spending limits beforehand.
  • Microsoft keeps marketing Copilot externally without restriction as a tool for every workforce.

Microsoft has set fixed upper limits for AI token consumption for its business units starting July 2026 and is making the cheaper model GPT-5.6 the internal standard. According to an internal memo from CoreAI chief Jay Parikh, the reason is a significant increase in costs since developer teams have increasingly been using agentic tools like GitHub Copilot. Individual engineers reportedly spent several thousand dollars per month on tokens previously.

Microsoft sets up token budgets for each business unit

Parikh, as Executive Vice President, leads the CoreAI division, which is responsible for Microsoft’s developer tools – including Copilot and GitHub – as well as a large part of the internal AI infrastructure. According to 404 Media, which first published the email, he sent it this week to engineering teams within the company. Each division will reportedly receive its own token budget as a target figure, and individual employees will also be able to track their consumption via an internal dashboard.

Parikh writes literally, “Tokenmaxxing is not what we are optimizing for” – consuming as many tokens as possible is not the goal. Instead, the workforce should focus on results that matter for customers and the business.

As another cost-saving measure, OpenAI’s GPT-5.6 becomes the default setting for internal requests because it is reportedly cheaper than previously used alternatives. Already in July, the company had begun redirecting some requests in Excel and Outlook to its own MAI models to cut costs for models from OpenAI and Anthropic. Both steps aim at the same goal: spending less money per completed task without cutting overall AI usage.

Costs soared despite cheaper token prices

Individual developers reportedly spent between several hundred and a few thousand dollars per month on tokens, according to the memo – independently unverified. The cause is primarily the shift from simple autocomplete suggestions to agentic tools that independently execute multiple work steps in a row, consuming a multiple of the tokens in the process. According to the magazine TheNextWeb, token prices have fallen by roughly 98 percent since late 2022, while many companies’ AI bills tripled over the same period – higher volume, in other words, fully eats up the savings per request.

Parikh stresses in the memo that the goal is not fewer tokens, but more impact per token. An anonymous Microsoft employee told 404 Media the move feels like an admission that the company can barely afford its own AI products internally. For workforces outside Microsoft, the case offers a sober takeaway: teams rolling out agentic coding tools should track cost per solved task rather than just the token price – otherwise similar surprises may follow, as they did for Microsoft’s own engineers. Microsoft CEO Satya Nadella had already warned of such a cost trap in AI usage back in July.

Other companies are cutting their teams’ AI spending too

Microsoft is not the first company to limit its employees’ AI usage. Tesla capped its employees’ AI spending at $200 per week back in July 2026, after internal leaderboards had previously fueled consumption. According to the same report, Uber, Meta, Amazon, and Walmart had already introduced comparable spending limits. TheNextWeb also reports that Adobe, Atlassian, and the bank Citi have cut their internal AI budgets as well.

Externally, Microsoft continues to promote Copilot without reservation as a tool every company should roll out to its entire workforce – the same message the company has used for months to market its own AI products. Internally, the opposite motto now applies: not more consumption, but more impact per token spent. According to TheNextWeb, the cost-cutting has not dented the business figures so far: Microsoft’s AI business keeps growing sharply, and the memo is aimed solely at internal cost control.

What matters now is whether the internal austerity squares with Microsoft’s external message that every company should deploy as many Copilot licenses as possible. If “impact per token” becomes the internal yardstick, it could become a model for other companies also struggling to keep their AI budgets in check. It also remains open how Microsoft intends to convince its own developer teams to use more sparingly a tool the company simultaneously markets as indispensable.

Frequently asked questions

What does the internal term "tokenmaxxing" mean?

The term describes excessive AI use in which employees consume as many tokens as possible without weighing the benefit of each request.

How much did individual Microsoft developers previously spend on AI?

According to the internal memo, individual employees' monthly spending ranged from several hundred to a few thousand dollars. These figures have not been independently confirmed.

Which AI model does Microsoft now use internally by default?

OpenAI's cheaper GPT-5.6 is now the default for internal requests instead of pricier alternatives.

Does the budget cap also apply to paying Copilot customers?

No, the limits apply only to Microsoft's own employees. The public Copilot offering continues to be marketed unchanged.

Which other companies have similarly limited their AI spending?

Tesla, Uber, Amazon, Walmart, Adobe, Atlassian, and the bank Citi have reportedly already introduced comparable spending controls.

Sources (3)
  1. 404 Media: Microsoft Tells Engineers 'Tokenmaxxing Is Not What We Are Optimizing For'
  2. TheNextWeb: Microsoft tells employees to stop tokenmaxxing, sets division-level AI budgets
  3. TechRadar Pro: 'Tokenmaxxing is not what we are optimizing for' – Microsoft tells engineer to calm down on AI usage

Your AI update for the work week

Once a week, the most important AI news – plus one practical tip to try right away. No spam, unsubscribe anytime.

← Back to the blog