The Models Are Getting Cheaper. Enterprise AI Bills Keep Rising.
.png)
On September 14, 2026, the temporary 50% boost Anthropic added to its weekly rate limits in May expires. Anthropic will replace it with a smaller, permanent increase of 25% over the original baseline.
Consider a baseline of 100. Before May, users had 100. For the past four months, they had 150. Going forward, they get 125. That amounts to a 25% increase over the original baseline and a 17% reduction from the capacity teams have been using. Anthropic acknowledged the reduction in its announcement, which was posted, pulled, and then reposted with clarification.
Anthropic extended the boost through mid-July, then July 19, August 19, and August 31, each time suggesting it might become permanent. Teams incorporated that capacity into their weekly workflows. After running for four months and receiving five extensions, the allowance had become part of how people worked, regardless of its temporary label.
The episode highlights the risk of dependency. When a company's daily knowledge work runs through a single closed AI vendor, that vendor controls its effective unit economics, throughput, and ability to scale internal AI workflows. A plan update quickly becomes a cost and capacity problem for the finance team.
Premium models have become the costly default for everyday work
Open models are improving fast, which is making the underlying intelligence cheaper. On many everyday knowledge-work tasks, their performance is close enough to frontier models that the difference in price deserves serious consideration during procurement.
Enterprise AI bills continue to climb because companies buy more than raw intelligence. They also buy a set of defaults: the model employees use for every request, the vendor their workflows depend on, the interface where work accumulates, and a pricing structure outside their control.
The costliest default is using the best available model for every task. Meeting summaries and high-stakes legal analyses call for different budgets. The same goes for a routine research brief, which often runs on the most expensive model simply because that model is the default button in the interface. Much of enterprise AI still works this way, applying one expensive model indiscriminately. As usage grows, bills rise even faster and common workflows become increasingly dependent on one vendor.
Claude Sonnet 5 is a strong model, and some work genuinely deserves premium capacity. Companies still need to decide whether it should automatically handle every piece of knowledge work.
DeepSeek V4 Flash costs up to 23x less while performing within a few points of Claude Sonnet 5
The most relevant comparison is DeepSeek V4 Flash against Claude Sonnet 5, the workhorse model many teams use across Claude.ai and the Claude API.
According to Artificial Analysis, DeepSeek V4 Flash costs $0.44 per million input tokens and $1.32 per million output tokens. Claude Sonnet 5 costs $2.00 and $10.00 for the same. DeepSeek is roughly 4.5x cheaper on input and 7.6x cheaper on output. On a per-task basis, the difference reaches about 23x: $0.22 per task versus $5.09.
Performance is much closer. On the Artificial Analysis Intelligence Index, Sonnet 5 leads 38 to 35, a real but modest edge. Sonnet 5 scores 1501 on GDPval-AA v2 compared with DeepSeek's 1456, while its AA-Briefcase score is 1355 against 1257. Sonnet remains stronger on raw intelligence, but DeepSeek comes within a few points on the work enterprises actually do while costing a fraction as much.
A serious AI strategy assigns models according to the task. For most organizations, the amount of work that needs a premium model is “less than we currently assume.” Lower-cost models can draft solid first versions, synthesize the right internal material, and prepare structured briefings with the right context. Using a premium closed model for all of that work wastes budget without improving quality control.
Company context helps lower-cost models handle more enterprise work
Context is the most important variable in enterprise AI, and raw model comparisons often leave it out. A model lacks knowledge of how your company approves a renewal, how your marketing team talks about the product, which documents are current, and where your regulatory team draws the line. Even powerful models have to guess when that information is missing. Give a cheaper model the right context and its performance improves dramatically on the work that matters to your company.
Most enterprise AI tasks require the reliable application of institutional knowledge, including approved messaging, team playbooks, prior decisions and corrections, current systems of record, and accumulated feedback that shows an agent what good work looks like. When every agent can access that context, and corrections from one interaction become durable knowledge for the next, employees can write shorter instructions. Agents can also handle harder, multi-step work while producing more consistent output across the team.
Context also makes switching models less disruptive. When value accumulates in individual prompts and a vendor-specific interface, moving to another model requires teams to rebuild how they work. Keeping context, workflows, governance, and agents in a platform layer above the model lets companies select the best model for each task without starting over. The company's knowledge becomes the source of its advantage, independent of any particular lab's pricing page.
A model-independent architecture reduces costs and vendor lock-in
Dependence on a closed vendor creates a structural imbalance. Employees adopt the tool, workflows begin to rely on it, and usage grows. The vendor can then change pricing, packaging, access, or capacity while customers face the cost of rebuilding how work gets done elsewhere. That sequence creates vendor lock-in, even when the change is presented as a plan improvement.
Companies can continue using Claude and other frontier models while choosing a different default for everyday work. A durable enterprise AI architecture routes each task to the right model, using lower-cost open models for high-volume, routine work and reserving premium models for cases where the additional quality justifies the additional cost.
This architecture keeps organizational context, documents, workflows, feedback, and permissions independent of the underlying model. Employees get one interface they will actually use, which reduces fragmented tools, shadow AI, and unmanageable spending. Leaders gain visibility into usage, cost, and agent performance, allowing adoption to build across the organization instead of resetting with every new hire.
elvex is built to provide this platform layer. DeepSeek V4 Flash is available on elvex today, with V4.1-Flash landing shortly. Along with model access, elvex provides organizational context that compounds, a unified interface employees actually use, and the flexibility to keep agents and workflows independent of any provider's pricing or roadmap. Enablement is built around the company's real workflows instead of a pile of licenses.
Review AI capacity before the next vendor change forces the issue
Companies should review their approach to AI capacity while they still have options, before another rate-limit update, packaging change, or price adjustment turns the issue into an emergency.
Open models are within a few points of the frontier workhorse on the tasks that define everyday enterprise productivity, and they cost a fraction as much. Supported by the right context, workflows, and governance, these models deliver results worth far more than their benchmark scores suggest.
Anthropic's rate-limit change gives enterprises a practical reason to reconsider their defaults. They now have an alternative to paying premium prices for every task and allowing one vendor to determine the cost and capacity of company-wide AI.
Transform your workflows today
Learn how we can help you modernize your business.

.avif)