TechSignal.news
Enterprise AI

Enterprise LLM Inference Prices Drop 43% in 10 Weeks to $1.16 Per Million Tokens

Average enterprise inference pricing fell to $1.16–$1.18 per million tokens in early August, down from $2.04 in May. The collapse forces buyers to rethink cost models, vendor selection, and multi-model routing strategies.

TechSignal.news AI4 min read

The Price Floor Just Moved

Enterprise LLM inference prices hit $1.16–$1.18 per million tokens between August 6–8, according to Jefferies research using Silicon Data. That is the lowest level recorded in 2026, down 43% from $2.04 in late May and 20% from $1.45 in late July. The driver is a global price war accelerated by Chinese open-source models, particularly DeepSeek, undercutting closed frontier APIs from OpenAI, Anthropic, and Google.

For buyers, this is not an incremental cost decline. It is a structural shift that changes how you model budgets, choose vendors, and think about lock-in. A 1-billion-token-per-month application now costs approximately $1,500 per month in raw model spend at $1.50 per million tokens, down from $2,040 in May. That difference compounds across large-scale deployments and makes multi-model routing economically mandatory, not optional.

Three Pricing Tiers Now Define Deployment Strategy

The market has stratified into three distinct bands. Open-weight models — Llama, Qwen, DeepSeek — are priced at $0.04–$0.30 per million input tokens when hosted on specialized inference platforms. Mid-tier models sit at $0.20–$1.50 per million tokens. Frontier closed models from OpenAI, Anthropic, and Google remain in the $5–$10 per million input tokens range, with output tokens costing $25–$50 per million.

BenchLM's frontier LLM price index dropped 88% from a March 2023 baseline of 100 to 12 today. A blended price across 10 flagship models fell to $4.39 per million tokens on August 1, down 3.9% from $4.57 in late February. Specific flagship pricing: GPT-5.6 Sol at $11.25 per million tokens, Claude Opus 5 at $10.00, Kimi K3 at $6.00, and Gemini 3.1 Pro at $4.50.

This creates clear decision points. Reserve frontier models for agentic or high-stakes tasks where quality and tooling justify the 2–3× price premium over mid-tier alternatives. Use mid-range models when quality is adequate and cost is sensitive. Route bulk, low-risk workloads to open-weight models at commodity pricing. A typical enterprise should now model 70–80% of requests at $1–$2 per million tokens and 20–30% at $5–$12 per million tokens for premium calls.

What This Means for Procurement and Risk

Buyers must now demand routing transparency and monthly cost reconciliation in contracts. Multi-model orchestration is the default deployment pattern, and price bands differ by an order of magnitude. Without visibility into which model handled which request, cost reconciliation becomes impossible and overbilling becomes inevitable.

Rapid price moves introduce cost-volatility risk. Contracts need explicit price-adjustment clauses and spending caps. If your vendor locks you into a fixed rate while the market drops 20% in a few weeks, you are paying a premium for stability that may not be worth it. Conversely, if you are on spot pricing without caps, a sudden reversal in pricing trends could blow your budget.

The rise of Chinese open-source models introduces a second layer of risk. Finance teams see DeepSeek and Qwen as low-cost alternatives. Regulatory and security teams must evaluate data residency, model provenance, and potential national security concerns, particularly for deployments in regulated industries or government-adjacent work. The unit economics are compelling, but the compliance and geopolitical risk may not be.

Google Undercuts OpenAI and Anthropic by Half

Google's Gemini 3.1 Pro at $4.50 per million tokens sits roughly halfway between the frontier pricing of GPT-5.6 Sol and Claude Opus 5 at $10–$11 per million tokens and the commodity open-weight tier. This is aggressive positioning. Google is betting that most enterprise workloads do not need the absolute highest-quality frontier model and that a 50% price cut will shift share.

For buyers, this creates leverage. If your use case does not require the cutting edge of reasoning or tool use, you now have a credible mid-tier alternative from a hyperscaler with existing enterprise relationships and compliance certifications. That pricing gap also gives you negotiating room with OpenAI and Anthropic — demand volume discounts or commit-based pricing that closes the gap, or route more traffic to Google.

What to Watch

Pricing will continue to fall as inference capacity scales and open-source models improve. The 88% decline from March 2023 is not done. Buyers should resist long-term pricing commitments unless they come with price-adjustment clauses tied to public indices or competitive benchmarks.

Watch for consolidation in the specialized inference layer. Companies building open-source-centric clouds are seeing investment (Groq raised $350 million in August), but the proliferation of inference providers also signals fragmentation. Winners will be determined by uptime, latency, and contract flexibility, not just raw pricing.

Finally, expect regulatory scrutiny of Chinese open-source models to intensify. If you are building pricing models that assume unrestricted access to DeepSeek or Qwen, have a fallback plan. The unit economics are real, but the geopolitical risk is not priced in.

LLMInferencePricingAI DeploymentOpen Source

Technology decisions, clearly explained.

Weekly analysis of the tools, platforms, and strategies that matter to B2B technology buyers. No fluff, no vendor spin.

More in Enterprise AI