TechSignal.news
SaaS Infrastructure

AWS Cuts OpenAI GPT-5.6 Pricing 80% as IBM Commits $240M to Cost-Optimized AI

Amazon slashed Bedrock pricing for GPT-5.6 Luna by 80% while IBM locked in a $240M infrastructure deal with Together AI targeting open-source model inference costs.

TechSignal.news AI4 min read

AWS undercuts Azure on OpenAI model pricing

Amazon cut pricing for OpenAI GPT-5.6 models in Bedrock by up to 80%, effective July 30. GPT-5.6 Luna now costs $0.20 per million input tokens and $1.20 per million output tokens—down from implied prior pricing near $1.00 and $6.00 respectively. The Terra variant dropped 20%.

The reduction creates a concrete wedge against Microsoft Azure OpenAI Service. An enterprise running 1 billion input tokens monthly on Luna now pays $200 versus roughly $1,000 before the cut. For organizations comparing token economics across clouds, Bedrock pricing on this SKU is now materially cheaper than equivalent Azure OpenAI consumption for the same model.

This matters because AI inference costs have blocked broader rollouts. Buyers who capped token volumes to stay within budget can now expand usage—more queries, additional business units—without raising spend. Finance teams modeling multi-cloud AI procurement must now itemize per-token pricing differences rather than assume parity between AWS and Azure for OpenAI access.

The risk: usage growth outpacing price cuts. An 80% reduction does not prevent total spend from rising if token consumption increases more than 5x. FinOps teams need guardrails on query volume expansion, especially in pilot programs where early adopters may treat cheaper tokens as license to remove usage constraints.

IBM locks $240M infrastructure deal targeting open-source inference

IBM and Together AI announced a multi-year agreement worth $240 million to deploy NVIDIA HGX B300 clusters on IBM Cloud for open-source AI model inference. The deal positions IBM Cloud as a specialist platform for enterprises seeking cost-efficient alternatives to proprietary frontier models.

Together AI operates multi-tenant inference workloads for open-source models like LLaMA variants. The $240M contract provides contracted compute capacity at scale, enabling Together AI to negotiate volume pricing with IBM and pass savings to enterprise customers through lower per-request or per-token rates. Specific pricing has not been disclosed, but the partnership explicitly targets buyers looking to reduce inference costs by avoiding proprietary model premiums.

This creates a credible low-cost lane for enterprises optimizing AI budgets. Organizations running high-volume inference workloads on open-source models can compare IBM Cloud + Together AI economics against AWS Inferentia, Google Cloud Vertex AI, and Azure GPU instances. The NVIDIA HGX B300 hardware targets inference specifically rather than training, reducing the hardware cost overhead passed through to customers.

For procurement teams, the IBM deal signals that open-source model hosting is attracting large committed infrastructure investments. Buyers evaluating build-versus-buy for AI inference can benchmark Together AI's managed service pricing against the fully loaded cost of operating their own GPU fleet. The $240M commitment also reduces platform risk—Together AI's infrastructure footprint is now backed by a multi-year IBM contract rather than relying solely on startup capital.

What benchmark data shows about wasted spend

Flexera's 2026 benchmark data quantifies where enterprises still overspend. The average organization wastes 30% of cloud budget on idle resources, overprovisioned instances, and unoptimized storage. VendorBenchmark's FinOps report shows organizations with active cost optimization programs reduce infrastructure spend by 23% annually, but only 34% of enterprises have formal FinOps processes in place.

The AWS price cuts and IBM infrastructure deal both target the same problem from different angles. AWS reduces unit cost for proprietary models, lowering the marginal cost of each inference request. IBM and Together AI reduce total cost by shifting workloads to open-source models on cost-optimized hardware. Both approaches require procurement teams to model total cost of ownership—not just list prices.

For buyers negotiating cloud contracts in Q4 2026, the implication is clear: per-token pricing for AI inference is now a competitive variable across hyperscalers. Enterprises should benchmark Bedrock, Azure OpenAI, and Google Vertex AI pricing using actual workload token volumes rather than relying on vendor estimates. For open-source models, Together AI on IBM Cloud creates a pricing anchor that AWS and Azure must compete against on cost-per-inference metrics.

What to watch

AWS's 80% price cut on GPT-5.6 Luna puts pressure on Azure to match or justify a premium. If Microsoft does not adjust pricing, enterprises will shift OpenAI workloads to Bedrock for cost arbitrage. Watch for Azure OpenAI price changes in the next quarter.

The IBM–Together AI deal tests whether enterprises will accept infrastructure lock-in for lower AI costs. If Together AI demonstrates materially lower inference costs on IBM Cloud, AWS and Google will need to either match pricing or offer migration incentives. Together AI's ability to publish transparent per-token benchmarks against hyperscaler alternatives will determine how much market share this deal captures.

FinOps teams should audit AI inference spend monthly. Token consumption grows faster than budget holders expect, and price cuts do not prevent overruns if usage is unmonitored. Set alerts when token volume increases exceed 2x month-over-month to catch runaway costs before they compound.

cloud-cost-optimizationai-infrastructureaws-bedrockibm-cloudfinops

Technology decisions, clearly explained.

Weekly analysis of the tools, platforms, and strategies that matter to B2B technology buyers. No fluff, no vendor spin.

More in SaaS Infrastructure