85% of Enterprise LLM Projects Fail — New Data Shows Why and How to Fix It
Recent enterprise adoption research quantifies LLM deployment failure rates and cost overruns, shifting budget priorities from model selection to data governance and monitoring infrastructure.
72% of Enterprises Increase LLM Spend Despite 85% Failure Rate
Kong Inc.'s 2025 enterprise generative AI report reveals that 72% of enterprises plan to increase LLM spending this year, with nearly 40% committing more than $250,000 to LLM initiatives. This spending increase occurs despite newly published data from Atlan showing that 85–95% of enterprise LLM projects fail to achieve their intended outcomes. The divergence between rising budgets and documented failure rates is forcing a fundamental shift in how enterprises approach LLM deployment — away from model experimentation and toward standardized governance, cost controls, and multi-model architectures.
The Kong data establishes a clear benchmark for serious LLM programs. Enterprises investing $250,000 or more are no longer funding ad-hoc pilots. Instead, they are building platforms that standardize data access, controls, and evaluation across multiple LLM endpoints. This architectural shift favors API gateway providers, data catalog platforms, and observability tools over single-model commitments. CIOs now have peer baseline numbers when presenting multi-model platforms and vector database budgets to CFOs.
Infrastructure Consumes 80% of Implementation Time, Not Model Choice
Atlan's enterprise guide to LLMs provides the mechanism behind the failure rates. Infrastructure and context engineering consume approximately 80% of implementation time in production LLM deployments, with model selection accounting for only the remaining 20%. More critically, 61% of teams skip data auditing and governance before selecting models, creating downstream failures that no amount of prompt engineering can fix.
A documented case from ZenML illustrates the cost consequences: an LLM agent loop generated $47,000 in unintended compute costs before a budget alert stopped it. This was not a model failure but a monitoring and governance failure. The incident type is now appearing in risk committee discussions as evidence for mandating cost caps, rate limits, and usage alerts in all LLM contracts before scaling pilots.
Atlan prescribes a data-first deployment sequence that directly contradicts the common practice of starting with model selection:
1. Govern and catalog data with documented ownership, lineage, and sensitivity classification 2. Choose model architecture based on privacy and latency requirements, not benchmark scores 3. Build a Retrieval-Augmented Generation layer with access-controlled retrieval 4. Implement governance and lineage tracking for AI use 5. Deploy with monitoring for response quality, latency, and retrieval relevance
This sequence shifts budget allocation from AI labs to data and platform teams. Enterprises are redirecting spending toward data cataloging, metadata and lineage tools, RAG infrastructure, and monitoring pipelines. Model providers — whether OpenAI, Anthropic, Google, or open-source stacks — are increasingly evaluated on native support for data governance integration, built-in monitoring APIs, and budget control features rather than tokens per second or MMLU scores.
Multi-Model Architectures Replace Single-Vendor Lock-In
The Kong report's emphasis on standardization drives a specific architectural pattern: multi-LLM deployments brokered through a common gateway with shared governance. Enterprises want standardized policy enforcement, routing, and observability for multiple model endpoints rather than committing to a single provider. This pattern favors API gateway and AI gateway vendors that can sit in front of multiple LLMs and provide centralized authentication, routing, and rate limits.
The architectural shift creates a new competitive landscape. Data catalog and lineage platforms such as Atlan, Collibra, and Alation gain leverage because LLM success is now framed as contingent on governed data assets and metadata. Vector databases like Pinecone and Weaviate become first-class components because retrieval accuracy is positioned as the quality lever. Cost and usage observability tools become mandatory to prevent runaway loops.
Model benchmark scores lose relevance as selection criteria. Enterprises evaluating vendors now ask: Does this platform integrate with our existing data catalog? Can we set cost caps and usage alerts? Does it support access-controlled retrieval? Can we route traffic across multiple models without rewriting application code?
What This Means for 2H26 Budgets and RFPs
The combination of Kong's spending data and Atlan's failure analysis creates a new baseline for vendor evaluation. Any platform pitching LLM deployment without concrete answers for cost governance, multi-model routing, and evaluation will be filtered out in RFPs. Finance and risk teams are demanding hard guardrails as part of the initial pitch, not as optional add-ons.
For enterprises in planning cycles, the 85–95% failure rate and $47,000 runaway cost example provide board-level justification for three changes:
- Mandating data governance platforms and observability as prerequisites before scaling LLM pilots - Requiring cost caps, rate limits, and usage alerts in all LLM contracts and statements of work - Shifting budget from model experiments to data cataloging, RAG infrastructure, and monitoring pipelines
The Kong report raises the procurement bar. Enterprises now have quantified peer benchmarks for spend levels and architectural patterns. Vendors without multi-model support, data governance integration, and cost controls will face longer sales cycles and increased scrutiny from finance teams holding them to the $250,000+ benchmark.
Technology decisions, clearly explained.
Weekly analysis of the tools, platforms, and strategies that matter to B2B technology buyers. No fluff, no vendor spin.
