SAN FRANCISCO — Hyperscale cloud providers pour billions into AI infrastructure, yet real-time inference costs remain stubbornly high. Meta, Microsoft and Amazon are all grappling with efficiency plateaus as they scale large language models for production workloads.
The problem traces to the limitations of general-purpose GPUs for the specific demands of inference—driving a quiet, multi-billion-dollar push into custom silicon and novel interconnects across the industry.
Meta Platforms, with its $60 billion AI infrastructure budget, allocates a significant portion to its custom MTIA inference chips. Alphabet's Google Cloud relies heavily on its Tensor Processing Units, now in their fifth generation, for both training and inference tasks.
These custom solutions aim to reduce the per-query cost of AI services. Cloud providers face pressure to maintain gross margins on AI offerings, which run at 20 percent to 30 percent—significantly below traditional SaaS multiples. Lower inference costs are critical for profitability.
Building proprietary chips creates a competitive moat. Amazon Web Services is developing its own Trainium and Inferentia ASICs to differentiate its cloud offerings and reduce reliance on external vendors.
The core problem lies in the unpredictable nature of inference workloads. Unlike training, which is batch-oriented, inference requires ultra-low latency and dynamic resource allocation, often with sparse data patterns. That makes traditional GPU architectures inefficient for many real-time applications.
Companies are pursuing diverse strategies: some focus on chip design, others on compiler optimization, and a few are exploring analog AI or optical computing for specific inference tasks. The result is a fragmented, complex landscape.
The capital allocation risk is substantial. Developing a custom ASIC costs hundreds of millions of dollars and takes years. A misstep can leave a company with obsolete hardware before it recoups its investment.
Nvidia, despite its dominant position in training GPUs, faces new competition in the inference segment. Its share of the total AI chip market, estimated at 80 percent, could erode if hyperscalers succeed with their internal solutions.
Nvidia shares traded at $190.01, down 3.6 percent, reflecting investor caution about future market segmentation. The broader Nasdaq fell 1.7 percent to 24,443.
Nvidia's data center revenue reached $26 billion in its most recent quarter, driven primarily by training chips. The inference market is projected to reach $50 billion by 2028.
The long-term winners will be those who can optimize their full stack—from silicon to software—for cost-effective AI inference. That requires deep vertical integration and sustained research and development investment.

