Gimlet Labs AI inference cloud technology is expanding after the company raised $300 million in Series B funding, valuing the San Francisco-based business at $3 billion. The platform is designed to distribute different parts of inference workloads across GPUs, CPUs, and specialized accelerators rather than forcing the entire workload onto one processor architecture.
Announced September 4, the round was led by Andreessen Horowitz and included participation from Sapphire Ventures, M12, Arm, Menlo Ventures, Factory, and several other investors. Gimlet says the financing brings its total capital raised to $392 million.
The more important infrastructure figure may be outside the funding round itself: Gimlet says it has added billions of dollars in contracted revenue, developed a gigawatt-scale data center pipeline, and is moving toward hundreds of megawatts of managed capacity.
Gimlet Labs AI Inference Cloud Challenges Homogeneous Infrastructure
Most large AI environments remain heavily centered on one primary accelerator architecture. Gimlet’s argument is that inference increasingly contains several computational phases with different bottlenecks, making one processor type inefficient for every stage.
The company says its software can divide workloads across different hardware architectures and assign each phase to the processor best suited to it.
That can include GPUs, CPUs, SRAM-oriented accelerators, and other purpose-built AI chips.
Gimlet works with hardware from NVIDIA, AMD, Intel, Arm, Cerebras, and d-Matrix, according to the company.
Inference Is Becoming A Data Center Scheduling Problem
The architecture reflects a broader shift in AI infrastructure.
Training historically drove much of the industry’s focus on large homogeneous GPU clusters. Inference creates a different optimization problem because workloads can include prefill, decode, retrieval, tool execution, multimodal processing, and other stages with very different compute and memory characteristics.
Gimlet’s approach is to treat those stages as a scheduling problem across heterogeneous infrastructure.
The attraction for data center operators is improved utilization. If specialized processors can perform selected phases more efficiently than general-purpose GPUs, operators may be able to extract more useful inference capacity from a fixed power envelope.
Hundreds Of Megawatts Changes The Scale Of The Company
Gimlet says it is now scaling toward hundreds of megawatts of managed infrastructure and has a gigawatt-scale data center pipeline.
Those figures are company disclosures rather than independently audited deployment totals, and a pipeline should not be treated as operational capacity.
Even so, they indicate that Gimlet’s ambitions extend well beyond selling inference software.
The company operates its own managed cloud while also offering its software for deployment in customer environments. That puts it closer to the emerging neocloud model, where software optimization, hardware procurement, networking, and physical data center capacity are increasingly managed as one platform.
Performance Claims Need Real-World Validation
Gimlet says its heterogeneous architecture can achieve gains of up to 10 times in throughput and interactivity for certain workloads.
That is a vendor performance claim, not evidence that every application will see the same result.
The economics will depend on workload composition, model architecture, interconnect performance, software overhead, accelerator pricing, utilization, and the ability to keep different processor pools balanced.
Heterogeneity can improve efficiency, but it can also create more complex scheduling, observability, failure handling, and capacity-planning requirements.
Why The $300 Million Funding Round Matters
The $300 million round shows that investors are increasingly willing to finance alternatives to the assumption that AI infrastructure must scale primarily by adding more identical GPUs.
For CIOs and infrastructure architects, that matters because inference economics may increasingly be determined by orchestration rather than accelerator performance alone.
If different chips can be pooled effectively, organizations could gain more flexibility in procurement and reduce dependence on one hardware architecture for every stage of an AI service.
The challenge is proving that those gains survive production complexity at data center scale.
What Happens Next For Gimlet Labs
Gimlet plans to use the new capital to expand its multi-silicon cloud and its team while increasing managed infrastructure capacity.
The next meaningful evidence will be customer deployments, independently reproducible performance results, and the amount of its stated data center pipeline that becomes operational capacity.
The underlying idea is worth watching regardless of the outcome: as AI inference expands, the industry may optimize not only which accelerator runs a model, but which accelerator handles each distinct stage.
Conclusion
The Gimlet Labs AI inference cloud represents a bet that heterogeneous computing can improve the economics and utilization of large-scale inference infrastructure.
If different chips can be pooled effectively, organizations could gain more procurement flexibility while extracting more work from constrained power and data center capacity. The operational challenge will be coordinating those resources without creating excessive scheduling, observability, networking, or reliability complexity.
As Gimlet converts its pipeline into operating capacity, customers will be watching whether its performance claims remain measurable at production scale. The broader lesson is clear: the future of AI inference may depend as much on intelligent orchestration as it does on any single accelerator.

