← The DataCenter Blog
Data Center AI

Gimlet Labs and Cerebras Plan 100MW of Wafer-Scale AI Inference

Cerebras SystemsSep 30, 2026

Gimlet Labs, an AI inference cloud startup backed by Andreessen Horowitz and Menlo Ventures, announced a partnership with Cerebras Systems (NASDAQ: CBRS) on September 28 to bring the chipmaker's wafer-scale compute into Gimlet Cloud. The companies plan to deploy 100 megawatts of Cerebras-powered inference capacity, with the first Cerebras-powered Gimlet Cloud datacenter expected to come online later this year.

Combining wafer-scale chips with GPUs

Gimlet Cloud combines Cerebras' Wafer Scale Engine with GPUs in what the companies call a disaggregated inference architecture, routing each phase of an AI model's execution to whichever hardware suits it best. The goal is speeds of up to 3,000 tokens per second for latency-sensitive agentic and real-time applications.

"Inference speed matters. It determines how productive AI can be," said Zain Asgar, co-founder and CEO of Gimlet Labs. Sean Lie, co-founder and CTO of Cerebras, said combining "the fastest tokens from Cerebras with the highest throughput GPUs delivers the best datacenter economics for everyone."

Building on an existing relationship

The partnership builds on joint customer engagements the two companies have run since last year, with an integrated solution already serving inference traffic in private deployments. Gimlet Labs will also serve as a launch partner for Cerebras' next-generation CS-4 hardware, with Gimlet Cloud customers expected to gain direct access to it in 2027.

Read more from Gimlet Labs.