· 6 min read

Gimlet Labs Raised $300M to Mix Nvidia, AMD, and Arm Chips in One Inference Cloud

Gimlet Labs raised $300 million in a Series B led by Andreessen Horowitz, pushing its valuation to $3 billion and its total funding to $392 million. New investors Arm and M12 joined the round alongside existing backers Sapphire Ventures, Menlo Ventures, and Factory. That's a big number for a company most solo operators have never heard of, solving a problem that most solo operators feel every single month without knowing its name: your inference costs are shaped less by how efficient the model is and more by which chip your provider happened to route your request to.

What "disaggregated inference" actually means

An AI model serving a request doesn't do one uniform kind of computation from start to finish. Inference has distinct phases, prompt processing and token generation being the big two, and those phases have genuinely different hardware needs: one is more compute-bound, the other more memory-bandwidth-bound. Most providers today run the whole request on a single type of chip anyway, because building infrastructure that intelligently splits phases across different hardware is hard, and because there hasn't historically been enough commercial pressure to solve it.

Gimlet's pitch is that they've built the orchestration layer to do exactly that: route different phases of an inference workload to whichever chip actually suits it, whether that's an Nvidia GPU, an AMD chip, an Arm-based processor, or a purpose-built AI accelerator from a company like Cerebras or d-Matrix. Gimlet says it works with all of the above. The claimed upside is better latency, higher throughput, and lower cost per token, because you're not paying premium GPU pricing for a workload phase that would run just as well, and cheaper, on different silicon.

Where the money is going

Beyond the raise itself, the more interesting signal is what Gimlet plans to build with it. Part of the funding goes toward growing the capacity of their existing serverless inference platform, with plans to add several hundred additional megawatts of computing power. The more unusual part is a stated plan to build a custom, inference-optimized server that skips the traditional motherboard entirely, designed to run outside conventional data centers. That's a genuinely different bet than "we'll rent more GPU capacity," it's a hardware company move wrapped inside what started as a software orchestration pitch.

Why this matters even if you'll never touch Gimlet's product

I want to be upfront that a solo operator reading this is extremely unlikely to ever directly integrate with Gimlet Labs. This isn't a tool you're going to add to your stack next to Supabase and Resend. The reason it's worth paying attention to anyway is that Gimlet is one of several well-funded companies now competing specifically on the axis of "make inference cheaper and faster by using hardware more intelligently," rather than the axis most of the public conversation focuses on, which is which lab ships the next flagship model. Competition on that infrastructure axis, if it actually delivers, is one of the mechanisms that keeps per-token API pricing trending down over time rather than up, which directly affects the margin on anything you're building with an AI feature baked in.

The a16z-led valuation jump, from a company that had raised a combined $92 million before this round to $392 million total, at a $3 billion valuation, tells you institutional money thinks this infrastructure layer is where real value gets captured over the next few years, not just at the model layer. That's a useful signal even if you never write a line of code against Gimlet's API.

The honest take

I'd push back on the implied inevitability in a lot of the coverage I read while researching this, the kind that treats "multi-silicon inference" as an obviously solved problem now that a well-funded startup says they've solved it. Orchestrating workloads intelligently across genuinely different hardware architectures, at the reliability and latency SLAs enterprise customers demand, is a hard distributed-systems problem, and a lot of well-funded infrastructure startups have promised exactly this kind of abstraction layer before and shipped something narrower than the pitch. Gimlet working with Nvidia, AMD, Intel, Arm, Cerebras, and d-Matrix as partners is a genuinely strong signal of industry buy-in, but partnership announcements and production-grade reliability across six different chip architectures under real customer load are not the same claim.

My actual read: this is worth tracking as one input into where inference pricing goes over the next year or two, not something to bet your unit economics on today. If you're pricing an AI feature into your product right now, price it against current API rates with a reasonable buffer, and treat any future price drops from this kind of infrastructure competition as upside, not as a plan.

Author

Sources

Stay in the Loop

Get new posts delivered to your inbox. No spam, unsubscribe anytime.

Newsletter coming soon. Set PUBLIC_CONVERTKIT_FORM_ID in .env to activate.

Related Posts