· 5 min read

Together AI Raised $800 Million at $8.3B to Commoditize Inference. If You're Selling 'Cheaper Inference,' You're Racing a $1.15B-ARR Company.

On July 1, Together AI closed an $800 million Series C at an $8.3 billion valuation. The round was led by Aramco Ventures with participation from Nvidia, Vista Equity Partners, General Catalyst, and a roster of other LPs that reads like a "we're hedging our bets on AI infrastructure" portfolio.

Here's what matters: Together AI's annual bookings crossed $1.15 billion last quarter. They're planning to grow capacity roughly 50-fold over the next five years. And their core business is renting GPU clusters and managing open-source model inference.

They're not training models. They're not selling a proprietary stack. They're selling "we have GPUs, we run your inference, we keep the cost down." Which is exactly the service a hundred indie consultants thought they could sell in 2024.

The inference arbitrage window is closing

If you've been consulting on "we'll run your model inference cheaper than the cloud vendor default," I want to be direct: you have 6 to 18 months before this margin goes away entirely.

Here's why. Open-source inference demand tripled in the last twelve months. Together AI and similar infrastructure plays have capital and runway to drop prices on compute until the arbitrage disappears. They're not trying to be profitable on inference margins; they're trying to own the infrastructure layer. The price will fall.

At the same time, cloud vendors (AWS, GCP, Azure) have noticed the same opportunity and are building commodity inference offerings too. You're not just racing Together AI; you're racing AWS's AI infrastructure team. And AWS has economics you can't match.

The play that looked good in 2024, "I'll manage your GPU cluster and handle the inference orchestration, and you'll save 30% on cloud costs," is now competing against $1.15B/year infrastructure companies with institutional backing.

Where the real money actually went

The capital flooding into inference infrastructure tells you something important about where the value actually sits. It's not in inference as a service. Inference is becoming a utility.

The money is moving to two places instead. First: the models themselves are becoming more specialized and efficient (smaller context windows, better quantization, domain-specific training). Second: the applications sitting on top of inference are where differentiation happens. The layer you can sell to an enterprise is "I've solved your use case with an AI system," not "I've rented you a GPU cheaper."

Together AI's $800M isn't a sign that inference arbitrage is the future. It's a sign that infrastructure is being consolidated so thoroughly that only well-funded players can win. Everyone else has to play a different game.

What you'd actually sell instead

If your consulting practice currently depends on inference margins, you need to reposition. Not because you did something wrong, but because the game changed.

Map your actual clients into two buckets. First: "they care about cost per inference." Those customers will commoditize toward the cheapest provider (Together AI, AWS, or a self-hosted setup). You're not going to beat that on price, and you don't want to race on price. Second: "they care about outcomes." They need AI to solve a specific business problem, and they don't actually know how to get there.

The second bucket is your real business. "Run inference 3% cheaper" is not a business. "I'll integrate your sales pipeline with AI scoring and you'll close 8% more deals" is. One is infrastructure arbitrage. The other is consulting.

The indie consultants who survive the next 18 months are the ones who move from "I run your inference" to "I solve your problem with AI, and here's the inference layer I picked to do it." The infrastructure becomes a detail you handle, not the thing you sell.

The honest risk

There's a chance Together AI stumbles. These mega-rounds have a track record of underperformance. The company could burn through its capital, lose customers to self-hosted setups, or fail to execute on the 50x capacity growth. It's possible.

But planning your business assuming Together AI fails is not a bet I'd recommend. It's better to plan assuming they succeed and adjust if they don't.

The other risk: if open-source model quality plateaus and everyone migrates back to closed-source APIs (Claude, GPT), the infrastructure layer becomes less valuable. But that's not the bet I see in the market right now. The bet is that open models are here to stay and infrastructure is becoming the game.

Author

Sources

Stay in the Loop

Get new posts delivered to your inbox. No spam, unsubscribe anytime.

Newsletter coming soon. Set PUBLIC_CONVERTKIT_FORM_ID in .env to activate.

Related Posts