Together AI Raised $800 Million at $8.3B Valuation. Open-Source Model Infrastructure. Meanwhile Inference Margins Are Collapsing. Here's Why That Round Works.
On July 6, Together AI closed $800 million in Series C funding at an $8.3 billion valuation. The core story: enterprise open-source model infrastructure. Companies run Llama, Mistral, and fine-tuned open models on Together's platform instead of paying per-token API fees to Anthropic or OpenAI.
This is a massive bet on a category (open-source model self-hosting) that seems to be losing. Inference API margins are collapsing: Claude dropped from $3/$15 to $2/$10 pricing in July alone. Llama is free. Yet Together just raised $800M and investors think that's a smart move.
Here's why both things are true at once.
The apparent contradiction
In June 2026, your inference options were:
- Claude API: $3 per million input tokens, $15 per million output tokens. Easy, no engineering.
- Together (Llama hosted): $0.50 per million tokens. Slightly more setup, but cheaper.
- Together (your own fine-tuned model): $0.80 per million tokens after training cost. Custom, potentially better output.
By July, Claude dropped to $2/$10. That's competitive with Together's base Llama pricing. So why does Together need $800M if the price is the same?
Because price is not the only variable. And the customers paying are not the same.
Three customer segments, three pricing models
The AI market is bifurcating into three customer archetypes:
Segment 1: API Users (70% of market)
Use case: "I need AI reasoning for my product. I don't want to operate infrastructure."
Vendors: Anthropic (Claude), OpenAI (GPT), Google (Gemini), Mistral API
Economics: Pay per token. Margins compress as price drops. Commodity market. No switching cost except integration time.
Where this wins: Everything else. Most products use Claude API or GPT.
Margin pressure: Real. Claude at $2/$10 means Anthropic fights on speed, quality, and features (computer use, vision, long context). Not pricing.
Segment 2: Self-Hosted Open Models (15% of market)
Use case: "I need to run inference at scale with specific optimization for my use case. I'll operate the infrastructure."
Vendors: Together AI, Anyscale (Ray), Hugging Face Inference API, Modal, AWS Bedrock (for fine-tuned models)
Economics: You pay for hosting + optimization + operations. The model is free or open-source.
Where this wins:
- Companies with 100B+ tokens/month inference (at that scale, the engineering cost is less than the API savings)
- Companies with specific model optimizations (quantization, distillation, fine-tuning)
- Companies with data that can't leave their infrastructure
- Companies in jurisdictions where API access is restricted
Margin structure: Together makes money on infrastructure hosting, optimization, support, and model-specific features. Not on tokens.
Segment 3: Fine-Tuned Proprietary Models (15% of market)
Use case: "Our proprietary data is our moat. We need models trained on it that we control."
Vendors: Enterprise fine-tuning (Anthropic Workbench, OpenAI Fine-tuning, together.ai Model Studio), custom training shops
Economics: High upfront cost (training), lower per-token cost (inference), but requires deep ML expertise.
Where this wins: Companies where the edge is proprietary knowledge (internal documentation, customer data, financial models, medical records).
Margin structure: Whoever owns the fine-tuned model owns the margin. Together captures the infrastructure cost.
Why Together's $800M round makes sense
Together is not betting that it will compete with Claude API on price. It's betting that Segment 2 (self-hosted infrastructure) and Segment 3 (fine-tuned models) represent a sustainable, growing business that doesn't depend on API pricing.
Here's the unit economics:
Segment 1 (Claude API):
- Customer pays: $2 per million input tokens
- Anthropic's gross margin: ~70%
- Your product margin: whatever's left after your costs
Segment 2 (Together):
- Customer pays: $X for hosting (based on GPU rental), $Y for optimization, $Z for support
- Together's gross margin: ~60% (hardware is expensive)
- Your product margin: whatever you save vs. API + what you gain from optimization
The hidden advantage: if you run 1 trillion tokens/month (a massive company), that's $2 million/month on Claude API, or $800K/month on Together infrastructure (including their 40% margin). The customer saves $1.2M/month. Together captures $800K/month. That's a sustainable business.
But it only works at scale.
The honest limitation: scale requirements
Together's business model requires customers to have high inference volume. If you're running 10 billion tokens/month (a mid-market company), the math is:
- Claude API: $20K/month
- Together self-hosted: $30K/month (GPU + ops overhead outweigh the per-token savings)
At that scale, you stay on Claude API.
Together's $8.3B valuation assumes a significant portion of the AI market shifts to Segment 2. That's a bet, not a certainty. But if the bet wins (if companies at scale decide proprietary fine-tuning or optimization is worth the engineering cost), Together owns the infrastructure layer.
Where Together wins (and loses)
Together wins if:
- Large enterprises (>$100M revenue) run in-house AI inference for proprietary data
- Inference volumes scale such that 30-50% of the market operates their own models
- Specialized model optimization (quantization, multi-token prediction, custom kernels) delivers measurable ROI
- Geopolitical fragmentation forces companies to run models outside public cloud
Together loses if:
- Inference API prices keep dropping (they will)
- Closed-source models (Claude, GPT) remain superior enough that the optimization advantage is negligible
- Most enterprise AI use cases don't require proprietary model training
- Open models asymptote in capability and don't improve enough to justify the ops cost
What I'd actually do
If you're building an AI product, decide which segment you're in:
Segment 1: API Users
- Positioning: "We bring domain expertise to Claude/GPT." You charge for integration, validation, support. Not for reasoning.
- Timeline: You're fast to market, slow to defensible margins.
- Escape velocity: Reposition to Segment 3 (proprietary data edge) before your customers realize they can run Claude themselves.
Segment 2: Self-Hosted Infrastructure
- Positioning: "We run your model at 40% the cost of API." This only works if your customer has 500B+ tokens/month.
- Timeline: You're slow to market, but you have durable margins.
- Escape velocity: Add support, fine-tuning services, or proprietary model optimization.
Segment 3: Fine-Tuned Proprietary Models
- Positioning: "We train models on your data that you control." This works for any company with a proprietary dataset and $100K+ budget.
- Timeline: You're slow to market, high touch, durable margins.
- Escape velocity: You are escape velocity. Model ownership is the moat.
Together's $800M bet is on Segment 2 growing faster than Segment 1. They might be right. But know which segment you're playing in, because the unit economics and timelines are completely different.
Author
Lukas
@lukcombinator