· 6 min read

The 'Use All the Tokens' Era Just Ended. Companies Are Clawing Back AI Spend, and That's the Service You Should Be Selling.

The 'Use All the Tokens' Era Just Ended. Companies Are Clawing Back AI Spend, and That's the Service You Should Be Selling.

For about two years the corporate directive on AI was "use more of it." CNBC has a name for the era (tokenmaxxing) where employers pushed developers and teams to lean on frontier models as hard as possible, results be damned, because being seen to adopt AI mattered more than the bill. That era just flipped. The same companies now want clearer ROI, tighter controls, and lower-cost alternatives. Uber reportedly put spending tiers on some AI tools starting at a $1,500-a-month base, with employees having to request more. Flo Crivello, the CEO of Lindy, switched his company entirely off Anthropic's Claude and moved 100% of traffic to DeepSeek's cheaper open-weight models. D.A. Davidson analyst Gil Luria warned that some of the labs' largest enterprise customers may start reining in out-of-control token spend.

If you're a solo operator, the instinct is to read this as bad news: the AI gold rush cooling off. I read it the opposite way. The market just repriced the valuable skill from "spend more" to "spend right," and spending right on AI is a thing solo operators are structurally better at than the companies now panicking about it.

Why you're already good at the thing they suddenly need

You've never had a tokenmaxxing budget. From day one your AI spend came out of the same pocket as your rent, so you learned the efficiency habits under duress. The data backs this up: for solo operators in the $10K–$50K/month revenue range, the typical AI-tool spend sits around $100 a month, and the resulting operating margins run high precisely because the cost base is disciplined. You route cheap models to the easy calls. You cache. You trim context because you watched a bloated prompt double your bill. You reach for the frontier model only when the task actually demands it.

That behavior was survival for you. It's now exactly what a mid-size company with a runaway Anthropic invoice is desperate to install. They spent two years optimizing for adoption and never built the muscle for restraint. You have nothing but that muscle.

The offer shapes that sell right now

I've been turning this into concrete engagements, and the framing that lands isn't "AI consulting." It's "I will lower this specific number on your bill." A few shapes that work:

An AI-spend audit. You go through a team's actual usage (which calls hit which models, how big the prompts are, what's cached and what isn't) and hand back a ranked list of where the money leaks. This is the wedge, because it produces a number the buyer can feel, and the number is almost always embarrassing.

A routing and caching retrofit. Most teams that grew up tokenmaxxing send everything to the top model. Splitting traffic so the easy 80% goes to a cheap model, caching repeated calls, and reserving the expensive model for the calls that measurably need it routinely cuts the bill by more than half without a quality drop anyone notices. That's a clean, scoped project with a result you can measure before and after.

A "same output, fewer tokens" engagement. Prompt slimming, context trimming, moving deterministic work out of the model entirely. Unglamorous, and it prints money.

The reason these convert is that the buyer's mood has changed. A year ago pitching "let me cut your AI usage" sounded like you didn't get it. Today it sounds like you're the only person in the room who does.

Don't oversell the efficiency gospel either

Here's where I'll argue against my own pitch. Efficiency taken too far is its own trap. Some of the tokenmaxxing was genuinely productive: teams did ship faster by leaning on strong models, and a client who cheeses out to the cheapest model for everything will produce worse work and blame the AI. If you sell "spend right" as "spend as little as possible," you'll deliver a cost cut and a quality regression, and you'll own both. The actual skill you're selling is judgment about which calls deserve the expensive model, not a blanket downgrade. Lindy moving 100% to DeepSeek is a real data point, but it's one CEO's bet, not proof that the cheap model wins every workload.

So the honest version of the service is: cut the waste, keep the calls that earn their cost, and be able to prove the quality held with an eval, not a vibe. Do that and you're not riding the end of a hype cycle: you're supplying the discipline the hype cycle forgot to build. The buyers just started paying for it.

Author

Sources

Stay in the Loop

Get new posts delivered to your inbox. No spam, unsubscribe anytime.

Newsletter coming soon. Set PUBLIC_CONVERTKIT_FORM_ID in .env to activate.

Related Posts