The US Government Is Now Reading State Audit Reports With ChatGPT, and Defunding Programs Based on What It Finds. If You Sell 'AI Document Review,' Read the Liability First.
The US Department of Health and Human Services is using ChatGPT and other large language models to review the audits that states and grantees file with the federal government, flagging chronic noncompliance, repeat deficiencies, material weaknesses, and delinquent filings, and then acting on the flags. In June it denied the annual recertification of Hawaii's Medicaid Fraud Control Unit, cutting off that unit's federal funding. This is a shift away from the decades-old "pay and chase" approach to fraud, toward "let a model read the paperwork and tell us where to stop paying."
Set aside whether you think this is good policy. As a solo operator who might sell, or already sells, AI document-review tooling, this is the most instructive deployment you'll see all year, because it shows both the size of the demand and the exact place the whole category gets dangerous.
What HHS is actually doing
The mechanics are mundane and that's the point. States and grantees file audits. There are mountains of them, they're dense, and reviewing them by hand is slow and expensive, which is why noncompliance historically slid through. HHS pointed LLMs at the pile to triage it: surface the audits with repeat deficiencies, the ones that are chronically late, the ones with material weaknesses, and route those for action. The model isn't inventing a new capability. It's doing document triage at a scale humans couldn't sustain, against documents with real money attached.
And the money is real. Combined criminal and civil recoveries from Medicaid fraud cases totaled close to $2 billion in fiscal 2025, and the broader program-integrity effort is aimed at waste estimated, by various accounts, in the tens to hundreds of billions a year. When the prize is that big, "AI reads the documents faster" stops being a productivity demo and becomes an enforcement strategy. The Hawaii recertification denial is the proof that flags turn into consequences.
The opportunity is obvious, and so is the trap
Here's the part for builders. The demand HHS is acting on exists everywhere there's a pile of documents and a decision riding on them: insurance claims, loan files, vendor compliance, contract review, due diligence, grant applications. "Point an LLM at the document pile and surface what matters" is a genuinely good product wedge, and a solo operator can build a sharp, narrow version of it for a specific industry. The government just validated the category at the highest possible stakes.
The trap is in the same sentence. The value of these tools is that they drive decisions, and the moment a model's read drives a decision with consequences, you've inherited liability you may not have priced in. If your tool flags a "material weakness" that isn't there and a client acts on it, that's not a cute hallucination, that's a defunded program, a denied claim, a killed deal, and a furious customer asking why your software said something false. The thing that makes the product valuable is the same thing that makes it dangerous.
Notice what HHS did and didn't do, because it's the tell. The reporting describes the LLMs as reviewing and flagging: triage. The consequential decisions still route through the agency's process. Even the government, with lawyers and a mandate, is using the model to find the needles, not to fire the gun. If the highest-stakes deployment in the country keeps a human and a process between the model's read and the irreversible action, your solo product selling into a regulated industry has no excuse to do less.
What I'd actually build
If I were building in this category, the order of operations would be the opposite of the demo instinct. The demo instinct is to show the model making the call: "watch it approve or reject this automatically." Resist that. Build the human-in-the-loop and the audit trail first, then sell the automation around it.
Concretely, that means three things. The model surfaces and explains, it doesn't decide: every flag comes with the passage it's based on, so a human can check the model's work in seconds instead of trusting it blind. Every decision keeps a trail: what the model saw, what it flagged, who reviewed it, what they did. That trail is what saves you and your client when someone challenges an outcome, and in a regulated industry someone will. And you make the consequential step a human action by design, not by accident: the software's job ends at "here's what to look at and why," and a person owns "here's what we're doing about it."
That's also how you charge more, not less. "Fully automated, trust us" is a race to the bottom and a lawsuit waiting to happen. "We make your experts faster and give you a defensible record of every decision" is what a regulated buyer actually pays for.
The honest counter-take
You could argue I'm being too cautious and that the buyers who want speed will route around the careful vendor to the one who promises full automation. That's partly true: there's always a customer who wants the cheap, fast, no-friction version, and someone will sell it to them. For low-stakes documents where a wrong flag costs nothing, more automation is genuinely fine, and bolting on review layers there is just friction that loses you the deal.
But the whole reason this category is lucrative is the high-stakes end, where the documents drive money and the wrong answer is expensive. That's exactly where the careful design isn't optional and where the full-automation vendor eventually detonates. HHS isn't running this on low-stakes paperwork: it's running it on audits that decide federal funding, and even there it kept the model in a triage role. Copy that. The opportunity is real and a solo operator can absolutely win in it. Just don't sell the one feature (fully automated consequences with no review layer) that turns your best customer into your worst legal problem.
This piece touches government fraud-enforcement actions; the specifics of any agency's process can change, so treat the policy details as of mid-June 2026 and verify against primary sources before relying on them.
Author
Lukas
@lukcombinatorSources
- HHS Uses AI to Audit Medicaid State Funding (DistilINFO)
- HHS Launches Crackdown Using AI To Detect Medicaid Fraud And Waste (WSJ, via Sahm Capital)
- CMS's two-pronged approach to crushing fraud, waste, abuse (Federal News Network)
- What to Know About Recent Federal Actions Involving State Medicaid Program Integrity (KFF)