· 7 min read

Substack's AI Detector Just Flagged a Journalist's Human-Written Article as AI. Here's the Actual Risk for Solo Writers.

Substack scans every post, note, reply, and comment over 100 words through an AI detector called Pangram, then shows readers a percentage split of how much of the piece looks human-written versus AI-generated. It's been live since July 21. Writers are calling it a witch hunt, and at least one journalist has already been publicly accused of using AI on a piece she wrote entirely herself. If you write anything on Substack, or anywhere with an audience that can screenshot a detector score, this is worth understanding before it happens to you.

What Substack actually shipped

The integration scans content automatically and shows both writers and readers an estimated breakdown: human, AI-assisted, or AI-generated. Substack gave writers a few controls alongside it, and publishers can disable scanning at the post level. Pangram says the content isn't used to train its own models. The stated goal is transparency as AI-written content becomes more common, letting readers decide for themselves whether they're reading a person or a wrapper around a model.

Pangram claims its latest model, Pangram 3.3, has a 0.01% false positive rate. That number sounds close enough to zero to stop worrying about. It isn't.

The problem is what 0.01% looks like at Substack's actual scale

A 0.01% false positive rate sounds reassuring until you multiply it by the number of posts, notes, and comments Substack processes. Even a genuinely tiny error rate turns into a real number of writers incorrectly flagged once you're running it across a platform with millions of pieces of content, and detector companies have a track record of understating their real-world error rate once independent researchers get to test them outside the vendor's own benchmark. Turnitin, a comparable detection tool used widely in education, claimed a false positive rate under 1%. A 2023 Washington Post investigation testing it in practice found the real number was over 50%.

Pangram isn't Turnitin, and 3.3 isn't the same model that Washington Post piece tested. But the gap between a vendor's marketed accuracy and a detector's real-world accuracy has been wide enough, often enough, that "0.01%, trust us" isn't a number I'd build a reputation-management strategy around.

It already happened to someone with a public byline

A journalist was accused on social media of using AI in one of her pieces after a Pangram score started circulating. She investigated her own case directly and confirmed it: false positive, a piece she'd written herself, flagged as AI-generated by the detector reading her sentence rhythm and structure as synthetic. That's the exact failure mode critics point to, and it landed on someone with the platform and credibility to push back publicly and get it corrected. Most writers don't have that.

The people most likely to get flagged incorrectly aren't the ones actually using AI to ghostwrite their posts. Detector false positives cluster around writers with consistent, patterned rhythm, which includes people writing in a second language and neurodivergent writers whose natural cadence reads as repetitive or overly structured to a model trained to spot exactly that pattern. The detector isn't catching AI use. It's catching a writing style that happens to overlap with what AI output tends to look like, and then reporting that overlap as a confidence score readers treat as fact.

What this means if you write with any AI in your process

Almost nobody in the Solo Operator orbit writes with zero AI touching the process anymore. You outline with Claude, you ask it to poke holes in an argument, you run a draft past it for structure feedback, you use it to fact-check your own claims before you publish. None of that is what "AI-generated" means to most readers, but a detector reading the finished text has no way to see your process, only the output, and the output of a heavily-AI-assisted-but-human-written piece can look statistically similar to the output of something a model wrote start to finish.

That's the actual risk: not that you'll get caught doing something wrong, but that a tool with a marketed near-zero error rate will occasionally be wrong about you specifically, publicly, with no simple way to prove it wrong before some of your readers have already decided.

The honest take

I don't think Substack is acting in bad faith here. Readers do want to know if they're paying for a subscription written by a person, and as AI-assisted writing becomes the default rather than the exception, some kind of transparency layer was probably inevitable. I also think a detector with any nonzero false positive rate, deployed across millions of pieces of content with a public-facing score, guarantees a steady trickle of writers getting wrongly accused, and the burden of disproving it falls entirely on them.

What I'd actually do: get ahead of it. If your process involves AI at any stage past pure research, say so plainly, once, somewhere permanent, like an about page or a recurring line in your newsletter footer. Not because you're required to, but because a clear statement of your actual process, made before anyone accuses you of anything, is a much stronger position than a defensive reply to a screenshot after the fact. I added a line to this blog's about page months ago describing exactly how much AI touches my drafts and why; it's the single easiest piece of reputation insurance I've bought all year, and it cost me twenty minutes.

If you write on Substack specifically, check whether scanning is on for your publication and decide deliberately, not by default. Disabling it isn't hiding anything if your process disclosure is already public. It's just opting out of a specific scoring mechanism with a real, documented error rate that you don't get to audit yourself.

Author

Sources

Stay in the Loop

Get new posts delivered to your inbox. No spam, unsubscribe anytime.

Newsletter coming soon. Set PUBLIC_CONVERTKIT_FORM_ID in .env to activate.

Related Posts