What a Better AI Detector Doesn't Fix
Pangram is more like fingerprints than DNA, but also, Substack isn’t a courtroom

Welcome back to The Third Hemisphere, where I try to make sense of how AI is reshaping work, thinking, and creativity, often by watching my own assumptions get upended.
If you were forwarded this and want to subscribe, click below. If you want to support a real human writing about AI, upgrade to paid.
First, apologies to readers who are sick of AI detection. Trust me, I am sick of it too, and pine for the halcyon days when there wasn’t AI to detect in the first place. I truly can’t wait to write about something uncontroversial and light-hearted next week, perhaps a hot take on the Fauci hearings.
Nevertheless, here we are. For anyone who missed it, I recently wrote about why “It’s OK to Change Your Mind About AI Detection.” The narrowish point I wished to convey is that Pangram is pretty damn good at detecting AI text under third-party testing conditions, and it was driving me crazy that serious people with large platforms wouldn’t acknowledge this basic empirical point.
But, after reading and engaging with many reactions, including a great discussion in the comment section of the last post, I’m realizing my last post came off to some readers as too promo for Pangram, or even naive, as though I hadn’t thought about the other issues at stake, or implications in play. This, of course, also drives me crazy, so here I am with a companion post. The point here, inspired by reader comments, is to lay out why I think even a quite good Pangram won’t fix a number of other important issues.
Efficacy vs Effectiveness
I should have been more precise in my last post, so let me borrow a distinction from the biomedical world, where I spend a lot of my time. In medicine, efficacy is how a treatment performs under the controlled conditions of a clinical trial: selected patients, careful protocols, close monitoring. Effectiveness is how it performs in the real world, where patients have three other conditions, or happen to be, um, women—you know, that minor subgroup the FDA actively barred from early-phase trials until 1993. The point is a drug can be efficacious in clinical trials and still fall short in the real world. It will not work has expected in some people, there will be unexpected side effects in others, etc. This is why FDA approval comes with strings attached. Manufacturers have to report adverse events for as long as a drug is on the market, and when a drug is cleared on more provisional evidence, through accelerated approval, the FDA can require confirmatory trials to verify the benefit is actually there.1
Using that vocabulary: Pangram is efficacious. That’s amazing! Under third-party testing conditions, it performs similarly to its maker’s claims. But efficacious is not the same as effective, and this is where I differ, a lot, from Pangram, whose position seems to be that its validation studies prove the tool “works” in the real world. What concerned me most about Substack’s “Against Claudefishing“ announcement was the absence of any reported mechanism for real-world monitoring. Sure, writers can dispute individual verdicts, but a dispute box is case-by-case recourse, and it’s unclear what even happens, if anything, behind closed doors of Pangram and Substack. To my knowledge, Pangram and Substack do not have a plan to estimate ongoing, aggregate error rates from actual deployment.
This is a major oversight. One particular concern of mine is that even though Pangram’s overall false positive rate is low, false positives might cluster in real-world subgroups not identified in the validation studies. Liam Dugan, a University of Pennsylvania graduate student whose PhD dissertation is on AI detection, told me when I spoke to him for a Slate article that “for most people, they might never, ever get a false positive. And for other people, the false positives are sort of disproportionately allocated on them because they just happen to write like AI.” Non-native speakers were the subgroup people rightly worried about early on, and there’s progress on that front. A paper published last month by researchers in Brussels took forty master’s theses submitted to their own faculty before 2019, every one written by a non-native English speaker, and submitted them to Pangram, which flagged none of them. This is progress! I’d still like to see more evidence before we put the non-native English speaker discrimination issue to rest, but even if we do, I doubt non-native English speakers are the only subgroup more likely to trip up detectors.
There may be clusters nobody has thought to look for yet—some writing whose combination of genre, topic, and style just by chance tends to inhabit spaces near the AI/human boundary. We don’t know exactly how Pangram performs on real people writing in 2026, in multiple languages, with all of the weird ways they use AI and have absorbed AI-isms. Pangram works a lot better than AI detectors used to, but it will still fail in unpredictable ways.
Bite marks, fingerprints, and DNA

The second issue a better detector doesn’t fix: people have different thresholds for how much false-accusation risk they’ll tolerate. This is fine! But it is a normative question, and accuracy data can’t settle it. (I mean, it is partly an accuracy question, in that we lack sufficient real-world data. But the efficacy data exists, and we can start making decisions with it.) For how to think about living with an imperfect but useful measurement, let me borrow from an area I’ve done a lot of previous reporting on: forensic science.
Some forensic science is genuinely junk science. Bite-mark analysis is theoretically unsound—there is no good evidence that human dentition is unique, or that skin records it faithfully—and when examiners have been put to the test, their error rates are far too high. It should not be allowed in court. Fingerprinting, you might be surprised to learn, is pretty good but has a noticeable error rate: in one of the better studies to date, FBI-affiliated examiners made false matches at a rate of about 1 in 1,000. DNA testing is excellent, though it too has a small but real error rate in real-world use, from contamination and sample mix-ups, and adjacent techniques like mixed-sample interpretation are dicier. Does this mean we throw out all of forensic science as junk? No. We accept that error rates exist and deal with them, in one of the highest-stakes setting society has.
Some people want perfection, or close to it: they want Pangram to be like DNA analysis. Others smear it as if it were bite marks. Really, Pangram is more like fingerprint analysis. Pretty good, not perfect. This gray area means people react very differently on a normative grounds. One commenter reacted this way:
One thing that intrigues me is why people are worried about false positives so much. Yes, it's a nuisance to be accused of something you are not. But outside of contexts where an AI accusation carries formal sanctions, like in academia, I feel like it's really a matter of taste (or editorial preference), especially in commercial settings like Substack. For the most part, AI or not is just another dimension publishers and readers are free to judge your writing on. And if they miss out on something great just because they prejudged you, that's on them.
Whereas for another commenter declared:
The facts are it's not 100% so innocent people get attacked and possibly ruined. That's enough.
These are both acceptable positions, but they are normative ones about what is an appropriate level of risk, who bears it, and so forth. My simple plea in the first post (and this isn’t about these commenters specifically, just in general) is that any normative argument about how to deploy AI detection starts from acknowledging the substantial empirical data we have about Pangram, not pretending it doesn’t exist or citing old studies.
Substack is not a courtroom
There is one big problem with both my clinical and legal analogies: Substack is not a court of law or a hospital. As Ruv Draba, who critically quoted my last post put it:
I think this is an important point. What makes a 1-in-1,000 fingerprint error tolerable in court is the adjudicating apparatus, however imperfect: rules of evidence, cross-examination, a judge deciding what a jury may hear, several levels of appeals.2 What makes releasing a drug into the population reasonable is its use is overseen and monitored by experts in relatively controlled medical settings. On Substack, the process is a single Pangram score, often poorly understood, and then potentially screenshotted and thrown into a feed. This commenter is correct that talking about the tool in the abstract, without the sociotechnological context of its deployment, while not exactly “misleading,” is probably insufficient.
A social problem, treated as a technical one
Which brings me to what I feel is the core problem with this partnership: It doesn’t actually solve, nor is it capable of solving, the issue. I don’t doubt the sincerity of the people behind it. Chris Best of Substack framed “Claudefishing” as a mismatch between what a reader expects and what they get, which is a very real problem. Max Spero of Pangram has been consistent that his tool should “never be the ending arbiter” but a starting point for a more thorough investigation. I believe both of them mean it. I also think both of them are wildly naive about how AI detection will function in the actual world we live in. In their vision, readers scan judiciously, click through to reports, understand the strengths and limitations of AI detection, keep up to date on the literature, carefully consider notions of authorship, etc. In the real world, people’s behavior will be driven by an interface that invites none of this nuance. I predict two spectacular kinds of failures, which I’ll illustrate with examples.
The first is on the writer’s side. People say they want less AI writing, but not everyone has considered the full range of AI use and what it can enable. I urge anyone who automatically recoils against AI use in writing to read Emma Klint. She’s a non-native English writer with a self-described neurodivergent brain, and she uses AI for everything she publishes:
I have strong verbal processing and low working memory. AI lets me use that verbal processing to compensate for what my working memory can’t hold. That helps me access thoughts that were already there, that I lost track of while I was still thinking them. Simple as that.
People often warn that using AI will “steal your voice,” but Emma describes using AI as the opposite:
When I started writing with AI, I wasn’t afraid of losing my voice, because I didn’t really have one. I had lots of thoughts and a point of view, sure, but not a writing voice. The process of writing through dialogue, and thinking with AI, is how I found it.
I’ll admit, sometimes in Emma’s writing I hear echoes of Claude, and aesthetically the English-speaking writer snob in me doesn’t like it. But I think it would be shallow of me to take that reaction very seriously. Emma’s posts clearly have a tremendous amount of thought put into them, and judging her valuable chronicling of AI use on the basis of a few linguistic tics strikes me as the wrong metric. As Emma herself put it: “I’m not arguing that Pangram will get the percentage wrong. Even a perfect score would answer the wrong question.” My concern is that the Pangram integration may discourage writers like Emma Klint from writing at all—and I don’t think that’s a net win.
One clever fix proposed in my comments is that Substack sidestep detect-and-punish with detect-and-deprioritize: let readers flip a switch that says “prioritize human writing in my feed,” so that a false positive might limit a writer’s algorithmic reach but spare them a public accusation. Of the design alternatives I’ve heard, this seems reasonable until I think of writers like Emma Klint. A writer like her would struggle to gain traction, without her knowing quite why. Writers who use AI to pump out half-baked content, on the other hand, would just rewrite until they slip past the filter. Users who flipped the “deprioritize” switch would then end up with a feed of human writing blended with undisclosed hybrid writing anyway. This deployment would screen out the Klints and yet reward the evaders.
The second failure is on the reader side. I think of a post from Marc Watkins, a lecturer of writing and rhetoric at the University of Mississippi, who chronicled a week using Pangram’s Chrome extension labeling every post in his social feeds. From his post, “How An AI Detector Made Me Trust People Less“:
I noticed my behavior change. Dramatically. I stopped interacting and reading posts and instead focused on labels. When we allow a company to place labels on our social media interactions, we cede some agency to opaque systems, ultimately giving them a great deal of power over our interactions.
Watkins is a sophisticated reader, aware of what these labels can and cannot claim. And yet, his behavior changed anyway.
Both of these failures reveal that the core problem not just technical but social: People want a determination on provenance, authenticity, and originality, but what they get is a determination on text. So, probably, all of this is part of a longer social process of redefining things like authorship and authenticity, which is really hard to do, and which no scan, however efficacious, will do for us.
So let me state my position as plainly as I can, since I failed to in my last post: Pangram is efficacious and people should stop making sloppy or outdated arguments to the contrary; and also: I think this version of the Pangram/Substack integration will do more harm than good, and I think Max and Chris are a bit naive about the social implications of integrating a powerful weapon in call-out culture. I think it is perfectly coherent to accept the empirical reality that Pangram works fairly well (with major limitations, of course) and normatively be against its deployment (for a variety of reasons).
In announcing the partnership, Best wrote that when readers have to wonder whether what they’re reading is real, it “undermines trust in authorship and threatens the livelihood of writers.” I agree with the problem, but not the solution. A Pangram score cannot differentiate a Claudefisher from a Klint, and a feed of color-coded metrics led Watkins to start trusting people less in a week. The Substack/Pangram partnership was built to restore trust between writers and readers, but the irony is, I suspect, that it will only serve to further fray it.
Those confirmatory trials are frequently late or never finished, which is its own scandal, but, hey, in principle it’s a good idea. (And, hey, it is improving.)
Again, reality asserts itself: in practice, courts have almost never excluded fingerprint evidence as unreliable, and juries almost never hear an error rate at all






“The Substack/Pangram partnership was built to restore trust between writers and readers, but the irony is, I suspect, that it will only serve to further fray it.”
Bingo 💯💯💯
Tim, my compliments on the rethinking and also the publishing of the rethinking. Those are rare here. I was also interested to note that you work in biomedical; I do a fair bit of informatic work with Australia's drug and devices regulator. I was glad to see efficacy vs effectiveness brought up.
I feel I also owe you a sort of opportunity-cost apology. I didn't realise you'd reflect so deeply, and to some extent I have let you down. I had an article sitting in the background which I could have offered you to read, but didn't. I hesitated because that reads as self-promotion on Substack, but this article talks quantitatively about risk to authors and also the risk of 'overfitting' -- which is typically what happens when a vendor tries to fine-tune a detector too high in machine learning. It comes from a place that has seen and deployed such detectors before. [https://substack.com/home/post/p-208793746]
And since I'm now offering you further information, I'd also invite you to revisit questions of authorship and motive. I have some links to help with that too.
Several writers have useful things to say about the notion of authorship. It has cultural and contextual character that has been oversimplified in the AI debate. Psychologist @Matt Grawitch has a good piece on substance vs translation [https://substack.com/home/post/p-183457261], and @nochka (Valentina) has an illuminating exposition from the world of perfume. [https://substack.com/home/post/p-208563772]. @Daniel P Hirschi explains why as someone who thinks outside English, machine translation is the *last step* -- the substantive authoring is done upstream, in a non-English language. [https://www.midliferegeneration.com/p/there-is-a-human-behind-these-words]
And in the light of what is known and reasonably *should* be known, I'd also invite you to revisit motive in both Substack and the Pangram vendor here. You granted pure motive to vendor and operator without ever checking the evidence, and that risks being a virtue-signal foreclosing a legitimate topic to examine.
Firstly, Pangram is accountable for what it does and doesn't disclose -- and it is not transparent with its data and independent assurance. If its AI-detector were a regulated medical product you'd want clear claims, indications, contraindications, and disclosure and independent assurance of what was tested on and how much of it. Pangram has published some results but there has been no independent assurance of them. Reliable disclosure is very much *not* what Pangram has done. It's hard to credit the claims without it. There's a strong commercial incentive to overpromise, commercialise on the fear here, then cash out as the fear ebbs.
Further, Pangram does *not* have to sell this product for uses that it knows or reasonably should know might bring its prospective customers and *their* customers into moral jeopardy. There are real questions to answer here given that the technology isn't as new as it looks -- those researching it know that it has been around since the 1990s, and that it has been continuously developing. We know a lot about how it works and fails. The information industry operates much like the Wild West, because regulators are still catching up. Asking questions is how we help it mature.
Secondly, there's no better time to examine motive than when norms are being established from a position of power and privilege, and money is at stake . It's fair to ask how well what is being introduced aligns with the stated benefits, what else might have been done, and what operator/vendor benefits are *also* being promoted without disclosure.
Chris Best has not being been entirely frank and transparent with his announcement. He singled out 'users worried about AI' and pitched this as service to them. But this service costs Substack money, and Substack's charity tends to be targeted toward the brand (e.g. Substack Defense), so this deserves further scrutiny.
In fact, he has another dog in this fight because AI undermines the referral algorithm which differentiates Substack. I wrote about that in a sidebar chat with a guitar-pedal manufacturer who is also on Substack: [https://substack.com/@ruvdraba/note/c-304432384] It explains how Substack's particular approach of disclosure-with-enforcement offers strategic benefit for Substack and how it's risk-shifting to authors in particular.
And then there's this curio. Apparently Substack wants to offer a two-tier service: bestsellers get sponsored AI access to reader reaction; everyone else gets Pangram, a head-pat and an invitation to subscribe. [https://substack.com/@ruvdraba/note/c-304949438]
In update, there's also this. At best, Chris isn't reading the room well. At worst, he's brand-building from both initiatives. [https://substack.com/@ruvdraba/note/c-305951030]
Tim, the issue here isn't entertaining people, nor (for me at least) clickbaiting off a hot topic. This topic falls into my bailiwick. Ethically I have a code of conduct which says I can't *not* comment, and commentary displaces other things I'd rather be doing.
The key issue is that this announcement was made overnight, was done *to* the community, not *with* the community, and it was billed as supporting a duty of disclosure that was never discussed, which still has not been introduced to the Terms of Service, and which was instantly enforced at an uncalibrated and unmanaged authors’ risk in a way guaranteed to immediately be both alarming and divisive.
We didn't choose the framing, but now the conversation deserves to run as long as it needs to. We can each form our own inferences, but we can't build effective consensus, community ethics nor reciprocal accountability from individual inference alone.
All the best, RD.