16 Comments
User's avatar
Jay Rooney's avatar

“The Substack/Pangram partnership was built to restore trust between writers and readers, but the irony is, I suspect, that it will only serve to further fray it.”

Bingo 💯💯💯

Ruv Draba's avatar

Tim, my compliments on the rethinking and also the publishing of the rethinking. Those are rare here. I was also interested to note that you work in biomedical; I do a fair bit of informatic work with Australia's drug and devices regulator. I was glad to see efficacy vs effectiveness brought up.

I feel I also owe you a sort of opportunity-cost apology. I didn't realise you'd reflect so deeply, and to some extent I have let you down. I had an article sitting in the background which I could have offered you to read, but didn't. I hesitated because that reads as self-promotion on Substack, but this article talks quantitatively about risk to authors and also the risk of 'overfitting' -- which is typically what happens when a vendor tries to fine-tune a detector too high in machine learning. It comes from a place that has seen and deployed such detectors before. [https://substack.com/home/post/p-208793746]

And since I'm now offering you further information, I'd also invite you to revisit questions of authorship and motive. I have some links to help with that too.

Several writers have useful things to say about the notion of authorship. It has cultural and contextual character that has been oversimplified in the AI debate. Psychologist @Matt Grawitch has a good piece on substance vs translation [https://substack.com/home/post/p-183457261], and @nochka (Valentina) has an illuminating exposition from the world of perfume. [https://substack.com/home/post/p-208563772]. @Daniel P Hirschi explains why as someone who thinks outside English, machine translation is the *last step* -- the substantive authoring is done upstream, in a non-English language. [https://www.midliferegeneration.com/p/there-is-a-human-behind-these-words]

And in the light of what is known and reasonably *should* be known, I'd also invite you to revisit motive in both Substack and the Pangram vendor here. You granted pure motive to vendor and operator without ever checking the evidence, and that risks being a virtue-signal foreclosing a legitimate topic to examine.

Firstly, Pangram is accountable for what it does and doesn't disclose -- and it is not transparent with its data and independent assurance. If its AI-detector were a regulated medical product you'd want clear claims, indications, contraindications, and disclosure and independent assurance of what was tested on and how much of it. Pangram has published some results but there has been no independent assurance of them. Reliable disclosure is very much *not* what Pangram has done. It's hard to credit the claims without it. There's a strong commercial incentive to overpromise, commercialise on the fear here, then cash out as the fear ebbs.

Further, Pangram does *not* have to sell this product for uses that it knows or reasonably should know might bring its prospective customers and *their* customers into moral jeopardy. There are real questions to answer here given that the technology isn't as new as it looks -- those researching it know that it has been around since the 1990s, and that it has been continuously developing. We know a lot about how it works and fails. The information industry operates much like the Wild West, because regulators are still catching up. Asking questions is how we help it mature.

Secondly, there's no better time to examine motive than when norms are being established from a position of power and privilege, and money is at stake . It's fair to ask how well what is being introduced aligns with the stated benefits, what else might have been done, and what operator/vendor benefits are *also* being promoted without disclosure.

Chris Best has not being been entirely frank and transparent with his announcement. He singled out 'users worried about AI' and pitched this as service to them. But this service costs Substack money, and Substack's charity tends to be targeted toward the brand (e.g. Substack Defense), so this deserves further scrutiny.

In fact, he has another dog in this fight because AI undermines the referral algorithm which differentiates Substack. I wrote about that in a sidebar chat with a guitar-pedal manufacturer who is also on Substack: [https://substack.com/@ruvdraba/note/c-304432384] It explains how Substack's particular approach of disclosure-with-enforcement offers strategic benefit for Substack and how it's risk-shifting to authors in particular.

And then there's this curio. Apparently Substack wants to offer a two-tier service: bestsellers get sponsored AI access to reader reaction; everyone else gets Pangram, a head-pat and an invitation to subscribe. [https://substack.com/@ruvdraba/note/c-304949438]

In update, there's also this. At best, Chris isn't reading the room well. At worst, he's brand-building from both initiatives. [https://substack.com/@ruvdraba/note/c-305951030]

Tim, the issue here isn't entertaining people, nor (for me at least) clickbaiting off a hot topic. This topic falls into my bailiwick. Ethically I have a code of conduct which says I can't *not* comment, and commentary displaces other things I'd rather be doing.

The key issue is that this announcement was made overnight, was done *to* the community, not *with* the community, and it was billed as supporting a duty of disclosure that was never discussed, which still has not been introduced to the Terms of Service, and which was instantly enforced at an uncalibrated and unmanaged authors’ risk in a way guaranteed to immediately be both alarming and divisive.

We didn't choose the framing, but now the conversation deserves to run as long as it needs to. We can each form our own inferences, but we can't build effective consensus, community ethics nor reciprocal accountability from individual inference alone.

All the best, RD.

Tim Requarth's avatar

I really appreciate this reply - just a quick note that I’m going to think about it, read the article you sent, and respond in kind when I return from weekend travel

Ruv Draba's avatar

Thank you for this response, Tim. Enjoy the travel, and have a safe trip.

Lakis Polycarpou's avatar

Thanks again for another well-reasoned piece. You have definitely given me some things to think about.

I have tended to be pretty pro-quality-AI-detection (which at this point seems to mean Pangram) but you make some good counter points.

I think it's important to discuss the whole topic in the context of the harms of rampant AI use on platforms like Substack. It not just the quality of the thinking, or the annoying Claude-voice. As Linda Caroll recently pointed out, there are currently "writers" on Substack who are feeding other people's posts into LMMs to auto-generate replies in an attempt to juice their visibility.

People say that we should judge posts and comments based on their quality, not their provenance. But how does that work when a single person with an LLM is able to produce more text in a minute that I can possibly read? Who is going to curate this material and separate the Emma Klints from the tsunami of junk food out there? The onus is entirely on the reader to filter the wheat from the chaff.

Without a tool like Pangram, LLMs create a perverse incentive structure. One the one hand, the algorithm rewards volume. On the other, many people (most?) say they prefer to read non-AI writing. So the incentive is to generate as much copy as possible and then gaslight readers by claiming no one can really recognize AI-generated text. Its these people who are the most outraged by the integration of Pangram—because they fear its accuracy, not it's failure rate.

(Its this rank dishonestly that infuriates me most. We are already a society drowning in bad faith and lies.)

You quote Marc Watkins on how Chrome made him trust people less. I would posit that real trust was broken the minute people started using LLMs to write for them. People who hate AI slop have become paranoid about everything they are reading, with or without Pangram. I would posit that this leads to a lot more "false positives" than Pangram. It also leads to stupid things like people deliberately altering their writing style to "sound less AI."

Finally, I have a bit of a separate point, not fully articulated, but one I'd like to wrestle with.

In spite of my ambivalence, I have used Claude a lot in the last couple of years for a variety of writing and non-writing tasks, though not for actual words on the page—I never liked what it did to my voice.

Pangram, however, has made me more anti-AI in general. If it's virtually impossible to find a false positive in any text written before 2021, what does that say about the nature of what LLMs are producing? It's starting to feel less and less like a creative mixing of human content and more like some kind of alien thing that at best "infects" our prose, and at worst infects our thinking.

In that context, having the ability to reliably separate generated writing from human written words seems increasingly critical.

Evan Maxwell's avatar

A partial like but a reply, as well, Lakis. Your last point is an overreaction, so far as I am concerned. I think you are arguing against some of the people who USE AI, not against the tool itself.

As a lifelong crime reporter, I have come to the conclusion that there are always people who game the system. Clever crooks (as opposed to straights) have been at work from the beginning to profit or at least entertain themselves by proving how smart they are. Look at the Internet and the billions of dollars that have been scammed from honest or gullible people over the last two decades.

And most of the scams are merely copies of con games that were wide-spread in the 19th and early 20th centuries.

We aren't going to prevent all those games run from call centers in Myanmar and we won't do away with people gaming the world with schemes that utilize AI to send emails or draft spurious posts. But the source of those schemes is human intelligence; it and shouldn't be laid off to AI as though if we did away with AI, the world would be all good again.

Look at "slop" as the equivalent of junk mail through the USPS. I would love to avoid having my tax dollars used to deliver material that goes directly into the waste, but I value AI to the extent that I don't see the virtue in trying so hard to correct the slop problem with a less-than-very-strong detection system, and as far as I can tell, Pangram isn't there yet.

Convince me that my Grammarly-assisted posts (changed to the minimal extent that a copy editor would be allowed to hook paragraphs and chase commas) will not be ostracized by lazy readers unwilling to read what I have to say. Then I will be more at ease with the segregation system that SubStack has put in place.

Darius Bacani's avatar

AI is a great assistant, but authenticity still wins. I also find Undetectable.ai's AI Image Detector useful for checking visuals before publishing or sharing them. In the end, your real experience is what stands out.

Gael MacLean's avatar

I like your posts Tim because they are conversations not edicts. I'm also glad you brought up Emma's post because Ai is a much bigger picture than these petty clickbait articles about policing our work. I don't mean yours as I clarified. I am really sick of the complete invasion of privacy this data mining economy has created. We are fighting Flock camera in our small town. As we all know what they say they are being used for and what they actually get used for are two different things. Companies lie to position themselves then maximize their profits by selling our data. Google now scans and stores all our emails for Ai 'training' purposes. No matter the intent, Ai is going rogue, look at OpenAi and Claude's recent breaches. Another Ai detector will come along that is even more accurate than Pangram then substack will be saddled with an albatross. And who is paying for Pangram? Paid subscriptions? Slop its slop. Human or Ai and I can make that assessment pretty quickly when I open a post. And I move on. I don't need Big Brother spying over my shoulder trying to influence my preferences. I read good, interesting, orignal, well written stories and I try and create the same. In protest we should all turn off Ai detection in our posts. And mind our own business.

Tim Requarth's avatar

Really appreciate this kind note and thoughtful reply. Agree that I probably give more credit to motive than some companies deserve….!

Gael MacLean's avatar

Tech startups: Here today—gone today.

tsunimee's avatar

"A social problem, treated as a technical one" 🎯 yes!! and I wrote recently about the why behind that choice. Every mechanism Substack could have built that would have actually cost them something never made it past the idea stage. What was built instead asks nothing of Substack at all. Pangram *raised* its prices a week after the integration launched.

Stephen Marshall's avatar

Interestingly there is quite a lot of behind the scenes work done to validate the statistics associated with DNA testing, including specific population studies to reduce the type of error you describe. We are very far away from doing that level of verification in AI detection and face the problem that human DNA doesn't change anything like as much as AI and the detectors themselves are nowhere as deterministic as the vendors would like you to believe.

Tim Requarth's avatar

Totally agree that in some ways even the analogy to fingerprints (let alone DNA) fails, and that detectors will be inherently unstable in performance in the real world. I guess the question for me is what to do, if anything, with a tool that is reasonably though not infallibly accurate on a narrow task of identifying AI-gen text? Perhaps a role as a signal in screening out slop accounts?

Stephen Marshall's avatar

Or we need to stop being as concerned as some seem to be currently about how things are generated and more focused on the utility of what we engage with. Verification of facts, of insight, and sensitivity to manipulation all seem to me to be more important than historical norms of production.

Evan Maxwell's avatar

Tim Requarth has his nerve. First he tells us to trust and embrace Pangram on Substack.

Then he says, "No, wait a minute."

I am joking here because Tim has performed a real service: He has thought and then rethought his views, taking into account issues that distinguish theoretical from real effects.

Pangram is pretty dang good, but by no means perfect. Neither, it turns out, are fingerprint classifiers. And while he doesn't exactly say this, he is aware of the damage that false scoring can have serious real-world effects.

In this whole argument, I am beginning to think about neurodivergency. AI seems to have tremendous value for some neurodivergents. That alone should make judgments about "Does he or doesn't she?"

Spurred on by Tim's new essay, I asked Gemini whether Pangram scores use of Grammarly as using AI. Turns out that the question is more compllcated than I thought. I cut and paste below the Gemini response. Take for what it is worth.

**Short answer:** It depends on *how* you use Grammarly. Basic proofreading typically won't trigger Pangram, but using Grammarly's AI generative features to rewrite paragraphs often will.

---

### How Grammarly Features Impact AI Detection

* **Basic Proofreading (Safe):** Standard spell-check, basic punctuation, and simple style fixes preserve your underlying sentence structures. AI detectors analyze statistical predictability (burstiness and perplexity); fixing typos doesn't alter those patterns enough to flag your work.

* **Heavy Generative Rewriting (High Risk):** If you use Grammarly’s "Rephrase," "Improve Tone," or AI sentence-rewriting tools, it strips away natural human voice. Pangram flags this because the output adopts the uniform, highly predictable phrasing typical of generative AI.

---

### The Pangram Problem: False Positives

Pangram claims high accuracy, but like all statistical AI detectors, it suffers from false positives.

Because AI tools are trained on clear, proper language, **formal, polished, and academic writing naturally triggers Pangram's AI detection**—even without heavy Grammarly edits. Academic terminology, varied vocabulary, and consistent sentence flow often get misclassified as machine-written.

### Best Practices to Protect Your Work

1. **Stick to granular edits:** Accept Grammarly’s typo and grammar suggestions individually rather than allowing one-click global "enhancements" or full paragraph rewrites.

2. **Keep Version History:** Work in Google Docs or Word with track changes or version history turned on. If Pangram falsely flags your paper, your step-by-step editing log serves as definitive proof of authentic human composition.

3. **Preserve your natural voice:** Resist the urge to "dumb down" formal phrasing just to satisfy a flawed detector; documented draft histories protect your original work better than self-censorship.

So, my friends, if you use Grammarly, read Gemini , think about the piece, and proceed accordingly.

Jay Rooney's avatar

Gonna disagree with Gemini on #3; alas, once people latch onto the red percentage number, they will have already made up their minds and no amount of logs or diffs will satisfy them. Writers will have no choice but to either dumb down their writing or leave in the borderline phrasing and thus risk bringing the Eye of Sauron upon them.