Pangram AI Detector on Novels: What 97% Really Proves

Someone in your writing group posts a screenshot at midnight. It's a detector readout with a big red number on it, and the book under the number is a novel everyone in the group has been talking about for a month. By breakfast the thread has three hundred replies, and half of them say the same thing: if it can happen to a book on the Goncourt list, it can happen to mine.
That fear is reasonable. But the case behind it is more complicated than the screenshot, and the number means something narrower than most people sharing it think. Here's what a high AI-detector score on a novel tells you, what it doesn't, and what to do if one gets pointed at your book.
What happened with the Goncourt and Pangram
Short answer: the Académie Goncourt did not pull Thélyson Orélien's novel because of one detector score. It pulled it after a detector result, a full-book retest, a translation test and separate plagiarism findings all pointed the same way — and even then the vote wasn't unanimous.
The book is C'était ça ou mourir, a novel about a young Haitian teacher who flees gang violence and makes his way to Canada, published by Boréal in Quebec and Grasset in France. On September 21, 2026, an X account called Balance ton Claude posted that the novel had been "written almost entirely by artificial intelligence," citing an analysis from the detector Pangram, according to Orélien's Wikipedia entry. The book had just won the Prix du roman Fnac and had sold roughly 35,000 copies in France, AFP reported via Geo News.
Orélien denied using AI and said the flagged rhythms came from Haitian and Caribbean language traditions. That's a serious argument — detectors have a documented history of misreading writers outside the statistical mainstream. A 2023 Stanford study found that popular GPT detectors "consistently misclassify non-native English writing samples as AI-generated."
Then the evidence stacked up. Radio-Canada ran the entire novel through Pangram and got 94% of the text flagged as AI-generated, while six other novels from the Goncourt selection came back 100% human. Researchers then translated passages into English to test the cultural-style defense; Pangram still flagged them, which it did not do for translations of known human text. Radio-Canada's fact-checking unit separately found that about half of a 2012 short story by Orélien had been lifted from Le Matin des magiciens, plus copied passages in several of his older journalism pieces.
On September 25, the Académie removed the book from its selection, calling it "very likely largely the product of artificial intelligence" and citing the plagiarism. The vote was nine to one. The dissenter, Pascal Bruckner, said he refused to take part in a "lynching" and called it the best novel of the season.
Hold onto that split.
What a "97% AI" score actually measures
A Pangram percentage is how much of the text the model classified as AI-like. It is not a probability that a machine wrote the book, and it is not a measure of how many words came from a model.
This is the part the screenshots flatten. Pangram's own model card for Pangram 4 is unusually candid about it. The tool reads a document in overlapping windows and labels each segment Human, AI-Assisted, or AI-Generated. The per-segment score is, in Pangram's words, "a continuous AI-involvement score, not the probability of the displayed discrete label." The confidence rating attached to each segment is "not a calibrated probability estimate."
So when a readout says 94% or 97%, the honest translation is: the model put an AI-style label on that share of the text. It is a description of what the prose looks like to a classifier, not a reconstruction of how the prose was made. French tech site Clubic, testing the same question after the Orélien story broke, put it plainly: a detector can flag something worth attention but doesn't reconstruct how a text was produced.
Here's the phrase worth keeping: a detector score is a smoke alarm, not a fire report. It tells you to go look. It doesn't tell you what's burning or whether someone just made toast.
Why the false-positive rate isn't the whole story
Pangram's headline numbers are genuinely low. The model card lists a false-positive rate of 0.0041% for English and 0.0026% for French — roughly one human document in 24,000 misflagged in English. That's part of why researchers interviewed by Radio-Canada took the Orélien result seriously. Marzena Karpinska told the broadcaster they could not say with 100% certainty it was AI, but there was "a very strong probability" of significant AI assistance.
Two things complicate the comfort, though.
First, those rates are measured on Pangram's evaluation sets, and a published novel is not a typical document. It's 80,000-plus words of one voice, often deliberately stylized, sometimes repetitive on purpose — a refrain, a narrator who circles the same memory. The model card itself says the tool works best on formal prose and that predictions depend on surrounding context.
Second, other detectors disagree with each other, sometimes violently. AFP ran the Orélien excerpts through seven other tools and got results ranging from entirely AI to entirely human. Clubic's own controlled test found the split even on texts with known origins: a human article heavily rewritten by ChatGPT scored 38% AI on Pangram and 98% human on GPTZero. Same text, opposite stories.
If two well-funded detectors disagree that much on one article, a screenshot of one tool on one chapter proves only that somebody ran a detector.
What made the Goncourt case different from a screenshot
What moved the Académie was convergence: several independent lines of evidence, including some that had nothing to do with detectors, all pointing the same direction.
Look at what was actually on the table by September 25:
- A full-book scan, not a cherry-picked chapter, run by a news organization rather than an accuser.
- A control group — six other novels from the same shortlist, same tool, all clean.
- A test designed specifically to check the author's defense (the translation experiment).
- Plagiarism findings in the author's earlier work, established by comparing text to text, which needs no AI detector at all.
Strip away the detector entirely and you'd still have a documented plagiarism problem. That's the opposite of what most writers fear: a stranger with a free account, one chapter, and a grudge.
Call it the convergence test: one number is an accusation; four independent lines of evidence are a case. Whether the Académie got the call right is a fair argument — Bruckner thought not, and Orélien still denies it. But "a detector score got a novel disqualified" isn't what happened.
We hit a related tension in AI slop vs. Daggermouth, where a detector put 60% on a book that had already sold.
What to do if a public score targets your book
Don't argue with the number. Produce the history the number can't see.
Here's a worked version. Say you're Delphine, and your self-published fantasy novel — 92,000 words, four years of evenings — has been out for six weeks. A reviewer posts a screenshot: chapter seven, 3,400 words, "97% AI." Your sales page picks up two one-star reviews that just say "AI slop" before lunch.
Her instinct is to run the chapter through five detectors and post whichever says human. That's the move that loses: it turns the argument into a contest of screenshots, which are weak evidence in both directions.
What she does instead:
- She pulls her drafting history. Chapter seven exists in eleven dated versions across fourteen months, including the one from March where Orsolya the ferrywoman still had a brother, and a margin note to herself about cutting him. No model generates a fourteen-month revision trail with a deleted sibling in it.
- She shows the scaffolding. Her notebook page mapping who knows about the forged seal in which chapter. The voice memo where she talked through the river-crossing scene on a walk. Beta-reader comments dated before publication, quoting lines that changed afterward.
- She discloses honestly what she did use. She used a model to brainstorm place names for the river towns and to catch comma splices. She says so, specifically. Vague denials invite suspicion; specific accounts end it.
- She asks for the full-book scan, not the chapter. If an accuser wants to rely on a detector, the Radio-Canada standard is the fair one: the whole manuscript, plus comparison books in the same genre.
- She posts once, calmly, and stops. One clear statement with the evidence, pinned. Then she gets back to book two.
Step one does the work, and only if the history exists. Most writers don't think about their drafting trail until someone demands it, and by then old versions are overwritten or stranded in a web app they cancelled two years ago. In our guide to proving you didn't use AI, we went through how to build that record before you need it — dated drafts, notes, and the kind of mess only a human leaves.
This is one place where your manuscript lives stops being an abstract preference. NovelMage is a desktop app that keeps your manuscript on your own machine rather than in a vendor's cloud, so your drafts and notes are files you control and can produce on demand, not something you have to export before a subscription lapses.
Should you run your own work through a detector?
Only if you're ready to treat the result as a question rather than a verdict — and never as your main defense.
A detector sometimes lights up on exactly the passages you let a model smooth too much, which is useful editing information. If a hand-drafted chapter scores high, your sentences may have drifted toward the flattest, most predictable rhythm available — the same problem we covered in how to use AI without losing your writing voice. Fix the prose because it's flat, not because a tool flagged it.
What you shouldn't do is rewrite every chapter until the number drops. That's writing for a classifier instead of a reader, and Clubic's test shows why it's a losing game: two respected tools can land 36 points apart on the same text. You'll end up sanding off whatever made your voice yours.
If you use AI deliberately — for brainstorming, for a stuck scene, for continuity checks — the stronger position is to own that choice and keep it on your terms. NovelMage runs local models through Ollama or LM Studio, or your own Claude, GPT or Gemini keys, for a one-time $99.99, so you decide scene by scene where a model touches the work and the record of that stays on your device.
Frequently Asked Questions
Can an AI detector prove a novel was written by AI?
No. Even Pangram, which publishes unusually low error rates, describes its score as an "AI-involvement" estimate, not a probability, and its confidence rating as uncalibrated. A high score is a reason to look harder. In the Goncourt case, the decision rested on a full-book scan, a control group, a translation test and independent plagiarism findings together.
What does "94% AI" mean on Pangram?
It means the model labeled roughly 94% of the text as AI-like. It doesn't mean there's a 94% chance a machine wrote it, and it doesn't mean 94% of the words came from a model. It describes how the prose reads to the classifier, not how it was produced.
The number is the start of a conversation
The Goncourt story will get retold as "a detector got a novel banned," and that retelling will make writers afraid of the wrong thing. The real lesson is narrower and more useful: a score can start an investigation, but only evidence a score can't produce — drafts, notes, dated history, text-to-text comparison — can finish one.
Keep your trail, and say plainly what you used. And if you want a writing setup where the manuscript, the drafts and the AI choices all stay on your own machine, you can download NovelMage and try it on your next book.