The short version: Claude and heavily-prompted ChatGPT drafts consistently score lowest on detectors like Originality.ai and GPTZero, and dedicated rewrite tools such as Undetectable.ai can push AI scores close to zero, but the editing you do after the draft matters more than which tool you started with. I’ve run this test myself with real client content and real detector scores, and the results surprised me more than they should have.
I ran the same brief through five tools. The scores were not close.
Back in early 2026 I had a client, a small B2B software company, who wanted proof that their blog content wouldn’t get flagged before it went anywhere near Google or a nervous editor’s desk. So I ran an experiment. I took one 900-word brief on “how to reduce customer churn” and generated it five different ways: a raw ChatGPT-4o draft with a bare prompt, the same prompt but heavily engineered with tone and structure instructions, a Claude 3.5 Sonnet draft, a Jasper draft using Boss Mode with the client’s brand voice document loaded in, and a version where I took the raw ChatGPT draft and ran it through Undetectable.ai’s rewrite function.
Then I ran all five through three detectors: GPTZero, Originality.ai, and Copyleaks. Here’s roughly what came back:
- Raw ChatGPT draft, bare prompt: 97% AI (Originality.ai), flagged as AI by all three tools
- ChatGPT with heavy prompt engineering (voice, sentence length variation, specific examples requested): 58% AI
- Claude 3.5 Sonnet, single clean prompt, no extra engineering: 41% AI
- Jasper Boss Mode with brand voice document: 34% AI
- Undetectable.ai rewrite of the raw ChatGPT draft: 3% AI
That last number is the one that makes people uncomfortable, and it should. A tool built for one purpose, sounding human to a machine, did that job better than any actual writing model did on its own. That’s not a loophole nobody’s found yet. That’s the entire business model of tools like Undetectable.ai, StealthGPT, and similar rewrite layers, and it’s worth knowing they exist even if you never touch one.
Claude tends to score lower than ChatGPT, and here’s the likely reason
Across dozens of client pieces since that first test, I’ve noticed Claude drafts land lower on detectors than default ChatGPT output more often than not. Anthropic trained Claude with a different emphasis on conversational variation and it tends to produce longer, more irregular sentences without being told to. GPTZero and Originality.ai both lean heavily on “perplexity” and “burstiness”, basically how predictable each next word is and how much sentence length varies. ChatGPT’s default writing style is smoother and more evenly paced, which is exactly what makes it feel polished and exactly what makes it easier to spot.
Jasper sits in the middle because most people using it are feeding it a brand voice document, past blog posts, tone notes, which pulls the output away from generic GPT phrasing. Gemini’s output, in my testing, behaves similarly to raw ChatGPT: clean, correct, and detectable.
None of this is fixed. These are all snapshots of models and detectors that update constantly. A tool that scores well against GPTZero this month could score badly next month once GPTZero retrains on newer outputs. Anyone telling you a specific tool is permanently “undetectable” is either selling you something or hasn’t checked lately.
The uncomfortable bit nobody wants to put in writing
Here’s what most posts on this topic won’t tell you plainly: AI detectors are not reliable enough to be trusted as a gate, in either direction. A widely cited Stanford study on stylometric analysis, the technique most detectors are built on, found detection tools misclassify non-native English speakers’ writing as AI-generated at strikingly high rates, in some tests over half the time, simply because non-native writers tend to use more predictable, less “bursty” sentence structures. Meanwhile, a well-edited AI draft with human input can score as fully human even though a machine wrote most of the sentences.
So the tool that’s “hardest to detect” isn’t really a writing model at all. It’s a human editor with twenty minutes and a red pen. I’ve taken 90%+ AI-flagged drafts down to single digits just by breaking up three sentences, adding one specific number, cutting a paragraph that restated the intro, and swapping two generic adjectives for something concrete. No rewrite tool needed. That’s the part that gets left out of most “best AI writing tools” roundups because it doesn’t sell software.
What moves the needle, in order of impact
If you want a draft that reads as human, whether or not you care about detectors, this is the order I’ve found matters most:
- Specificity. Real numbers, named tools, actual dates. AI models default to vague ranges and soft claims. Replace “many businesses” with “62% of small firms” or a named example and both readers and detectors treat the text differently.
- Sentence length variation. Force yourself to mix a six-word sentence with a thirty-word one. AI models rarely do this unless prompted, and detectors are built almost entirely around spotting when they don’t.
- Removing the AI’s favourite connective tissue. Words like “furthermore,” “moreover,” “in today’s fast-paced world,” and “it’s important to note” appear constantly in raw AI output and barely ever in how people talk.
- First-person detail. A specific memory, a client name, a dollar figure from a real project. Models can’t invent this convincingly, and when you add it yourself, both readers and machines register it as human.
- Structural imperfection. Real writers repeat a word they already used two sentences ago. Real writers start a paragraph with “and” or “but.” Perfectly balanced grammar is, ironically, one of the biggest tells.
If you’re already prompting ChatGPT to sound more human rather than editing after the fact, you’ll get partway there, but prompting alone rarely closes the gap the way a proper edit pass does.
Step by step: how to draft something that won’t get flagged for the wrong reasons
- Write your own outline first, in your own words, even if it’s rough bullet points. This gives the model your actual structure instead of a generic one.
- Generate the draft with a specific prompt that includes your real examples, not just a topic.
- Read it aloud. Anywhere you stumble or a sentence feels too smooth, that’s usually where the AI cadence is strongest.
- Cut every instance of “it’s worth noting,” “in conclusion,” and “overall.” These phrases do almost no work and appear disproportionately in generated text.
- Add one thing only you know: a client result, a number from your own data, a mistake you made. This single step does more for detection scores than any paraphrasing tool.
- Run it through one detector, not five. If you’re bouncing between GPTZero, Originality.ai, and Copyleaks trying to satisfy all three, you’ll edit the personality out of the piece chasing a number that moves for no clear reason.
Why detection scores matter less than people think, and when they matter a lot
For most blog content, detection scores are a vanity metric. Google has said repeatedly it doesn’t penalise content for being AI-assisted, it penalises content for being unhelpful, thin, or written purely to game rankings. I’ve had posts score 40% AI on Originality.ai and outrank competitors’ fully human, badly written pieces without any trouble. If your content answers the question and includes real detail, the detector score is close to irrelevant for search.
Where it matters a great deal: academic submissions, some publishing contracts, agency work with clients who have explicit no-AI policies, and platforms that actively penalise flagged content. This is also where the ethics get murky, because a student submitting a Claude draft rewritten through Undetectable.ai has done something different from a marketer polishing a ChatGPT draft for a blog. Same tools, very different stakes.
This is also why the wider conversation around AI content and platform monetisation rules keeps shifting. YouTube, and increasingly other platforms, care less about whether AI touched the content and more about whether it’s disclosed and whether it’s original enough to be worth the audience’s time. Detection tools built for text don’t map neatly onto that standard at all.
The real problem detection scores can’t fix
Here’s the thing I keep coming back to with clients: even a 0% AI score doesn’t fix content that reads like everyone else’s. I’ve reviewed drafts that pass every detector clean and still say absolutely nothing memorable, because the writer optimised for burstiness and sentence variation instead of having an actual point of view. If you want to know why that happens so often, I wrote a longer piece on why AI-written marketing sounds like everyone else’s, and the short version is that most people prompt for tone but never feed the model a genuine opinion to work from. A tool can help you sound human. It cannot give you something worth saying.
Agencies feel this pressure hardest because they’re producing volume for multiple clients at once. If you’re managing this at scale, it’s worth reading how other teams are structuring the process in how agencies are using AI for content creation without losing the specificity that makes each client’s content theirs.
My honest recommendation
If you need a starting point today: Claude for first drafts, because it needs less correction to sound less mechanical. A genuine human edit pass, every single time, no exceptions. And I’d steer clear of dedicated “humanising” rewrite tools for anything client-facing or professional, not because they don’t work, they clearly do based on my own numbers, but because they optimise for fooling a machine rather than for saying something true and specific. Those aren’t always the same goal, and chasing the wrong one costs you the thing that keeps readers around.
Frequently asked questions
Which AI writing tool has the lowest AI detection score by default?
In my own testing, Claude drafts consistently scored lower on Originality.ai and GPTZero than default ChatGPT drafts, without any extra prompting for tone or sentence variation, largely because Claude naturally produces more irregular sentence lengths.
Do rewrite tools like Undetectable.ai work?
Yes, in my test a raw ChatGPT draft flagged at 97% AI dropped to around 3% after running through Undetectable.ai’s rewrite feature, but the output still needed a human fact-check pass because rewriting for detection doesn’t fix accuracy or add genuine substance.
Can editing an AI draft myself lower its detection score without any special tool?
Yes, and often more effectively than a paraphrasing tool. Adding specific numbers, varying sentence length, cutting generic connector phrases, and inserting a real personal detail can take a draft from over 90% AI-flagged down to single digits with about twenty minutes of manual editing.
Should I worry about AI detection for blog content that’s just for SEO?
Generally no. Google has said it doesn’t penalise content for being AI-assisted, only for being unhelpful or thin, so detection scores matter far more for academic work, publishing contracts, or platforms with explicit no-AI policies than for ordinary marketing content.