Updated September 3, 2026
The open benchmark for AI humanizers.
AI Humanizer Benchmark tests the major AI humanizers against GPTZero, Originality.ai, Copyleaks, Winston AI, ZeroGPT, QuillBot, and Grammarly. Every monthly cycle runs the same 33 texts through every tool and measures meaning preservation and readability alongside bypass rate. We publish every prompt, raw output, detector verdict, and the scoring script on GitHub, so the rankings can be reproduced.
Currently leading
September 2026Leaderboard
One score per tool, measured across seven detectors and seven writing categories. How scores are computed
- Last tested
- 2026-09-03T16:11:14.775Z
- Sample size
- 33
- Methodology
- v1.0.0
| Rank | Humanizer | Penalties | Trend | Last tested | ||||
|---|---|---|---|---|---|---|---|---|
| 1– | 84.90 | 86.3 | 84.8 | 77.3 | none | today | ||
| 2– | 84.80 | 85.3 | 87.6 | 75.7 | none | today | ||
| 3– | 81.60 | 90.4 | 76.2 | 72.4 | -1.0 | today | ||
| 4– | 78.30 | 79.9 | 75.7 | 76.7 | none | today | ||
| 5– | 76.10 | 82.6 | 70.8 | 74.7 | none | today | ||
| 6– | 73.80 | 73.2 | 80.8 | 63.8 | none | today | ||
| 7– | 71.90 | 74.6 | 70.1 | 82.7 | -3.0 | today | ||
| 8– | 68.50 | 70.7 | 77.3 | 74.0 | -5.0 | today | ||
| 9– | 66.80 | 86.5 | 65.3 | 73.2 | -11.0 | today | ||
| 10– | 39.00 | 3.7 | 59.2 | 81.4 | -4.0 | today | ||
| 11– | 38.20 | 0.2 | 89.6 | 71.7 | -12.0 | today |
11 humanizers · 7 detectors · scores out of 100, higher is better. Click a column to sort; hover a penalty for the breakdown.
View cycle dataHow scores are computed
See the full methodologyOverall score formula
How often the rewritten text is classified as human by the detector panel.
How closely the rewrite preserves the original text's meaning.
Whether the output is clear, fluent, natural writing.
Whether performance holds across all seven writing types or varies widely between them.
Penalties (right) subtract from the composite, and the result never goes below zero.
Possible penalties
- Identical to input
The text came back essentially unchanged; no meaningful rewriting occurred.
-2.0-10.0 - Refusal
The tool declined to process the text, typically due to its own content filter.
-1.0-10.0 - Meaning drift
The rewrite departed substantially from the original's meaning.
-1.0-10.0 - Length inflation
The output ran far longer than the input, diluting the AI signal with added words rather than removing it.
-1.0-10.0 - Length deflation
The output ran far shorter than the input; content was dropped rather than rephrased.
-1.0-10.0
Maximum possible total penalty: -50.0 (off 100).
How it works
- 01
Generate the test texts
Fresh AI-written texts are generated each cycle across seven writing categories (essays, emails, blog posts, marketing copy, and more), locked in before any tool runs.
- 02
Run every humanizer on them
Every humanizer rewrites all 33 identical texts on its own default settings and plan: exactly what a normal user gets, not a hand-tuned best case.
- 03
Score against 7 detectors
Each output faces GPTZero, Originality.ai, Copyleaks, Winston AI, ZeroGPT, QuillBot, and Grammarly, every score recorded as returned, with no rescaling.
- 04
Grade meaning and readability
Each rewrite is also scored for how much of the original meaning survives, via semantic similarity, and for readability, rated by a language model.
- 05
Publish all the data
The full raw test data (inputs, outputs, per-detector verdicts, and the scoring code) ships openly and permanently on GitHub with every monthly cycle.
The detector panel
Seven commercial AI detectors, each with an equal vote in the bypass score. Passing a single detector is weak evidence; the panel measures against all seven. Each detector page ranks every humanizer against that detector alone.
GPTZero
Widely deployed in education; often the first check a text faces
Originality.ai
The strictest detector in our panel; standard among publishers
Copyleaks
Enterprise detection used across education and hiring
Winston AI
Commercial detector issuing document-level verdicts
ZeroGPT
Free browser detector with very large consumer reach
QuillBot
Detection built into the QuillBot writing suite
Grammarly
Detection inside Grammarly's writing products
+Suggest a detectorKnow one we should test against?Why this exists
Most rankings of AI humanizers are commissioned: affiliate listicles, vendors reviewing themselves, claims with no test procedure behind them. This benchmark replaces that with measurement: a standing monthly test where every tool faces the same texts, the same detectors, and the same formula, and where every number traces back to a raw input and output anyone can inspect.
Read the full storyBy use case
The overall ranking weighs everything equally. For a specific need, like the best AI humanizer for essays or performance against one detector, these reweighted views answer it directly.
Best for Students
Ranked with essay performance counting double, for coursework that has to pass a classroom detector check.
Best for Essays
Which tools keep an essay's argument intact while passing as human; essay categories count double here.
Best for GPTZero
Every tool ranked on its GPTZero results alone; nothing else in the score.
Best for Originality.ai
Ranked on Originality.ai alone, the strictest detector in our panel.
Best for Turnitin
Turnitin cannot be tested directly, so this ranks tools against our strictest detector as the closest proxy.
Best for SEO
For content that must pass both editors and detectors; blog and marketing categories count double.
Best for Academic writing
For theses, papers, and applications; formal writing categories weigh double in this ranking.
Best for Marketing
Landing pages, product copy, campaign emails; persuasive categories count double here.
Best for Free
Only the tools with a genuinely usable free tier, ranked by the same overall score.
Frequently asked questions
What is an AI humanizer?+
An AI humanizer rewrites AI-generated text so that AI detectors classify it as human-written: restructuring sentences, varying rhythm, and adjusting vocabulary while preserving the meaning. Many tools degrade the writing in the process, which is why this benchmark measures meaning preservation and readability alongside bypass rate.
How are the tools tested?+
Once a month, every humanizer rewrites the same set of freshly generated texts spanning essays, emails, blog posts, and more, each tool on its default settings. Every output is then scored by seven commercial AI detectors and measured for meaning preservation and readability.
What does the score out of 100 mean?+
It combines four measurements: detector bypass (42%), meaning preservation (32%), readability (16%), and consistency across writing types (10%). Quality failures such as meaning drift, padded or truncated output, and refusals subtract penalty points. A tool cannot rank highly by beating detectors at the expense of the text.
Which AI detectors do you test against?+
GPTZero, Originality.ai, Copyleaks, Winston AI, ZeroGPT, QuillBot, and Grammarly: the detectors that students, teachers, editors, and hiring teams use in practice. Every humanizer page breaks results down against each one, and every detector has its own bypass ranking.
How often do the rankings change?+
A full cycle runs every month with newly generated test texts, so no tool can tune against a known test set. Rankings move as tools update their models. Past cycles remain archived, so any tool's trajectory can be traced.
Do tools pay to be listed or ranked?+
No. Nothing on this site is sponsored, no ranking position can be bought, and there are no affiliate links. A tool's position is determined by its measured results and nothing else.
How is bias handled?+
Some tools in these rankings are made by the benchmark's operator. They run through the identical pipeline as every other tool, with the same prompts, settings, detectors, and scoring code. All raw data is published, so any special treatment would be visible in the detector verdicts themselves.
Can I check the results myself?+
Yes, at two levels. The quick check: paste any published output into the detector yourself and compare verdicts. The full check: clone the public data repository and run one command that recomputes the entire leaderboard from the raw inputs, outputs, and detector scores.