Two benchmarks can test the exact same tool against the exact same detector and report wildly different pass rates, not because either one made an error, but because passing was never actually defined the same way in both tests. This single detail, what counts as a pass, turns out to matter as much as the raw detector score itself.
Once this distinction is visible in one specific case, it becomes much harder to read any pass-rate claim at face value without first asking what threshold actually produced it, a habit worth carrying into every future comparison.
Two Different Bars, Reported Side by Side
Phrasly’s September 2026 benchmark made this distinction explicit in how it reported StealthGPT’s results. Counting any output Pangram rated at least 50 percent human as a pass, a looser threshold, StealthGPT’s Super model passed on 53 percent of the 72 texts tested. Counting only outputs rated entirely human, a fraction_human of exactly 1, a stricter threshold, just 19 of the 72 outputs, about 26 percent, met that bar.
Both of those numbers describe the same 72 outputs from the same tool tested the same day. The difference between 53 percent and 26 percent comes entirely from where the line for passing gets drawn, not from any difference in what was actually measured underneath.
A gap this size, roughly double depending solely on where the threshold sits, shows just how much a single, unstated definitional choice can shape the final headline figure a reader actually sees.
Why This Distinction Gets Lost in Marketing Claims
A vendor advertising a pass rate rarely specifies which of these two definitions, or some other threshold entirely, the figure is based on. StealthGPT’s own documentation cites an 89 percent bypass rate from an internal test, without the same explicit breakdown between a lenient and a strict passing bar that Phrasly’s benchmark provides for the same tool.
This matters because a reader comparing two advertised numbers from two different vendors has no way to know whether both vendors are using the same definition of passing, unless each one discloses its threshold explicitly. Two tools could have genuinely identical underlying detector scores and still report very different headline pass rates, purely because one vendor chose a stricter bar for what counts as a pass and the other did not.
The incentive here is not hard to understand. A vendor choosing between two accurate ways to describe the same underlying result has an obvious reason to lead with whichever framing produces the larger number, even without any intent to mislead a reader who never thinks to ask which definition was used. None of this requires bad faith on anyone’s part, a company publishing its most favorable accurate number is a normal marketing choice, which is exactly why the burden falls on the reader to ask which number is being shown.
What a Careful Comparison Actually Requires
Reading past the headline number to find the actual definition behind it is the only way to compare two pass rates meaningfully.
Before comparing pass rates across different tools or benchmarks, it helps to check:
- Whether a pass means any positive human rating or specifically a full, entirely human rating
- Whether the detector model and version are the same across every number being compared
- Whether the underlying sample size and text type are similar enough to compare fairly
- Whether the raw, per-output scores are available to verify which threshold the summary figure actually reflects
Reporting Multiple Thresholds Is a Sign of Good Methodology
A benchmark that reports both a lenient and a strict measure side by side, as Phrasly’s did for average human score alongside the share of entirely human outputs, gives a reader more honest information than one that reports only a single pass rate with its threshold left unstated. The gap between the two figures is itself informative, a tool with a wide gap between its lenient and strict pass rates is producing a lot of borderline results, while a tool with a narrow gap is producing results that are either clearly human or clearly not, with little in between.
That borderline-result pattern is itself worth paying attention to, since a tool that frequently lands in the ambiguous middle is a fundamentally less predictable choice than one that reliably lands clearly on one side or the other, even if their average scores happen to look similar at a glance. Predictability matters in practice as much as the average score does, since a tool whose results cluster tightly around either outcome gives a user a much better sense of what to expect on their own next piece of writing.
The full breakdown behind both thresholds, for every tool tested, is available in the Phrasly Benchmark, which is exactly the level of detail worth looking for before treating any single advertised pass rate as directly comparable to another.
The Definition Behind the Number Matters as Much as the Number
A pass rate without a disclosed definition of passing is an incomplete piece of information, regardless of how impressive or unimpressive the figure looks on its own. Asking what counts as a pass, before asking how high the pass rate is, turns out to be the more useful first question every time.
That single habit, pausing to ask what passing actually means before reacting to the number itself, catches more misleading comparisons than any amount of scrutiny applied to the headline figure alone ever could.
For more on how AI detection thresholds and scoring methods vary across tools, further reading on the Phrasly blog covers the underlying research for anyone comparing competing pass-rate claims across different tools and detectors.
