Detector Checker AI Detector Benchmark Results

First-party evaluation • 880 texts • 440 source-matched human–AI pairs • English • public RAID train subset • no adversarial attack

Detector Checker’s AI detector benchmark achieved 87.3% strict balanced accuracy (95% paired cluster-bootstrap interval: 85.1%–89.4%) on a pre-specified set of 880 texts from the public RAID train_none file. The set contained 440 human texts and 440 AI-generated texts paired by source across eight writing domains and eleven historical generator labels.

The human false-positive rate was 7.5% (33/440; 95% Wilson interval: 5.4%–10.3%). The AI false-negative rate was 11.1% (49/440; 95% Wilson interval: 8.5%–14.4%). A further 3.4% of all results were indeterminate. We retained indeterminate results and technical failures in the strict metric denominators rather than removing them.

These figures describe this tested sample and detector version. They are not a guarantee for every document, language, model, or editing condition, and they do not prove who wrote a text.

Try the AI detector · Read the methodology · Download the benchmark files

Results at a glance

Measure Result What it means
Texts 880 440 human and 440 AI-generated texts
Source-matched pairs 440 Each AI text was paired with a human source/control from the same source task
Strict balanced accuracy 87.3% Average of Human→Lower and AI→Higher; indeterminate and failed results are not correct
95% interval for strict balanced accuracy 85.1%–89.4% Paired nonparametric cluster bootstrap on pair_id, 5,000 replicates
Human false-positive rate 7.5% 33 of 440 human texts received a Higher-signal result
AI false-negative rate 11.1% 49 of 440 AI texts received a Lower-signal result
Indeterminate rate 3.4% 30 of 880 results were indeterminate
Coverage 96.6% 850 of 880 results were Lower or Higher rather than indeterminate/failed
Technical failure rate 0.0% 0 final failures in 880 calls; three calls succeeded after one retry
Ordinal AUROC 0.964 Ranking discrimination only; the score was not treated as a probability

How to read this page: “Lower signal” and “Higher signal” are detector outcomes, not proof of human or AI authorship. The 0–100 score is ordinal, not a calibrated probability. See how to interpret Detector Checker results and false-positive limitations.

The complete pre-specified result

The headline result uses the full set selected before testing: 352 analysis pairs plus 88 sealed_audit pairs, for 440 pairs in total. The model, configuration, score bands, metric definitions, and sample assignment were frozen before endpoint calls. No model or threshold change was made during the run.

Outcome matrix

Rows are the known source label. Columns are the category emitted by the product backend.

Ground truth Lower signal Indeterminate Higher signal Technical failure Total
Human 395 (89.8%) 12 (2.7%) 33 (7.5%) 0 (0.0%) 440
AI 49 (11.1%) 18 (4.1%) 373 (84.8%) 0 (0.0%) 440
Confusion matrix for Detector Checker’s 880-text AI detector benchmark, showing Lower, Indeterminate, Higher, and failed outcomes for human and AI texts.

Primary and supporting metrics

Metric Estimate 95% interval Numerator / denominator Interval method
Strict balanced accuracy 87.3% 85.1%–89.4% 0.5 × (395/440 + 373/440) Paired cluster bootstrap
Human→Lower rate 89.8% 86.6%–92.3% 395/440 Wilson
AI→Higher rate 84.8% 81.1%–87.8% 373/440 Wilson
False-positive rate, `P(Higher Human)` 7.5% 5.4%–10.3% 33/440
False-negative rate, `P(Lower AI)` 11.1% 8.5%–14.4% 49/440
Overall indeterminate rate 3.4% 2.4%–4.8% 30/880 Wilson
Coverage 96.6% 95.2%–97.6% 850/880 Wilson
Selective balanced accuracy 90.3% 88.3%–92.2% Determinate outcomes only Paired cluster bootstrap
Ordinal AUROC 0.964 0.953–0.975 440 human + 440 AI scores Paired cluster bootstrap
Technical failure rate 0.0% 0.0%–0.4% 0/880 Wilson

Selective balanced accuracy is shown only with coverage. It excludes indeterminate outputs, so it is not the headline measure. AUROC measures whether AI texts tend to receive higher scores than human texts; it does not show probability calibration.

Sealed-audit verification

We reserved 88 matched pairs—176 texts—as sealed_audit. That subset produced 88.1% strict balanced accuracy (95% paired cluster-bootstrap interval: 83.0%–92.6%). Its false-positive rate was 5.7% (5/88; 95% Wilson interval: 2.5%–12.6%), and its false-negative rate was 11.4% (10/88; 95% Wilson interval: 6.3%–19.7%).

The sealed subset supports the direction and approximate magnitude of the full-set result. It is still part of the 880-text headline set, however, so it is not an external replication or statistically independent of the headline figure.

Sealed-audit outcome matrix

Ground truth Lower signal Indeterminate Higher signal Technical failure Total
Human 82 (93.2%) 1 (1.1%) 5 (5.7%) 0 (0.0%) 88
AI 10 (11.4%) 5 (5.7%) 73 (83.0%) 0 (0.0%) 88

Full set and sealed subset compared

Metric Full pre-specified set Sealed-audit subset
Texts / pairs 880 / 440 176 / 88
Strict balanced accuracy 87.3% (85.1%–89.4%) 88.1% (83.0%–92.6%)
Human false-positive rate 7.5% (5.4%–10.3%) 5.7% (2.5%–12.6%)
AI false-negative rate 11.1% (8.5%–14.4%) 11.4% (6.3%–19.7%)
Indeterminate rate 3.4% (2.4%–4.8%) 3.4% (1.6%–7.2%)
Coverage 96.6% (95.2%–97.6%) 96.6% (92.8%–98.4%)
Selective balanced accuracy 90.3% (88.3%–92.2%) 91.1% (86.8%–94.8%)
Ordinal AUROC 0.964 (0.953–0.975) 0.968 (0.940–0.989)
Final technical failures 0/880 0/176

Results by writing domain

Each domain contains 55 human texts and 55 AI texts. These subgroup estimates are exploratory. Their intervals are wider than the overall interval, and small differences between rows should not be treated as settled rankings.

Domain Human n AI n Strict balanced accuracy (95% cluster-bootstrap interval) Human FPR AI FNR Overall indeterminate
Abstracts 55 55 92.7% (88.2%–97.3%) 0/55 (0.0%) 6/55 (10.9%) 1.8%
Books 55 55 94.5% (90.0%–98.2%) 4/55 (7.3%) 1/55 (1.8%) 0.9%
News 55 55 90.9% (85.5%–95.5%) 5/55 (9.1%) 1/55 (1.8%) 3.6%
Poetry 55 55 82.7% (76.4%–89.1%) 3/55 (5.5%) 16/55 (29.1%) 0.0%
Recipes 55 55 70.9% (62.7%–79.1%) 7/55 (12.7%) 9/55 (16.4%) 14.5%
Reddit 55 55 88.2% (82.7%–93.6%) 7/55 (12.7%) 5/55 (9.1%) 0.9%
Reviews 55 55 92.7% (88.2%–97.3%) 2/55 (3.6%) 4/55 (7.3%) 1.8%
Wikipedia 55 55 85.5% (78.2%–92.7%) 5/55 (9.1%) 7/55 (12.7%) 3.6%

Recipes had the lowest strict balanced accuracy in this sample and the highest indeterminate rate. Poetry had the highest observed AI false-negative rate. Recipes and Reddit shared the highest observed human false-positive rate. These are descriptive findings, not evidence that subject matter alone caused the errors.

Strict balanced accuracy, human false-positive rate, and AI false-negative rate with confidence intervals across eight tested writing domains.

Results by historical generator label

Every generator row below contains 40 AI texts and 40 source-matched human controls. The names are historical RAID corpus labels. They do not establish performance on every current or future product carrying a similar name.

The human FPR in each row is calculated on that row’s source-matched human controls. Those texts were not generated by the named model; the row pairing is used to control source/task differences.

RAID label Strict balanced accuracy (95% cluster-bootstrap interval) AI→Higher AI→Lower / FNR AI indeterminate Human-control FPR
chatgpt 92.5% (86.2%–97.5%) 36/40 (90.0%) 3/40 (7.5%) 1/40 (2.5%) 1/40 (2.5%)
gpt4 80.0% (70.0%–88.8%) 31/40 (77.5%) 5/40 (12.5%) 4/40 (10.0%) 5/40 (12.5%)
gpt3 91.3% (85.0%–96.2%) 36/40 (90.0%) 3/40 (7.5%) 1/40 (2.5%) 0/40 (0.0%)
gpt2 85.0% (76.2%–92.5%) 33/40 (82.5%) 6/40 (15.0%) 1/40 (2.5%) 4/40 (10.0%)
llama-chat 90.0% (82.5%–96.2%) 37/40 (92.5%) 2/40 (5.0%) 1/40 (2.5%) 4/40 (10.0%)
mistral 86.2% (78.8%–92.5%) 34/40 (85.0%) 4/40 (10.0%) 2/40 (5.0%) 2/40 (5.0%)
mistral-chat 95.0% (90.0%–98.8%) 38/40 (95.0%) 1/40 (2.5%) 1/40 (2.5%) 2/40 (5.0%)
mpt 87.5% (80.0%–93.8%) 33/40 (82.5%) 5/40 (12.5%) 2/40 (5.0%) 3/40 (7.5%)
mpt-chat 93.8% (88.8%–98.8%) 38/40 (95.0%) 1/40 (2.5%) 1/40 (2.5%) 3/40 (7.5%)
cohere 72.5% (62.5%–81.2%) 23/40 (57.5%) 15/40 (37.5%) 2/40 (5.0%) 5/40 (12.5%)
cohere-chat 86.2% (78.8%–92.5%) 34/40 (85.0%) 4/40 (10.0%) 2/40 (5.0%) 4/40 (10.0%)

At 40 AI texts and 40 human controls per row, these estimates are useful diagnostics—not precise current-model rankings. In particular, the result for chatgpt is not a claim about every version of ChatGPT, and the result for gpt4 is not a claim about every GPT-4-family model.

Lower, Indeterminate, Higher, and technical-failure shares for AI texts across eleven historical RAID generator labels.

Historical model mapping documented by RAID

RAID’s generation code maps the corpus labels as follows. This table documents the historical corpus; Detector Checker did not regenerate these samples.

RAID corpus label Generator mapping in RAID generation code Important qualification
chatgpt gpt-3.5-turbo-0613 Historical API model ID, not current ChatGPT as a whole
gpt4 gpt-4-0613 Historical API model ID
gpt3 text-davinci-002 in generation/model.py RAID’s current README describes this row as text-davinci-003; because the public documentation conflicts, we retain the corpus label and do not claim the ambiguity is independently resolved
gpt2 gpt2-xl Historical open model ID
llama-chat meta-llama/Llama-2-70b-chat-hf Chat-tuned historical model
mistral mistralai/Mistral-7B-v0.1 Base model
mistral-chat mistralai/Mistral-7B-Instruct-v0.1 Instruction-tuned model
mpt mosaicml/mpt-30b Base model
mpt-chat mosaicml/mpt-30b-chat Chat-tuned model
cohere command through the completion/generation path API alias rather than an immutable dated revision
cohere-chat command through the chat path API alias rather than an immutable dated revision

Sources: RAID paper, RAID repository, RAID generation mapping, and RAID dataset card.

Claude and Gemini were not tested

This benchmark does not contain documented Claude or Gemini outputs. It therefore provides no Claude-specific or Gemini-specific accuracy rate.

The Claude detector page and Gemini detector page should link to this disclosure and describe product use and evidence gaps accurately. These RAID results must not be copied into those pages as if they were Claude or Gemini test results. A defensible model-specific claim would require a separate pre-specified evaluation using exact, dated model IDs and matched human controls.

The same limitation applies to any untested current model family or version.

Results by text length

Length bins are descriptive, and the long-text bin is small. Bins are assigned from each text’s own length, so the human and AI counts can differ slightly inside a bin.

Length bin Human n AI n Strict balanced accuracy (95% cluster-bootstrap interval) Human FPR AI FNR Overall indeterminate
Short 152 153 90.2% (86.6%–93.2%) 4.6% 11.1% 2.0%
Medium 269 267 85.1% (82.0%–88.0%) 9.7% 12.0% 4.1%
Long 19 20 94.7% (86.8%–100.0%) 0.0% 0.0% 5.1%

The long-text row is too small to support a strong comparative claim. Across the complete sample, human texts ranged from 84 to 451 words (mean 238.77; median 232.5), and AI texts ranged from 85 to 445 words (mean 240.04; median 233).

Greedy and sampling results

The AI side of the sample contained 220 greedy outputs and 220 sampling outputs, each paired with a human control.

Decoding label Human controls AI texts Strict balanced accuracy (95% cluster-bootstrap interval) Human FPR AI FNR Overall indeterminate
Greedy 220 220 89.3% (86.4%–92.0%) 7.3% 7.3% 3.4%
Sampling 220 220 85.2% (82.0%–88.4%) 7.7% 15.0% 3.4%

The higher observed FNR for sampling is descriptive within this selected corpus. Generator family, source, and other generation settings are also involved, so this table should not be read as a causal experiment on decoding alone.

Score distribution and overlap

Human texts had a median score of 4 and AI texts had a median score of 98. The corresponding means were 13.0 and 82.9. Within the 440 source-matched pairs, the AI score was higher in 429 pairs, tied in one pair, and lower in 10 pairs. Paired concordance was 97.6% (95% cluster-bootstrap interval: 96.1%–98.9%).

The distributions still overlapped: a human text reached 98, and an AI text reached 1. This is why a detector score should not be used as standalone proof of authorship or as the sole basis for a consequential decision.

Normalized distribution of Detector Checker ordinal scores for 440 human and 440 AI texts, with Lower, Indeterminate, and Higher score bands.

The score is not a probability

Detector Checker’s score is an ordinal signal. A higher score means the analyzed text showed more of the patterns used by this detector; it does not mean there is an equivalent percentage probability that AI wrote the text. Use the result as one input to review, not as standalone proof of authorship.

Published integer-score bands at the time of testing were:

Boundary-label contract finding

The product backend emitted the category presented to the user. We used that emitted category for the performance tables above. We separately compared it with the category implied by the returned integer score and published bands.

Three of 880 records (0.34%) differed at a band boundary. Exactly five core results had a returned integer score of 40 or 50; three of those five disagreed with the published band:

Sample ID Ground truth Returned integer score Backend-emitted category Category implied by published integer band
RAID-news-llama-chat-05-H Human 40 Lower Indeterminate
RAID-news-gpt2-02-H Human 50 Indeterminate Higher
RAID-recipes-gpt2-03-A AI 50 Indeterminate Higher

This could occur if a higher-precision internal value is classified before the score is rounded, or if the display and classification layers use different rules. The benchmark data alone cannot establish the cause. The score and label contract should be unified and boundary-tested; any corrected detector release should receive a new version and a new benchmark run.

Technical reliability and request latency

All 880 planned calls produced a final valid response with a unique request ID. Three calls required one retry; no call required more than two attempts. There were no final technical failures.

Request statistic Result
Final valid responses 880/880
Calls retried once 3
Median elapsed time 3.62 seconds
p90 6.36 seconds
p95 7.56 seconds
p99 11.72 seconds
Minimum 1.47 seconds
Maximum 28.81 seconds

These timings describe this benchmark runner and test window, not a service-level guarantee for every user, region, or load condition.

Distribution of elapsed request time for all 880 benchmark calls, with median and p95 markers.

Methodology

Test data

We selected a deterministic, source-matched subset from the public RAID train_none.csv file—the labeled training partition without RAID adversarial transformations. We did not use RAID’s hidden-label test partition, and this is not an official RAID leaderboard score.

Design item Value
Corpus RAID train_none.csv
Source URL https://dataset.raid-bench.xyz/train_none.csv
Source ETag "a3b0661bddab1f7e63499e8f309c5eb4-8"
Source Last-Modified header Tue, 04 Jun 2024 20:20:08 GMT
Deterministic selection seed 20260827
Languages English only
Adversarial attack label none
Human texts 440
AI texts 440
Source-matched pairs 440
Domains Abstracts, books, news, poetry, recipes, Reddit, reviews, Wikipedia
Historical generator labels 11
Pairs per generator × domain cell 5
analysis split 352 pairs / 704 texts
sealed_audit split 88 pairs / 176 texts

The human and AI members of each pair share the source identifier. Pairing helps reduce differences in topic and source task, but it does not remove every difference between human and generated texts.

Live detector execution

The 880 texts were sent sequentially to the public production detector endpoint used by the website, with analysis_level=standard and language=auto. The benchmark runner used a three-second pause between calls and stored the final score, emitted classification, detected language, request ID, response time, and all detector version keys. Nonces and other request secrets are not included in public artifacts.

The core run began at 2026-08-26T22:31:08.030392+00:00 and ended at 2026-08-27T00:16:58.306864+00:00.

Frozen detector version

Every successful result contained the same version values; no version drift was observed.

Version field Exact value
Release ID rel-2026-08-26-00ad353f8f70
Engine engine-2026-08-26
Model combiner_v12_human_correction_p04_r82rescue
Aggregation ensemble-2026-07-22
Configuration 2026-07-22
Document thresholds thresholds-2024.1
Sentence thresholds sentence-thresholds-2024.1
Analysis schema detector-checker-benchmark-analysis-1.1

Metric rules

The four possible analysis states were Lower, Indeterminate, Higher, and Technical failure. A recognized backend classification_key was authoritative because it represented the product outcome. The published integer score bands were used only as fallback and as a product-contract diagnostic.

The primary metric was:

Strict balanced accuracy = 0.5 × [P(Lower | Human) + P(Higher | AI)]

Indeterminate and technical-failure results remained in their class denominators and were never counted as correct. Technical failures were excluded from AUROC only when no score existed. The score was treated as ordinal—not as a calibrated probability.

We used 95% Wilson score intervals for binomial rates. Composite and rank metrics used a paired nonparametric cluster bootstrap on pair_id, with 5,000 replicates and base seed 20260827. Pair-level resampling preserves the matched human–AI design.

For a fuller explanation, see our evaluation methodology.

Limitations

  1. This is a first-party evaluation. Detector Checker selected the subset, ran its own production detector, and analyzed the results. Publishing the protocol, hashes, and sample-level metadata improves auditability but does not make the study independent.
  2. This is a public labeled training subset, not a hidden test. We used a deterministic subset of RAID train_none, not RAID’s hidden-label test set and not an official leaderboard submission.
  3. Training overlap has not been independently ruled out. We cannot independently establish that none of these public RAID records, or closely related records, influenced detector development or tuning. The result therefore does not prove out-of-training generalization.
  4. The full-set headline includes the sealed subset. The 176-text sealed audit is a verification slice inside the 880-text result, not an external or statistically independent replication.
  5. English, unattacked text only. The core accuracy result does not establish performance on Arabic, other languages, paraphrased text, translated text, adversarial attacks, or substantial human editing. See supported languages for product support, which should not be confused with benchmarked accuracy.
  6. Historical generator labels. The tested corpus represents historical models and API aliases. Results must not be generalized to every current model sharing a brand name.
  7. Claude and Gemini are absent. This benchmark establishes no Claude- or Gemini-specific rate.
  8. Subgroups are small. Domain rows contain 55 texts per truth class; generator rows contain 40 per truth class; the long-text bin is smaller still. The subgroup tables are exploratory.
  9. Multiple comparisons. We did not adjust subgroup intervals for simultaneous comparisons. The highest or lowest row should not be promoted as a universal ranking.
  10. Ground truth is inherited from RAID. Detector Checker did not independently recreate or manually verify authorship for every source row.
  11. The score is not calibrated probability. AUROC and separation do not justify interpreting a score of 80 as an 80% probability of AI authorship.
  12. A boundary-label inconsistency was observed. Three returned boundary scores did not agree with the published integer band. The emitted backend category was used for the primary metrics.
  13. Source-text republication is restricted in our public package. We publish IDs, metadata, input hashes, results, and methods—not complete RAID source texts.
  14. Results are version-specific. A material change to preprocessing, model, aggregation, thresholds, or category logic requires a new benchmark version.

For practical guidance, read AI detector limitations and false positives.

What this benchmark supports—and what it does not

Supported wording:

On our pre-specified, source-matched English subset of the public RAID train_none file, Detector Checker release rel-2026-08-26-00ad353f8f70 achieved 87.3% strict balanced accuracy (95% paired cluster-bootstrap interval: 85.1%–89.4%). The human false-positive rate was 7.5%, and the AI false-negative rate was 11.1%.

Unsupported wording includes:

Download the benchmark package

The public benchmark package contains sample-level outcomes without corpus text, the pre-specified manifest, analysis summary, protocol, charts, and file-integrity checksums. It excludes request IDs, full corpus text, sentence content, cookies, nonces, authorization data, and other request secrets. For live-detector data handling, read the privacy policy.

Download the public benchmark package (ZIP)

References

Version history

Page date Benchmark release Detector release Change
27 August 2026 benchmark-raid-clean-v1 rel-2026-08-26-00ad353f8f70 Initial publication of the 880-text first-party RAID train-subset evaluation

Future versions must remain separately accessible or be clearly archived. A result should never be silently replaced after a detector change.

Frequently asked questions

How accurate is Detector Checker’s AI detector?

On the full pre-specified 880-text sample described here, Detector Checker release rel-2026-08-26-00ad353f8f70 achieved 87.3% strict balanced accuracy, with a 95% paired cluster-bootstrap interval of 85.1%–89.4%. This applies to the tested English public-RAID subset, not every possible document.

What was the false-positive rate?

The human false-positive rate was 7.5%: 33 of 440 human texts received a Higher-signal result. The 95% Wilson interval was 5.4%–10.3%. A detector result should not be the sole basis for a disciplinary, employment, legal, or academic decision.

What was the false-negative rate?

The AI false-negative rate was 11.1%: 49 of 440 AI texts received a Lower-signal result. The 95% Wilson interval was 8.5%–14.4%.

What happened to indeterminate results?

We kept them visible. Thirty of 880 results—3.4%—were indeterminate. They were not counted as correct in strict balanced accuracy. Accuracy among determinate results is shown only beside 96.6% coverage.

Is the 0–100 score a probability?

No. It is an ordinal detector signal, not a calibrated probability and not proof of authorship.

Does this benchmark test ChatGPT?

It includes the historical RAID label chatgpt, mapped in RAID’s generation code to gpt-3.5-turbo-0613. That does not establish performance on every current or future ChatGPT model. See the ChatGPT detector page for scope-specific guidance.

Does this benchmark test Claude or Gemini?

No. The selected RAID subset contains no documented Claude or Gemini outputs, so this page establishes no Claude- or Gemini-specific accuracy rate.

Is this an independent or hidden benchmark?

No. It is a first-party test on a deterministic subset of the public labeled RAID training file. It is not RAID’s hidden-label test set, an official RAID leaderboard submission, or an independent audit.

Can this result prove who wrote a text?

No. The benchmark contains false positives, false negatives, and indeterminate outcomes. Detector results should be combined with context, provenance, revision history, and human review.

Try the AI detector · Read the methodology