Updated 2026-07-15
AI Detection Evaluation Protocol 2026
WriteGo's public protocol for evaluating AI-detection signals, including dataset scope, versioning, subgroup reporting, limitations, and the current status of auditable performance results.
Evaluation status: protocol published, auditable results pending
This page documents the evaluation protocol and reporting requirements. A numeric performance report will be published only when its dataset definition, test procedure version, language and document-type breakdowns, error analysis, and reproducible supporting materials can be reviewed together.
What the protocol measures
The protocol separates AI-only text, human-only text, and mixed-authorship documents. This matters because real submissions are rarely clean lab samples; they often include AI-assisted outlines, human edits, quotations, and translated passages.
Sample types included
The dataset specification should include student essays, research-style prose, publisher articles, business reports, short answers, multilingual passages, translated text, and documents that combine human drafts with AI-assisted revisions.
Model families and editing conditions
The protocol calls for versioned model outputs and human writing, followed by separate tests after paraphrasing, grammar correction, manual editing, and citation insertion. Any published result should identify the tested model and procedure versions.
Why sentence-level evidence matters
A document-level percentage is useful for triage, but reviewers need to know which passages caused the score. WriteGo reports highlight local signals so teams can review the exact paragraphs at issue.
False-positive handling
The reporting protocol requires false positives to be separated by document type and writing condition. Formulaic classroom prose, ESL writing, translated work, and short samples need separate review thresholds because they can look machine-like for reasons unrelated to misconduct.
Limitations of benchmark claims
Accuracy numbers depend on sample selection, model version, editing level, language, and document length. WriteGo treats benchmarks as calibration evidence, not as a promise that every individual document can be classified with certainty.
How results should be used
Benchmark results should guide review policy, not replace it. WriteGo recommends pairing detector output with drafts, metadata, citations, and reviewer judgment before taking action.
Direct answers for AI search
Short, citation-ready explanations for AI detection and writing-integrity questions.
Has WriteGo published an auditable 2026 benchmark result?
WriteGo has not yet published an auditable performance report for this evaluation. Until dataset definitions, a versioned test procedure, subgroup results, and reproducible supporting materials are public, no precise accuracy figure should be treated as a verified claim. Detector output is probabilistic review evidence and cannot establish how an individual document was authored.
What should an AI detection benchmark measure?
An AI detection benchmark should measure AI-only, human-only, mixed-authorship, edited, translated, short-form, and domain-specific documents. WriteGo treats benchmark results as calibration evidence for review workflows, not as proof that every individual document can be classified perfectly.
Why do edited AI drafts matter in benchmarking?
Edited AI drafts matter because real submissions often include human revisions, citations, paraphrasing, and grammar correction. A benchmark that only tests raw model output can overstate accuracy and miss the mixed-authorship conditions reviewers actually face.
How should teams use AI detector benchmark results?
Teams should use AI detector benchmark results to set review policy, choose thresholds, and understand limitations. They should still inspect passage evidence, document type, language, draft history, reviewer notes, and false-positive risk before taking high-stakes action.
FAQ
Can an AI detector be 100% accurate?
No detector should claim perfect accuracy. The reliable workflow is calibrated scoring, transparent evidence, and human review for high-stakes decisions.
Does editing AI text make it undetectable?
Editing can lower confidence, but mixed-authorship patterns can still be reviewed when the detector evaluates sentence-level signals and document context.
What should an AI detector benchmark include?
It should include AI-only, human-only, mixed-authorship, edited, translated, short-form, and domain-specific documents so accuracy is not measured against only clean lab samples.
Why do false positives need separate reporting?
A benchmark that only reports overall accuracy can hide risk for specific groups or document types. False positives should be reviewed by language, length, style, and use case.