Built to explain, not just detect
GBT Zero started from a simple observation: a score without a reason is nearly useless in any review process that requires evidence. Our detection methodology is built to surface the why behind every result — not just a number, but the specific signals that produced it.
The tool is free for basic use and trusted by educators, editors, and content teams who need more than a binary verdict. This page explains the detection approach, who built it, and what it is — and isn’t — designed to do.
How GBT Zero reads AI-generated text
Large language models generate text by selecting statistically likely tokens — a process that leaves measurable traces in the output. GBT Zero quantifies four of those traces independently, then combines them into a calibrated detection score. Using four separate signals reduces false positives caused by any single measure behaving unusually in formal or non-native English writing.
Perplexity
Perplexity measures how predictable the word choices in a passage are. Because language models select high-probability tokens by design, AI-generated text tends to score low on perplexity — meaning the word choices are less surprising than typical human writing. Human authors make idiosyncratic, lower-probability word selections that raise the perplexity of their text.
Low perplexity → predictable token selection → AI signalBurstiness
Burstiness captures variance in sentence length across a passage. Human writing naturally alternates between short declarative sentences and longer, clause-heavy constructions. AI-generated text tends toward uniform sentence length — a rhythmic regularity that stands out when measured statistically. Low burstiness is a consistent and reliable signal across AI models.
Low variance in sentence length → AI rhythm signalLexical diversity
Lexical diversity is calculated as the ratio of unique word types to total words in a passage (the type-token ratio). AI models tend to reuse common vocabulary within a passage, reducing this ratio compared to human writing of similar length and register. Paraphrased AI content often preserves the low lexical diversity of the original output.
Type-token ratio below baseline → vocabulary reuse signalSentence entropy
Sentence entropy applies information-theoretic analysis at the sentence level, measuring the information density and structural variation within individual sentences. AI-generated paragraphs frequently show lower entropy than human prose of comparable length because they avoid structural irregularities that carry information. This signal is particularly useful for detecting lightly edited AI output.
Shannon entropy below baseline → structural uniformity signalThe four signals are weighted and combined using a calibration model trained on a balanced corpus of human-written and AI-generated text from multiple sources. Weights are updated periodically as new model output becomes available. A single elevated signal is treated as weak evidence; consistent elevation across all four signals produces a high-confidence detection result.
What guides how we build this tool
Detection without explanation creates more disputes than it resolves. Every decision in GBT Zero’s design reflects three core commitments.
Show your working
Every detection result displays the four individual signal bars — not just the combined score. Users can see which signals drove a result and by how much, making the output reviewable rather than opaque. This is especially important for contested cases where the score alone would not be defensible.
Scores should mean something
A score of 75% should mean approximately 75% of similar texts with these signal characteristics were AI-generated in calibration testing — not just “high probability.” We avoid inflated confidence claims and document where the detection model performs less reliably, including on short texts and heavily edited output.
Results require interpretation
GBT Zero is a measurement tool, not a verdict. The detection score is one input into a review process, not the conclusion of one. We design the interface to surface context — signal breakdowns, confidence qualifiers, and length warnings — because that context is what makes the score actionable rather than just alarming.
Built for anyone who needs evidence, not just a score
GBT Zero is designed for professional and academic contexts where a detection result has real consequences and needs to be defensible. The tool’s transparency features — signal bars, confidence indicators, sentence-level breakdown in the full report — exist because our users need to be able to explain a result, not just cite it.
Educators and academic institutions
Review student submissions before grading. Use the signal breakdown to understand what specifically triggered a result before making an academic integrity determination. GBT Zero supports the review process — the judgment remains with the educator.
Editors and publishers
Verify contributor work before publication. Catch AI-assisted content that passes surface-level review but shows consistent statistical patterns across multiple detection signals. The full report’s sentence-level breakdown identifies which specific passages warrant closer attention.
Content and editorial teams
Audit AI-assisted drafts against internal editorial standards. Maintain consistency in voice and authenticity across output at volume. GBT Zero gives team leads a structured way to evaluate submissions without reading every word of every draft individually.
Who built GBT Zero
GBT Zero is developed by a small team with backgrounds in computational linguistics, machine learning, and statistical language modeling.
Computational linguist with a background in distributional semantics and authorship attribution. Sarah leads the research behind GBT Zero’s signal design — particularly the perplexity and lexical diversity models. Her prior work focused on measuring statistical divergence between human and model-generated text in multilingual corpora.
Machine learning engineer specializing in anomaly detection and signal processing. Marcus leads calibration of GBT Zero’s multi-signal scoring engine against current language model output. He manages the corpus used to weight and validate the detection model, and oversees updates as new AI systems are released.
What GBT Zero cannot do
Every detection tool has documented failure modes. Understanding ours helps you use GBT Zero’s results appropriately — especially in high-stakes situations.
- Short texts (under 100 words) produce unreliable scores. The statistical signals GBT Zero measures require sufficient word count to be meaningful. Results for very short submissions should be treated as indicative, not definitive.
- Heavily paraphrased or human-edited AI content scores lower than unmodified output. Significant post-generation editing can reduce the statistical traces that detection relies on.
- Formal human writing — particularly from non-native English speakers — can produce elevated scores. Structural regularity and lower lexical diversity in ESL academic writing overlap with AI-generation patterns. The signal breakdown provides context for interpreting unexpected results.
- No AI detector achieves 100% accuracy. GBT Zero should be used as one input in a review process that includes human judgment — not as the sole basis for an academic integrity determination or editorial decision.
- Detection coverage is updated periodically but may lag behind the very latest AI model releases. Newly released models may score differently than established systems until calibration is updated.
If you have questions about a specific detection result or use case, the contact page is the right place to reach the team.
Try the detector
Free detection score and four signal bars — no registration required.