Product methodology
How the GPTHuman AI Detector Produces Its Score
A first-party explanation of the current detector workflow, score bands, language signals, data path, and limits you should understand before acting on a result.
Reviewed by GPTHuman Editorial Team ·
- Minimum input
- 50 words
- Shorter submissions are rejected before analysis.
- Verdict threshold
- 70+
- Applied to the AI-associated or human-associated score.
- Possible verdicts
- 3
- AI-associated, human-associated, or mixed.
The method in plain language
GPTHuman uses a language model to assess writing patterns and return a structured estimate. It does not observe the document’s real creation history.
The detector sends the submitted passage with a fixed analysis instruction to a configured AI provider. The model reviews the passage for linguistic signals associated with generated and human writing, then returns two complementary scores, one overall verdict, and sentence-level labels.
This implementation is not a locally calculated perplexity or burstiness classifier. Those concepts can explain some detector research, but the current GPTHuman workflow does not expose or calculate either metric. Read the separate perplexity and burstiness guide for general background.
What happens when you run a check
The current request path separates text processing from the quota record shown in account history.
Your browser submits the passage
The detector accepts plain text through the tool form. The server requires at least 50 words and rejects empty, oversized, or malformed requests.
The server validates the request
Input type, size, tool name, and word count are checked before any model call. A quota reservation records operational usage for the request.
A configured provider evaluates the text
The submitted passage and detector instruction are sent to an available AI provider. Provider and model availability can change the route used for a request.
The server validates the response shape
The response must contain numeric AI and human scores, an allowed verdict, and a sentence array before it is returned to the interface.
The result appears in your browser
GPTHuman returns the analysis without writing the submitted text or generated result to the application’s usage-history table. See Text and usage processing for the policy details.
Language signals the detector is instructed to review
No single feature decides the result. The current instruction asks the model to weigh several kinds of evidence together.
- Sentence variety
- Changes in sentence length and structure, including repeated openings or unusually even paragraph construction.
- Vocabulary range
- Whether word choice is varied and context-specific or relies heavily on predictable, generic phrasing.
- Idiomatic usage
- Expressions and phrasing that fit the language and context, without assuming that fluency proves human authorship.
- Specificity
- Concrete examples, details, constraints, and qualifications compared with broad claims that could fit almost any topic.
- Personal voice
- Subjective perspective or situated experience when the genre would reasonably contain it. Its absence is not automatically an AI signal.
- Formulaic patterns
- Stock transitions, excessive hedging, repeated sentence shapes, generic examples, and mechanically balanced sections.
How the current score bands map to a verdict
The two headline scores are instructed to total 100. The interface should be read as a risk-style estimate, not a calibrated probability about a person.
| Score condition | Displayed verdict | Responsible interpretation |
|---|---|---|
| Human-associated score is 70 or higher | Human | The passage more strongly matches the instructed human-associated patterns. |
| Neither score reaches 70 | Mixed | The available signals do not support either strong label under the current threshold. |
| AI-associated score is 70 or higher | AI | The passage more strongly matches the instructed AI-associated patterns. |
A displayed “80% AI” result should not be translated into “an 80% chance this person used AI.” The number is a model output under the current instruction, provider, text, and threshold. GPTHuman has not published a calibration study that supports that personal-probability interpretation.
How to use sentence-level labels
Sentence labels are pointers for close reading, not instructions to rewrite every highlighted line.
- Check whether highlighted sentences repeat the same opening, use generic transitions, or lack the detail promised by the surrounding section.
- Compare the sentence with drafts, notes, citations, and document history before making an authorship claim.
- Keep exact facts, names, dates, quotations, URLs, and technical terms intact when revising.
- Do not add errors, random synonyms, or awkward fragments merely to move a detector score.
- For a disputed result, preserve the exact passage, date, screenshot, tool, and any visible settings before retesting.
If verified human work is flagged, follow the evidence-preservation and appeal workflow in AI detector false positives.
Known limits and failure conditions
Detection becomes less dependable when the input provides little signal or falls outside the conditions the model handles well.
- Short text
- The tool accepts 50 words, but the current instruction treats passages under 100 words as lower-signal evidence.
- Mixed authorship
- A human-edited AI draft or AI-edited human draft may not fit a clean binary origin label.
- Language and genre
- Performance can vary across languages, formal templates, code, quotations, reference lists, and specialized writing.
- Provider changes
- Configured providers, model versions, and service availability can change. GPTHuman does not currently expose a public detector version identifier with each result.
- Repeat runs
- Model-based analysis can vary between runs. A second score is another model output, not independent evidence of provenance.
- Consequential decisions
- Grades, employment actions, accusations, or publication decisions require stronger evidence and a fair human review process.
Accuracy claims and evaluation status
GPTHuman does not publish a detector accuracy percentage because the planned benchmark has not yet produced auditable results.
The 500-sample test protocol is a preregistration-style plan with blank result tables. It defines sample groups, thresholds, error metrics, stress tests, and reporting requirements, but it is not a completed benchmark.
- No accuracy, precision, or false-positive percentage is claimed on this methodology page.
- A future result must identify the frozen dataset, detector configuration, test date, exclusions, failed runs, and uncertainty intervals.
- Any material change to the detector instruction or score rule should update this page’s reviewed date and version note.
Frequently asked questions
Does GPTHuman calculate perplexity or burstiness for my text?
No. The current GPTHuman detector uses a language-model assessment guided by a fixed list of linguistic signals. The interface does not calculate or expose a perplexity or burstiness metric.
What is the minimum text length?
The server requires at least 50 words. Passages under 100 words still contain less evidence, so their result deserves extra caution.
Does a 90% AI score prove AI use?
No. It is an estimate produced for the submitted text under the current model instruction and provider. It does not prove who authored the document or whether a policy was broken.
Does GPTHuman store detector submissions in account history?
The application usage log stores operational fields such as account or guest identifier, tool, word count, and timestamp. It is not designed to store the submitted text or detector result. The configured AI provider and hosting infrastructure may process request data under their own terms.
Has GPTHuman completed the 500-sample accuracy study?
No. The published article is the test protocol, not a results report. GPTHuman should not publish benchmark percentages until the tests, data, analysis, and limitations are available for review.
Sources and product evidence
Product-specific statements were checked against the current implementation. External sources provide context; they are not copied into this guide.
- GPTHuman AI Disclosure — product processingHow AI-backed tools and the rule-based Readability Scorer process text.
- GPTHuman 500-sample detector benchmark protocolA published method with intentionally blank result tables.
- GPT detectors are biased against non-native English writersPrimary research illustrating why subgroup testing and human review matter.
Related resources
- GPTHuman AI DetectorRun a text check and keep this methodology beside the result.
- AI detector false positivesA fair review workflow for verified human writing that is flagged.
- Ethical Use PolicyAcceptable and unacceptable uses of detector output.
- Text and usage processingWhat the application sends, stores, and does not store.
Use the score as a starting signal
Run a check when it helps you review a passage, then combine the result with drafts, sources, revision history, and a fair human assessment.
Try the AI Detector