Skip to content

Product methodology

How the GPTHuman AI Detector Produces Its Score

A first-party explanation of the current detector workflow, score bands, language signals, data path, and limits you should understand before acting on a result.

Reviewed by GPTHuman Editorial Team ·

Minimum input
50 words
Shorter submissions are rejected before analysis.
Verdict threshold
70+
Applied to the AI-associated or human-associated score.
Possible verdicts
3
AI-associated, human-associated, or mixed.

The method in plain language

GPTHuman uses a language model to assess writing patterns and return a structured estimate. It does not observe the document’s real creation history.

The detector sends the submitted passage with a fixed analysis instruction to a configured AI provider. The model reviews the passage for linguistic signals associated with generated and human writing, then returns two complementary scores, one overall verdict, and sentence-level labels.

This implementation is not a locally calculated perplexity or burstiness classifier. Those concepts can explain some detector research, but the current GPTHuman workflow does not expose or calculate either metric. Read the separate perplexity and burstiness guide for general background.

What happens when you run a check

The current request path separates text processing from the quota record shown in account history.

  1. Your browser submits the passage

    The detector accepts plain text through the tool form. The server requires at least 50 words and rejects empty, oversized, or malformed requests.

  2. The server validates the request

    Input type, size, tool name, and word count are checked before any model call. A quota reservation records operational usage for the request.

  3. A configured provider evaluates the text

    The submitted passage and detector instruction are sent to an available AI provider. Provider and model availability can change the route used for a request.

  4. The server validates the response shape

    The response must contain numeric AI and human scores, an allowed verdict, and a sentence array before it is returned to the interface.

  5. The result appears in your browser

    GPTHuman returns the analysis without writing the submitted text or generated result to the application’s usage-history table. See Text and usage processing for the policy details.

Language signals the detector is instructed to review

No single feature decides the result. The current instruction asks the model to weigh several kinds of evidence together.

Sentence variety
Changes in sentence length and structure, including repeated openings or unusually even paragraph construction.
Vocabulary range
Whether word choice is varied and context-specific or relies heavily on predictable, generic phrasing.
Idiomatic usage
Expressions and phrasing that fit the language and context, without assuming that fluency proves human authorship.
Specificity
Concrete examples, details, constraints, and qualifications compared with broad claims that could fit almost any topic.
Personal voice
Subjective perspective or situated experience when the genre would reasonably contain it. Its absence is not automatically an AI signal.
Formulaic patterns
Stock transitions, excessive hedging, repeated sentence shapes, generic examples, and mechanically balanced sections.

How the current score bands map to a verdict

The two headline scores are instructed to total 100. The interface should be read as a risk-style estimate, not a calibrated probability about a person.

Current GPTHuman detector score bands and verdict rules
Score conditionDisplayed verdictResponsible interpretation
Human-associated score is 70 or higherHumanThe passage more strongly matches the instructed human-associated patterns.
Neither score reaches 70MixedThe available signals do not support either strong label under the current threshold.
AI-associated score is 70 or higherAIThe passage more strongly matches the instructed AI-associated patterns.

A displayed “80% AI” result should not be translated into “an 80% chance this person used AI.” The number is a model output under the current instruction, provider, text, and threshold. GPTHuman has not published a calibration study that supports that personal-probability interpretation.

How to use sentence-level labels

Sentence labels are pointers for close reading, not instructions to rewrite every highlighted line.

  • Check whether highlighted sentences repeat the same opening, use generic transitions, or lack the detail promised by the surrounding section.
  • Compare the sentence with drafts, notes, citations, and document history before making an authorship claim.
  • Keep exact facts, names, dates, quotations, URLs, and technical terms intact when revising.
  • Do not add errors, random synonyms, or awkward fragments merely to move a detector score.
  • For a disputed result, preserve the exact passage, date, screenshot, tool, and any visible settings before retesting.

If verified human work is flagged, follow the evidence-preservation and appeal workflow in AI detector false positives.

Known limits and failure conditions

Detection becomes less dependable when the input provides little signal or falls outside the conditions the model handles well.

Short text
The tool accepts 50 words, but the current instruction treats passages under 100 words as lower-signal evidence.
Mixed authorship
A human-edited AI draft or AI-edited human draft may not fit a clean binary origin label.
Language and genre
Performance can vary across languages, formal templates, code, quotations, reference lists, and specialized writing.
Provider changes
Configured providers, model versions, and service availability can change. GPTHuman does not currently expose a public detector version identifier with each result.
Repeat runs
Model-based analysis can vary between runs. A second score is another model output, not independent evidence of provenance.
Consequential decisions
Grades, employment actions, accusations, or publication decisions require stronger evidence and a fair human review process.

Accuracy claims and evaluation status

GPTHuman does not publish a detector accuracy percentage because the planned benchmark has not yet produced auditable results.

The 500-sample test protocol is a preregistration-style plan with blank result tables. It defines sample groups, thresholds, error metrics, stress tests, and reporting requirements, but it is not a completed benchmark.

  • No accuracy, precision, or false-positive percentage is claimed on this methodology page.
  • A future result must identify the frozen dataset, detector configuration, test date, exclusions, failed runs, and uncertainty intervals.
  • Any material change to the detector instruction or score rule should update this page’s reviewed date and version note.

Frequently asked questions

Does GPTHuman calculate perplexity or burstiness for my text?

No. The current GPTHuman detector uses a language-model assessment guided by a fixed list of linguistic signals. The interface does not calculate or expose a perplexity or burstiness metric.

What is the minimum text length?

The server requires at least 50 words. Passages under 100 words still contain less evidence, so their result deserves extra caution.

Does a 90% AI score prove AI use?

No. It is an estimate produced for the submitted text under the current model instruction and provider. It does not prove who authored the document or whether a policy was broken.

Does GPTHuman store detector submissions in account history?

The application usage log stores operational fields such as account or guest identifier, tool, word count, and timestamp. It is not designed to store the submitted text or detector result. The configured AI provider and hosting infrastructure may process request data under their own terms.

Has GPTHuman completed the 500-sample accuracy study?

No. The published article is the test protocol, not a results report. GPTHuman should not publish benchmark percentages until the tests, data, analysis, and limitations are available for review.

Sources and product evidence

Product-specific statements were checked against the current implementation. External sources provide context; they are not copied into this guide.

Use the score as a starting signal

Run a check when it helps you review a passage, then combine the result with drafts, sources, revision history, and a fair human assessment.

Try the AI Detector