Checkmark Plagiarism Logo
Checkmark Plagiarism
Menu
Back to Learning
DetectionHow It Works~6 min read

What an AI Score Can — and Can't — Tell You

How Checkmark's AI detection presents its results: passage-level findings, a calibrated confidence gradient instead of a fake-precision number, N/A guardrails for short texts, and process evidence to check the score against.

The Checkmark Plagiarism Team
What an AI Score Can — and Can't — Tell You

We build an AI detector, so take this from the people with every incentive to oversell one: an AI score, by itself, should never end a conversation about a student’s work. The interesting question is what a score can responsibly do — and how a report has to be built for it to do that. This post explains the choices behind Checkmark’s AI detection.

What the detector actually measures

AI detectors are pattern classifiers. They measure how strongly a passage resembles the statistical patterns of machine-generated prose versus human drafting. That resemblance is real, measurable signal — and it is also probabilistic. Fluent human writers sometimes pattern like machines; lightly edited AI output sometimes patterns like people. Any tool that turns this into a confident yes/no is misrepresenting what it knows.

Choice one: passage-level, not paper-level

A single document-wide percentage invites the worst misreading — “this essay is 62% AI” — when the true situation is usually “these three paragraphs look AI-like and the rest doesn’t.” Checkmark underlines specific passages in the essay, each with its own evidence card in the report sidebar. Teachers see which parts, not just how much, which changes what they can do next: read the flagged paragraphs, compare them with the student’s usual voice, and check how they were written.

Checkmark report sidebar showing AI Detection cards with confidence sliders alongside paste and plagiarism findings

Choice two: a gradient, not a decimal

Each AI card shows a confidence slider between typical human writing patterns and typical AI patterns. We use a gradient deliberately: a slider communicates “strong signal, weigh it” where a number like 87.3% communicates a precision the underlying science does not have. And the card carries a permanent disclaimer, verbatim: “Typical AI writing pattern versus typical human writing styles. Do not solely rely on this score to determine AI authorship.” That sentence is not legal boilerplate. It is the correct way to use every AI detector on the market, including ours.

Choice three: N/A beats a guess

Short texts do not contain enough signal for reliable AI classification. Below roughly 150 words, Checkmark’s AI tile reads N/A — with a tooltip saying exactly why — rather than reporting a number we would not defend. The same guardrail applies to our transcription and uncited-source analyses at their own thresholds. If a detector never tells you “not enough data,” it is telling you something else: that it will always produce a number, whether or not the number means anything.

Choice four: a second, independent signal

The strongest check on an AI score is not a better AI score — it is evidence of a different kind. Checkmark’s reports pair every AI finding with writing-process analysis: paste events with the original pasted text preserved, transcription patterns, and a Playback that replays the session keystroke by keystroke. The combinations are what carry meaning:

  • AI-flagged passage + arrived in one paste + no drafting: a clear picture worth a conversation.
  • AI-flagged passage + visible drafting, revision, and typo-fixing over 40 minutes: most likely a false positive — and the student has the receipts.
  • Low AI score + heavy transcription: the detector was never the right question; the process was.

This is also the honest answer to humanizer tools. Rewriting AI output can move the AI score; it cannot retroactively create a writing session that never happened.

What this means in practice

For teachers: treat the AI tile as a reason to look, never as a verdict. Open the flagged passages, check the process evidence, and use the report’s private flag statuses (Flagged, Resolved, Not Flagged) to track follow-ups — students never see the flags.

For students: your drafting history is your best protection. Writing in Google Docs or a tracked editor means a false flag can be answered with the strongest evidence there is — the visible record of you doing the work.

More on the mechanics: AI Writing Detection and Writing Process Analysis. Or run your own text through the demo and read the report it produces.

What an AI Score Can — and Can't — Tell You