The Glossary
The terms behind AI detection and humanizers, each in a sentence or two — plain, honest, no jargon for its own sake.
- AI detectorn.
A tool that reads finished text and estimates how likely it is that a language model produced it. It judges the words alone — never the writing process — so its output is a probability, not proof.
Read the full entry →- AI-generated / AI-refined / Human-writtenn.
The three-way split a modern detector reports instead of one number. AI-generated means the text reads model-authored; AI-refined means a human wrote the substance and a model polished the surface; human-written means it reads person-drafted. The three add up to 100.
Read the full entry →- AI humanizern.
A tool that rewrites text to read more like natural human writing — varying sentence rhythm, cutting machine tells, calibrating confidence — while preserving the meaning. A legitimate writing aid when the ideas being expressed are genuinely the writer's own.
Read the full entry →
- Burstinessn.
The variation in sentence length and complexity across a passage. Human writing is bursty — long sentences next to short ones; unedited model output is often uniform. Low burstiness is treated as an AI signal.
Read the full entry →
- Calibrationn.
How well a detector's scores match reality — whether the texts it calls 80% AI really are AI about 80% of the time. A poorly calibrated detector can be confidently wrong, which is why a stated confidence level matters as much as the score.
Read the full entry →- Calibration curven.
A plot that compares a detector's stated confidence against the true rate of correctness across thousands of labelled texts. A well-calibrated detector hugs the diagonal — when it says 70%, it's right about 70% of the time. A curve that bows away from the diagonal is the visual signature of a detector that bluffs.
Read the full entry →
- Driftn.
What happens to a detector's accuracy as the underlying language models keep improving. A detector trained on last year's GPT output silently gets worse on this year's, because the gap between model writing and human writing has narrowed. The fix is continual re-evaluation, not a one-time benchmark.
Read the full entry →
- Em-dash overusen.
A recurring surface tell of unedited AI prose — long sentences punctuated by parenthetical asides set off with em-dashes, often three or four to a paragraph. Em-dashes are not the problem; the rhythm of using them as a default connective is. Read aloud and you can hear it.
Read the full entry →
- False positiven.
When a detector flags genuinely human writing as AI. Polished essayists and non-native English speakers are the most common victims, because formal, fluent prose can share surface features with model output.
Read the full entry →- False negativen.
When a detector misses AI-written text and scores it as human. Lightly edited or paraphrased model output is the usual cause — small changes can be enough to drop a statistical detector's confidence.
Read the full entry →
- Genericismn.
Prose that could describe almost anything — comprehensive solutions, robust frameworks, significant challenges. The most consistent surface signal of AI text, because models reach for the high-probability noun. The fix is specificity: replace the generic phrase with the one fact that earns the claim.
Read the full entry →
- Hallucinationn.
When a language model states something false or invented with full confidence — a fake citation, a wrong date, a non-existent fact. A reason to verify model output rather than trust it, especially for sources and figures.
Read the full entry →- Hedgingn.
Calibrating the strength of a claim with phrases like "this suggests", "in most cases", "arguably". Skilled human writers hedge the claims they are less sure of and commit hard to the ones they will defend; AI text often states everything at the same flat level of certainty.
Read the full entry →
- Large language model (LLM)n.
The kind of AI system behind tools like ChatGPT, Claude and Gemini. Trained on large amounts of text to predict likely continuations, it can draft, rewrite and answer — and, increasingly, write in ways close to human prose.
Read the full entry →
- Paraphrase launderingn.
The practice of feeding a passage from a source into a model and treating the reworded output as your own writing about the source. It is not paraphrase — it is hidden quotation, because the engagement with the source has been outsourced. The honest version reads the source, closes it, and writes from memory and notes.
Read the full entry →In the journalWriting an academic paper with AI in the room- Perplexityn.
A measure of how predictable a text is to a language model. Model-written text tends to choose high-probability, unsurprising words and so has low perplexity; human writing is usually less predictable.
Read the full entry →- Plagiarismn.
Presenting someone else's words or ideas as your own. Distinct from AI detection — plagiarism checking compares text against existing published sources, while AI detection estimates whether a machine wrote it.
Read the full entry →- Promptn.
The instruction given to a language model to produce output. The same model can produce very different text depending on how it's prompted.
Read the full entry →In the journalWhat is AI detection?
- Recalln.
Of all the truly AI-written texts a detector is shown, the fraction it correctly identifies. A detector can have high accuracy but low recall — missing most AI text — if it sets its threshold cautiously. Recall and false-positive rate together describe a detector far better than a single accuracy number.
Read the full entry →In the journalFalse positives: when a detector is wrong- Registern.
How formal or informal a piece of writing is. Skilled writers modulate register within a single piece — a careful technical sentence next to a plain one — while unedited model output tends to lock into one register and hold it. Variation in register is one of the strongest human signals on the page.
Read the full entry →- ROC curven.
Receiver operating characteristic curve — a plot of a detector's true-positive rate against its false-positive rate as the decision threshold is moved. The area under the curve summarises how well the detector separates AI from human text overall, independently of any one threshold choice.
Read the full entry →
- Specificityn.
In writing, the willingness to reach for the concrete fact rather than the generic noun. "Handles 14 currencies" beats "comprehensive solution". The single largest difference between AI prose and good human prose. In detection statistics, the same word has a separate technical meaning — the fraction of human texts correctly identified as human.
Read the full entry →- Sycophancyn.
The tendency of a language model to agree, flatter and enthuse rather than push back or judge. It shows up as cover letters that praise the company indiscriminately, essays that endorse every position they cite, and rewrites that soften every sharp claim. Real writing has moments of judgement; sycophantic prose does not.
Read the full entry →In the journalEditing a cover letter the AI helped you draft
- Tokenn.
The unit a language model reads and writes in — roughly a word or part of a word. Model usage and limits are usually measured in tokens rather than words.
Read the full entry →
- Voicen.
The sense, on the page, that a particular person — with a particular history, taste and temperament — chose these words. Voice shows up in tiny preferences a model cannot reproduce: which metaphors a writer reaches for, which jokes they would never make, the specific words they use when tired. The most durable human signal in any text.
Read the full entry →
- Watermarkingn.
A technique where an AI provider subtly biases its model's word choices so its own output can later be identified. Promising in theory, but it only works for participating providers and survives editing poorly — so it isn't a general detection solution.
Read the full entry →
Want the longer version? The Journal has full explainers on how AI detection works.