Key takeaway: AI writing detection produces false positives and false negatives that no institution should treat as evidence. The durable answer is verifiable process, not a probability score.
Why it matters: Universities, publishers and content teams are all outsourcing an accountability decision to a tool that can't carry it. That's a governance problem, not a detection problem.
What happened with AI writing detection
An academic who supervises masters theses ran an experiment: use large language models to help write a paper, then see whether AI writing detection tools could catch it. The honest conclusion was that AI-assisted writing is easy to produce and genuinely hard to prove after the fact.
You can read the original account in the Opinio Juris piece on AI-assisted writing and detection. It's a candid tour of a gap most institutions would rather not look at directly.
The academic literature agrees. A 2026 study in the International Journal for Educational Integrity evaluated the reliability of detection tools in higher education and found their outputs inconsistent enough to be risky as evidence, while guidance from university libraries now openly documents both false positives and false negatives.
Source: International Journal for Educational Integrity (Springer), 2026
Most people will conclude the detectors just need to get better
The consensus reaction is a technical one: the models are immature, vendors will improve accuracy, and in a year or two we'll have detectors that reliably separate human from machine. Buy the better tool, tighten the threshold, wait for the next release.
It's a reasonable instinct, and it isn't entirely wrong — models do improve. But it assumes the problem is precision. It isn't. Even a very accurate detector hands you a probability, and a probability is not a finding you can put in front of a student, an editor or a regulator.
We think you're solving a detection problem that's actually a workflow problem
In our experience building agents, the moment you rely on a black-box score to make an accountability decision, you've already lost the argument. A number that says "82% likely AI" can't be cross-examined. It can't show its reasoning. It can't be appealed. That's not a stronger tool away from being fixed — it's the wrong shape of evidence.
This is our long-standing view: the workflow is the IP. AI output on its own is difficult to attribute and, in most jurisdictions, difficult to protect. What you can attribute, defend and stand behind is the human-authored process that produced it — the drafts, the prompts, the decisions, the review steps and the person who signed off.
Chasing detection is a defensive posture. You're forever one model release behind the thing you're trying to catch, running an unwinnable arms race against your own students, writers or contributors. Building a verifiable process is an offensive one. You stop asking "can I prove this was AI?" and start asking "can this person show me how they got here?"
That reframing changes what you buy. It's the difference between a detector that guesses and an agent workflow built for education that captures provenance as work happens — versions, contributions and checkpoints — so authorship is a record, not a hunch. The same logic protects a publisher's copy desk or a marketing team's blog: originality you can evidence beats originality you assert.
We build AI systems for a living and we'll say it plainly: a tool that produces an unaccountable score is worse than no tool, because it launders a guess into a verdict. If your integrity policy would collapse the first time someone demanded to see the working, it was never a policy — it was a bet on a vendor.
What this means for marketing teams
Content teams face the same authenticity question as universities, minus the tribunal. Here's the practical version:
- Stop treating detector scores as pass/fail. If you run one at all, cap its role at "prompt a conversation" — never a rejection — and document that rule this quarter.
- Capture provenance by default: keep briefs, drafts and edit history for every published piece, so any article can be reconstructed in under 10 minutes.
- Write a one-page AI-use disclosure standard for contributors and freelancers, and make it a contractual expectation, not a suggestion.
- Audit your last 20 published pieces for a named human owner and a review step. Anything without both gets a process, not a scolding.
- If you want help designing that provenance workflow rather than buying another scanner, our team can walk through your content pipeline.
Frequently asked questions
Are AI writing detection tools reliable enough to accuse someone?
No. Independent studies and university guidance document that AI detectors produce both false positives and false negatives, so their scores should never be treated as proof of misconduct on their own.
What's a better alternative to AI detection for academic integrity?
Focus on verifiable process. Ask students and writers to retain drafts, notes and version history so authorship can be demonstrated directly, rather than inferred from a probability score.
Can AI-generated writing be copyrighted?
In most jurisdictions, purely AI-generated output isn't protectable. The defensible layer is the human-authored workflow around it — the decisions, edits and structure a person contributes.




