Picture a student who spent a week on her essay. She wrote every sentence herself, in English, her second language. A few days later, she gets an email: the AI detector flagged her work, and she has to explain herself.
For the student, it’s stressful. For the lecturer, it’s a dead end. The detector gives a percentage, not proof, and no one can say for sure who wrote the essay.
That’s the core problem with AI detection. It tries to catch AI after the fact, while AI can already write almost any assignment. The more useful question comes earlier: what should students be able to do on their own, and where does AI actually belong?
So, how do you decide?
This article explains why AI detectors can’t solve academic integrity, how to redesign assessment around what really matters, and how the same logic applies when AI moves beyond the classroom.
What Are AI Detectors
An AI detector is a tool that estimates whether a piece of text was written by a person or generated by AI. It analyzes patterns in the writing, such as how predictable the word choices are, and returns a probability score. In higher education, AI detectors are mostly used to check student essays and assignments.
The key word is estimates. An AI detector doesn’t find evidence in the text. It makes a statistical guess about how the text was produced.
How AI Detectors Analyze Text
Large language models generate text by predicting the most likely next word. As a result, AI-generated writing tends to be smooth, consistent, and statistically predictable. Detectors are built to spot that predictability.
Most tools combine a few signals:
- Predictability of word choices (perplexity): how easy it is for a language model to guess each next word. Low perplexity suggests AI-generated text.
- Variation in sentence rhythm (burstiness): people tend to mix long and short sentences, while AI output is often more uniform.
- Trained classifiers: machine learning models trained on large sets of human-written and AI-generated text, which learn to tell the two apart.
The output is usually a percentage or a label, such as “likely AI-generated.” Some tools also highlight the sentences they consider suspicious.
How Universities Use AI Detectors
In many universities, AI detection isn’t a separate step. It’s built into tools lecturers already use. Turnitin, for example, offers AI writing detection as part of its Similarity Report, available to institutions with a Turnitin Originality license. Through integrations with learning platforms such as Canvas, Moodle and Blackboard, lecturers can see an AI score right next to the plagiarism check.
A typical workflow looks simple. A student submits an assignment through the learning platform. The system runs the check. The lecturer sees a score and decides whether to follow up.
That simplicity is exactly why detectors spread so quickly on campus. And the pressure is real: in a national survey by AAC&U and Elon University, 73% of faculty said they had personally dealt with academic integrity issues involving students’ use of AI. But a quick fix for a real problem is also where the problems start.
Why AI Detectors Fail in Higher Education
According to Inside Higher Ed, at least a dozen universities, including Northwestern, Georgetown and New York University, have disabled Turnitin’s AI detection software. Their reasons come down to one thing: detectors can’t reliably tell human writing from AI-generated text. In practice, that shows up in four ways.
1. Accuracy: Detectors Get It Wrong in Both Directions
AI detectors flag human writing as AI-generated, and they miss text that AI actually wrote. Turnitin’s own detector turned out to have a much higher false-positive rate than the company originally suggested.
The results aren’t stable either. Earlier in this article, we showed the ZeroGPT report for the introduction of one of our blog posts. When we ran a shorter fragment of the same text, the tool highlighted different sentences as AI-generated, including ones it hadn’t flagged the first time.

ZeroGPT report for a shorter fragment of the same text. Note the sentences highlighted this time and the “Make it Human” offer below the result.
Look closely at the report, and you’ll notice something else. Right below the result, next to the words “Using AI Text? It Can Be Detected,” there’s a button: “Make it Human.” The same service that checks text for AI also offers a way to make AI-generated text harder to detect. Indiana University’s teaching center points to a growing body of online tools designed to make AI-generated text undetectable, and some of them sit on the same page as the detector.
That creates an uneven playing field. Students who write honestly risk being flagged, while those who know how to hide AI use often aren’t. So the tool can end up catching the wrong people and missing the right ones.
2. Fairness: Non-Native English Writers Pay the Price
Detectors look for writing that seems predictable: simple sentence structures, common word choices, a steady rhythm. You could see it in our own test. The sentences ZeroGPT highlighted were the most standard ones, like a list of familiar terms such as “declining GPAs, attendance patterns, financial aid status.” The trouble is that many people write like that for reasons that have nothing to do with AI.
Students writing in their second language often rely on clearer, more standard phrasing. So do students taught to write in a structured, formulaic way. Indiana University’s guidance for faculty notes “a significant number of false positives, especially with text generated by non-native speakers of English.”
For universities with large international or online student populations, this is more than a technical flaw. It’s an equity problem built into the grading process.
3. Evidence: A Detector Score Isn’t Proof
Imagine a lecturer bringing an integrity case to a disciplinary committee. The only evidence is a report saying the essay is “78% likely AI-generated.” What does that number actually prove?
Very little, and even the companies behind these tools say so. Turnitin states that the percentage on its AI writing indicator “should not be used as the sole basis for action or a definitive grading measure by instructors.” ZeroGPT goes further. The “Instructions for Educators and Evaluators” printed under every result say: “These outcomes should not be utilized to penalize students.”
Universities take the same view. MIT’s Committee on Discipline doesn’t consider detector output alone sufficient evidence in academic integrity cases. And that makes sense: when a system makes a judgment nobody can explain or verify, it’s hard to defend (we explored this problem in more depth in our article on explainable AI).
A detector score can start an investigation. It can’t finish one.
4. Privacy: Student Work Leaves the University
Every time a lecturer pastes an essay into a free online detector, that text goes to a third-party company. It may include the student’s name, personal experiences, or original research. And once it’s uploaded, the university no longer decides what happens to it.
That’s why Indiana University’s guidance says instructors should not upload student work to any AI-detection tools “due to privacy and intellectual property concerns.” As the same guidance puts it, “we have no control over how the companies use submitted work.”
For universities, this is also a compliance question. Student work often contains personal data, and in regions with strict data protection rules, such as the EU under GDPR, sending it to an unvetted external service can create legal risk on top of the academic one.
A tool meant to protect academic integrity shouldn’t put student data at risk.
What AI Detection Costs Universities
Together, these problems create a cost that’s harder to measure. When any assignment might be checked by an algorithm, students and lecturers start treating each other with suspicion. MIT’s committee on AI in education describes exactly this: professors feeling pressured to use unreliable detection tools, and a climate of mutual suspicion between students and instructors. Students may start writing to avoid a flag instead of writing to learn, and lecturers spend time investigating instead of teaching.
And even a perfect detector wouldn’t change the underlying situation. The same MIT report found that AI can already produce credible responses to almost any written assignment in the undergraduate curriculum, from essays to math proofs to code. When AI can complete an assignment on its own, the question is no longer how to catch it. It’s what that assignment should measure in the first place.
How to Redesign Assessment When AI Can Do the Assignment
MIT’s committee on AI in education suggests an order worth following. Before designing AI-aware assessments, “instructors should reconsider their goals for student learning in every subject they teach.” Schools and universities around the world are already testing what that looks like.
1. Decide What Students Must Be Able to Do Without Help
Take a first-year course where students write essays based on sources. AI can already produce a polished essay on almost any topic. So what should a student leaving this course still be able to do on their own? Maybe it’s judging whether a source is credible, or building an argument and defending it when someone pushes back.
It’s a question faculty are already asking themselves. In a national survey by AAC&U and Elon University, 90% said AI will diminish students’ critical thinking skills. Defining which thinking students still have to do themselves is the first step to protecting it.
2. Design Assessment Around That Skill
Denmark offers a clear example. The Danish Ministry of Education decided that upper-secondary students must verbally defend the major written assignment they complete at home. Teachers ask about the arguments, sources, methodology and conclusions in the paper. The written work still counts, but understanding is checked in conversation.
That’s the same direction the MIT report recommends for universities: oral exams, semester portfolios, and out-of-class assignments paired with in-class conversations. These formats sound old-fashioned, but they assess the thinking behind the work.
3. Decide Where AI Belongs
The University of Sydney turned this step into a simple framework: a “two-lane approach” to assessment.
- Lane 1 (secure): in-person assessment, such as interactive oral exams, in-class tasks and supervised tests. This is where the university needs a trustworthy judgment of what a student can do.
- Lane 2 (open): assessment where AI use is supported, for example to brainstorm, suggest structure, find counterarguments or give feedback.
The United Arab Emirates University takes a similar route in its student AI policy. Instructors assign each task one of three roles for AI:
- Prohibited: for foundation skills and in-class exams, among others.
- Assistive: AI supports the work, for example in brainstorming, proofreading or testing code.
- Integral to learning: using AI is part of the skill students are building.
The same policy still requires faculty to use AI detection tools, a sign of how many universities are caught between the old approach and the new one.
The rules differ from course to course. What they share is that each one follows from a learning goal, not from what a detector can catch.
Planning Where AI Fits at Your University?
AI detectors promised a simple answer to a hard question. They can’t deliver it, and universities that keep relying on them risk flagging honest students while missing the real issue. When AI can complete almost any written assignment, protecting academic integrity starts with deciding what students must still do on their own, and building assessment around it.
At DLabs.AI, we specialize in AI for higher education. We build AI solutions for universities, run AI training for faculty and staff, and help teams decide where AI belongs in teaching and assessment. If you’re planning to implement AI thoughtfully and responsibly, in a way that supports both students and faculty, contact us.
Artykuł Why AI Detectors Fail in Higher Education and What to Do Instead pochodzi z serwisu DLabs.AI.


