Insights & Perspectives

How a treadmill saved my academic integrity: A personal insight into the growing role (and risk) of AI detectors in academic assessment

Brianne Lindsay

After eight years away from formal education, I recently returned to postgraduate study. Having just submitted my first assignment, I received an unexpected notification—my work had been flagged by the university’s AI detection software for suspected use of generative AI.

I was shocked. How on earth was I meant to prove I had written this assignment myself? I spent the evening upset and panicking slightly, before reminding myself that sleeping on it was probably a good idea—and that I’d need a plan of attack in the morning.

While the assignment was entirely my own—reflecting both the course content and my professional, real-world experience—the detection system identified my style of writing and level of synthesis as ‘AI-generated.’

In true academic fashion, I began researching the issue myself. I came to realise that characteristics such as high-level synthesis, a polished tone, and advanced writing structure may have contributed to the false positive. While I was ultimately able to defend my work, the experience highlighted some significant concerns around the use of AI detection tools in higher education—particularly their reliability, potential for harm, and the broader implications for learning, trust, and fairness.

The rise of artificial intelligence has transformed the world as we know it over the past few years. It’s reshaping how we do business, how we work, how we connect with one another—and how we learn.

This shift has raised important questions about how institutions uphold academic integrity. This has led to the rapid adoption of AI detection tools in higher education institutions, such as universities and polytechnics, in an attempt to flag work produced by generative AI and protect academic integrity—a move driven by good intentions, but not without consequences.

However, the use of these AI detectors brings significant risks and ethical challenges, most prominently, false positive results and the impact this has on learners who are being accused of ‘cheating’.

In my case, the assignment topic—and the lens I brought to it—aligned closely with areas I have extensive real-world experience in. Therefore, I felt I was able to demonstrate a level of synthesis and critical thought beyond the learning from the course. This, combined with my experience as a writer and editor, meant my assignment had a polished feel—it was clear, well-structured, and demonstrated advanced critical thinking.

The kicker? Research by Elkhatat et al. (2023) found that AI detection tools have been proven to associate high-quality, human-written content as AI-generated, particularly work with advanced synthesis, structure, and technical polish.

The AI detection software that my education institution uses claims to be 98% accurate. However, their own AI scientist has been publicly quoted as acknowledging that the results should be interpreted with a ‘grain of salt,’ noting that human judgement is essential when reviewing flagged content. Numerous academic studies have demonstrated the risk of false positives (particularly in strong writers and non-native speakers) (Liang et al., 2023; Elkhatat et al., 2023), a lack of transparency in how the algorithm works (Coffey, 2024), and a significant risk of undermining falsely-accused learner trust and motivation (Fowler, 2023).

Collectively, these studies acknowledge one crucial insight—these tools cannot replace critical human judgement. They highlight the need for due process by education institutions, human interpretation, and contextual understanding of the reports. Higher education institutions should reflect on exactly how these tools are used, and ensure that teaching staff and other faculty members are aware of their limitations and how to effectively use them to support judgements of academic integrity—not to deliver the final verdict. AI detector reports must be used as the start of dialogue with a learner—not the basis for an accusation of ‘cheating’.

And how did a treadmill save me, you may be asking?

While trawling through my version history to demonstrate the development of my assignment, I stumbled across a note I’d left to myself mid-draft:
“Reference here – from that article I read on the treadmill, I can’t remember it atm haha.”

That purely human, messy, unfiltered moment turned out to be the piece of evidence that confirmed my assignment was, in fact, my own.

So from now on, I just might do all my readings while walking on a treadmill—just in case.

Coffey, L. (2024, February 9). Professors cautious of tools to detect AI-generated writing. Inside Higher Ed. https://www.insidehighered.com/news/tech-innovation/artificial-intelligence/2024/02/09/professors-proceed-caution-using-ai 

Elkhatat, A. M., Elsaid, K., & Almeer, S. (2023). Evaluating the efficacy of AI content detection tools in differentiating between human and AI-generated text. International Journal for Educational Integrity, 19, 17. https://doi.org/10.1007/s40979-023-00140-5 

Fowler, G. A. (2023, April 3). We tested a new ChatGPT detector for teachers. It flagged an innocent student. The Washington Post. https://www.washingtonpost.com/technology/2023/04/01/chatgpt-cheating-detection-turnitin/ 

Liang, W., Yuksekgonul, M., Mao, Y., Wu, E., Zou, J. (2023). GPT detectors are biased against non-native English writers. Patterns, 4(7), 100681. https://doi.org/10.1016/j.patter.2023.100779 

Discover more from Base Insights

Subscribe now to keep reading and get access to the full archive.

Continue reading