When the Accuser Is a Chatbot

Vision. Innovation. Competitiveness. Growth.

When the Accuser Is a Chatbot

Spread the love

What a Brazilian Plagiarism Scandal Teaches Us About AI and Academic Integrity

eyesoneurope

In mid-2023, a story out of Brazil went viral for an unexpected reason. It wasn’t about a student caught cheating with ChatGPT. It was the opposite: a professor had used ChatGPT itself to decide that a student’s final paper was AI-generated — and got it wrong.

A university student in Brasília had her uncle’s undergraduate thesis flagged by the review board after a professor ran it through ChatGPT and asked whether the tool had written it. ChatGPT, ever confident, said yes. The accusation nearly cost him his grade — until he found a way to fight back: he ran one of the professor’s own published articles through the same chatbot, and ChatGPT just as confidently declared that it, too, had generated the professor’s writing. The case went public, racked up millions of views, and cracked open a much bigger problem hiding underneath it.

Once the story spread, more students came forward with similar complaints. A computer science student in Rio de Janeiro described his study group receiving a zero after a professor accused them of using AI — with no evidence offered beyond the accusation itself. He and his groupmates approached the professor, the course coordinator, and the university ombudsperson, without success. Meanwhile, in a São Paulo school, a Portuguese teacher had better luck: she used ChatGPT to check essays from two seventh-graders she’d grown suspicious of, and when she revealed what she’d done, one student admitted to using the tool. But her method worked more because she knew her students’ writing intimately from years in the classroom — not because the chatbot was actually a plagiarism detector.

Why “asking the chatbot” doesn’t work

It’s worth being precise about why this keeps failing, because the mechanics matter for anyone designing academic policy.

ChatGPT generates text by predicting likely words based on patterns learned during training — it doesn’t retain a searchable archive of everything it has ever generated for anyone. Computer scientist Lucas Lattari, who has studied this issue, points out that once a model is trained and released, no new interactions get folded back into its knowledge until an entirely new version is built. So when a professor asks the tool “did you write this?”, there is no database being checked and no memory being consulted. The model is simply generating another guess, in the same confident tone it uses for everything else — including for content it demonstrably did not write.

That last point was easy to demonstrate. Fact-checking outlet Aos Fatos ran an experiment: they fed ChatGPT two paragraphs from a 1998 Brazilian edition of Era dos Extremos, the historian Eric Hobsbawm’s landmark 1994 book — text that existed a full 24 years before ChatGPT was released — and asked whether it had been AI-generated. ChatGPT answered yes, attributing the passage to itself.

Eric Hobsbawm “Era dos Extremos”

The outlet then tried the reverse experiment: they took a genuine ChatGPT-written passage and ran it through several dedicated AI-detection tools. Most of them — tools built specifically for this purpose — wrongly concluded the text was human-written. Only one, OpenAI’s own experimental classifier, got it right, and even that tool carried its own disclaimer that it should not be trusted for high-stakes decisions and performs worse on text from children or non-English writers.

The uncomfortable conclusion, echoed by researchers and privacy experts alike, is that reliable AI-detection may not be a problem technology solves anytime soon — and treating any current tool as a verdict-machine is itself a kind of malpractice.

The real damage isn’t cheating — it’s false accusation

It’s tempting to file this under “growing pains of new technology” and move on. But the stakes here are not abstract. Students described real consequences: a zeroed group assignment, a risk of failing a semester, months added to a degree — all triggered by a tool that cannot actually do what it was asked to do, with no corroborating evidence and, in at least one case, no real path to appeal.

Rafael Zanatta

Rafael Zanatta, director of the Brazilian NGO Data Privacy Brasil, frames this as less a story about tempted students and more a story about institutional unpreparedness. Chatbots are designed to produce confident, fluent answers regardless of whether those answers are true, which is exactly the design trait that misleads an unprepared evaluator. The risk compounds in institutions where teachers are under-resourced, unsupported, and working without any real infrastructure for handling AI-related integrity questions — they reach for the tool that’s in front of them, because nothing better has been given to them.

That is a fixable problem. It just isn’t fixed by asking a chatbot to judge itself.

A better dividing line: process, not prose

If output can no longer reliably tell you whether AI was involved, the sensible move is to stop trying to reverse-engineer authorship from the finished text and instead look at the process that produced it. That reframing does double duty: it protects innocent students from false accusations, and it gives real structure to what “legitimate AI use” actually looks like.

Signs of legitimate, productive AI use:

  • The student can explain and defend every claim, structure choice, and piece of evidence in their own words, unprompted.
  • There’s a visible trail — drafts, notes, an outline, a history of revisions — showing the work developing over time.
  • AI was used the way a tutor or research assistant is used: to brainstorm angles, stress-test an argument, clarify a confusing concept, or clean up grammar — not to originate the core ideas or analysis.
  • The student discloses AI use where it’s relevant, rather than hiding it, because the institution has made disclosure normal and safe rather than an automatic confession of guilt.
  • The final work reflects the student’s own voice, prior skill level, and specific course context — things a generic AI-generated answer rarely captures.

Signs actually worth investigating (carefully, and with due process):

  • No supporting drafts, notes, or research trail exist at all — the work appears fully formed.
  • The student cannot explain their own reasoning, terminology, or sources when asked directly, in conversation.
  • The content strays into material never covered in class, or performs at a level sharply inconsistent with the student’s earlier, verified work.
  • There’s a pattern across multiple submissions of generic phrasing, oddly uniform structure, or the same distinctive errors also seen in classmates’ work.

Notice what’s missing from the second list: “a chatbot said so.” That’s deliberate. Every one of these signals depends on a human evaluator who knows the student, the course, and the work — not on outsourcing judgment to the same tool the policy is trying to police.

What this means in practice

For students: the safest and most genuinely useful habit is to keep a visible paper trail of your own thinking — save drafts, keep notes on what you asked an AI tool and why, and make sure you can talk through your work’s reasoning without the page in front of you. That habit protects you twice over: it’s what real learning looks like, and it’s your best defense if you’re ever wrongly accused.

For institutions: stop treating any chatbot — including the very tool students might be using — as a verdict machine. Build assessment around process (oral check-ins, drafts, revision history) rather than a single final artifact judged in isolation. Give faculty real training and real policy, so they’re not left, unsupported, to invent their own detection methods under pressure. And build a genuine appeals process before a false accusation costs a student a grade, a semester, or their trust in the institution.

As Zanatta put it: students determined to cheat will keep finding creative ways to do it, whatever policy exists. The more urgent fix isn’t a better detector — it’s an institutional culture that treats AI as a subject for open, honest conversation with students rather than forbidden fruit to be sniffed out after the fact. The Brasília case didn’t prove that Brazilian students were cheating with AI. It proved that punishing suspicion instead of evidence is a far easier trap to fall into — for institutions as much as for anyone else.

eyesoneurope

 

Leave a Reply

Your email address will not be published. Required fields are marked *