AI visibility is the practice of recording how a student worked with AI — every prompt, every response, every revision — so an educator can assess the process rather than estimate whether the finished text was machine-written.
A detector puts you in an argument with your student. A record puts you in a conversation.
A record of how the work was done. Three things you need in order to grade it, none of which a probability about the finished text can give you.
Which sentence, when it arrived, and what came before it. To grade a piece of writing you have to know how it was made — something you read in a record, not something you judge from the finished text.
Feedback only teaches if the student can look at the same thing you looked at. A transcript of their own session is something to talk about together; it belongs to them as much as to you.
Anything that arrives after submission is a grade. You can see a student stuck, or leaning on the AI, while there is still time to say something. That is the part that changes what they learn.
Schools that stop relying on detectors usually move to one of three approaches. They answer different questions, and most teachers end up combining them.
They are not two versions of the same tool. They ask different questions, and only one of the answers has a next step attached.
| AI detection | AI visibility | |
|---|---|---|
| The question | Might this text have been written by AI? | How did this student work? |
| What you get back | A probability about the finished text | The sequence — prompts, replies, revisions, in order |
| Where it looks | At the artifact, after the fact | At the working, as it happens |
| What you can do next | Raise it with the student, or let it go | Teach, grade against a rubric, or ask a better question |
| What the student sees | A verdict about them | Their own working, which they can learn from |
| What a parent or an appeal panel sees | A number they are asked to trust | The exchange itself — what was asked, what came back |
| Where it struggles | Hardest on students writing in a second language | Only covers work done inside the workspace |
Accuracy figures are mostly vendor-reported and conditional on how much of a document gets flagged, so the average rate is the less useful number. What matters is who the errors fall on. When researchers tested seven commercial detectors on essays by non-native English writers, every one of them misclassified that writing as AI-generated — while classifying native-speaker writing correctly.
The likely cause is mechanical rather than malicious: a writer with a smaller working vocabulary produces more predictable text, and predictability is what these tools measure. The researchers cautioned against using them in educational settings for exactly this reason.
Where most students are writing in their second language, that failure mode lands on the same people repeatedly. Visibility has no equivalent, because it reads what happened rather than estimating it from the prose.
Liang et al., Patterns (Cell Press), 2023 — US eighth-grade essays vs TOEFL essays.
AI detectors promised to settle this and did not. Here is what a record genuinely cannot do.
A student can draft somewhere else and paste it in. What the record shows is a page that arrived with no working behind it — a signal worth raising, though not proof.
Cognity does not label a student dishonest and does not produce a score you are meant to act on. It shows what happened. The judgement stays with the teacher, where it belongs.
Work done in another tab is not in the record. What Cognity can tell you is what happened in the task you set, which is the part you are grading.
Answered straight, including the ones where the honest answer is not the one we would prefer.