Back to Insights
Insights15 min read

AI Detectors for Teachers: What Districts Bought and What the Research Shows

Abbas Khan
Abbas KhanSeptember 18, 2026
AI Detectors for Teachers: What Districts Bought and What the Research Shows

Last updated: September 19, 2026

Quick Answer

Schools are adopting AI fast and buying it barely at all. The uptake sits one layer below the purchasing office. On a teacher card, or in a free browser tab. The most widely deployed AI in American classrooms is the detector, pointed at students rather than used by them, and it is the category almost nobody formally procured.

  • In Civic IQ spend records, districts bought AI detectors for as little as $156, and one charter paid for GPTZero on a credit card.
  • Turnitin appears in 13,037 agency records. GPTZero appears in 80, Copyleaks in 237. Most teachers did not choose their detector.
  • On humanized AI text, Pangram held 92.5% accuracy. Turnitin managed 50%, Copyleaks 22.5%, GPTZero 2.5%.
  • Honest AI editing gets flagged at 38 to 80%. Text run through a humanizer stays flagged under 4%.

Almost every article about AI in schools is an opinion. This one has purchase orders under it.

This is the companion to our wider look at how fast districts buy AI compared with how fast they write rules. That piece covers the whole category. This one covers the part pointed at students.

Schools are adopting AI fast and procuring it barely at all. The uptake happens one layer below the purchasing office. On a teacher card. Inside a free Chrome extension. Through a browser tab nobody signed a contract for.

And the single most widely deployed kind of classroom AI is the one pointed at students rather than used by them. The AI detector.


What the purchase records actually show

Start with what districts pay for detection. The amounts explain the governance gap on their own.

District Tool Amount Record date
Simi Valley Unified, CA GPTZero Premium, annual $1,344 Jun 2024
Pennsauken Township Board of Education, NJ GPTZero subscription $1,920 Jul 2024
Freire Charter School, PA GPTZero, paid by credit card $156 Mar 2025
Marion Independent School District, IA Turnitin Originality add-on $1,283 Nov 2024
Santa Clara Unified, CA Copyleaks renewal $3,000 Jan 2025
New Haven Unified, CA Annual plagiarism detection $9,273 Oct 2024

Individual purchase records surfaced by Civic IQ. Amounts are as recorded.

Every figure there sits far below any bid threshold. Most district thresholds start around $25,000.

The Freire record is the whole argument in one line. A school bought an AI detector on a credit card. No bid. No board vote. No privacy review. None was required at $156.

Most teachers did not choose their detector

Turnitin appears in 13,037 agency records in our data. GPTZero appears in 80. Copyleaks appears in 237.

That ratio is the market. Districts bought Turnitin for plagiarism years ago. The AI feature arrived in April 2023 as a switch someone flipped. The incumbent won by already being there.


Do AI detectors for teachers actually work?

The honest answer has changed. Most of what ranks for this question is three years out of date in one direction, or vendor copy in the other.

The 2023 evidence was damning

Weber-Wulff and colleagues tested 14 tools in the International Journal for Educational Integrity. Their verdict: neither accurate nor reliable. All scored below 80%. Only five cleared 70%. Paraphrased AI text was caught 26% of the time.

Liang and colleagues, in Patterns, found seven detectors mislabelled over half of TOEFL essays as AI. The average false positive rate was 61.22%. The same tools scored near-perfectly on US 8th-grade essays.

The causal test is the part that matters. Richer vocabulary in the TOEFL essays dropped the false positive rate to 11.77%. Simpler word choice pushed mislabelling of real 8th-grade essays from 5.19% to 56.65%.

The detectors were reading linguistic simplicity. They were not reading authorship.

OpenAI’s own classifier caught 26% of AI text while falsely flagging 9% of human text. It was withdrawn in July 2023 for low accuracy.

The 2026 evidence is split, and the tool matters

A June 2026 head-to-head in the same journal tested GPTZero, Pangram, Copyleaks and Turnitin. The good news: all tools correctly found fully human texts, and false positives were rare across the board.

The bad news is sensitivity.

Tool Accuracy on humanized AI text
Pangram 92.5%
Turnitin 50%
Copyleaks 22.5%
GPTZero 2.5%

So the claim that detectors do not work is no longer true as stated. What is true is narrower, and worse for schools. The detector that performs best in the literature is the one almost no district has. The three most teachers use are the three that fail on exactly the text a cheating student would produce.

The finding that survives any accuracy improvement

Researchers at Notre Dame tested what happens to students who follow the rules. Light edits stand in for the rule-following AI help most districts now allow. Those were flagged at 38 to 80%. Text run through a humanizer tool stayed flagged under 4%.

Their conclusion, verbatim: honest AI editing results in a higher sanction risk than humanizer-assisted evasion.

Better models do not fix that. It describes what the tool rewards. A detector that punishes disclosure and rewards concealment is not a weak integrity instrument. It is an inverted one.

See which districts are buying AI tools right now, and what they paid
Civic IQ tracks 100,000+ agencies, $14T+ in tracked spend, 1.5M+ documents per month, all 50 states.

See Civic IQ →


Is 20% AI detection bad?

This gets asked constantly, and the honest answer is that Turnitin decided you should not be asking it.

In March 2023 the company claimed a false positive rate under 1%. Two months later its chief product officer revised that publicly: real-world use was yielding different results from the lab.

Turnitin tested 800,000 pre-ChatGPT papers. Below 20% AI writing, it found more false positives. The under-1% claim was re-scoped to documents above 20%. A new figure appeared: a sentence-level false positive rate of about 4%.

By July 2024 the company stopped showing the low band at all. Its own own guidance states that scores in the 1% to 19% range get no score and no highlights, to avoid false positives.

So a 20% score is the floor of what Turnitin is willing to display, precisely because everything below it was too unreliable to show. It is the least meaningful number in the product. It is not a threshold of guilt.

The underlying problem is real. Turnitin’s own data shows roughly 15% of essays submitted between October 2025 and February 2026 were more than 80% AI-written. At launch it was 3.3%. The worry is justified. The instrument is the issue.


The free detector problem is a purchasing problem

Thousands of Americans a month search for a free AI checker. That demand is the clearest signal in this market. Teachers need a tool. They do not have one. They will not wait for the district to buy one.

Follow what that means in contract terms.

A free web detector is not on an approved vendor list. It has no data privacy agreement. It has had no security questionnaire, no FERPA review, no state privacy alliance. A teacher pastes in a student essay and that work leaves district control. It goes to a vendor whose training practices, retention policy and home country nobody has checked.

The Education Department’s Office for Civil Rights has already written down how this goes wrong. Its guidance on avoiding discriminatory use of AI opens with a scenario. A teacher runs book reports through a free online service that claims it can spot generative AI. The service has a low error rate for native English speakers and a high one for everyone else. It flags the class’s only two English Learners. Neither used AI.

The framing is possible national-origin discrimination under Title VI. Given the search volume, that is not a warning about some future risk. It describes standard practice.

The instrument meant to catch this does not reach it. The National Data Privacy Agreement is the dominant K-12 contract form, with more than 222,000 agreements signed. It reached version 2.2 in November 2025 with no AI provisions at all. Even a perfect agreement would not help, because a free browser tool never touches one.


The rules exist. Almost none of them bind.

State education agencies have been clear that detectors should not be trusted. Almost none of that guidance has any force.

AI for Education’s tracker counted 35 states plus Puerto Rico with official K-12 AI guidance as of August 2026. More than a third of states discourage relying solely on detection tools.

Some are blunt. North Carolina says AI detectors have proven undependable, and should never be the only factor in deciding whether a student cheated. Massachusetts tells districts to stop relying on them, calling them often inaccurate.

Guidance is not law. State guidance also tends to skirt the cheating question. It tells districts not to use detectors, without telling them what to do instead.

Three 2026 statutes changed that, and only one bites

  • Maryland SB 720, effective June 1, 2026, requires every local school system to adopt an AI policy within 120 days of state guidance, name a central-office AI coordinator, and use state rubrics when selecting AI systems.
  • Idaho S1227, effective July 1, 2026, requires vendors to disclose whether products use generative AI and give assurances on data protection and algorithmic transparency.
  • Oklahoma SB 1734, the Responsible Technology in Schools Act, passed 42 to 0 and 89 to 0. It requires human review of AI outputs and annual parent disclosure of every AI vendor in use. It also bars AI as the primary basis for grading, discipline, placement, promotion or retention.

Oklahoma never uses the word detector. But that clause is the only enacted US law that limits detector-driven punishment. It arrived through a purchasing statute rather than an academic integrity one.


Districts are regulating the wrong side of the desk

Look at what the biggest districts did in 2026 and the asymmetry is hard to miss.

New York City adopted a moratorium on student-facing generative AI from pre-K through 8th grade. It added screen-time caps and a few high school pilots. Los Angeles Unified restricted generative AI for all students, all grades, for 2026-27.

Both policies are detailed about what students may use. Neither says anything about what a teacher may run a student essay through.

That is the governance gap in one sentence. Districts write rules about student access. Meanwhile the AI that can fail a student, alter a transcript and trigger a discipline hearing sits on the other side of the desk. Unregulated, and mostly unpurchased.

Students noticed before districts did. RAND’s September 2025 survey found 54% of students using AI for school, and half of students worried they would be falsely accused of using it. Only 45% of principals reported any AI policy.


What the courts have actually said

Three cases define the record, and together they say something narrower than either side claims.

The student win was procedural

Newby v. Adelphi University is the case people cite. A professor reported a Turnitin AI score of 100%. The student produced two other detectors, both returning 0%. The court annulled the integrity violation as arbitrary and capricious, ordered the record expunged and denied the university’s motion to dismiss.

Read the holding carefully. The defect was procedural. The university ignored contrary evidence and denied a real appeal. The court did not rule that detectors are invalid. Schools may keep using them. What they may not do is treat a score as a closed record.

The two cases that do not say what people think

The Hingham High School case is cited everywhere as the AI false-accusation case. It is not. The student admitted using generative features, and the work contained made-up citations. The dispute was about how harsh the discipline was. The school won.

Kato v. Palo Alto Unified was the real test. It involved a 76% Turnitin score and a demand that the district stop treating scores as final. It was dropped in September 2026. The district reported no settlement, no payment, and no change to any record or policy.

Net: no US court has held that relying on a detector score violates a K-12 student’s rights. What courts have held is that ignoring the student’s rebuttal does. The exposure is procedural, and it is entirely avoidable.

Higher education is moving faster

Berkeley opted out within weeks of launch. Vanderbilt disabled the detector in August 2023. At a 1% false positive rate, it calculated, around 750 student papers could have been wrongly labelled. Yale and Pittsburgh followed.

Washington State University went further in February 2026. It cancelled its Turnitin AI detection contract outright and published the math. Roughly 1,485 of 148,547 papers were likely falsely flagged. About a third of integrity cases resolved not responsible where detection was the only evidence.

That one matters because it is a contract cancellation with published false-positive math. A purchasing decision rather than a pedagogical opinion.


What is replacing detection

Vendors have read the room. The category is pivoting from judging finished text to recording how text gets written.

Turnitin Clarity shows a student’s drafting process. It captures pasted text, typing patterns and draft history, with video playback. Uptake is unclear, and the company has published no numbers. Colorado Boulder ended its pilot, saying it did not show enough to justify wider use.

Grammarly Authorship sorts text into typed, pasted, AI-generated or AI-modified. In March 2026 it added agent-level tagging. Then in June 2026 the parent company bought GPTZero. Worth saying plainly. The leading standalone detector is now a feature inside a writing assistant, made by a company whose own product generates the text being detected.

Brisk Teaching’s Inspect Writing is the most honest of the lot. It declines to produce a score at all, showing a factual record of how the document was produced instead.

Watermarking is not the answer

Google’s SynthID portal launched in May 2025 promising text detection in the coming weeks. Sixteen months later it accepts only images, video and audio. Google watermarks its own model output and gives educators no way to check it.

OpenAI has never shipped text watermarking. It cites tamper-fragility, and the risk of stigmatising AI use by non-native English speakers. C2PA does not support plain text at all. And a May 2025 attack paper strips watermarks from seven schemes at near-total success, for under a dollar per million tokens.

Keystroke analysis is the live research frontier. Its own authors are cautious. The lead researcher says it is not ready to be a stand-alone arbiter of misconduct, and raises privacy and consent issues. That is the right instinct. Every one of these process tools is keystroke-level surveillance of minors. None has been through a review designed to ask that question.

The quieter answer is assessment redesign. One London law school moved from take-home essays to invigilated exams and watched first-class grades rise from 13% to over 45%.


What a district should actually require

Detection is where the risk clusters. It is the only AI that produces action against a student. Here is a usable standard, in the order the questions matter.

  1. Name the tool, or ban the category. An approved list with one named detector beats a policy that says use your judgment. That one produces forty free web tools with no agreement behind them.
  2. Write the rebuttal process before the purchase order. The liability is procedural. A student must be able to see the score, respond, and submit contrary evidence that gets weighed on the record. That costs nothing. It closes the only risk any court has found.
  3. Require the false positive rate in writing, broken out by English Learner status, with the study attached. The market leader will not be able to supply a current number. That itself is information.
  4. Prohibit the score as sole evidence. This is now statutory in Oklahoma and recommended by more than a third of states. Put it in the contract, and not only the handbook.
  5. Ask the surveillance question about process tools. Draft replay and keystroke telemetry are more intrusive than a detector rather than less. They need an agreement covering behavioural data, a retention limit, and a parent disclosure.
  6. Budget for the free-tool problem. Teachers will find a tool whether you give one or not. The cheapest risk control in this category is giving them a vetted one.

The signal underneath all of this

Step back from detection and the pattern holds to every category of school AI.

Google turned Gemini on by default for students of all ages in August 2026. OpenAI’s ChatGPT for Teachers is free through June 2028. It reached more than 300,000 educators by August 2026. Microsoft bundles Copilot into its office suite. Huge volumes of classroom AI are arriving through contracts nobody renewed, and switches nobody flipped.

Formal purchasing is not slow because districts are careless. It is slow because it was built for a world where technology arrived as a purchase. A thing with a price, a signature and a renewal date. AI arrives as a default setting. Or a free tier. Or a $20 charge on a teacher card.

The purchase orders are not the market. They are the thin slice of the market that still looks like buying.

For anyone selling into K-12, the implication is the same. The buying signal is no longer the RFP. It is the policy uptake deadline. The board agenda item. The state rubric requirement. The superintendent memo naming a tool. Maryland districts are adopting AI policies now, and Oklahoma districts have until 2027-28. Every one of those policies will name approved vendors and exclude others. Almost none of it will be competitively bid.

That is where the rules are being written. It is just not where anyone is looking for them.

The buying signal is the policy deadline rather than the RFP
Civic IQ surfaces board agendas, policy adoptions and purchase records before a bid ever posts.

Start Free →

Open RFPs related to this topic

Live opportunities surfaced by Civic IQ as of September 2026. Status changes daily.

Track every RFP in your category with Civic IQ →


Frequently asked questions

What AI checker do teachers use?

Most commonly Turnitin, because districts already license it for plagiarism and the AI feature was added to the existing product in 2023. Teachers without district access typically use GPTZero, Copyleaks, ZeroGPT or a free web tool. Between roughly a third and two-thirds of secondary teachers use a detector regularly, depending on the survey.

How do teachers check for AI?

Three ways. Running text through a detector. Comparing the work to a student’s known writing. And increasingly, inspecting the writing process itself through revision history or replay tools. The third is where the market is moving, because it produces a record rather than a probability.

Is 20% AI detection bad?

20% is the lowest score Turnitin will display. It is the floor precisely because everything below it produced too many false positives to show. Turnitin’s own own guidance says scores from 1% to 19% are suppressed. A 20% result is the least informative number in the product, not evidence of misconduct.

Can teachers detect ChatGPT without a tool?

The research says no. A 2024 study found experienced teachers correctly found only 37.8% of AI-generated texts, worse than chance, and were overconfident in their judgments. A PLOS ONE study that slipped fully AI-written answers into real university exams found 94% went undetected.

What is the best AI detector for teachers?

By the independent reviewed evidence, Pangram. A June 2026 comparison found it held 92.5% accuracy on humanized AI text where Turnitin managed 50%, Copyleaks 22.5% and GPTZero 2.5%. But no detector should be the sole basis for a discipline ruling, and Oklahoma has made it unlawful to use AI as the primary basis for grading or discipline.

Are free AI detectors safe for teachers to use?

Not from a compliance standpoint. A free web detector has no data privacy agreement with the district, so student work leaves district control to an unvetted vendor. The Education Department’s Office for Civil Rights uses exactly this scenario as its lead example of potentially discriminatory AI use under Title VI.

Related reading: our analysis of the largest K-12 edtech vendors by spend, our look at the ESSER budget cliff, and our guide to selling into school districts.

Purchase figures come from individual records surfaced by Civic IQ from agency purchase orders and checkbook data. Detector accuracy figures come from the cited reviewed studies rather than vendor materials. Civic IQ tracks 100,000+ agencies, $14T+ in tracked spend, 1.5M+ documents per month, all 50 states.

Abbas Khan

Written by

Abbas Khan

Bring us your territory.
We'll show you what is forming.

B2G and SLED sales intelligence. Surface government procurement signals from 100,000+ state, local, and education agencies months before the RFP.

Try Civic IQ for free