The six week pattern
A school runs an AI workshop. Attendance is good, feedback is positive, several teachers try things the following week.
Six weeks later, practice looks much as it did before.
This is common enough to be predictable, and it is usually not a reflection of the trainer or the staff. It is a consequence of what was taught.
Three kinds of AI literacy
The 124 study review identifies three distinct things people mean when they say AI literacy.
Functional literacy. How to operate the tool. Which button, which prompt, which feature.
Critical literacy. How to question it. How to oversee output, judge it against a standard, and decide when not to use it at all.
Indirectly beneficial. What learning to work with AI teaches you about your own thinking and reasoning.
Almost all school AI training delivers the first and stops.
There are understandable reasons. Functional content is easy to demonstrate, easy to structure into a ninety minute session, and produces visible activity in the room. Critical literacy is slower, less demonstrable, and requires the trainer to know the participants' actual work.
But functional literacy alone changes very little, for a reason that is worth stating plainly.
A teacher who can operate ChatGPT but has never been shown how to interrogate it is not better equipped than a teacher who never opened it. She is faster at producing something that looks finished.
The national data supports this. CENTA's 2025 survey of over 5,000 Indian educators found 67 percent rating their AI skills as above average, while only 57 percent could identify a basic AI misconception. Confidence is running ahead of capability, which is exactly what training that stops at operation produces. The full picture is in what the data actually says about AI use by teachers in India.
The specific failure this creates
Two failure modes appear in staff rooms, and they look opposite while sharing a cause.
Some teachers over-trust. The output arrives formatted, confident and grammatical, so it is accepted. An assessment stating that water evaporates only at 100 degrees goes to forty children, not because the teacher does not know the science, but because the output did not look like something requiring a check.
Others avoid entirely. Sensing that they cannot evaluate what they are being given, they decline to use it, which is a rational response to an unfamiliar instrument with no calibration.
Both come from the same place: not knowing what to do with professional judgment when AI is in the room.
The correction worth making
It is often said that you cannot distinguish a good AI output from a flawed one by looking at it.
That is not quite true, and the imprecision matters, because if nobody can tell then there is no point training anyone.
The accurate statement is that you cannot tell if you have never been trained for it and have not practised. Once you have, you can, quite quickly, and it stops feeling like additional work.
That is the entire difference between a workshop and a method.
What effective training does differently
It starts from what teachers already have. An experienced teacher already defines outcomes before planning, already anticipates misconceptions, already evaluates her materials against a standard and revises them. Those are the same judgments AI use requires. Effective training makes that transfer explicit rather than teaching the judgments from scratch.
It uses the teacher's real work. Not an invented example about the solar system, but the assessment she is genuinely setting on Thursday. Transfer from artificial examples is weak, and teachers know it.
It teaches evaluation before generation. If a teacher can reliably identify why an output is inadequate, the generation problem largely solves itself. The reverse is not true.
It makes rejection normal. The single most useful behaviour change is teachers feeling licensed to say a draft is not good enough. That is a cultural intervention as much as a technical one, and it usually needs a school leader to model it visibly.
It ends with something reusable. A session that produces one working setup holding the teacher's own standards outlasts a session that produces notes.
What school leaders can do without technical knowledge
Three actions require no understanding of how a language model works.
Ask to see the prompt, not just the output. When a teacher shows you AI generated material, ask what she requested and what she changed. The answer reveals whether judgment was applied.
Reject weak AI output publicly. If staff see leadership decline AI generated work, they learn they are permitted to.
Hold the standard that already existed. Not a special AI standard, the same one. "Would this have been good enough if a teacher had written it by hand?" is sufficient as a complete test.
A way to measure where you are
Before commissioning further training, ask staff to score themselves on six questions:
Do you define the specific outcome before prompting, or just the topic?
Do you ask AI what you might be missing, before asking it to produce?
Does your instruction carry your constraints, format and range of learners?
Do you evaluate output against criteria set in advance?
When it is wrong, do you give numbered corrections rather than "make it better"?
Do you have a rule for when something is finished?
Most experienced teachers score two. Not because the judgment is absent, but because nobody indicated it applied here.
Those six questions are not arbitrary. They are the first six steps of the IMPACTS framework, the method I teach through Credo Learnings, phrased as a diagnostic. Identify, Mine, Prompt, Assess, Calibrate, Tune. A seventh step, Systematise, covers what happens after a teacher gets the first six right and wants to stop rebuilding the same thing each term.
That result tells you what to commission next, with more precision than any vendor conversation.