Back to blog
Implementation

How to Design a Six-Week Pilot for a New AI Tool in K-12

Project planning timeline chart on paper with pencil

The most common failure mode for AI tool pilots in K-12 schools is not that the tool does not work. It is that the pilot was designed in a way that makes it impossible to know whether the tool worked. An administrator approves a three-month open-ended trial, a handful of teachers try it inconsistently, there is no pre-measurement and no defined success criteria, and at the end of the semester the question "should we expand this?" has no data to draw on.

A well-designed six-week pilot is more useful than a three-month open-ended trial in most cases. It has enough duration to get past the initial learning curve but short enough that teacher attention does not drift. It is tight enough to run with controlled conditions. And six weeks maps cleanly onto a curriculum unit in most K-12 schedules, which makes it possible to design the pilot around a specific instructional context with measurable outcomes.

Before week one: the three decisions that determine everything

Before the pilot starts, three decisions need to be made explicitly. First, what is the specific question the pilot is trying to answer? "Does this tool work?" is not a specific question. "Does this tool help teachers identify students with prerequisite gaps in fraction operations before the algebra unit begins, and does acting on that information reduce the rate of foundational errors in the first two weeks of algebra instruction?" is a specific question.

Second, who will champion the pilot? A successful pilot in a school typically runs through one or two teachers who have both the conviction to try something new and the professional standing to report results credibly to their colleagues. A pilot that is mandated from above without a committed classroom teacher is unlikely to generate useful implementation data, because the teachers will comply minimally rather than engage seriously.

Third, what does a positive result look like, and what does it need to look like to justify expansion? Defining success criteria before the pilot starts is essential because school environments are noisy -- lots of things change during a six-week period -- and it is easy to selectively interpret ambiguous results as confirming whatever you wanted to believe. If you agreed upfront that the pilot succeeds if three or more teachers report meaningfully changing their instructional approach based on the tool's output, you have a clear bar to evaluate against.

Week one: setup and baseline measurement

The first week should be almost entirely setup and baseline measurement. Do not expect teachers to use the tool and generate usable results in week one. The IT integration may still be settling, teachers are learning the interface, and students need to understand what they are being asked to do. Trying to rush pedagogical application into the first week creates implementation noise that confuses the data.

Baseline measurement in week one means collecting the pre-pilot data you will compare against at the end. For a diagnostic-focused tool, this might mean: recording what the teacher's current assessment of each student's readiness is (their subjective impression), and optionally running a traditional pre-assessment of the relevant prerequisite concepts. This baseline lets you compare the tool's gap detection output to the teacher's existing knowledge, which is one of the most useful comparisons a pilot can generate.

Weeks two through five: structured implementation

The core of the pilot should be structured enough to generate consistent data but not so scripted that it prevents teachers from exercising judgment about how to use the tool's output. A useful structure: teacher runs the diagnostic sequence with students in weeks two and three, reviews the gap summary, and implements one specific instructional response (a targeted small-group session, a modified warm-up sequence, or a different practice set assignment) for students flagged as having gaps. Weeks four and five, the unit advances and the teacher notes whether flagged students show the expected difficulty or not.

Keeping a simple log matters more than most teachers expect. "Student X was flagged for fraction division gap on Feb 12. I ran a targeted 10-minute group session on Feb 14. Student X's error rate on fraction-dependent problems in the next unit was [lower / about the same / higher] than expected." This is not rigorous research data, but it is the kind of documentation that makes the end-of-pilot debrief actually informative rather than impressionistic.

Week six: debrief and decision

The pilot debrief should happen in week six, while the experience is fresh and before the next curriculum unit creates new demands on teacher attention. The debrief should cover: what did the teacher do differently because of the tool's output (as distinguished from what they would have done anyway), were the gaps the tool flagged consistent with their own observations, and what would need to change for them to use this tool as a regular part of their practice?

The debrief output should feed directly into a decision: expand, pause, or stop. Expand means commit to broader rollout with a defined plan. Pause means there is interest but unresolved issues (usually technical integration or workload) that need to be addressed before expansion. Stop means either the tool did not demonstrate the value that justified the investment or the fit with this school's context was poor -- both legitimate outcomes that are more useful than an inconclusive open-ended trial.

A pilot that ends with a clear decision, even if that decision is "not the right tool for us," was a well-designed pilot. A pilot that ends with "let's keep exploring" without any clear next step has not generated the decision-forcing information that makes a pilot worth running.

See gap detection in action

Join the early-access program and run a pilot with your classroom or program.