Watching how a learner reasons through a problem rather than only whether the answer came out right, and acting on what that shows.
The instrument exists and runs daily. It is the maths run in NeuroCraft, where a timer and a cell-by-cell decomposition make the interior of a solution observable. This document sets out what that instrument captures, which parts of NeuroCraft qualify and which do not, and what has to be built next.
Students arriving in engineering classes are increasingly able to produce an answer and increasingly unable to say what operation the answer required. The two capacities have come apart. A student who can obtain a result without being able to name the step that produced it cannot check that result, cannot repair it when it breaks, and cannot recognise when a machine has produced something plausible and wrong.
The AI-in-research work in the adjacent programme measured a version of this directly. A 1.5-billion-parameter model deciding a bursary case end to end scored 55.6 per cent. The same model, used to extract features which a two-parameter decision rule then applied, scored 93.3 per cent. The gain came from naming the operation and handing it to the method that suits it, and the model never reached for the smaller method unless a person integrated it. The reasoning instruction set generalises that finding into a taxonomy of operations, each paired with the classical algorithm that discharges it, with the language model repositioned as orchestrator rather than authority.
The same argument applies to teaching, and it is more urgent there. If competence in a graduate is the ability to name the operation, select the method, and certify the result, then teaching has to be able to observe those three things separately. A final mark observes none of them. It records that something worked or did not, and it is silent on which step failed, and therefore silent on whether the failure is the student's or the course's.
That gap is the reason for the programme. The deployment horizon matters as well: the silicon refresh runs on roughly a five-year cycle and an academic generation on roughly ten, so the students now in first year pass through the system once. Instrumentation that arrives after they graduate arrives for someone else.
The framework will eventually touch the wider estate, and planning as though that were imminent would misrepresent the state of things. For a long time yet this surfaces in one place, which is NeuroCraft, and within NeuroCraft in one surface, which is the maths run.
The maths run is the only place where the two conditions hold together. A problem is decomposed into cells that each carry a declared expectation, and a clock plus an assistance ladder record what it cost the learner to satisfy each one. Everything else in NeuroCraft assesses the answer. Assessing the answer is useful and it is not what this programme is about.
Concentrating on one surface is also the right development order. The maths run is where the capture format, the diagnostic reading and the adult-facing view can be worked out against a real learner at high frequency. Extending a format that has been proven on one surface is ordinary work; extending one that has not been is speculation multiplied by six.
Twenty step-decomposed question classes cover written arithmetic, equivalent fractions, addition over a lowest common denominator, decimals, mixed numbers, long division and word problems, presented in three panes of a single run. Each class declares its cells, the expression each cell must satisfy, and a feedback line naming the misconception that produces the wrong value. Problems are posed by a seeded constraint solver rather than authored one at a time, and the constraints force the teaching moment deliberately: a carry that must happen, a borrow, a zero that must be placed, a denominator that has to be inferred rather than read.
Guided mode then walks the algorithm cell by cell. Every entry is judged as it is typed and recorded with its timing. The assistance ladder prices each answer at five points, falling to three after a first miss and one after a second, where the first hint states the observation, the second states the instruction, and neither reveals the number. A timeout counts as a miss. Per-attempt telemetry records credit banked against credit available, misses separated into wrong entries and expiries, and misses attributed to the individual cell.
The consequence is that a wrong answer resolves into a location and a cost. A learner who reaches the right result on e1d unaided and needs both hints every time on e2d is telling the system something specific, which is that scaling a fraction is understood when the multiplier is given and not understood when it has to be inferred. No mark carries that.
| Capability | State | Mechanism | Role in this programme |
|---|---|---|---|
| Per-child identity | LIVE | Email-less child accounts under a guardian; team-around-the-child scoping | Every attempt is attributable without profiling children to a third party. |
| Attempt records | LIVE | Durable per-learner rows through the learning service | Cross-device truth from which mastery and gaps are computed. |
| Mastery and decay | LIVE | Three-correct streak to secure, seven-day decay, review woven back into the plan | The short return of the loop, at skill granularity. |
| Difficulty laddering | LIVE | Three constraint levels per generator, selected by mastery state, with carer override | Adaptation acts on the generator rather than on a bank of fixed items. |
| Session pulse | LIVE | Thirty-second beat while the tab is visible; focus gaps by construction | Separates disengagement from difficulty, which matters greatly for an ADHD learner. |
| Prescription and nudge | LIVE | Adult sets today's plan; the learner's app opens it within one beat | Human judgement enters the loop without leaving the instrument. |
| Gap aggregation | LIVE | Credit share, hint load, timeout load and worst cells per skill | First diagnostic output. Currently per learner. |
| Curriculum coverage | LIVE | CAPS topic map per grade and subject against what is planned | Anchors components to a public curriculum rather than to private labels. |
| Opportunity-indexed log | REQUIRED | Ordered per-cell transactions carrying a practice count | Prerequisite for the long return. Cannot be reconstructed later. |
| Population curves | REQUIRED | Error rate against opportunity, per component, across learners | The long return itself. |
The honest summary is that the forward path of the loop and the short return are built and working, and the long return is not yet possible because of one gap in how the record is stored.
NeuroCraft carries several assessment approaches that were built for different reasons and at different times. Only one of them unpacks the interior of a solution. The rest are sound as practice and weak as instrumentation, and the distinction is worth stating precisely rather than being blurred by the fact that all of them produce attempt records.
| Surface | Decomposition | What is recorded | Diagnostic reach |
|---|---|---|---|
| Maths run 20 grid classes |
PER CELL | Every cell entry, timed, with hint level, credit banked against available, and misses split into wrong entries and expiries | Names the step that failed and what it cost. The only surface that supports the framework as it stands. |
| Meta-question scaffolds 5 of 39 definitions |
SUB-QUESTIONS | One answer per sub-question, with timing | Locates failure to a stage of a word problem. Shallower than a cell, because the sub-question is itself answered in one move. |
| Value explorers 15 definitions |
CRITERIA | Pass or fail against declared criteria on a live model | The criteria are already named and separable. Which criterion failed is computed and then discarded rather than stored. |
| Multiple choice 39 definitions |
NONE | Submitted answer string and correctness | Answer level. Every distractor already carries a named misconception, which the record does not yet resolve. |
| Matching, ordering, yes-or-no, fill-in | NONE | Submitted answer and correctness | Answer level throughout. |
| Reasoning probes 15 definitions |
NONE | Choice of explanation after a correct answer | Asks the right question, which is whether the learner can say why. Stores it as one more right-or-wrong. |
| Lesson and webbook quizzes | NONE | Correctness, and in one path not even the submitted answer | Answer level. Suited to checking that a chapter was read. |
| Speed run | NONE | Answer and response time against a personal baseline | Measures fluency, deliberately. Fluency is a real construct and it is not understanding. |
Two of these could move up without redesigning anything, which makes them worth doing before any new surface is built.
The multiple-choice definitions already attach a named misconception to each distractor, explaining exactly what wrong procedure produces that option. The submitted answer is already stored. Resolving the stored answer back to the distractor that produced it would convert thirty-nine definitions from right-or-wrong into misconception frequencies across a cohort, which is a genuine diagnostic reading obtained from data already on disk. This is the cheapest real gain available in the system.
The explorers already evaluate declared criteria and already know which one failed. Storing the failing criterion rather than only the verdict costs one field.
Neither upgrade makes those surfaces equivalent to the maths run, because neither observes a learner mid-procedure. Both would make them contribute to the same diagnostic picture instead of only to a score.
The first learner does not look forward to the system and has to be pressured into using it. That observation outranks every measurement result in this document, because a pressured learner produces avoidance data rather than learning data, and an instrument nobody opens measures nothing.
The measurement half was built carefully and the reasons to return were not. What exists is an assessment instrument with reward tokens attached, and the faults are specific rather than a general shortage of polish.
| Element | State | Fault, and what replaced it |
|---|---|---|
| Clock | CORRECTED 3 AUG | Was a bar that drained, reddened, pulsed an alarm and ended in a failure chime, which is a bomb rather than suspense. Now a bonus window: it cools green to amber, lapses on a level two-note settle, and its expiry pays nothing extra while taking nothing away. |
| Economy | CORRECTED 3 AUG | Was a worth of five visibly reduced to three then one, with a red minus flying off the total. Now every correct entry pays a base point, a streak multiplier rides on top inside the window, and a miss resets the streak without touching anything banked. Loss aversion is removed by having nothing to lose. |
| Reward schedule | CORRECTED 3 AUG | Was fixed and perfectly predictable. Now carries an occasional unannounced bonus, because predictable reward is the weakest schedule available and more so where the dopamine response tracks prediction error. |
| Persistence | OPEN | A run ends with a congratulation and a points total that resets, so nothing survives the session. Wanted: something that accumulates and is worth returning to, such as a collection, a world, or a character that grows on the strength of work done. |
| Agency | OPEN | A fixed ladder of five question types, two each. Stages can be skipped and nothing meaningful is chosen. Wanted: real choices with real consequences, including what to work on and what to risk. |
| Voice | OPEN | Every line is earnest and instructional, with no character, no humour and no irony. Wanted: a companion with attitude, which a child will work to impress or to provoke. Irony also lets a correction land without sounding like one. |
| The adult | OPEN | Appears as prescriptions, nudges and a gap report, which is presence as surveillance. Wanted: presence as company. Something done together, or something the learner can show, beats something the learner is watched doing. |
Two further points are structural rather than cosmetic. Cell-by-cell written arithmetic is the most laborious part of school mathematics; it is the right thing to instrument and the wrong thing to make the whole experience. And every interaction in the current run is evaluated, so there is no unjudged space anywhere in the system, which is difficult for any learner and harder for an inattentive one.
Adding points to an activity a learner already enjoys can reduce the enjoyment once the points stop, which is the overjustification effect. The correction is to make the activity itself the attraction and to let the scoring stay quiet, rather than to increase the reward when engagement falls.
Cell-by-cell decomposition is one way to make reasoning observable and it should not be the only one. The candidates below are chosen against three conditions at once: each teaches something the current run does not, each is more enjoyable than filling in a grid, and each still exposes the learner's process rather than only an answer. The third condition is what keeps them inside this programme rather than being decoration.
| Method | What it teaches | Why a learner returns to it | What it instruments |
|---|---|---|---|
| Error hunt BUILT 3 AUG |
Recognising a wrong procedure, which is a different capacity from executing a right one | Catching someone out is inherently satisfying, and the learner judges rather than performs | Which misconceptions are recognised and which pass unnoticed |
| Estimation brackets | Number sense before procedure: nearer ten or nearer a hundred | Rounds last seconds, carry no procedure, and are hard to fail badly | The bracket chosen, which exposes magnitude reasoning directly |
| Build the question | Working backwards from a result to a problem that produces it | Many answers are right, so it cannot be failed in the ordinary way | The construction strategy chosen, which is richer than a correct answer |
| Predict then check | Committing to an expectation before evidence arrives | Being right about a prediction is more satisfying than being right about a sum | Prediction error, which is both the learning moment and the datum |
| Teach the character | Naming why a step works, in order to explain it to someone else | Helping rather than being tested, with a character who visibly improves | Whether the learner can name the operation, which is the capacity the programme is actually about |
| Race the ghost | Nothing new; it re-frames practice already being done | Competing against a replay of one's own earlier run rather than a clock | Pace against a personal baseline, adjusting by construction rather than by rule |
A worked solution is presented containing exactly one mistake and the learner taps the line that breaks. It scored highest against all three conditions and cost the least to build, so it went first.
It is a mode of the existing question classes rather than a new one. Setting hunt on any of the twenty grid types renders its posed problem as a completed working, with the entries filled in and one value spoilt by a transform drawn from a registry of plausible slips: out by one, out by ten, doubled, halved, digits swapped. Lines become the answer, and scoring judges which line was chosen rather than which digits were typed. No new content is authored, because the correct working is exactly what the generator already computes.
Two properties matter more than they look. The spoilt value is chosen only from cells the pose actually displays, since several classes declare more cells than any one pose shows, and spoiling a hidden one would produce an unanswerable question. The fill-in step labels are also suppressed, because they name which part of a line was already worked out and would point straight at the spoilt cell.
The lowest-common-denominator variant produces the sharpest version of the task. A learner sees a stated denominator of forty while every line beneath it correctly uses twenty, so the mistake is found by noticing that a line does not agree with the rest of the working. That is internal-consistency reasoning, and it is much closer to what checking a machine's output requires than any single execution step.
The clock, the badges and the adaptive pacing exist and are pointed the wrong way. The clock should set a bonus window rather than a deadline, so the personal baseline that already drives it becomes a source of upside instead of pressure. Badges should be collectible rather than earned, with some rare and some ironic, since a badge that arrives unexpectedly is worth more than one whose conditions were published in advance. The assistance ladder should keep the price visible while ceasing to subtract, because the diagnostic signal comes from what help was taken rather than from the deduction.
Every component of this approach has a literature behind it, and the programme is stronger for saying so precisely rather than claiming invention.
| Thread | Origin | What it supplies |
|---|---|---|
| Model tracing | Anderson, ACT-R Cognitive Tutors; Carnegie Learning | Following a solution path against a rule model of the skill. Establishes the step as the level at which tutoring produces its effect. |
| Knowledge tracing | Corbett & Anderson 1995; Piech et al. 2015 | Estimating acquisition of each component, updated on every observation. NeuroCraft's mastery and decay rules are a simple instance. |
| Assistance-based scoring | Feng, Heffernan & Koedinger 2009 | Hints and attempts predict achievement better than binary correctness. Justifies the credit ladder as measurement rather than decoration. |
| Component model discovery | Cen, Koedinger & Junker 2006 | Fitting curves per component to find where the decomposition itself is wrong. The formal method behind the long return. |
Two results are worth carrying explicitly. VanLehn's 2011 review found that systems intervening at the step level produced effect sizes close to human tutoring, of the order of d = 0.76 against 0.79, while systems responding only to final answers produced roughly d = 0.31. Separately, Koedinger's group has repeatedly taken a component whose curve refused to flatten, split it on the evidence, redesigned the teaching around the split, and measured improved learning. The teaching-gap claim therefore has an established method and a demonstrated payoff, and this programme should adopt that method rather than reinvent it.
One terminology point. The phrase "deep learning monitoring" collides with Deep Knowledge Tracing, which denotes the neural-network family inside this exact field, and with deep learning in the Marton and Säljö sense of an approach to study. Either reader would misfile the work. Step-level diagnostic learning is the working title; process-level learning analytics suits institutional audiences; instrumented curriculum suits the teaching-gap half.
Two features of the maths run are prior art and should be presented as such: a hint ladder with declining credit has equivalents in ALEKS and ASSISTments, and misconception-named feedback descends from Brown and Burton's work on procedural bugs. Three appear to be less well covered, and the third of them is a direction rather than something already demonstrated.
In the standard pipeline, items are authored, a matrix mapping items to components is built by hand, curve analysis later shows the mapping to be wrong, and the items are re-authored. Building that mapping is the acknowledged bottleneck of the field, to the point that 2025 work applies language models to cluster items into components after the fact.
In NeuroCraft the mapping requires no discovery. A question class such as uiEquivFractions declares a cell e2d whose criterion states that the numerator moved from a to a·k, so the multiplier is k and it applies to the denominator. The label, its semantics, its misconception line and its difficulty ladder are created with the item. When population data names e2d as the point where learning stalls, the finding points at an editable generator, so the discovery-to-redesign loop can run continuously instead of once per research project.
Prior art: Learning Factors Analysis performs the discovery. KCluster and related 2025 work attacks the same bottleneck post hoc.
This one is a direction rather than a result, and it is recorded here so the capture format does not foreclose it. Components in the literature are almost always domain-specific skills, such as converting a fraction or balancing an equation. The adjacent reasoning instruction set defines twenty-five operations in ten families, each paired with the classical algorithm that discharges it and each carrying a precondition and a certificate. Those operations are candidate components at a level above the domain, and they are the level the motivation section cares about.
The arithmetic cell records that a learner cannot apply a multiplier to a denominator. An operation label would record that the same learner cannot recognise when a constraint must be propagated, which is the part that would transfer to reactor design. Whether the operation layer can be made observable at all is an open question, and it stays downstream of proving the domain layer on one surface.
Prior art: component models are domain-specific by convention. Transfer across domains is well studied; instrumenting a shared operation vocabulary underneath two domains appears uncommon.
The finest grain in DataShop, the field's main step-level repository, is the step. A cell within a written algorithm sits below that and separates a positional error from an arithmetic one. Separately, assistance scores are conventionally computed after the fact and stay invisible during the task; showing the price before the learner chooses turns hint-taking into an elicited confidence judgement in addition to a behavioural trace.
Prior art: step-level logging is standard, cell-level is not. Certainty-based marking is the nearest relative of visible pricing, without step-level tracing.
Confirmed rather than anticipated. The first learner has to be pressured into opening the system, which makes voluntary use the binding constraint on the whole programme and every accuracy figure conditional on it. Voluntary opening should be tracked as an outcome in its own right, and no measurement result should be reported without it.
The help-seeking literature treats hint avoidance as a maladaptive behaviour on equal footing with hint abuse, and a visible price selects for avoidance in anxious or perfectionist learners. The signature is long silences ending in timeouts rather than early hint use. If it appears, hold the price above zero or price both hints equally, which keeps the signal and removes the penalty gradient.
Written algorithms decompose cleanly into cells, which is why the build started there. Reasoning operations do not, and the risk is that the system measures procedural fluency accurately while the capacity the motivation section is actually about goes unmeasured. W3 exists to confront this, and if the operation vocabulary cannot be made observable, the programme should say so rather than substitute the easier measurement.
Curve fitting requires a population. The N = 1 site produces design signal and affective constraints and cannot produce a teaching-gap finding. Single-case methodology should govern any claim made from it, and the cohort sites need to come online early.
Minor learners and university students require clearance and informed consent before records are used for research, and clinical records in NeuroPlay stay outside this programme's data boundary entirely. Clearance should be in place before the cohort sites begin recording for research purposes.
The spine consists of three things: content that declares its steps, a record that keeps every attempt against those steps, and a diagnosis that separates the learner's gap from the teaching's. Any surface carrying content can in principle carry it, and several in the estate eventually should.
Undergraduate reaction engineering is the most interesting of them, because the decomposition is already explicit in how the subject is taught. Sizing a reactor proceeds through the mole balance, the rate law, the stoichiometric relation between concentration and conversion, combination, evaluation and a dimensional check. A failed sizing problem currently scores as one event when it could be any of those, and separating them across a cohort would say something about the course rather than about the students. The ScholarCloud class anchor would carry the telemetry, Publon Press would need exercises that declare steps at authoring time rather than in code, and a credential backed by step evidence would assert which operations a graduate demonstrated and under how much assistance. NeuroPlay shares the guardian and team-around-the-child spine and already observes process rather than outcome, so the method transfers even though its clinical records stay behind their own boundary.
None of that is scheduled, and it should not be until the reading works on one surface with one learner. The order matters more than the ambition.
bash build-local.sh · effect sizes in the theory section are quoted approximately from VanLehn (2011) and should be checked against the source before external use.