Evaluating the Steps in Understanding
NeuroCraft · step-level diagnostic learning

Evaluating the Steps in Understanding

Watching how a learner reasons through a problem rather than only whether the answer came out right, and acting on what that shows.

The instrument exists and runs daily. It is the maths run in NeuroCraft, where a timer and a cell-by-cell decomposition make the interior of a solution observable. This document sets out what that instrument captures, which parts of NeuroCraft qualify and which do not, and what has to be built next.

Opened 2 August 2026
Lead R. Rawatlal
Site NeuroCraft, maths run
Status Forward path live · long return blocked on capture
MOTIVATIONWhy this is being built, ahead of whether it publishes.

Reason for the work

Students arriving in engineering classes are increasingly able to produce an answer and increasingly unable to say what operation the answer required. The two capacities have come apart. A student who can obtain a result without being able to name the step that produced it cannot check that result, cannot repair it when it breaks, and cannot recognise when a machine has produced something plausible and wrong.

The AI-in-research work in the adjacent programme measured a version of this directly. A 1.5-billion-parameter model deciding a bursary case end to end scored 55.6 per cent. The same model, used to extract features which a two-parameter decision rule then applied, scored 93.3 per cent. The gain came from naming the operation and handing it to the method that suits it, and the model never reached for the smaller method unless a person integrated it. The reasoning instruction set generalises that finding into a taxonomy of operations, each paired with the classical algorithm that discharges it, with the language model repositioned as orchestrator rather than authority.

The same argument applies to teaching, and it is more urgent there. If competence in a graduate is the ability to name the operation, select the method, and certify the result, then teaching has to be able to observe those three things separately. A final mark observes none of them. It records that something worked or did not, and it is silent on which step failed, and therefore silent on whether the failure is the student's or the course's.

A mark reports that a student failed. It does not report what they cannot do, and it never reports what the course failed to teach.

That gap is the reason for the programme. The deployment horizon matters as well: the silicon refresh runs on roughly a five-year cycle and an academic generation on roughly ten, so the students now in first year pass through the system once. Instrumentation that arrives after they graduate arrives for someone else.

SCOPEOne site, one surface, for a good while yet.

Where the work sits

The framework will eventually touch the wider estate, and planning as though that were imminent would misrepresent the state of things. For a long time yet this surfaces in one place, which is NeuroCraft, and within NeuroCraft in one surface, which is the maths run.

The maths run is the only place where the two conditions hold together. A problem is decomposed into cells that each carry a declared expectation, and a clock plus an assistance ladder record what it cost the learner to satisfy each one. Everything else in NeuroCraft assesses the answer. Assessing the answer is useful and it is not what this programme is about.

Concentrating on one surface is also the right development order. The maths run is where the capture format, the diagnostic reading and the adult-facing view can be worked out against a real learner at high frequency. Extending a format that has been proven on one surface is ordinary work; extending one that has not been is speculation multiplied by six.

Content generator declares each step Learner work every entry, timed, with its hint history Record per learner, per component Diagnosis learner gap or teaching gap prescription · what this learner does next revision · what every learner is taught next
Closed loop, two returns. The forward path is instrumentation. The value sits in the two returns, which carry the same signal to different places. The short return adapts what one learner practises next. The long return changes the content itself, and it only becomes readable once enough learners have crossed the same step. Systems that carry only the short return are adaptive practice. Carrying both makes the curriculum an object under measurement.
INSTRUMENTThe maths run. Live and carrying real attempts.

What the maths run captures

Twenty step-decomposed question classes cover written arithmetic, equivalent fractions, addition over a lowest common denominator, decimals, mixed numbers, long division and word problems, presented in three panes of a single run. Each class declares its cells, the expression each cell must satisfy, and a feedback line naming the misconception that produces the wrong value. Problems are posed by a seeded constraint solver rather than authored one at a time, and the constraints force the teaching moment deliberately: a carry that must happen, a borrow, a zero that must be placed, a denominator that has to be inferred rather than read.

Guided mode then walks the algorithm cell by cell. Every entry is judged as it is typed and recorded with its timing. The assistance ladder prices each answer at five points, falling to three after a first miss and one after a second, where the first hint states the observation, the second states the instruction, and neither reveals the number. A timeout counts as a miss. Per-attempt telemetry records credit banked against credit available, misses separated into wrong entries and expiries, and misses attributed to the individual cell.

The consequence is that a wrong answer resolves into a location and a cost. A learner who reaches the right result on e1d unaided and needs both hints every time on e2d is telling the system something specific, which is that scaling a fraction is understood when the multiplier is given and not understood when it has to be inferred. No mark carries that.

Records, adaptation and the adult view

CapabilityStateMechanismRole in this programme
Per-child identityLIVEEmail-less child accounts under a guardian; team-around-the-child scopingEvery attempt is attributable without profiling children to a third party.
Attempt recordsLIVEDurable per-learner rows through the learning serviceCross-device truth from which mastery and gaps are computed.
Mastery and decayLIVEThree-correct streak to secure, seven-day decay, review woven back into the planThe short return of the loop, at skill granularity.
Difficulty ladderingLIVEThree constraint levels per generator, selected by mastery state, with carer overrideAdaptation acts on the generator rather than on a bank of fixed items.
Session pulseLIVEThirty-second beat while the tab is visible; focus gaps by constructionSeparates disengagement from difficulty, which matters greatly for an ADHD learner.
Prescription and nudgeLIVEAdult sets today's plan; the learner's app opens it within one beatHuman judgement enters the loop without leaving the instrument.
Gap aggregationLIVECredit share, hint load, timeout load and worst cells per skillFirst diagnostic output. Currently per learner.
Curriculum coverageLIVECAPS topic map per grade and subject against what is plannedAnchors components to a public curriculum rather than to private labels.
Opportunity-indexed logREQUIREDOrdered per-cell transactions carrying a practice countPrerequisite for the long return. Cannot be reconstructed later.
Population curvesREQUIREDError rate against opportunity, per component, across learnersThe long return itself.

The honest summary is that the forward path of the loop and the short return are built and working, and the long return is not yet possible because of one gap in how the record is stored.

AUDITWhich NeuroCraft surfaces evaluate steps, and which evaluate answers.

Assessment surfaces measured against the framework

NeuroCraft carries several assessment approaches that were built for different reasons and at different times. Only one of them unpacks the interior of a solution. The rest are sound as practice and weak as instrumentation, and the distinction is worth stating precisely rather than being blurred by the fact that all of them produce attempt records.

SurfaceDecompositionWhat is recordedDiagnostic reach
Maths run
20 grid classes
PER CELL Every cell entry, timed, with hint level, credit banked against available, and misses split into wrong entries and expiries Names the step that failed and what it cost. The only surface that supports the framework as it stands.
Meta-question scaffolds
5 of 39 definitions
SUB-QUESTIONS One answer per sub-question, with timing Locates failure to a stage of a word problem. Shallower than a cell, because the sub-question is itself answered in one move.
Value explorers
15 definitions
CRITERIA Pass or fail against declared criteria on a live model The criteria are already named and separable. Which criterion failed is computed and then discarded rather than stored.
Multiple choice
39 definitions
NONE Submitted answer string and correctness Answer level. Every distractor already carries a named misconception, which the record does not yet resolve.
Matching, ordering, yes-or-no, fill-in NONE Submitted answer and correctness Answer level throughout.
Reasoning probes
15 definitions
NONE Choice of explanation after a correct answer Asks the right question, which is whether the learner can say why. Stores it as one more right-or-wrong.
Lesson and webbook quizzes NONE Correctness, and in one path not even the submitted answer Answer level. Suited to checking that a chapter was read.
Speed run NONE Answer and response time against a personal baseline Measures fluency, deliberately. Fluency is a real construct and it is not understanding.

Raising the second tier cheaply

Two of these could move up without redesigning anything, which makes them worth doing before any new surface is built.

The multiple-choice definitions already attach a named misconception to each distractor, explaining exactly what wrong procedure produces that option. The submitted answer is already stored. Resolving the stored answer back to the distractor that produced it would convert thirty-nine definitions from right-or-wrong into misconception frequencies across a cohort, which is a genuine diagnostic reading obtained from data already on disk. This is the cheapest real gain available in the system.

The explorers already evaluate declared criteria and already know which one failed. Storing the failing criterion rather than only the verdict costs one field.

Neither upgrade makes those surfaces equivalent to the maths run, because neither observes a learner mid-procedure. Both would make them contribute to the same diagnostic picture instead of only to a score.

FINDINGThe instrument works. The learner avoids it.

Voluntary use as a precondition

The first learner does not look forward to the system and has to be pressured into using it. That observation outranks every measurement result in this document, because a pressured learner produces avoidance data rather than learning data, and an instrument nobody opens measures nothing.

The measurement half was built carefully and the reasons to return were not. What exists is an assessment instrument with reward tokens attached, and the faults are specific rather than a general shortage of polish.

Faults in the present design

ElementStateFault, and what replaced it
Clock CORRECTED 3 AUG Was a bar that drained, reddened, pulsed an alarm and ended in a failure chime, which is a bomb rather than suspense. Now a bonus window: it cools green to amber, lapses on a level two-note settle, and its expiry pays nothing extra while taking nothing away.
Economy CORRECTED 3 AUG Was a worth of five visibly reduced to three then one, with a red minus flying off the total. Now every correct entry pays a base point, a streak multiplier rides on top inside the window, and a miss resets the streak without touching anything banked. Loss aversion is removed by having nothing to lose.
Reward schedule CORRECTED 3 AUG Was fixed and perfectly predictable. Now carries an occasional unannounced bonus, because predictable reward is the weakest schedule available and more so where the dopamine response tracks prediction error.
Persistence OPEN A run ends with a congratulation and a points total that resets, so nothing survives the session. Wanted: something that accumulates and is worth returning to, such as a collection, a world, or a character that grows on the strength of work done.
Agency OPEN A fixed ladder of five question types, two each. Stages can be skipped and nothing meaningful is chosen. Wanted: real choices with real consequences, including what to work on and what to risk.
Voice OPEN Every line is earnest and instructional, with no character, no humour and no irony. Wanted: a companion with attitude, which a child will work to impress or to provoke. Irony also lets a correction land without sounding like one.
The adult OPEN Appears as prescriptions, nudges and a gap report, which is presence as surveillance. Wanted: presence as company. Something done together, or something the learner can show, beats something the learner is watched doing.

Two further points are structural rather than cosmetic. Cell-by-cell written arithmetic is the most laborious part of school mathematics; it is the right thing to instrument and the wrong thing to make the whole experience. And every interaction in the current run is evaluated, so there is no unjudged space anywhere in the system, which is difficult for any learner and harder for an inattentive one.

Extrinsic reward can displace interest

Adding points to an activity a learner already enjoys can reduce the enjoyment once the points stop, which is the overjustification effect. The correction is to make the activity itself the attraction and to let the scoring stay quiet, rather than to increase the reward when engagement falls.

METHODSSix candidates. Cells are one method, not the method.

Teaching methods under consideration

Cell-by-cell decomposition is one way to make reasoning observable and it should not be the only one. The candidates below are chosen against three conditions at once: each teaches something the current run does not, each is more enjoyable than filling in a grid, and each still exposes the learner's process rather than only an answer. The third condition is what keeps them inside this programme rather than being decoration.

MethodWhat it teachesWhy a learner returns to itWhat it instruments
Error hunt
BUILT 3 AUG
Recognising a wrong procedure, which is a different capacity from executing a right one Catching someone out is inherently satisfying, and the learner judges rather than performs Which misconceptions are recognised and which pass unnoticed
Estimation brackets Number sense before procedure: nearer ten or nearer a hundred Rounds last seconds, carry no procedure, and are hard to fail badly The bracket chosen, which exposes magnitude reasoning directly
Build the question Working backwards from a result to a problem that produces it Many answers are right, so it cannot be failed in the ordinary way The construction strategy chosen, which is richer than a correct answer
Predict then check Committing to an expectation before evidence arrives Being right about a prediction is more satisfying than being right about a sum Prediction error, which is both the learning moment and the datum
Teach the character Naming why a step works, in order to explain it to someone else Helping rather than being tested, with a character who visibly improves Whether the learner can name the operation, which is the capacity the programme is actually about
Race the ghost Nothing new; it re-frames practice already being done Competing against a replay of one's own earlier run rather than a clock Pace against a personal baseline, adjusting by construction rather than by rule

Error hunt, as built

A worked solution is presented containing exactly one mistake and the learner taps the line that breaks. It scored highest against all three conditions and cost the least to build, so it went first.

It is a mode of the existing question classes rather than a new one. Setting hunt on any of the twenty grid types renders its posed problem as a completed working, with the entries filled in and one value spoilt by a transform drawn from a registry of plausible slips: out by one, out by ten, doubled, halved, digits swapped. Lines become the answer, and scoring judges which line was chosen rather than which digits were typed. No new content is authored, because the correct working is exactly what the generator already computes.

Two properties matter more than they look. The spoilt value is chosen only from cells the pose actually displays, since several classes declare more cells than any one pose shows, and spoiling a hidden one would produce an unanswerable question. The fill-in step labels are also suppressed, because they name which part of a line was already worked out and would point straight at the spoilt cell.

The lowest-common-denominator variant produces the sharpest version of the task. A learner sees a stated denominator of forty while every line beneath it correctly uses twenty, so the mistake is found by noticing that a line does not agree with the rest of the working. That is internal-consistency reasoning, and it is much closer to what checking a machine's output requires than any single execution step.

Reworking the elements already present

The clock, the badges and the adaptive pacing exist and are pointed the wrong way. The clock should set a bonus window rather than a deadline, so the personal baseline that already drives it becomes a source of upside instead of pressure. Badges should be collectible rather than earned, with some rare and some ironic, since a badge that arrives unexpectedly is worth more than one whose conditions were published in advance. The assistance ladder should keep the price visible while ceasing to subtract, because the diagnostic signal comes from what help was taken rather than from the deduction.

0 0.2 0.4 0.6 1 4 8 12 practice opportunity on the knowledge component error rate teaching gap acquired
Reading the long return. Error rate against practice opportunity, per component, aggregated over learners. A component that is correctly specified and adequately taught decays towards zero. A component whose curve stays flat has not been learned by repetition, which means either the step decomposition is wrong or the instruction attached to it does not work. Both readings are actionable and neither is visible from marks.
THEORYFour established threads the build draws on.

Theory the build rests on

Every component of this approach has a literature behind it, and the programme is stronger for saying so precisely rather than claiming invention.

ThreadOriginWhat it supplies
Model tracingAnderson, ACT-R Cognitive Tutors; Carnegie LearningFollowing a solution path against a rule model of the skill. Establishes the step as the level at which tutoring produces its effect.
Knowledge tracingCorbett & Anderson 1995; Piech et al. 2015Estimating acquisition of each component, updated on every observation. NeuroCraft's mastery and decay rules are a simple instance.
Assistance-based scoringFeng, Heffernan & Koedinger 2009Hints and attempts predict achievement better than binary correctness. Justifies the credit ladder as measurement rather than decoration.
Component model discoveryCen, Koedinger & Junker 2006Fitting curves per component to find where the decomposition itself is wrong. The formal method behind the long return.

Two results are worth carrying explicitly. VanLehn's 2011 review found that systems intervening at the step level produced effect sizes close to human tutoring, of the order of d = 0.76 against 0.79, while systems responding only to final answers produced roughly d = 0.31. Separately, Koedinger's group has repeatedly taken a component whose curve refused to flatten, split it on the evidence, redesigned the teaching around the split, and measured improved learning. The teaching-gap claim therefore has an established method and a demonstrated payoff, and this programme should adopt that method rather than reinvent it.

One terminology point. The phrase "deep learning monitoring" collides with Deep Knowledge Tracing, which denotes the neural-network family inside this exact field, and with deep learning in the Marton and Säljö sense of an approach to study. Either reader would misfile the work. Step-level diagnostic learning is the working title; process-level learning analytics suits institutional audiences; instrumented curriculum suits the teaching-gap half.

CLAIMSSeparated from what is already established.

Distinctive elements

Two features of the maths run are prior art and should be presented as such: a hint ladder with declining credit has equivalents in ALEKS and ASSISTments, and misconception-named feedback descends from Brown and Burton's work on procedural bugs. Three appear to be less well covered, and the third of them is a direction rather than something already demonstrated.

Generative content carries its own component model

In the standard pipeline, items are authored, a matrix mapping items to components is built by hand, curve analysis later shows the mapping to be wrong, and the items are re-authored. Building that mapping is the acknowledged bottleneck of the field, to the point that 2025 work applies language models to cluster items into components after the fact.

In NeuroCraft the mapping requires no discovery. A question class such as uiEquivFractions declares a cell e2d whose criterion states that the numerator moved from a to a·k, so the multiplier is k and it applies to the denominator. The label, its semantics, its misconception line and its difficulty ladder are created with the item. When population data names e2d as the point where learning stalls, the finding points at an editable generator, so the discovery-to-redesign loop can run continuously instead of once per research project.

Prior art: Learning Factors Analysis performs the discovery. KCluster and related 2025 work attacks the same bottleneck post hoc.

Reasoning operations as a second component layer

This one is a direction rather than a result, and it is recorded here so the capture format does not foreclose it. Components in the literature are almost always domain-specific skills, such as converting a fraction or balancing an equation. The adjacent reasoning instruction set defines twenty-five operations in ten families, each paired with the classical algorithm that discharges it and each carrying a precondition and a certificate. Those operations are candidate components at a level above the domain, and they are the level the motivation section cares about.

The arithmetic cell records that a learner cannot apply a multiplier to a denominator. An operation label would record that the same learner cannot recognise when a constraint must be propagated, which is the part that would transfer to reactor design. Whether the operation layer can be made observable at all is an open question, and it stays downstream of proving the domain layer on one surface.

Prior art: component models are domain-specific by convention. Transfer across domains is well studied; instrumenting a shared operation vocabulary underneath two domains appears uncommon.

Sub-step forensics and visibly priced assistance

The finest grain in DataShop, the field's main step-level repository, is the step. A cell within a written algorithm sits below that and separates a positional error from an arithmetic one. Separately, assistance scores are conventionally computed after the fact and stay invisible during the task; showing the price before the learner chooses turns hint-taking into an elicited confidence judgement in addition to a behavioural trace.

Prior art: step-level logging is standard, cell-level is not. Certainty-based marking is the nearest relative of visible pricing, without step-level tracing.

SEQUENCEW1 is time-sensitive. The rest follow evidence.

Work packages

  1. Opportunity-indexed capture in the maths run. Record each cell attempt as a sequenced row carrying a count of how many times that learner has met that component. Substrate for everything below, and it cannot be retrofitted onto rolled-up records, so it comes first regardless of what else is running.
  2. Turn the clock and the economy around. DONE 3 AUG Bonus window in place of a deadline, a streak that builds in place of a worth that is deducted, and occasional unpredictable bonuses. Smallest change with the largest effect on whether the system is opened at all, and it touches the rig rather than any generator.
  3. Error hunt. DONE 3 AUG The first genuinely new method, generated from misconceptions the generators already declare. Instruments recognition rather than execution.
  4. Voice and collection. A companion with attitude carrying the feedback lines, and something that accumulates across sessions. Addresses the two absences the learner is responding to, which are the lack of anything to come back for and the lack of anyone to come back to.
  5. Single-learner shakedown. Run the reworked instrument daily. Watch for help-seeking suppression, confirm the telemetry separates a comprehension gap from a fluency gap, and treat voluntary opening as the primary outcome measure rather than accuracy.
  6. Distractor and criterion decode. Resolve stored answers back to the named misconception on multiple choice, and store the failing criterion on explorers. Cheap, uses data already on disk, and brings the second-tier surfaces into the same diagnostic picture.
  7. Diagnostic reading for the adult. Turn the gap output into something a parent or teacher acts on without interpretation: which step, whether comprehension or fluency, and what to do next. The instrument is only as good as this view.
  8. Coverage of the maths run. Extend cell decomposition to the arithmetic still assessed at answer level, so a learner's whole maths path is observable rather than only the fractions and decimals stretch.
  9. Cohort curves. Once more than one learner is on the instrument, fit error rate against opportunity per component. Components whose curves do not flatten become candidates for splitting or for redesign of the teaching attached to them.
  10. Operation vocabulary. Map the reasoning instruction set's operations onto the generators so each cell carries a domain component and an operation label. This is the bridge to any other subject, and it is deliberately downstream of proving the reading on one.
RISKFour. The first is live now.

Risks

Measuring an instrument the learner avoids

Confirmed rather than anticipated. The first learner has to be pressured into opening the system, which makes voluntary use the binding constraint on the whole programme and every accuracy figure conditional on it. Voluntary opening should be tracked as an outcome in its own right, and no measurement result should be reported without it.

Priced assistance may suppress help-seeking

The help-seeking literature treats hint avoidance as a maladaptive behaviour on equal footing with hint abuse, and a visible price selects for avoidance in anxious or perfectionist learners. The signature is long silences ending in timeouts rather than early hint use. If it appears, hold the price above zero or price both hints equally, which keeps the signal and removes the penalty gradient.

Instrumenting what is easy to instrument

Written algorithms decompose cleanly into cells, which is why the build started there. Reasoning operations do not, and the risk is that the system measures procedural fluency accurately while the capacity the motivation section is actually about goes unmeasured. W3 exists to confront this, and if the operation vocabulary cannot be made observable, the programme should say so rather than substitute the easier measurement.

Single-learner data carries no power for component discovery

Curve fitting requires a population. The N = 1 site produces design signal and affective constraints and cannot produce a teaching-gap finding. Single-case methodology should govern any claim made from it, and the cohort sites need to come online early.

Consent and ethics precede population use

Minor learners and university students require clearance and informed consent before records are used for research, and clinical records in NeuroPlay stay outside this programme's data boundary entirely. Clearance should be in place before the cohort sites begin recording for research purposes.

LATERReal, and not yet scheduled.

Where this goes eventually

The spine consists of three things: content that declares its steps, a record that keeps every attempt against those steps, and a diagnosis that separates the learner's gap from the teaching's. Any surface carrying content can in principle carry it, and several in the estate eventually should.

Undergraduate reaction engineering is the most interesting of them, because the decomposition is already explicit in how the subject is taught. Sizing a reactor proceeds through the mole balance, the rate law, the stoichiometric relation between concentration and conversion, combination, evaluation and a dimensional check. A failed sizing problem currently scores as one event when it could be any of those, and separating them across a cohort would say something about the course rather than about the students. The ScholarCloud class anchor would carry the telemetry, Publon Press would need exercises that declare steps at authoring time rather than in code, and a credential backed by step evidence would assert which operations a graduate demonstrated and under how much assistance. NeuroPlay shares the guardian and team-around-the-child spine and already observes process rather than outcome, so the method transfers even though its clinical records stay behind their own boundary.

None of that is scheduled, and it should not be until the reading works on one surface with one learner. The order matters more than the ambition.

ACCESSReachable now.

External resources

  • PSLC DataShop (Carnegie Mellon). Public step-level datasets with curve tooling and component model comparison built in. Most useful as the export format to target and as analysis tooling that need not be rebuilt.
  • Benchmark datasets. ASSISTments 2009 and 2015, EdNet, and the Eedi data from the NeurIPS 2020 Education Challenge. The Eedi set carries misconception labels on mathematics items and is the closest public analogue to what NeuroCraft generates.
  • Modelling libraries. pyKT for benchmarking knowledge tracing models and pyBKT for the classical Bayesian formulation, so the predictive claim can be tested against published baselines.
  • Communities. Educational Data Mining, Learning Analytics and Knowledge, and AIED. A South African foundation-programme application would be unusual in these venues.
SOURCESVerified August 2026.

References

  1. Brown, J. S. & Burton, R. R. (1978). Diagnostic models for procedural bugs in basic mathematical skills. Cognitive Science, 2(2).
  2. Corbett, A. T. & Anderson, J. R. (1995). Knowledge tracing: modeling the acquisition of procedural knowledge. User Modeling and User-Adapted Interaction, 4(4).
  3. Doignon, J.-P. & Falmagne, J.-C. (1999). Knowledge Spaces. Springer.
  4. Cen, H., Koedinger, K. & Junker, B. (2006). Learning Factors Analysis. Intelligent Tutoring Systems.
  5. Feng, M., Heffernan, N. & Koedinger, K. (2009). Addressing the assessment challenge with an online system that tutors as it assesses. UMUAI, 19(3).
  6. VanLehn, K. (2011). The relative effectiveness of human tutoring, intelligent tutoring systems, and other tutoring systems. Educational Psychologist, 46(4).
  7. Koedinger, K., Corbett, A. & Perfetti, C. (2012). The Knowledge-Learning-Instruction framework. Cognitive Science, 36(5).
  8. Piech, C. et al. (2015). Deep Knowledge Tracing. NeurIPS.
  9. Liu, Z. et al. (2022). pyKT: a Python library to benchmark deep learning based knowledge tracing models. arXiv:2206.11460
  10. Choi, Y. et al. (2020). EdNet: a large-scale hierarchical dataset in education. arXiv:1912.03072
  11. KCluster (2025). An LLM-based clustering approach to knowledge component discovery. arXiv:2505.06469
  12. PSLC DataShop. Carnegie Mellon University. pslcdatashop.web.cmu.edu
  13. Marton, F. & Säljö, R. (1976). On qualitative differences in learning. British Journal of Educational Psychology, 46.
  14. Rawatlal, R. (2026). Reasoning instruction set: operations, algorithms and certificates. Papers/ai-in-research-cybernetic/. Internal.
  15. Rawatlal, R. (2026). Representation over scale. Papers/ai-in-research-cybernetic/. Internal.
Papers/step-level-diagnostic-learning · artefact.html is the source of truth · rebuild the standalone local copy with bash build-local.sh · effect sizes in the theory section are quoted approximately from VanLehn (2011) and should be checked against the source before external use.