A student studying with Bloom is not reading a generated explanation and hoping it was right. Every question, every wrong-answer response, and every decision about whether an idea has landed is written ahead of time, approved by a person, and graded by the server — and each of those choices comes from a finding somebody published, not from a product meeting.
Two facts underneath all five
A study session makes no live AI calls. The questions, the wrong choices, and the explanation a student reads when they miss one are written in batches, checked when they are generated, and approved by a person in Bloom's concept library before any student sees them. The documented failure of a live model inside an assessment loop is a confident wrong explanation, and the way to not have one is to not ask.
The server is the only authority on a right answer. Bloom writes the session's plan before the first question renders and stores it. What the student's browser reports is believed about one thing only — which choice was tapped. Which choice was correct, how hard the question was, and what a wrong choice reveals are all read back from the stored plan, so nothing a browser sends can move a student's record.
1. Mastery, not attendance
Bloom, B. S. (1984). Mastery learning asks students to advance on demonstrated understanding rather than on the calendar. That is a claim about pacing, and pacing is something software can actually hold.
Every student has a state for every topic: unknown → shaky → solid → mastered. Reaching solid takes two correct answers in a row on questions of real difficulty; reaching mastered takes a clean exit ticket at the end of a session, and only that. A finished session is not a passed one: each session stores its own verdict — how much of the practice was correct, whether the exit ticket landed — and offers another pass instead of “move on” when the run was weak.
Nothing a student says about themselves moves a state. Fluency during study is a poor predictor of what will be there next week (Koriat & Bjork, 2005), so re-reading a topic, replaying an explanation, or feeling ready changes nothing in the record — only answered questions do.
Evidence is deliberately asymmetric. One lucky answer is not mastery and one slip is not ignorance, so a state falls one rung at a time and never falls back to unknown. Unknown means there is no evidence yet — never that the student got something wrong.
2. Spacing beats cramming
Ebbinghaus (1885/1913); Cepeda, Pashler, Vul, Wixted & Rohrer (2006). The same practice spread over days produces more retention than the same practice in one sitting. It is one of the oldest and best-replicated findings in the field, and it is almost never built into the tools students use.
Every topic a student has touched carries a date it comes due again: a shaky topic in one day, a solid one in three, a mastered one in seven — widening to three weeks only after a mastered topic has survived a review. Topics that are due come back oldest-first, at the start of the next session, before anything new is taught.
The intervals are fixed rather than fitted to each student, and that is a choice we will defend. A per-student forgetting curve needs far more data per topic than a school year produces, and an interval nobody can explain is an interval nobody can check. What is personal is which topics come due, and in what order.
3. Testing is a learning event
Roediger & Karpicke (2006). Retrieving something from memory strengthens it more than re-reading it does. Being asked is not only the measurement — it is a large part of the instruction.
So a session opens by asking before it teaches: up to two topics that are due, one question each, retrieved cold. And it closes with an exit ticket held out of the practice pool — a question the student did not just see — so that what it measures is whether the idea landed, not whether the last five minutes are still in working memory.
4. Practice should be hard enough to be diagnostic
Bjork & Bjork (2011). Conditions that make practice feel harder — and slower — often produce better long-term learning than the smooth ones. Easy practice is pleasant and tells you very little.
Difficulty steps up after two correct answers in a row and down after two wrong, and a session is filled across difficulty bands rather than run from the bottom, so it cannot settle onto easy questions. The exit ticket is the topic's hardest question. And easy correct answers, on their own, never promote a student: the ladder in section 1 only counts evidence at difficulty that was worth something.
5. A wrong answer is a belief, not a miss
Chi, Feltovich & Glaser (1981). Novices and experts fail differently, and a wrong answer usually reveals a specific, nameable idea the student is holding — not an absence of one.
Every topic in Bloom carries a named list of misconceptions, each with the belief stated plainly and a written counter to it, and every wrong choice on every question is mapped to one of them. Choose it, and the student is shown that misconception's own approved explanation — resolved on the server, the same words for every student, written and reviewed by a person rather than generated on the spot. Which misconception it was is written into the student's record, so the topic can come back aimed at the belief rather than the question.
What one session looks like
Sweller (1988); Wood, Bruner & Ross (1976); National Research Council (2000). Working memory is the binding constraint, support belongs at the step the learner cannot yet take alone, and instruction has to start from what the student already knows.
A Bloom session has hard ceilings, chosen as product rather than as a performance guard: up to two warm-up topics at one question each, up to six practice questions, one exit ticket, and one topic taught. Worked examples come before practice and are broken into steps rather than handed over as an answer, and each topic opens with a question the student answers before the explanation is revealed — learners who explain something to themselves first learn more than learners who are simply told (Chi, de Leeuw, Chiu & LaVancher, 1994). After the exit ticket the session shows what moved on the student's map and stops. There is no infinite feed.
Which topic gets taught is gated on prerequisites: the topic Bloom teaches next is the earliest one whose prerequisites are all at least solid and which the student has not yet made solid themselves. A student new to a course can map their strengths first — at most ten questions across a course, framed as a map and never as a test, and it never lowers a state the student already earned.
What we do not claim
- No two-sigma result. Bloom (1984) describes human one-to-one tutoring. We cite it as the design target that mastery-based pacing aims at. We have not run an efficacy study, we have no control group, and we publish no learning-gain number.
- No score projection. Mastery states are never converted into a predicted SAT, ACT, PSAT or AP score, or into a GPA forecast. Take the official practice test for the diagnosis; Bloom is for the weeks in between.
- No ability claim, and no ranking. A state is a record of this student's evidence against these questions. It is not an ability estimate, not a percentile, and we never tell a family where their child sits against other students.
- Not a diagnosis. Nothing in the Learning Lab is a screening or clinical instrument for a learning difference, and we do not present it as one.
- Parents see effort, not mistakes. Sessions this week, streak, how many topics sit at each state — never the question missed and never the misconception's name. Counselors see the same summary and nothing more.
- One live AI lane, clearly marked. A student can ask for a concept to be explained another way, and that is a live model call. It is locked to the topic, nothing it says is graded, and it is never the source of a question or an answer key.
Sources
- Bjork, E. L., & Bjork, R. A. (2011). Making things hard on yourself, but in a good way: Creating desirable difficulties to enhance learning. In Psychology and the real world (pp. 56–64). Worth.
- Bloom, B. S. (1984). The 2 sigma problem: The search for methods of group instruction as effective as one-to-one tutoring. Educational Researcher, 13(6), 4–16.
- Cepeda, N. J., Pashler, H., Vul, E., Wixted, J. T., & Rohrer, D. (2006). Distributed practice in verbal recall tasks: A review and quantitative synthesis. Psychological Bulletin, 132(3), 354–380.
- Chi, M. T. H., Feltovich, P. J., & Glaser, R. (1981). Categorization and representation of physics problems by experts and novices. Cognitive Science, 5(2), 121–152.
- Chi, M. T. H., de Leeuw, N., Chiu, M.-H., & LaVancher, C. (1994). Eliciting self-explanations improves understanding. Cognitive Science, 18(3), 439–477.
- Ebbinghaus, H. (1913). Memory: A contribution to experimental psychology (H. A. Ruger & C. E. Bussenius, Trans.). Teachers College. (Original work published 1885.)
- Koriat, A., & Bjork, R. A. (2005). Illusions of competence in monitoring one's knowledge during study. Journal of Experimental Psychology: Learning, Memory, and Cognition, 31(2), 187–194.
- National Research Council. (2000). How people learn: Brain, mind, experience, and school (Expanded ed.). National Academies Press.
- Roediger, H. L., III, & Karpicke, J. D. (2006). Test-enhanced learning: Taking memory tests improves long-term retention. Psychological Science, 17(3), 249–255.
- Sweller, J. (1988). Cognitive load during problem solving: Effects on learning. Cognitive Science, 12(2), 257–285.
- Wood, D., Bruner, J. S., & Ross, G. (1976). The role of tutoring in problem solving. Journal of Child Psychology and Psychiatry, 17(2), 89–100.