The learning science behind Study and Learn
In June 2026 Microsoft published a white paper on the science behind its Study and Learn agent. This walkthrough reads it closely: what the design gets right, what the evidence does and does not yet show, and how to talk about the tool with students without overclaiming. About fifteen minutes.
Two claims, one word apart
A tool can be built on good science and still not be proven to work. Those are two separate claims, and Study and Learn sits squarely between them.
The document
The white paper is version 1.0, dated June 2026, and is authored entirely by Microsoft staff. It is the vendor describing its own product, which is normal for a design paper and worth holding in mind throughout.
The distinction
Read the paper closely and two different kinds of claim run through it. One is about how the tool was made. The other is about what it does to learning. They are easy to blur together, and keeping them apart is the single most useful thing a teacher can do with this document.
“Designed on the science” and “proven to improve learning” are not the same sentence. The rest of this deck keeps them apart.
Sort the claim
Each line below is drawn from Microsoft’s own material. Decide whether it is a claim about the design, or a claim about a measured learning result. The verdict explains what the evidence behind it actually is.
The four pillars, and the science under each
The design is genuinely thoughtful. The question is not whether the pillars name real learning science; they do. It is whether a chat agent re-creates the conditions those studies measured, or only resembles them.
The frame Microsoft opens with
The two-sigma effect has never been reliably replicated at that size, and it required mastery-based testing as well as tutoring. Across ninety-six tutoring studies reviewed by von Hippel (Education Next, 2024), none reached two sigma; the meta-analytic effect for real tutoring is around 0.37 of a standard deviation (Nickow, Oreopoulos and Quan). Invoking two sigma to motivate an AI agent is rhetorically strong and empirically unsupported.
The four pillars
The paper calls these “architectural constraints, evaluated in every interaction”. Keep that phrase in view; the agent’s restraint is a set of instructions layered on a general model, not a different kind of machine. Explore each pillar below.
Inspect a pillar
The category question
Notice what the strongest caveats have in common. Scaffolding and the zone of proximal development describe responsive human interaction: contingent support from someone who genuinely models the learner’s state of mind. A language model predicts the next word; it simulates that contingency rather than possessing it.
This matters because the figures often cited are for human and step-based tutoring. VanLehn’s 2011 review corrected the folklore: human tutoring sits around an effect size of 0.79, well-designed step-based computer tutors around 0.76, and answer-giving systems around 0.31. Good structured tutoring approaches human tutoring; answer-giving, which is what an unguarded chatbot resembles, is the weakest form.
The pillars rest on real research. Whether an agent re-creates those conditions, or only resembles them, is an empirical question the paper does not settle.
What the independent evidence shows
There is real, independent evidence on AI tutors. It is genuinely mixed, it turns almost entirely on design, and none of it tested Study and Learn itself.
A result worth sitting with
Bastani and colleagues (PNAS, 2025) gave 994 high-school students a plain answer-giving chatbot. Grades jumped while they had it, then fell below the students who never used it once it was removed. A guardrailed version that withheld answers erased the harm, though it produced no lasting gain either. This is the clearest picture we have of the crutch effect, and of design deciding the outcome.
The studies, and what they actually tested
The direction of the design is supported in general. The specific product is untested.
Where does it sit?
Education uses a standard evidence ladder, the ESSA tiers used by the What Works Clearinghouse. Click each tier to test whether Study and Learn’s current evidence reaches it.
What the paper offers as evidence
Satisfaction is a weak, sometimes negative, guide to learning. Students often rate answer-giving tools highest precisely because they reduce effort; Roediger and Karpicke called this the illusion of competence, where fluency feels like mastery. The paper also notes a Digital Safety Board sign-off, which is a privacy, security and responsible-AI review; a compliance check, not a measure of teaching effect.
What this means for teaching staff
None of this makes Study and Learn a bad tool. Used as designed, it is a capable study companion. The task is to match what we say about it to what we actually know.
The honest status
Four positions worth taking
The tool can be genuinely useful and genuinely unproven at the same time. Both are true, and the advice above holds them together.
Microsoft’s own caution
That candour is the right note, and it is the note to carry into the workshop. The gap is not between an honest paper and a dishonest one; it is between a design rationale and proof of outcome. This deck simply keeps that gap visible.
Why this matters
Carry this back to the workshop’s hands-on Copilot section, where the same tool is tested from the other side, and to the sections on cognitive outsourcing and honest uncertainty.
→ Hands on: Copilot Study and Learn
→ Cognitive outsourcing
→ What we’re not sure about
Sources used. Microsoft, “Learning by design: Learning science foundation of the Study and Learn Agent”, v1.0, June 2026 (vendor white paper). First-principles appraisal of the Study and Learn efficacy evidence, 2026 (project source; references dated to August 2026). Key studies named in the deck: Bastani et al., PNAS 122(26), 2025; Kestin et al., Scientific Reports 15:17458, 2025; Wang et al., Tutor CoPilot, Stanford, 2024; VanLehn, Educational Psychologist 46(4), 2011; Sinha and Kapur, Review of Educational Research 91(5), 2021; Roediger and Karpicke, Psychological Science 17(3), 2006; von Hippel, Education Next, 2024; Nickow, Oreopoulos and Quan (tutoring meta-analysis). Evidence tiers: ESSA / What Works Clearinghouse. Figures for the satisfaction survey are Microsoft’s own and are reported as vendor assertions.