NTell World Ink.AI Workshop Press M for menu
An interactive walkthrough · for teaching staff

The learning science behind Study and Learn

In June 2026 Microsoft published a white paper on the science behind its Study and Learn agent. This walkthrough reads it closely: what the design gets right, what the evidence does and does not yet show, and how to talk about the tool with students without overclaiming. About fifteen minutes.

← Back to the workshop
Module 01

Two claims, one word apart

In this module: the distinction the whole reading turns on

A tool can be built on good science and still not be proven to work. Those are two separate claims, and Study and Learn sits squarely between them.

The document

In June 2026 Microsoft published “Learning by design: Learning science foundation of the Study and Learn Agent”. It is careful, well-referenced, and unusually candid in places. It also makes one move worth watching: it describes in detail how the agent was designed, and invites you to read that design as a reason to trust the outcome.

The white paper is version 1.0, dated June 2026, and is authored entirely by Microsoft staff. It is the vendor describing its own product, which is normal for a design paper and worth holding in mind throughout.

The distinction

Read the paper closely and two different kinds of claim run through it. One is about how the tool was made. The other is about what it does to learning. They are easy to blur together, and keeping them apart is the single most useful thing a teacher can do with this document.

A design claim
“Built using learning-science principles.”
A statement about the ingredients and intentions: what the designers drew on, what behaviours they told the agent to follow. This is what the white paper actually substantiates.
An efficacy claim
“Improves learning outcomes.”
A statement about measured results: that students who use it learn more, or remember longer, than students who do not. This is a different claim, and it needs different evidence.

“Designed on the science” and “proven to improve learning” are not the same sentence. The rest of this deck keeps them apart.

Sort the claim

Each line below is drawn from Microsoft’s own material. Decide whether it is a claim about the design, or a claim about a measured learning result. The verdict explains what the evidence behind it actually is.

Module 02

The four pillars, and the science under each

In this module: what Microsoft built, and how solid the research beneath it is

The design is genuinely thoughtful. The question is not whether the pillars name real learning science; they do. It is whether a chat agent re-creates the conditions those studies measured, or only resembles them.

The frame Microsoft opens with

The paper opens on Bloom’s “Two Sigma” result: a tutored student outperformed about 98 per cent of classroom peers. It is a powerful frame for a personal tutor at scale. It is also the weakest link in the citation set.

The two-sigma effect has never been reliably replicated at that size, and it required mastery-based testing as well as tutoring. Across ninety-six tutoring studies reviewed by von Hippel (Education Next, 2024), none reached two sigma; the meta-analytic effect for real tutoring is around 0.37 of a standard deviation (Nickow, Oreopoulos and Quan). Invoking two sigma to motivate an AI agent is rhetorically strong and empirically unsupported.

The four pillars

Microsoft frames the agent around four design pillars: adaptive scaffolding, productive struggle, active learning, and application and transfer. Each is a set of behaviours the agent is instructed to follow, and each is tied to a body of research.

The paper calls these “architectural constraints, evaluated in every interaction”. Keep that phrase in view; the agent’s restraint is a set of instructions layered on a general model, not a different kind of machine. Explore each pillar below.

Inspect a pillar

The category question

Notice what the strongest caveats have in common. Scaffolding and the zone of proximal development describe responsive human interaction: contingent support from someone who genuinely models the learner’s state of mind. A language model predicts the next word; it simulates that contingency rather than possessing it.

This matters because the figures often cited are for human and step-based tutoring. VanLehn’s 2011 review corrected the folklore: human tutoring sits around an effect size of 0.79, well-designed step-based computer tutors around 0.76, and answer-giving systems around 0.31. Good structured tutoring approaches human tutoring; answer-giving, which is what an unguarded chatbot resembles, is the weakest form.

The pillars rest on real research. Whether an agent re-creates those conditions, or only resembles them, is an empirical question the paper does not settle.

Module 03

What the independent evidence shows

In this module: the causal studies, and where the product sits on an evidence ladder

There is real, independent evidence on AI tutors. It is genuinely mixed, it turns almost entirely on design, and none of it tested Study and Learn itself.

A result worth sitting with

baseline: never used the tool grades up +48% while the tool is there grades down −17% after it is taken away

Bastani and colleagues (PNAS, 2025) gave 994 high-school students a plain answer-giving chatbot. Grades jumped while they had it, then fell below the students who never used it once it was removed. A guardrailed version that withheld answers erased the harm, though it produced no lasting gain either. This is the clearest picture we have of the crutch effect, and of design deciding the outcome.

The studies, and what they actually tested

Bastani et al. 2025PNAS, field RCTAnswer-giving AI flatters practice, then harms durable skill once removed; guardrails that withhold answers prevent the harm. The design, not the tool, decides.
Kestin et al. 2025Scientific ReportsA custom, guardrailed AI tutor beat in-class active learning at Harvard, roughly doubling gains. Single elite course, short duration, and immediate rather than delayed measures.
Wang et al. 2024Tutor CoPilot, StanfordLifted topic mastery by four percentage points, and more for weaker tutors. Crucially it assists a human tutor; it is not a student-facing agent working alone.
Study and LearnMicrosoft, 2026No independent, peer-reviewed study of this product exists as at August 2026. Every study above tested something else.

The direction of the design is supported in general. The specific product is untested.

Where does it sit?

Education uses a standard evidence ladder, the ESSA tiers used by the What Works Clearinghouse. Click each tier to test whether Study and Learn’s current evidence reaches it.

Click a tier to test it.

What the paper offers as evidence

The white paper’s own evidence is nine rounds of expert review and an anonymous survey of about 300 users, who rated features from 4.0 down to 3.56 out of five. That establishes that people found it usable and liked it. It does not establish that they learned more.

Satisfaction is a weak, sometimes negative, guide to learning. Students often rate answer-giving tools highest precisely because they reduce effort; Roediger and Karpicke called this the illusion of competence, where fluency feels like mastery. The paper also notes a Digital Safety Board sign-off, which is a privacy, security and responsible-AI review; a compliance check, not a measure of teaching effect.

Module 04

What this means for teaching staff

In this module: how to use and describe the tool without overclaiming

None of this makes Study and Learn a bad tool. Used as designed, it is a capable study companion. The task is to match what we say about it to what we actually know.

The honest status

Promising by design, unproven in outcome. The design sits on the favourable side of the evidence; the product has not been measured. That is the accurate position as at August 2026, and it is a defensible one to hold in front of a class or a course team.

Four positions worth taking

Separate the two claimsIn anything you tell students or write into assessment design, treat “built on learning science” and “improves learning” as different statements. Only the first is established for this product.
Do not cite the safety reviewThe Responsible-AI and safety sign-off says the tool is safe to deploy. It says nothing about whether it teaches better, so it should never stand in for efficacy evidence.
Aim students at the stuck momentPoint them to the agent after a genuine first attempt, when they are stuck, which is where productive-struggle conditions hold. That is different advice from “use it to study”.
Measure it locally before claimingIf CDU wants to claim a benefit, it needs a small local study. That means a comparison group, a real learning measure, and a delayed test after the tool is withdrawn, following the Bastani design.

The tool can be genuinely useful and genuinely unproven at the same time. Both are true, and the advice above holds them together.

Microsoft’s own caution

To its credit, the paper says much of this itself. It warns against a century of overpromising, from radio and film to the personal computer and the MOOC, each announced as the end of the tutoring gap. It calls this release “a first version, a foundation, not a finished product”, designed to be tested and refined as efficacy research comes in.

That candour is the right note, and it is the note to carry into the workshop. The gap is not between an honest paper and a dishonest one; it is between a design rationale and proof of outcome. This deck simply keeps that gap visible.

Why this matters

A well-designed tool and a proven tool look identical from the outside, right up until a later moment asks the student to recall, explain, or transfer what they were supposed to have learned. Teaching staff are the ones standing at that later moment. Knowing which claim you are relying on is what lets you set assessment, and advise students, on solid ground.

Carry this back to the workshop’s hands-on Copilot section, where the same tool is tested from the other side, and to the sections on cognitive outsourcing and honest uncertainty.

Hands on: Copilot Study and Learn
Cognitive outsourcing
What we’re not sure about

Sources used. Microsoft, “Learning by design: Learning science foundation of the Study and Learn Agent”, v1.0, June 2026 (vendor white paper). First-principles appraisal of the Study and Learn efficacy evidence, 2026 (project source; references dated to August 2026). Key studies named in the deck: Bastani et al., PNAS 122(26), 2025; Kestin et al., Scientific Reports 15:17458, 2025; Wang et al., Tutor CoPilot, Stanford, 2024; VanLehn, Educational Psychologist 46(4), 2011; Sinha and Kapur, Review of Educational Research 91(5), 2021; Roediger and Karpicke, Psychological Science 17(3), 2006; von Hippel, Education Next, 2024; Nickow, Oreopoulos and Quan (tutoring meta-analysis). Evidence tiers: ESSA / What Works Clearinghouse. Figures for the satisfaction survey are Microsoft’s own and are reported as vendor assertions.