The model writes, the service checks.
How MathTrail lets a task reach a child only once it has passed 10 checks, and what the checks do not see.
The chat's modelwrites
taskA fence 12 m long has a post every 3 m, both ends included. How many posts are there?
solver12 ÷ 3 + 1 = 5 → C
the task and its solverrefused, sent back
MathTrailchecks
10 checksthe solver runs twice
one right option · words the grade can read · no repeat of earlier tasks
passed every check
The childsolves
the answer sealed until the child answers
- 603 reference tasks
- 10 checks
- 17 topics
- 20 traps
- Grades 1–6
- open code, MIT
Four things the paper shows
The chat's model writes
Every task is written by the model of the chat a family already uses, in the child's language and around a story close to the child. MathTrail calls no model of its own. It is a small deterministic service that keeps nothing of a child between requests.
- The chat's modelwrites
- MathTrailchecks
- The childsolves
The solver runs twice
With the task, the model writes a short program that works the answer out and names the letter of the right option. The service runs it, moves the letters 2 places on, so that A becomes C, and runs it again. A program that works the answer out finds it in its new place, and one that wrote the letter by hand does not. The second run comes once the first picks a single option.
First run
- A3
- B4
- C5
- D6
- E12
→ C
Second run, letters moved
- A6
- B12
- C3
- D4
- E5
→ E
A program that wrote “C” by hand points at 3 in the second run, and the task is refused.
The answer is sealed
The answer, the solution, the explanations and the solver are encrypted and kept in the child's profile in the parent's Google Drive. Until the child answers, they are in nothing the service hands the card or the chat's model, and in none of its logs.
What the child sees
the question · the drawing · the options · the hint
Sealed until the answer
the answer · the solution · the explanations
XChaCha20-Poly1305 · bound to the child and the task
What the seal does not reach is the chat itself. The model that wrote the task knows its answer, and the chat's own record of the tools it called shows it to an adult who opens it.
A little harder, but within reach
The student model predicts how likely a child is to solve a task, and the service asks for the one whose chance is nearest 77.5%, the middle of the corridor from 70% to 85%. A new child first takes a trial series of 5 tasks, which finds where to start.
P = c + (1 − c) · σ(θ + δt − β)
- c = 0.2, one of 5 guessed
- the corridor aimed at
The student model against its goals
How closely the service follows a child, how well it keeps tasks within reach, how surely it declares a topic mastered, and what the child is shown late in a run. The numbers are the bench's, on simulated children, computed from the code this site is built from.
commit 7108ac9406e6 · 1,000 children × 200 answers seed 20261001
The paper reports the rule as it stood at the commit its numbers were computed on, 52ce86908135. This page follows the service as it runs now, so the two differ.
- the service now
- 95% interval
- the goal
- the ceiling, an estimate that knows where the child stands
How the estimate follows a child
The lag behind a child who learns
logits, less is better · the best possible is 0.00
Baseline
0.540.52–0.56
Tasks in the corridor 70%–85%
a child who learns · a share of tasks, more is better
Baseline
37.5%36.9%–38%
Not caught up after a jump
a share of children, less is better
Not reached
56.8%53.8%–59.8%
How mastery is declared
Masteries declared falsely
a share of the masteries declared, less is better
Reached
3.6%3.1%–4.2%
Masteries declared falsely, topics far apart
a share of the masteries declared, less is better
Reached
5.9%5.2%–6.6%
Answers until a mastery is declared
answers, less is better
Baseline
4.64.5–4.8
Answers until a mastery is declared, topics far apart
answers, less is better
Baseline
4.94.7–5.1
What the child is shown late in a runanswers 150–200 · the goal is no more than in answers 6–20
How far the rating moves
points on most answers, less is better
Reached
4140–41
How often the rank changes
changes a hundred answers, less is better
Reached
1.61.4–1.9
The rating, a child who learns
points on most answers, less is better
Reached
4040–41
The rank, a child who learns
changes a hundred answers, less is better
Reached
1.71.4–1.9
For comparison, with no goal
The error of the estimate after 200 answers
a child who stays put · logits, less is better
0.480.48–0.49
Tasks in the corridor 70%–85%
a child who stays put · a share of tasks, more is better
42.8%42.3%–43.3%
- Reached
- The whole interval lies on the better side of the goal.
- Not reached
- The whole interval lies on the worse side of the goal.
- Baseline
- The goal is drawn from the service's own number, so it is the bar for the next rule rather than a test of this one.
Live data
Whether the model's promises come true for real children, one month at a time. The numbers come from the log of answers, with no names and no task's text.
Coming
Promised → came true
Came true − promised, by the number of answers
The share right on the rule's tasks
The numbers come once a whole month is counted, early in the month after it.
A range is shown only when it holds 10 children or more and 30 answers or more, and every count in it is rounded to a multiple of 5. A smaller range is left out rather than folded in with another.
Where the reference tasks came from
Ideas are not protected by copyright, and texts are. So from a book only a task's idea and its answer are taken, and the wording, the story, the options and the traps are written anew.
In the public domain The copyright in these books has run out. They belong to everyone, and anyone may use them for free.
In all, 603 reference tasks each with a solver and a trap behind every wrong option
- 200 · grades 1–2
- 250 · grades 3–4
- 153 · grades 5–6
The style follows the Soviet collections. A reference task has no field for its source, so the author of a task's idea cannot be named.
The whole paper
Its methods, tables, protocols, limitations and ethics. The PDF comes once the paper names its authors and its archive.