MathTrail

The model writes, the service checks.

How MathTrail lets a task reach a child only once it has passed 10 checks, and what the checks do not see.

Code and data

The chat's modelwrites

taskA fence 12 m long has a post every 3 m, both ends included. How many posts are there?

solver12 ÷ 3 + 1 = 5 → C

the task and its solverrefused, sent back

MathTrailchecks

10 checksthe solver runs twice

one right option · words the grade can read · no repeat of earlier tasks

passed every check

The childsolves

  1. A3
  2. B4
  3. C5
  4. D6
  5. E12

the answer sealed until the child answers

Four things the paper shows

  1. The chat's model writes

    Every task is written by the model of the chat a family already uses, in the child's language and around a story close to the child. MathTrail calls no model of its own. It is a small deterministic service that keeps nothing of a child between requests.

    1. The chat's modelwrites
    2. MathTrailchecks
    3. The childsolves
  2. The solver runs twice

    With the task, the model writes a short program that works the answer out and names the letter of the right option. The service runs it, moves the letters 2 places on, so that A becomes C, and runs it again. A program that works the answer out finds it in its new place, and one that wrote the letter by hand does not. The second run comes once the first picks a single option.

    First run

    1. A3
    2. B4
    3. C5
    4. D6
    5. E12

    → C

    Second run, letters moved

    1. A6
    2. B12
    3. C3
    4. D4
    5. E5

    → E

    A program that wrote “C” by hand points at 3 in the second run, and the task is refused.

  3. The answer is sealed

    The answer, the solution, the explanations and the solver are encrypted and kept in the child's profile in the parent's Google Drive. Until the child answers, they are in nothing the service hands the card or the chat's model, and in none of its logs.

    What the child sees

    the question · the drawing · the options · the hint

    Sealed until the answer

    the answer · the solution · the explanations

    XChaCha20-Poly1305 · bound to the child and the task

    What the seal does not reach is the chat itself. The model that wrote the task knows its answer, and the chat's own record of the tools it called shows it to an adult who opens it.

  4. A little harder, but within reach

    The student model predicts how likely a child is to solve a task, and the service asks for the one whose chance is nearest 77.5%, the middle of the corridor from 70% to 85%. A new child first takes a trial series of 5 tasks, which finds where to start.

    P = c + (1 − c) · σ(θ + δt − β)

    • c = 0.2, one of 5 guessed
    • the corridor aimed at

The student model against its goals

How closely the service follows a child, how well it keeps tasks within reach, how surely it declares a topic mastered, and what the child is shown late in a run. The numbers are the bench's, on simulated children, computed from the code this site is built from.

commit 7108ac9406e6 · 1,000 children × 200 answers seed 20261001

The paper reports the rule as it stood at the commit its numbers were computed on, 52ce86908135. This page follows the service as it runs now, so the two differ.

  • the service now
  • 95% interval
  • the goal
  • the ceiling, an estimate that knows where the child stands

How the estimate follows a child

  1. The lag behind a child who learns

    logits, less is better · the best possible is 0.00

    Baseline

    goal ≤ 0.24

    0.540.52–0.56

  2. Tasks in the corridor 70%–85%

    a child who learns · a share of tasks, more is better

    Baseline

    goal ≥ 43.9%ceiling 61.1%

    37.5%36.9%–38%

  3. Not caught up after a jump

    a share of children, less is better

    Not reached

    goal ≤ 25%

    56.8%53.8%–59.8%

How mastery is declared

  1. Masteries declared falsely

    a share of the masteries declared, less is better

    Reached

    goal ≤ 20%

    3.6%3.1%–4.2%

  2. Masteries declared falsely, topics far apart

    a share of the masteries declared, less is better

    Reached

    goal ≤ 20%

    5.9%5.2%–6.6%

  3. Answers until a mastery is declared

    answers, less is better

    Baseline

    goal ≤ 6.9

    4.64.5–4.8

  4. Answers until a mastery is declared, topics far apart

    answers, less is better

    Baseline

    goal ≤ 7.4

    4.94.7–5.1

What the child is shown late in a runanswers 150–200 · the goal is no more than in answers 6–20

  1. How far the rating moves

    points on most answers, less is better

    Reached

    goal ≤ 73

    4140–41

  2. How often the rank changes

    changes a hundred answers, less is better

    Reached

    goal ≤ 3.8

    1.61.4–1.9

  3. The rating, a child who learns

    points on most answers, less is better

    Reached

    goal ≤ 73

    4040–41

  4. The rank, a child who learns

    changes a hundred answers, less is better

    Reached

    goal ≤ 3.6

    1.71.4–1.9

For comparison, with no goal

  1. The error of the estimate after 200 answers

    a child who stays put · logits, less is better

    0.480.48–0.49

  2. Tasks in the corridor 70%–85%

    a child who stays put · a share of tasks, more is better

    ceiling 60.7%

    42.8%42.3%–43.3%

Reached
The whole interval lies on the better side of the goal.
Not reached
The whole interval lies on the worse side of the goal.
Baseline
The goal is drawn from the service's own number, so it is the bar for the next rule rather than a test of this one.

Live data

Whether the model's promises come true for real children, one month at a time. The numbers come from the log of answers, with no names and no task's text.

Coming

Promised → came true

Came true − promised, by the number of answers

The share right on the rule's tasks

The numbers come once a whole month is counted, early in the month after it.

A range is shown only when it holds 10 children or more and 30 answers or more, and every count in it is rounded to a multiple of 5. A smaller range is left out rather than folded in with another.

Where the reference tasks came from

Ideas are not protected by copyright, and texts are. So from a book only a task's idea and its answer are taken, and the wording, the story, the options and the traps are written anew.

In the public domain The copyright in these books has run out. They belong to everyone, and anyone may use them for free.

In all, 603 reference tasks each with a solver and a trap behind every wrong option

  1. 200 · grades 1–2
  2. 250 · grades 3–4
  3. 153 · grades 5–6

The style follows the Soviet collections. A reference task has no field for its source, so the author of a task's idea cannot be named.

The whole paper

Its methods, tables, protocols, limitations and ethics. The PDF comes once the paper names its authors and its archive.

Code on GitHub