Altamens

Altamens Forschung · Open-data reanalysis · Protocol

How Much Can the Question Mix Change a 30-Second Mental-Math Score?

A preregistered reanalysis of an independent public arithmetic dataset: can difficulty-balanced question sets make a 30-second correct-count score more stable — or does item mix barely matter?

Altamens Research10 Min. Lesezeit

Zusammenfassung

This study asks whether two equally long arithmetic tests can produce meaningfully different scores for the same person simply because their questions differ. We use public, independently collected response-time and accuracy data from 59 US undergraduates completing 256 addition and subtraction verification trials each. Item difficulty is estimated on training participants, then held-out participants' responses are used to simulate repeated 30-second forms.

We compare naive random forms with forms balanced on operation, truth status, problem size, predicted response time and predicted error probability. The primary outcome is within-person form-to-form variation in correct-count score. This is a measurement study, not a validation of Altamens, a clinical assessment, or a source of population norms.

Kernaussagen

  • A 30-second correct-count score reflects three things at once: the person's performance, the questions they happened to receive, and the test conditions.
  • We reanalyse an independent public dataset — 59 US undergraduates, 256 arithmetic-verification trials each (15,104 planned participant-trials) — released under CC BY 4.0.
  • Simulated 30-second forms will compare naive random question mixes against forms balanced on cross-validated item difficulty, calibrated only on training participants.
  • The primary outcome is within-person form-to-form variation in correct-count score. A null or contrary result will be published.
  • This is a measurement-transparency study. It cannot validate Altamens as a clinical assessment, create population norms, or show that practice improves general cognition.

1The measurement question

A 30-second test has no room to average over hundreds of items. A handful of unusually slow questions can consume a large share of the available time. That means a correct-count score can reflect both a person's performance and the particular questions selected for that attempt.

The central question is narrow and answerable: in a fixed 30-second arithmetic sprint, how much can item composition alone change a person's correct-count score — and can difficulty-balanced forms reduce that form-to-form variation?

Answering it honestly requires separating four ideas that a single reported score collapses together:

  1. The person's performance — their speed and accuracy on that occasion.
  2. The questions shown — operation, operand size, truth status, answer plausibility and calibrated item difficulty.
  3. The test conditions — time limit, order, interface, device, latency, distractions and practice.
  4. The reported score — a correct count and accuracy generated by all three factors above.

2Why arithmetic questions differ in difficulty

The best-replicated reason is the problem-size effect: larger arithmetic problems reliably take longer to answer and produce more errors than smaller ones, a finding studied and modelled for over three decades.3 The source dataset for this study replicated it — response times tended to increase with problem size for both addition and subtraction.1

  • Operation — subtraction and addition items are not interchangeable in speed or error profile.
  • True/false decision demands — verifying a displayed answer engages different processes than producing one, and implausible false answers can be rejected faster than near-miss answers.4
  • Speed–accuracy tradeoff — a person can convert accuracy into attempted items and back; correct count alone hides which strategy they used.
  • Order, practice and fatigue — where an item lands in a session affects the response it gets.

If items differ this much, a short form containing more subtraction items or larger operands may be slower than another form of equal duration — which raises the practical question this study is built around: did the person's performance change, or did the form change?

3An independent open dataset

We will use the arithmetic-verification task from Bye, Harsch and Varma (2022), an independently collected, peer-reviewed dataset released openly under CC BY 4.0.1,2 Participants judged whether single-digit addition and subtraction equations were true or false, responding as quickly and accurately as possible. False answers generally differed from the correct answer by 2, with specific exceptions to avoid negative results or trivially rejectable values.

Abbildung 1Try the task format yourself
3 + 4 = 7
9 − 4 = 5
4 + 2 = 8
8 + 6 = 14
8 − 3 = 7
7 − 2 = 5
9 + 7 = 14
2 + 6 = 8
9 − 6 = 1
8 − 5 = 3
7 + 9 = 18
6 − 2 = 6

These items are illustrative examples constructed from the source task's documented stimulus rules — they are not drawn from the dataset, and your responses are timed locally in your browser only. Nothing is recorded or sent anywhere.

Twelve example equations in the source study's verification format: single-digit addition and subtraction, with false displayed answers offset from the correct answer by 2. Judge each as true or false and see your own response times — a feel for why some items consume more of a 30-second budget than others.

  1. 62 undergraduates initially recruited.
  2. Two excluded for prior pilot participation; one excluded for accuracy considerably below 90%.
  3. Final sample: 59 US undergraduates, ages 18–23.
  4. 256 arithmetic-verification trials per participant — 15,104 planned participant-trials in total.
  5. Minimum threshold to proceed: at least 50 participants retaining at least 200 valid trials each, and at least 10,000 valid participant-trials overall.
Abbildung 2Participant and trial flow
  1. Recruited62

    US undergraduates, ages 18–23

  2. Excluded: prior pilot participation−2
  3. Excluded: accuracy considerably below 90%−1
  4. Final sample59

    One arithmetic-verification CSV per participant, public on OSF

  5. Planned participant-trials15,104

    59 participants × 256 trials (realised set: 129 true, 127 false)

  6. Preregistered minimum to proceed≥ 50 · ≥ 10,000

    At least 50 participants with ≥ 200 valid trials each, and ≥ 10,000 valid participant-trials after exclusions

Recruitment, exclusions and planned trial counts as reported in the source paper, plus the preregistered minimum data thresholds this reanalysis must meet before proceeding. These are published facts about the dataset, not results of this study.

Privacy handling is fixed before ingestion: the participant date field is dropped immediately; survey, ACT/SAT, health, admissions and demographic fields are never used; raw files stay in a restricted analysis workspace; and only aggregate tables and reproducible code are published — no participant-level rows, even under pseudonyms.

4Preregistered hypothesis

Primary hypothesis: naively sampled forms containing more large-operand, subtraction, or otherwise slower and more error-prone items will produce fewer correct responses and greater within-person form-to-form score variation than forms balanced using item difficulty estimated only from training participants.

This is a preregisterable hypothesis, not a finding. A null or contrary result must be published. Balanced forms may fail to reduce variance if:

  • individual item-difficulty rankings differ substantially across people;
  • the item pool is too small;
  • most variation is caused by person-level inconsistency rather than item mix;
  • the balancing variables omit important sources of difficulty;
  • the 30-second simulation is insensitive to the available difficulty range.

5Simulating thousands of 30-second forms

The simulation replays real observed behaviour into a time budget. In plain English:

  1. Learn which items tend to be slower or more error-prone using training participants only.
  2. Build naive and balanced item sequences — never using the evaluation participant's own data to calibrate difficulty.
  3. Replay each held-out participant's observed item times and accuracy into a 30-second budget: an item counts only if its response finishes in time; correct completed items score one point; incorrect completed items consume time and score zero.
  4. Repeat at least 1,000 naive and 1,000 balanced forms per participant per evaluation repeat, until Monte Carlo error is negligible relative to participant-bootstrap uncertainty.
Abbildung 3How a question mix consumes a 30-second budget
Faster item · 900 msSlower item · 2400 ms
23 items fit in 30 s

Illustrative mechanism only, with hypothetical durations. Whether real item mixes move scores this way — and whether balanced forms reduce the variation — is exactly what the frozen analysis will test. No expected scores or accuracy values are shown because none have been estimated yet.

A mechanism demonstration of the replay simulation: each block is one item, and an item counts only if its response finishes inside the budget. Drag the slider to change the share of slower items and watch how many items fit in 30 seconds. The two durations are hypothetical round numbers chosen to illustrate the mechanism — they are not measured response times, model estimates or study results.

Two naive comparators are reported: an unconstrained random draw from the realised item pool, and a blueprint-matched draw that preserves broad source-task proportions (operation, truth status) without balancing calibrated difficulty. The blueprint-matched version is the stronger comparator: it asks whether balancing predicted difficulty adds value beyond simple content quotas. Balanced forms match a target blueprint across operation, truth status, problem-size bins, predicted response-time bins and predicted error-probability bins — the stratified-assignment logic studied in randomly-equivalent-forms research.5

Primary estimand: for each held-out participant, the standard deviation of correct-count scores across naive forms versus across balanced forms — reported as absolute SD difference, SD ratio, variance ratio and percent variance reduction. The term “reliability improvement” will not be used unless a recognised reliability coefficient is actually estimated and justified.

Secondary outcomes include mean correct count, mean accuracy, attempted items, the 5th–95th percentile score range, the probability that two random forms differ by at least 1, 2 or 3 correct answers, the speed–accuracy frontier, and sensitivity to 10-, 20-, 45- and 60-second limits.

6Analysis plan and safeguards

Response times are modelled with participant- and item-aware mixed models on the log scale, with operation, truth status, a problem-size spline, answer distance, trial position and prespecified interactions; accuracy uses a mixed-effects logistic model with the same structure. Estimated differences and uncertainty intervals are reported rather than p-values alone.

  • Item difficulty is two-dimensional — every item gets a cross-validated predicted response time and a predicted error probability; form construction preserves both rather than collapsing to one composite.
  • Cross-validation splits by participant, never by trial — repeated five-fold participant-level cross-validation with published seed handling; a held-out participant's own performance never shapes the difficulty used to build their forms.
  • Prespecified exclusions — the primary analysis keeps finite positive response times at or below 4 seconds (correct and incorrect, because both consume sprint time); a full sensitivity analysis repeats everything up to the original task timeout.
  • Uncertainty — participant-level nonparametric bootstrap (2,000 samples, preserving participant clustering), with 95% intervals and full distributions where possible.
  • Robustness — early-trial exclusions, addition-only and subtraction-only, true-only and false-only, mean/median/robust summaries, and hypothetical per-item transition-time overheads of 0–500 ms, labelled as hypothetical interface overhead rather than measured latency.

7What we will and will not claim

The claims boundaries are fixed now, before results exist, in line with good testing practice.6 After analysis, with numbers attached, it will be safe to say whether question mix changed simulated 30-second scores in this dataset, whether balanced forms reduced within-person form variation, and whether the result survived the prespecified sensitivity analyses.

8What this means for a 30-second Altamens score

The practical lesson is not that a 30-second score is meaningless. It is that it should be interpreted as an informal performance snapshot. Repeating the test under similar conditions, and using well-balanced item forms, can make comparisons more informative. The source task differs from Altamens, so its numerical estimates will not be transferred directly as Altamens norms or error bands.

9Limitations

  1. The sample contains only 59 US undergraduates aged 18–23.
  2. The source task is true/false verification; Altamens uses four answer choices.
  3. Source response times do not include measured Altamens device, rendering, touch or transition latency.
  4. The 30-second forms are simulations, not prospectively administered alternate forms.
  5. Each source item was observed within a long session, so order, practice, fatigue and context may affect responses.
  6. The 4-second cutoff and other preprocessing choices can affect estimates.
  7. Item difficulty can vary between people; a group-calibrated “hard” item is not equally hard for everyone.
  8. The item pool is limited to the source study's addition/subtraction design.
  9. The analysis cannot produce national, age, education, clinical or IQ norms.
  10. It cannot show that training improves processing speed, learning, driving, work performance or everyday functioning.
  11. It does not validate the Altamens test as a medical, educational or psychometric instrument.
  12. The originality check behind this protocol was a targeted preliminary search, not a systematic review.

Quellen

  1. 1.Bye JK, Harsch RM, Varma S Decoding Fact Fluency and Strategy Flexibility in Solving One-Step Algebra Problems: An Individual Differences Analysis. Journal of Numerical Cognition, 8(2), 281–294, 2022. Source paper for the arithmetic-verification dataset. https://doi.org/10.5964/jnc.7093
  2. 2.Bye JK, Harsch RM, Varma S Open Science Framework project: data, codebooks and analysis files. OSF, 2022. CC BY 4.0; raw files at osf.io/j79a6, codebooks at osf.io/kgd4c. https://osf.io/vx2rh/
  3. 3.Ashcraft MH, Guillaume MM Mathematical Cognition and the Problem Size Effect. Psychology of Learning and Motivation, Vol. 51, 121–151, 2009. https://doi.org/10.1016/S0079-7421(09)51004-3
  4. 4.Faulkenberry TJ A Single-Boundary Accumulator Model of Response Times in an Addition Verification Task. Frontiers in Psychology 8:1225, 2017. https://doi.org/10.3389/fpsyg.2017.01225
  5. 5.Liao C-W, Livingston SA Examining an Alternative to Score Equating: A Randomly Equivalent Forms Approach. ETS Research Report RR-08-14, 2008. Stratified random assignment on content and predicted difficulty. https://doi.org/10.1002/j.2333-8504.2008.tb02100.x
  6. 6.American Educational Research Association, American Psychological Association, National Council on Measurement in Education Standards for Educational and Psychological Testing. AERA, 2014. https://www.testingstandards.net/