Altamens

Altamens Research · Open-data reanalysis

How Stable Is Your Digit Span?

A 1,433-person test of scoring rules and immediate repeatability: the simplest all-trial score was the most repeatable, and our prespecified partial-credit score was the least — the opposite of what we predicted.

Altamens ResearchUpdated August 12, 202613 min read

Abstract

A digit-span result is usually reported as a single number, but that number depends on how trials are converted into a score as well as on the person being tested. We reanalysed the public Music Ensemble digit-span trial file (CC BY 4.0), in which adults completed a visual forward digit-span task twice in immediate succession using different fixed sequence sets. After prespecified cleaning, 38,054 substantive trials from 1,433 participants entered the primary analysis. Three scoring rules were defined before outcomes were inspected: longest passed span, total fully correct sequences, and serial-position partial credit restricted to sequence lengths attempted in both blocks. The primary estimand was ICC(2,1), a two-way random-effects, absolute-agreement, single-measure intraclass correlation, with 5,000 participant-bootstrap resamples.

Fully correct sequences were the most repeatable (ICC 0.611, 95% CI 0.566–0.651), followed by longest passed span (0.535, 0.478–0.586). Common-length partial credit was near zero (0.050, −0.009–0.114), a difference of −0.485 versus longest span (−0.556 to −0.407). Using all attempted lengths raised partial credit to 0.209 (0.145–0.270) but left it far below both simpler scores, and an edit-distance variant behaved almost identically. The prespecified hypothesis that partial credit would improve repeatability was therefore contradicted. A score can carry more numerical detail without carrying more stable between-person signal.

Key findings

  • In 1,433 adults who completed two consecutive forward digit-span blocks, the total number of fully correct sequences was the most repeatable score: ICC(2,1) = 0.611 (95% bootstrap CI 0.566–0.651).
  • Conventional longest-passed span was moderately repeatable: ICC(2,1) = 0.535 (0.478–0.586). The fully-correct count beat it by +0.076 (joint bootstrap CI +0.052 to +0.102).
  • The prespecified serial-position partial-credit score, computed over lengths attempted in both blocks, was the worst: ICC(2,1) = 0.050 (−0.009–0.114), or −0.485 versus longest span (−0.556 to −0.407). Our hypothesis was contradicted, not merely unsupported.
  • Only 33.0% of participants recorded the identical longest span twice; 45.8% moved by exactly one level and 21.2% by two or more. Mean absolute change was 0.960 span levels.
  • Block 2 was slightly higher on every score (longest span +0.199 levels, 95% CI +0.130 to +0.267), but block order is confounded with sequence set, so no cause can be identified.
  • This measures immediate same-session repeatability in adults aged 18–30. It is not a week-to-week stability estimate, a clinical threshold, a population norm, or evidence that training improves memory.

1The measurement question

Digit span looks like the simplest test in cognitive psychology. A sequence of digits is presented, you reproduce it in order, and the sequences get longer until the task becomes too hard. The result is usually reported as a single fact: “my digit span is seven.”

But that number does not come from the person alone. It also comes from the rule used to turn a set of trials into a score.

Suppose the target sequence is 5 2 9 4 7 1 and somebody responds 5 2 9 4 8 1. Five of six positions are correct. A strict pass/fail rule treats that response exactly like a completely wrong one. A partial-credit rule keeps the information that five positions were right.

Our research question was narrow and answerable: if the same person completes two consecutive versions of a forward digit-span task, which scoring rule produces the most repeatable result?

Our prespecified hypothesis was that partial credit would smooth the abrupt thresholds created by maximum-span scoring and therefore produce higher block-to-block reliability — an expectation consistent with the general argument for partial-credit span scoring in the working-memory literature.6 It did not. The partial-credit score was by far the least repeatable of the three.

2The task, and what “repeatability” means here

Participants completed a visual forward digit-span task. Digits were displayed one at a time. The task began with short sequences and presented two trials at each sequence length. If at least one of those two trials was recalled completely correctly, the sequence length increased by one. The task stopped after two incorrect trials at the same length.

Crucially, participants completed the whole task twice in immediate succession, using different fixed sequence sets. That gives two measurements per person under highly similar conditions, which is exactly what a repeatability question needs.

The adaptive stopping rule matters for the analysis, because it creates structural missingness: harder sequence lengths are never presented once a participant stops. A participant who does better in one block automatically meets more difficult sequences in that block, so any score computed over “all trials the person saw” is partly a score computed over a harder test. Digit-span methodology work has long noted that administration and stopping details shape the resulting score.7

3The dataset — and how to download it

We used the public Music Ensemble dataset: a standardized in-lab battery of musicianship, cognition and personality measures collected from 1,438 adults aged 18–30 across 35 research units in 16 countries.1 The project is released openly on OSF under CC BY 4.0.

Prespecified cleaning removed the familiarization block (block 0, 2,866 rows) and 50 substantive trials whose response field contained an empty array. There were zero missing participant/block/trial identifiers and zero duplicate participant × block × trial records. Timestamps and experimenter identifiers were discarded before analysis, and no country, site or demographic group is ranked anywhere in this report.

Figure 1Inclusion flow: from the raw OSF file to the primary dataset
  1. Raw digit-span trial rows on OSF40,970

    data/raw-after-selection/digit_span.csv · 16 variables · one familiarization block plus two substantive blocks

  2. Removed: familiarization block (block 0)−2,866

    Practice trials, excluded from every score

  3. Missing identifiers and duplicate records0

    No missing participant/block/trial IDs and no duplicate participant × block × trial rows were found

  4. Removed: empty responses−50

    Substantive trials whose response field was an empty array; a sensitivity analysis retains them as zero-credit attempts

  5. Primary substantive trials38,054

    1,433 participants, every one with identifiable observations from both substantive blocks

  6. Protocol-defined minimum to proceed — met≥ 1,200

    At least 1,200 participants with two identifiable blocks and parsable stimulus/response strings

Every exclusion applied to the public digit-span trial file, in the prespecified order. The protocol required at least 1,200 participants with two identifiable blocks before outcomes could be examined; 1,433 participants qualified, so the analysis proceeded.

We did not take the dataset's correctness field on trust. Exact accuracy was independently reconstructed from the raw stimulus and the raw participant response, then compared with the supplied flag: source/recomputed disagreements = 0. The published correctness field matched exact stimulus-response comparison perfectly, which is a good sign for the file's integrity and means our scores are built on verified trial outcomes.

4Three prespecified ways of scoring the same test

All three scores were derived independently within each block, from the same cleaned trials, using definitions written down before any outcome was inspected.

  1. Longest passed span — the longest sequence length at which at least one of the two trials was recalled perfectly. This is the conventional, intuitive result (“span = 7”), but it compresses an entire test into one threshold: a single successful trial can move somebody a whole level.
  2. Fully correct sequences — the total count of completely correct trials in the block. This is also the block scoring used by the original Music Ensemble study, and it uses the full set of pass/fail outcomes while staying trivially easy to explain.
  3. Serial-position partial credit — the proportion of presented digit positions recalled with the correct digit in the correct location. A six-digit sequence with five correctly positioned digits scores 5 / 6 = 83.3%. In the primary analysis this was computed only over sequence lengths the participant attempted in both blocks, precisely to stop adaptive stopping from comparing an easier test with a harder one.

5How repeatability was estimated

The primary reliability statistic was ICC(2,1): a two-way random-effects, absolute-agreement, single-measure intraclass correlation.3,4

Absolute agreement is the important choice. An ordinary correlation can stay high even when everybody's score shifts systematically between blocks, because it only cares about preserved ranking. ICC(2,1) penalises that kind of disagreement, which is exactly what you want when the practical question is “would this person get the same number again?”

For every scoring rule we also computed Spearman rank stability, mean absolute block difference, within-person standard deviation, the average Block 2 − Block 1 difference, and Bland–Altman bias and 95% limits of agreement.5 Confidence intervals came from 5,000 participant-bootstrap resamples, and scoring-rule contrasts were computed within the same bootstrap draws so that the difference between two ICCs carries its own interval. Because participants were recruited at 35 research units, every headline coefficient was additionally re-estimated under a site-cluster bootstrap.

6Results: the simplest all-trial score won

Primary result: the count of fully correct sequences was the most repeatable score (ICC 0.611, 95% CI 0.566–0.651), longest passed span was moderately repeatable (0.535, 0.478–0.586), and the prespecified common-length partial-credit score was effectively unrepeatable (0.050, −0.009–0.114).

Figure 2Immediate repeatability by scoring rule
Fully correct sequencesprimary0.61 (0.570.65)
Longest passed span0.54 (0.480.59)
Serial-position partial credit (common lengths)0.05 (-0.010.11)
-0.010.170.350.530.72

Two-way random-effects, absolute-agreement, single-measure ICC. Hover a row for focus. The vertical grey line marks zero.

ICC(2,1) with 95% participant-bootstrap intervals (5,000 resamples), N = 1,433. Higher is more repeatable. Spearman rank stability followed the same ordering: 0.605, 0.535 and 0.064 respectively. The partial-credit interval crosses zero, so this analysis cannot distinguish its between-person signal from none at all.

The margin between the two simple scores was small but consistent: the fully-correct count exceeded longest passed span by +0.076, with a joint participant-bootstrap 95% CI of +0.052 to +0.102. That is worth stating plainly, because the better score was not the more sophisticated one. It was the source task's ordinary count of correctly recalled sequences.

The partial-credit result went the other way, and hard. Partial credit minus longest span was −0.485 (95% bootstrap CI −0.556 to −0.407). The interval sits far below zero, so the prespecified hypothesis was contradicted rather than merely unsupported.

Average performance was slightly higher in the second block on all three scores. Longest span rose from 6.548 ± 1.373 to 6.747 ± 1.401 (mean change +0.199, 95% CI +0.130 to +0.267). Fully correct sequences rose from 9.890 ± 2.443 to 10.273 ± 2.442 (+0.382, 95% CI +0.272 to +0.496). Partial credit rose from 0.876 ± 0.070 to 0.894 ± 0.074 (+0.0179, 95% CI +0.0128 to +0.0229).

Figure 3Block 1 versus Block 2 means
Block 1Block 2

Longest passed span (levels)

6.5
6.7

Fully correct sequences (count)

9.9
10.3

Mean scores in the first and second consecutive block for the two count-based rules (span levels and sequence counts). Partial credit shifted on its own scale, from 0.876 to 0.894. Because the blocks were always administered in the same order with different fixed sequence sets, block order is confounded with item set and no cause can be identified.

7Why more detailed scoring did not help

A partial-credit score preserves more information at the level of the individual trial. That is not the same thing as preserving stable differences between people, and this dataset separates the two ideas cleanly.

Restricted to lengths attempted in both blocks, the partial score concentrated most participants at high positional accuracy — mean 0.876 in Block 1 and 0.894 in Block 2, with standard deviations of only 0.070 and 0.074. The between-person spread available to be reproduced was therefore small. At the same time, the exact positions in which errors fell varied considerably between the two sequence sets. Most of the extra numerical resolution was error position, and error position did not repeat.

There is a second, structural reason the common-length restriction hurt. It deliberately throws away the lengths a participant reached in only one block — which is precisely where much of the between-person difference in ability lives. That is why relaxing the restriction improves the coefficient substantially, as the next section shows, without ever bringing it close to the simple scores.

8How much a span actually moved

Longest span is the easiest result for a person to understand, so its instability is the most practically relevant finding here. Across 1,433 participants tested twice within one session, only a third reproduced the same number.

Figure 4Change in longest passed span between two consecutive blocks
Identical span in both blocksprimary33.00%
Changed by exactly one level45.80%
Changed by two or more levels21.20%
0%13%25%38%50%

Descriptive percentages from this same-session experiment. Not norms, and not thresholds for interpreting any individual result.

Percentage of the 1,433 participants by size of block-to-block change in longest passed span. Direction was 38.7% improved, 28.3% declined and 33.0% unchanged. The mean absolute change was 0.960 span levels — in practical terms, about one digit-span level.

This means a change such as 7 → 8 is entirely compatible with ordinary immediate repeat variability. It should not be read as evidence that the person's underlying memory capacity improved.

Bland–Altman limits of agreement make the same point in units of each score.5 For longest span, bias was +0.199 with 95% limits of −2.408 to +2.805 span levels. For fully correct sequences, bias was +0.382 with limits of −3.802 to +4.567 sequences. For partial credit, bias was +0.018 with limits of −0.177 to +0.212. Averages hide large individual movements: the mean shift is small, and the individual range is not.

9Sensitivity and validation checks

The central result should not depend on one arbitrary implementation of partial scoring, so we ran the prespecified alternatives: an edit-distance credit instead of strict positional credit, and all attempted lengths instead of common lengths only.

Figure 5Partial-credit variants against the simple scores
Fully correct sequences (reference)primary0.61 (0.570.65)
Longest passed span (reference)0.54 (0.480.59)
Partial credit, all attempted lengths0.21 (0.140.27)
Edit-distance credit, all attempted lengths0.21 (0.140.27)
Partial credit, common lengths (primary)0.05 (-0.010.11)
Edit-distance credit, common lengths0.05 (-0.010.11)
-0.010.170.350.530.72

The vertical grey line marks zero. Both common-length intervals include zero.

ICC(2,1) with 95% participant-bootstrap intervals for four partial-credit implementations, shown against the two simple rules. Using all attempted lengths roughly quadruples partial-credit reliability, and the edit-distance variant is almost indistinguishable from strict positional credit — but no partial-credit version approaches longest span, let alone the fully-correct count.

Even in its best specification, partial credit remained far behind: all-attempted partial credit minus longest span was −0.327 (95% CI −0.400 to −0.248). The conclusion does not depend on which partial-credit definition we use.

  • Empty responses — the 50 empty-array trials were excluded in the primary analysis. Retaining them instead as genuine zero-credit attempts moved the common-length partial-credit ICC from 0.0498 to 0.0503. The conclusion is unchanged, so the exclusion rule is not driving the null.
  • Independent correctness recomputation — exact accuracy rebuilt from raw stimulus and response disagreed with the dataset's supplied flag on 0 trials.
  • Site-cluster bootstrap — resampling whole research units rather than individuals gave 95% ICC intervals of 0.479–0.582 for longest span, 0.572–0.646 for fully correct sequences and −0.019–0.120 for partial credit. The ordering survives respecting the dataset's clustered structure.
  • Group labels — all scoring was computed independently of musician/nonmusician status, site and country, and no group comparison or ranking was performed.

10What a digit-span score should be taken to mean

The most useful reading of these findings is that a digit-span result is not simply a fact about a person's memory. It is jointly produced by the participant, the particular sequences presented, the adaptive stopping rule and the scoring system. Change the last of those alone and immediate repeatability ranged from 0.611 to 0.050 in the very same trials.

Maximum span compresses a whole test into one threshold. Partial credit retained more trial detail but, here, mostly retained detail that did not survive a change of sequence set. The fully-correct sequence count was the best compromise available in this dataset: it uses the complete set of pass/fail trial outcomes while avoiding the single-threshold behaviour of maximum span.

The communication implication is just as important as the scoring one. A result is best described as performance under the current testing conditions, not as a fixed measurement of somebody's cognitive capacity.

11What we do and do not claim

The claims boundaries were fixed before results existed, in line with good testing practice.8 With the numbers attached, the supported claims are:

  • In this dataset, three defensible scoring rules applied to identical trials produced very different immediate repeatability.
  • The count of fully correct sequences was the most repeatable score tested (ICC 0.611), modestly but consistently ahead of longest passed span (0.535).
  • The prespecified common-length serial-position partial-credit score was not more repeatable than longest span; it was dramatically worse (0.050), and the hypothesis was contradicted.
  • Only about a third of participants reproduced the same longest span in two consecutive blocks, with a mean absolute movement near one span level.

12Limitations

  1. Both blocks occurred in the same laboratory session, so this estimates immediate repeatability, not stability over days, weeks or months.
  2. The two blocks used different fixed sequence sets and were always administered in the same order, so block order is confounded with item set.
  3. Adaptive stopping creates structural missingness: harder trials are never presented after a participant reaches their stopping point, which is exactly what the common-length restriction tried — imperfectly — to handle.
  4. The common-length restriction discards the lengths reached in only one block, and that choice is itself part of why the primary partial-credit coefficient is so low; the all-attempted variant is reported alongside it.
  5. The sample is adults aged 18–30 recruited for a musician/nonmusician research programme, not a population-representative normative sample, and residual site differences may remain despite the cluster bootstrap.
  6. The source task is a visual forward digit span starting at two digits; the Altamens Working Memory Test differs in interface, device and starting length, so these coefficients are not Altamens reliability estimates.
  7. Forward digit span is a narrow component of short-term memory and should not be equated with working memory as a whole or with general cognitive ability.
  8. Fifty empty-response trials required a handling rule; both rules were run and agreed, but the choice was still ours.
  9. The study was protocol-guided, with scoring rules and estimands frozen before outcomes were inspected, but it was not formally preregistered in a public registry before execution.

References

  1. 1.Talamini F, Grassi M, Altoè G, et al. Music Ensemble: a large dataset on musicianship, cognition, and personality in musicians and nonmusicians. Scientific Data 13, 473, 2026. Dataset descriptor, cited for methods only (article is CC BY-NC-ND). https://doi.org/10.1038/s41597-026-06654-0
  2. 2.Music Ensemble consortium Music Ensemble — Open Science Framework project (trial-level data). OSF · DOI 10.17605/OSF.IO/Y97T3, 2026. CC BY 4.0; digit-span trial file at data/raw-after-selection/digit_span.csv. https://osf.io/y97t3/
  3. 3.Shrout PE, Fleiss JL Intraclass correlations: uses in assessing rater reliability. Psychological Bulletin 86(2), 420–428, 1979. Definition of the ICC(2,1) form used here. https://doi.org/10.1037/0033-2909.86.2.420
  4. 4.Koo TK, Li MY A Guideline of Selecting and Reporting Intraclass Correlation Coefficients for Reliability Research. Journal of Chiropractic Medicine 15(2), 155–163, 2016. https://doi.org/10.1016/j.jcm.2016.02.012
  5. 5.Bland JM, Altman DG Statistical methods for assessing agreement between two methods of clinical measurement. The Lancet 327(8476), 307–310, 1986. https://doi.org/10.1016/S0140-6736(86)90837-8
  6. 6.Conway ARA, Kane MJ, Bunting MF, Hambrick DZ, Wilhelm O, Engle RW Working memory span tasks: A methodological review and user's guide. Psychonomic Bulletin & Review 12(5), 769–786, 2005. Reviews all-or-nothing versus partial-credit span scoring. https://doi.org/10.3758/BF03196772
  7. 7.Woods DL, Kishiyama MM, Yund EW, et al. Improving digit span assessment of short-term verbal memory. Journal of Clinical and Experimental Neuropsychology 33(1), 101–111, 2011. https://doi.org/10.1080/13803390903424355
  8. 8.American Educational Research Association, American Psychological Association, National Council on Measurement in Education Standards for Educational and Psychological Testing. AERA, 2014. https://www.testingstandards.net/