What the research actually shows

This page is about the evidence for training attention, including the parts of it that do not support training attention. It exists because we would rather write down where the evidence is thin than write around it.

Does training attention work?

Train a task and you get faster and more accurate at that task. That much is well established, in the cognitive training literature and in the meditation one.

Whether it reaches beyond the trained task is a separate question, and it is the one the field has spent decades failing to settle.

What do the largest reviews find?

Simons et al. (2016) reviewed the whole field and found extensive evidence for change on the trained task, less for tasks that resemble it, and little for everyday life.

One finding in that review deserves more attention than it gets: the size of the reported effects was inversely related to the quality of the study design. The weaker the design, the larger the apparent effect.

Sala et al. (2019) went further and pooled meta-analyses rather than studies. Once placebo effects and publication bias are accounted for, the far-transfer effect and its variance both come out at zero. Their summary of the field: "a field in search of a phenomenon."

What happens when the control group also expects to benefit?

The effects shrink, and often they disappear entirely. A control group that does something plausible and expects it to help is the hardest test in this field, and it is the test most of the studies did not run.

Sala et al. (2019) is the clearest statement of what survives it. On tasks that resemble the trained one, something real. Beyond that, nothing that can be told apart from zero.

Is meditation different?

It has the more direct evidence for attention specifically, and the shape of that evidence is worth being exact about.

Jha et al. (2015) compared mindfulness training that emphasised doing the practice against training that emphasised explaining it, across a stretch of high demand. The practice-focused group held steady. The explanation-focused group and the untrained group did not.

Read that carefully, because it is the part that gets rewritten in the retelling: nobody climbed. The effect is protective — attention holding under load rather than rising. That is the honest shape of most of what exists here, and it is not what this industry sells.

The meditation literature has its own critics on the same grounds. Van Dam et al. (2018) catalogued definitional confusion, weak measurement, and effects that shrink against active controls.

How do you measure sustained attention?

The task we use comes from Esterman et al. (2013). Images fade into one another over about eight tenths of a second instead of appearing. Something that appears suddenly takes your attention for you; something that fades in does not, so what is left is attention you are supplying yourself.

Nine in ten images are one category and you respond to every one. One in ten is the other category and you do nothing. After a minute of pressing, not pressing is the hard part.

The load is in the withholding, not in the seeing. Telling the two categories apart is meant to be easy, and a task that made the seeing harder would be measuring eyesight.

How much can one person read into their own score?

Less than most apps imply. Hedge, Powell and Sumner (2018) named this the reliability paradox: a task can separate groups of people cleanly and still track one individual poorly, because the very consistency that makes it robust leaves little room for one person to vary. Across seven classic tasks they found reliabilities ranging from zero to .82.

The consequences are practical rather than academic. A block has to run the same length every time, because blocks of different lengths are not comparable. The first sessions cannot count, because everyone climbs at the start through familiarity alone. And a single pair of sessions is not a result.

Does difficulty have to adapt to the person?

Apparently not. von Bastian and Eschen (2016) put 130 adults through 20 sessions over four weeks and compared difficulty that adapted to the participant, difficulty that varied at random, and difficulty the participant chose. The gains on the trained task were clear. The differences between the three conditions were not there.

Their conclusion was that exposure to varying levels of difficulty is sufficient on its own — which makes adaptivity a reasonable way of putting a task somewhere sensible for somebody, rather than the active ingredient it is usually sold as.

Their mechanism is where our thresholds come from: difficulty adjusted once per session, up at 80 per cent accuracy or above, down below 60 per cent. Those numbers are theirs rather than ours, and that is the entire point of using them. A difficulty rule quietly tuned until more people succeed has stopped measuring anything.

What about phones and short-form video?

This is the claim everyone in this category wants to make, and there is not much underneath it. There is little direct trial evidence in heavy short-form-video users specifically.

The idea is plausible. It is not established. So we do not say it.

Then why build any of this?

Because the parameters are real even where the promises are not. A task built to a published design, at a length that never varies, with thresholds taken from a paper instead of chosen, produces a number that means something. Very little in this category is built that way.

Whetstone is our attempt at it. We have run no trial of Whetstone. Measuring something carefully is not the same as having shown that it works, and we will not dress the first up as the second.

References

Every study below informed how something in this app is built, or what is said above about what the research found. None of them is a study of Whetstone.

  1. Esterman, M., Noonan, S. K., Rosenberg, M., & DeGutis, J. (2013). In the zone or zoning out? Tracking behavioral and neural fluctuations during sustained attention. Cerebral Cortex, 23(11), 2712–2723.The task the check-in is built from, and the reason its images fade rather than appear.
  2. Hedge, C., Powell, G., & Sumner, P. (2018). The reliability paradox: Why robust cognitive tasks do not produce reliable individual differences. Behavior Research Methods, 50(3), 1166–1186.Why a task can be robust across groups and still say little about one person, and why nothing is read into a single pair of sessions.
  3. von Bastian, C. C., & Eschen, A. (2016). Does working memory training have to be adaptive? Psychological Research, 80(2), 181–194.Where the 80 per cent and 60 per cent thresholds come from, and the finding that adaptive difficulty was no different from difficulty that merely varied.
  4. Jha, A. P., Morrison, A. B., Dainer-Best, J., Parker, S., Rostrup, N., & Stanley, E. A. (2015). Minds "at attention": Mindfulness training curbs attentional lapses in military cohorts. PLOS ONE, 10(2), e0116889.The clearest attention-specific meditation finding, and the reason it is described here as protective rather than as a gain.
  5. Simons, D. J., Boot, W. R., Charness, N., Gathercole, S. E., Chabris, C. F., Hambrick, D. Z., & Stine-Morrow, E. A. L. (2016). Do "brain-training" programs work? Psychological Science in the Public Interest, 17(3), 103–186.The field-wide review, and the finding that weaker study designs produced larger apparent effects.
  6. Sala, G., Aksayli, N. D., Tatlidil, K. S., Tatsumi, T., Gondo, Y., & Gobet, F. (2019). Near and far transfer in cognitive training: A second-order meta-analysis. Collabra: Psychology, 5(1), 18.What is left of far transfer once placebo effects and publication bias are accounted for, measured across the whole field at once.
  7. Van Dam, N. T., van Vugt, M. K., Vago, D. R., Schmalzl, L., Saron, C. D., Olendzki, A., et al. (2018). Mind the hype: A critical evaluation and prescriptive agenda for research on mindfulness and meditation. Perspectives on Psychological Science, 13(1), 36–61.Why the same caution applies to the meditation literature as to the cognitive training one.

Last written .