EVIDENCE REVIEW · MOTIVATION AND BEHAVIOUR DESIGN

Flow, dopamine, and stop signals — what the evidence actually supports

Macheng Shen × agent workflow · 2026-08-02 · cognitive state: speculative for the synthesis, survived for the meta-analytic results it rests on

Why this review exists. Almost every claim used to justify gamified interfaces, streaks, points, adaptive difficulty and screen-time limits is either (a) a well-replicated result being applied in the wrong direction, or (b) a plausible deduction with no controlled test behind it. Sorting those two piles is the whole content here.

Four findings worth the click, if you read nothing else.
1. The same meta-analysis that shows tangible rewards undermine intrinsic motivation (d ≈ −0.28 to −0.40) shows positive informational feedback enhancing it (d ≈ +0.31 to +0.33). One sign flip separates the two, and it is the most actionable number in this field.
2. Adaptive difficulty is the most expensive flow condition to build and has the weakest and most contested evidence of any of them.
3. Heart-rate variability from consumer hardware cannot tell flow from mild stress: Apple Watch SDNN against a chest strap gives roughly 29% mean absolute percentage error.
4. The largest coercive stop-signal interventions ever run — a six-hour national gaming curfew, and mandated playtime quotas — produced measured effects near the floor of measurement, and the most-cited study claiming quotas fail has a design flaw that means it could not have detected success either.

中文摘要

本文是一份证据审查,对象是「游戏化界面 / 心流设计 / 防沉迷」这一堆常被引用的主张。核心结论:(1) 多巴胺编码的是「想要」和「意外」,不是「快乐」和「奖励」 —— 多巴胺耗竭到 99% 的大鼠仍保留正常的享乐反应;完全被预测到的奖励不产生相位反应,因此「完全透明可控」与「多巴胺上有冲击力」在机制上是互斥的。(2) 可变比率(抽卡)最上瘾,而且它的收入结构会自己招供:开箱消费前 5%(月消费 >100 美元)贡献约一半收入,其中约三分之一符合问题赌博标准;消费与问题赌博相关 ρ=0.34,与收入相关 ρ=0.02 且不显著 —— 它抽的是脆弱性的租,不是财富的租。(3) 动态难度调节是最贵、证据最弱的一条:最佳受控实验(n=221)发现表现显著提升,而心流、愉悦、挑战、焦虑全部无显著差异;但 2024 年的系统综述报告多数研究得到正面结果 —— 诚实的读法是异质性加发表偏倚,不是已判死。(4) 别用生理信号测心流:消费级 HRV 的 SDNN 误差约 29%。(5) 停止信号里,弹窗是表演,强制不可用才有效:赌博场景弹窗只让不到 1% 的人停下,加了规范性与自评内容后从 0.67% 升到 1.39%;而 60 分钟后强制休息 15 分钟能显著延长后续自愿停顿,90 秒近乎无用。韩国「灰姑娘法」六小时全国宵禁,实测每天减少约 3.65 分钟上网,两年后归零,对网络成瘾与睡眠时长无影响,该法 2021 年废止。全文引用 30+ 条文献,逐条核实;并公开更正我们自己初稿里的三处引用错误。

1 · Dopamine: wanting is not liking, and surprise is not reward

Wanting ≠ liking. Rats with up to 99% depletion of accumbens and neostriatal dopamine retain normal hedonic taste reactions, normal learning of new hedonic values, and normal pharmacological enhancement of pleasure [11]. Dopamine is necessary for wanting an incentive, not for liking it. This makes "dopamine hit = pleasure" a category error at the level of design, not just of vocabulary: sensitised wanting with flat liking is the signature shape of addiction — want more, enjoy no more. Any mechanism that makes someone compulsively check but leaves them unsatisfied is manufacturing exactly that gap.

Honesty note: Berridge's strong position — that dopamine has essentially no role in reward learning — is a minority view, and he says so himself in the debate paper [12]. Most of the field holds a dual-role account. Nothing below depends on the strong version.

Dopamine codes prediction error. Midbrain dopamine neurons fire above baseline for better-than-predicted outcomes, remain at baseline for fully predicted rewards, and dip below it for omitted ones; during learning the response migrates backwards from the reward onto the cue that predicts it [13][14].

The corollary designers do not want to hear. A fully transparent, fully predictable reward schedule produces, by construction, no phasic dopamine response. "Transparent and controllable" and "dopaminergically punchy" are in genuine tension, and a design claiming both is usually smuggling hidden variance in somewhere.

[our speculation] The escape route we find most plausible is to put the uncertainty in the world model rather than the payout table — let the prediction error come from whether the thing you tried actually worked. That is mechanically coherent and completely untested: nobody has shown that naturally occurring real-world prediction error is dense enough to feel like anything.

2 · Variable ratio: maximally engaging, and the revenue structure confesses

Two independent syntheses link loot-box spending to problem gambling: a meta-analysis finds r = 0.26 across 15 studies (r = 0.37 after trim-and-fill) with gambling symptomatology and r = 0.25 across 7 studies with excessive gaming [15]; a systematic review and meta-synthesis reaches the same qualitative conclusion [16]. A large survey found effect size η² = 0.054 for loot-box spend against problem gambling, versus η² = 0.004 for other in-game purchases [18], and adolescent data show the association strengthens when loot-box contents are time-limited [19].

But the sharpest evidence is the shape of the revenue. Pooling 7,767 purchasers across six open datasets: the top 5% of spenders (over $100/month) generate roughly half of all loot-box revenue; about a third of that group screens positive for problem gambling; spending correlates with problem gambling at ρ = 0.34 — and with income at ρ = 0.02, not significant [17]. The business model extracts rent from vulnerability, not from wealth.

The honest limit: all of this is cross-sectional. The engagement inference is safe; the causal harm inference is not — selection (problem gamblers choosing loot boxes) is not excluded by any current design.

[deduction, not measured] Variable ratio and flow are not two settings of one dial; they are opposed targets. Flow requires clear goals, immediate deterministic feedback, and perceived control. A variable-ratio schedule is definitionally the insertion of uncontrollable randomness between action and feedback, which moves attention from the task to the payout. We rate this at P ≈ 0.60 rather than higher precisely because we could not find a controlled experiment introducing random rewards into a cognitive-output task and measuring flow and output quality. If you know of one, please write.

3 · The most expensive flow condition has the weakest evidence

Of the standard flow conditions — clear goals, immediate feedback, challenge–skill balance, freedom from distraction, sense of control — the one that costs the most to engineer is challenge–skill balance via dynamic difficulty adjustment, because it needs a live skill estimator.

The best-controlled test we found: n = 221, three conditions (DDA, fixed, incremental). The DDA group performed significantly better — and showed no significant difference in enjoyment, flow, challenge, or anxiety [20]. The authors themselves recommend future work using flow as an outcome, because theirs came out null.

A correction to our own draft, and it cuts against the conclusion we liked. Our working note cited a paper as a "review with mixed results" supporting the null. It is not a review: it is a single uncontrolled deep-learning DDA study reporting ~90% high enjoyment and immersion — i.e. evidence pointing the other way [21a]. The actual systematic review of this literature reports that most studies find significant DDA effects on enjoyment, flow, motivation and immersion [21b].

So the honest statement is: heterogeneous results, plausible publication bias, and essentially all of it in action and casual games rather than cognitive work. P(DDA does not reliably improve flow) drops to ≈ 0.45. The build decision nonetheless survives at ≈ 0.75, for a different reason: even under the optimistic reading the effect sits in a different task domain, and a user-facing intensity selector captures most of the value at a fraction of the cost.

The cheaper substitute has better evidence. Self-determination-theory work on games (four studies) finds in-game autonomy and competence satisfaction predict enjoyment, preference, and post-play wellbeing [2]. Handing difficulty selection to the user buys challenge–skill balance without an inference engine — and abrupt automatic adjustment is precisely what erodes the perceived autonomy that the same framework says is doing the work. [our speculation] that the substitution is sufficient; the SDT premise itself is not speculative.

4 · Do not measure flow with physiology

A systematic review of 48 studies using heart-rate variability in educational contexts found HRV has only moderate concurrent validity as a stress measure, unclear validity as an attention measure, and inconsistent correlation with performance [34].

Consumer hardware is worse than the reputation suggests. Apple Watch Series 9 and Ultra 2 against a Polar H10 chest strap, 39 participants, ~300 morning readings: SDNN gives MAPE = 28.9% (95% CI 26.2–31.6), MAE 20.5 ms, limits of agreement roughly −54 to +37 ms — while resting heart rate from the same device gives MAPE 5.9% [33a]. A separate validation against 3-lead ECG found near-perfect agreement for R–R interval (MAPE ≈ 1%) but only moderately acceptable agreement for the derived indices, and only at rest [33b]. Note also that Apple reports SDNN while most consumer devices report RMSSD, so cross-device comparison is invalid before you even get to accuracy.

What to use instead [our speculation]: behavioural proxies that cost nothing and interrupt no one — uninterrupted work-block length, application-switching rate (fragmentation means flow has already broken), edit or commit cadence — plus a validated flow scale [5] administered weekly rather than per session, since a per-session questionnaire is itself an interruption.

Symmetry, because this is where reviews usually cheat: we are confident physiology fails (P ≈ 0.85). We are much less confident these behavioural proxies succeed — they have no validation against a flow criterion either. The defensible position is "both are weak; one is free and non-intrusive."

5 · The one sign flip that matters most

The 128-experiment meta-analysis on extrinsic rewards and intrinsic motivation reports, on free-choice behaviour: engagement-contingent rewards d = −0.40, completion-contingent −0.36, performance-contingent −0.28. In the same analysis, positive informational feedback goes the other way: d = +0.33 free-choice and +0.31 self-reported interest [4]. Task-noncontingent and unexpected rewards were not shown to undermine.

Read that as a design rule: adding reward tokens to an activity someone is already intrinsically motivated by has a measured negative expected value; telling them informatively how it went has a measured positive one. The two interventions look similar on a screen and point in opposite directions.

The gamification literature is consistent with this once you look at what it measured.

A second correction to our own draft. We wrote that a 15-week study showed intrinsic motivation "declining over time" in a gamified environment. That is not what it found. Intrinsic motivation initially decreased and then rose again without exceeding baseline, which the authors attribute to a familiarisation effect; and the same environment supported psychological needs in some students while thwarting them in others [9]. The finding is ambivalence, not monotonic decline, and the ambivalence is more useful than the version we had written down.

6 · Work produces more flow than leisure

The classic experience-sampling result: flow-like states occurred in about 54% of work-time samples versus about 17–18% of leisure samples, while people simultaneously reported higher motivation during leisure [1]. Work already supplies clear goals, feedback, and challenge–skill matching; unstructured leisure supplies none of them.

Two honest qualifications. The sample was 78 adults over one week — small for an ESM study, and the finding is often cited as though it were large. And the mechanism it supports is not "work is good" but "the conditions are what matter, and one context happens to supply them" — which is precisely why importing game structure into work has a worse expected value than importing it into the unstructured side.

7 · Stop signals: what is theatre and what actually works

This is the section where the received wisdom is most wrong in both directions.

7.1 Messages are theatre

In real casino operator data, a pop-up appearing after 1,000 consecutive slot plays caused fewer than 1% of players to stop — 45 of 4,205 qualifying sessions terminated at exactly 1,000, which was nine times the pre-intervention rate and still under one percent [26]. Adding normative and self-appraisal content, across 1.6 million sessions, doubled it: from 0.67% to 1.39% [27]. That is a 98.6% ignore rate for the enhanced version.

7.2 Enforced unavailability works, and the dose matters

A randomised real-world experiment with 21,129 players and 156,989 mandatory breaks triggered after 60 minutes of continuous play compared 90-second, 5-minute and 15-minute forced breaks, with and without personalised feedback. The 15-minute break produced the longest subsequent voluntary pause; personalised feedback added nothing [28a]. An earlier study of 90-second forced session termination on video lottery terminals (n = 7,190) found no significant effect at all [28b].

The brake has to be "you temporarily cannot", not "you are being told" — and the dose has to be long enough that resuming is a decision rather than a reflex.

7.3 Self-control tools work for about three weeks

Systematic review and meta-analysis of digital self-control tools: a significant reduction in unwanted use across seven field experiments (≈0.5 SD), and no long-term effect — because such tools perform self-monitoring without building new habits [29]. Most of the studies ran about 21 days, and most used within-subjects designs that can inflate apparent effectiveness.

7.4 National curfews and quotas: near the measurement floor, and contested

South Korea's "shutdown law" banned under-16s from online games between midnight and 6 a.m. from November 2011; it was upheld by the Constitutional Court in 2014 and abolished in August 2021. A difference-in-differences analysis on ~244,000 adolescents found internet use fell by 3.65 minutes per day in 2012 and 3.20 in 2013, then became non-significant in 2014 and 2015 — with no effect on internet addiction and no effect on sleeping hours [31]. An independent evaluation on the same survey family found mixed-sign, effectively null results, including a sleep increase of about 1.5 minutes [32].

A six-hour nightly national ban bought roughly three and a half minutes a day, and the effect decayed to nothing within two years.

The most mis-cited study in this whole area, and we want to be precise about it. A widely shared paper analysing over seven billion hours of telemetry reports no credible evidence that Chinese playtime mandates reduced heavy gaming, with accounts becoming 1.14× more likely to play heavily afterwards [30]. Three qualifications travel with it and are usually dropped: Honest summary: it is not established that quotas work, and it is equally not established that they fail. Our working draft put P ≈ 0.80 on "hard quotas get routed around". We have revised that to 0.50.

8 · A discriminator we propose, clearly labelled as unvalidated

Putting the above together, here is the test we would apply to any feedback mechanism to ask whether it supports engagement or manufactures compulsion. All three conditions required.

  1. Output-coupled. The event is causally downstream of a real change in the world — one that stays meaningful if you delete the interface entirely.
  2. Schedule-transparent. Given the action, the timing and content of the feedback are predictable. No variable ratio anywhere.
  3. Terminating. A state exists in which there is genuinely nothing more to do today, and it is reached on most days.

This is a synthesis, not a tested construct. It has no validation data behind it, and we are publishing it as a hypothesis to be attacked rather than as a finding. The obvious way to attack it: build two variants differing only in condition (2), and measure output quality rather than session length.

9 · Calibrated confidence

ClaimP(mechanism)P(useful when building)What would change it
Dopamine encodes wanting and prediction error, not pleasure; fully predicted rewards produce no phasic response0.950.45The mechanism is not in doubt. The practice number is low because whether real-world prediction-error density is sufficient to substitute for engineered variance is entirely untested — someone shipping a non-random design that still feels alive after eight weeks would move it
Variable-ratio schedules are associated with problematic use at clinically relevant magnitude0.850.80A longitudinal or instrumented design showing the correlation reverses direction (selection rather than causation). All current evidence is cross-sectional
Variable ratio actively destroys flow (not merely coexists with it)0.600.70A controlled experiment on a cognitive-output task comparing deterministic against variable reward, with flow and output quality as outcomes. None exists — this is a deduction from flow's definition and is priced as one
Dynamic difficulty adjustment does not reliably improve flow0.450.75 (as a build decision)The mechanism is genuinely contested — one strong controlled null against a review reporting mostly positive results. The build decision is robust anyway. Change it if a controlled experiment on cognitive work uses flow as the primary outcome and finds a significant positive effect
Consumer physiological signals cannot validly detect flow0.850.65A published, cross-validated model on desktop cognitive work using consumer hardware that separates flow from mild stress
Message-based stop signals are near-useless; enforced unavailability at ≥15 min works0.850.75Both halves come from gambling populations, where base rates of compulsion are far higher than in ordinary work. A study in a non-clinical population finding materially higher compliance for messages, or finding rebound use after forced breaks
National playtime quotas get routed around0.50Revised down from 0.80. The evidence is contested in both directions; see §7.4. A well-identified study with age data would move this a long way in whichever direction it lands
Adding reward tokens to an already-intrinsically-motivated activity is net negative; informational feedback is net positive0.900.85The best-evidenced claim here — 128 experiments, both signs measured in one analysis. Note the scope limit: task-noncontingent and unexpected rewards were not shown to undermine, so a non-performance-linked nudge remains on the table where the bottleneck is initiation rather than motivation quality

References

Every reference was checked against a publisher record, PubMed, or an author copy, and checked against what it actually reports. Two sources in our working draft were dropped as unusable and one was affirmatively mis-described; those corrections are in the body above rather than hidden here.

  1. Csikszentmihalyi, M. & LeFevre, J. (1989). Optimal experience in work and leisure. Journal of Personality and Social Psychology 56(5):815–822. doi:10.1037/0022-3514.56.5.815 — n = 78 adults, one week.
  2. Ryan, R. M., Rigby, C. S. & Przybylski, A. K. (2006). The motivational pull of video games: a self-determination theory approach. Motivation and Emotion 30(4):344–360. doi:10.1007/s11031-006-9051-8
  3. Player Experience of Need Satisfaction (PENS) — instrument derived from [2]; cite as a measure, not as evidence.
  4. Deci, E. L., Koestner, R. & Ryan, R. M. (1999). A meta-analytic review of experiments examining the effects of extrinsic rewards on intrinsic motivation. Psychological Bulletin 125(6):627–668. doi:10.1037/0033-2909.125.6.627
  5. Rheinberg, F., Vollmeyer, R. & Engeser, S. (2003). Die Erfassung des Flow-Erlebens (Flow Short Scale). In Stiensmeier-Pelster & Rheinberg (eds), Hogrefe, 261–279.
  6. Sailer, M. & Homner, L. (2020). The gamification of learning: a meta-analysis. Educational Psychology Review 32(1):77–112. doi:10.1007/s10648-019-09498-w
  7. Mekler, E. D., Brühlmann, F., Tuch, A. N. & Opwis, K. (2017). Towards understanding the effects of individual gamification elements on intrinsic motivation and performance. Computers in Human Behavior 71:525–534. doi:10.1016/j.chb.2015.08.048
  8. Hanus, M. D. & Fox, J. (2015). Assessing the effects of gamification in the classroom. Computers & Education 80:152–161. doi:10.1016/j.compedu.2014.08.019 — see also the 2018 corrigendum, Computers & Education 127:298.
  9. van Roy, R. & Zaman, B. (2019). Unravelling the ambivalent motivational power of gamification: a basic psychological needs perspective. International Journal of Human-Computer Studies 127:38–50.
  10. Pickal, A. J., Stadler, M., Sailer, M., Bai, S., Ninaus, M., Greiff, S., Becker, N. & Koch, M. (2026). The winner takes it all — effects of leaderboard-based feedback on cognitive performance and motivation. Learning and Individual Differences 126:102836. doi:10.1016/j.lindif.2025.102836 (open access)
  11. Berridge, K. C. & Robinson, T. E. (1998). What is the role of dopamine in reward: hedonic impact, reward learning, or incentive salience? Brain Research Reviews 28(3):309–369. doi:10.1016/S0165-0173(98)00019-8
  12. Berridge, K. C. (2007). The debate over dopamine's role in reward: the case for incentive salience. Psychopharmacology 191(3):391–431. doi:10.1007/s00213-006-0578-x — explicitly a debate position paper.
  13. Schultz, W., Dayan, P. & Montague, P. R. (1997). A neural substrate of prediction and reward. Science 275(5306):1593–1599. doi:10.1126/science.275.5306.1593
  14. Schultz, W. (2016). Dopamine reward prediction error coding. Dialogues in Clinical Neuroscience 18(1):23–32. doi:10.31887/DCNS.2016.18.1/wschultz
  15. Garea, S. S., Drummond, A., Sauer, J. D., Hall, L. C. & Williams, M. N. (2021). Meta-analysis of the relationship between problem gambling, excessive gaming and loot box spending. International Gambling Studies 21(3):460–479. doi:10.1080/14459795.2021.1914705
  16. Spicer, S. G., Nicklin, L. L., Uther, M., Lloyd, J., Lloyd, H. & Close, J. (2022). Loot boxes, problem gambling and problem video gaming: a systematic review and meta-synthesis. New Media & Society 24(4):1001–1022. doi:10.1177/14614448211027175 — a meta-synthesis, not a pooled meta-analysis.
  17. Close, J., Spicer, S. G., Nicklin, L. L., Uther, M., Lloyd, J. & Lloyd, H. (2021). Secondary analysis of loot box data: are high-spending "whales" wealthy gamers or problem gamblers? Addictive Behaviors 117:106851. doi:10.1016/j.addbeh.2021.106851
  18. Zendle, D. & Cairns, P. (2018). Video game loot boxes are linked to problem gambling: results of a large-scale survey. PLoS ONE 13(11):e0206767 — see the 2019 correction and the 2019 replication.
  19. Zendle, D., Meyer, R. & Over, H. (2019). Adolescents and loot boxes: links with problem gambling and motivations for purchase. Royal Society Open Science 6(6):190049
  20. Robb, N. & Zhang, B. (2022). Performance-based dynamic difficulty adjustment and player experience in a 2D digital game: a controlled experiment. Acta Ludologica 5(1):4–22 — n = 221; conducted via Mechanical Turk, with author-flagged data-reliability limits.
  21. [21a] Romero-Mendez, E. A., Santana-Mancilla, P. C., Garcia-Ruiz, M., Montesinos-López, O. A. & Anido-Rifón, L. E. (2023). The use of deep learning to improve player engagement in a video game through a dynamic difficulty adjustment based on skills classification. Applied Sciences 13(14):8249. doi:10.3390/app13148249 — a single uncontrolled study, not a review.
    [21b] Mortazavi, F., Moradi, H. & Vahabie, A.-H. (2024). Dynamic difficulty adjustment approaches in video games: a systematic literature review. Multimedia Tools and Applications.
  22. Adamczyk, P. D. & Bailey, B. P. (2004). If not now, when? The effects of interruption at different moments within task execution. CHI '04, 271–278. doi:10.1145/985692.985727 — note the authors report their own time-on-task and resumption-lag results as inconsistent with prior work.
  23. Iqbal, S. T. & Bailey, B. P. (2008). Effects of intelligent notification management on users and their tasks. CHI '08, 93–102 — models detect breakpoints reasonably well but struggle to differentiate their type.
  24. Iqbal, S. T. & Bailey, B. P. (2010). Oasis: a framework for linking notification delivery to the perceptual structure of goal-directed tasks. ACM TOCHI 17(4), Article 15. doi:10.1145/1879831.1879833 — coarser breakpoints correspond with successively larger reductions in interruption cost.
  25. Altmann, E. M. & Trafton, J. G. (2004). Task interruption: resumption lag and the role of cues. Proceedings of the 26th Annual Conference of the Cognitive Science Society.
  26. Auer, M., Malischnig, D. & Griffiths, M. D. (2014). Is "pop-up" messaging in online slot machine gambling effective as a responsible gambling strategy? Journal of Gambling Issues 29. doi:10.4309/jgi.2014.29.3
  27. Auer, M. & Griffiths, M. D. (2015). Enhanced pop-up messaging study — 1.6 million sessions, two matched random samples of 800,000; 1.39% vs 0.67% session termination.
  28. [28a] Hopfgartner, N., Auer, M., Santos, T., Helic, D. & Griffiths, M. D. (2021). The effect of mandatory play breaks on subsequent gambling behavior among Norwegian online sports betting, slots and bingo players. Journal of Gambling Studies. doi:10.1007/s10899-021-10078-3 — 21,129 players, 156,989 breaks.
    [28b] Auer, M. et al. (2019) — 7,190 VLT players, 90-second forced termination, no significant effect.
  29. Monge Roffarello, A. & De Russis, L. (2023). Achieving digital wellbeing through digital self-control tools: a systematic review and meta-analysis. ACM TOCHI 30(4), Article 53. doi:10.1145/3571810
  30. Zendle, D., Flick, C., Gordon-Petrovskaya, E., Ballou, N., Xiao, L. Y. & Drachen, A. (2023). No evidence that Chinese playtime mandates reduced heavy gaming in one segment of the video games industry. Nature Human Behaviour 7(10):1753–1766. doi:10.1038/s41562-023-01669-8 — read with the three qualifications in §7.4.
  31. Choi, J., Cho, H., Lee, S., Kim, J. & Park, E.-C. (2018). Effect of the online game shutdown policy on internet use, internet addiction, and sleeping hours in Korean adolescents. Journal of Adolescent Health 62(5):548–555. doi:10.1016/j.jadohealth.2017.11.291
  32. Lee, C., Kim, H. & Hong, A. (2017). Ex-post evaluation of illegalizing juvenile online game after midnight: a case of shutdown policy in South Korea. Telematics and Informatics 34(8):1597–1606. doi:10.1016/j.tele.2017.07.006
  33. [33a] O'Grady, B., Lambe, R., Baldwin, M., Acheson, T. & Doherty, C. (2024). The validity of Apple Watch Series 9 and Ultra 2 for serial measurements of heart rate variability and resting heart rate. Sensors.
    [33b] Apple Watch Series 6 versus 3-lead ECG validation, n = 78. Sensors 25(8):2380 (2025).
  34. (2024). The validity of heart rate variability (HRV) in educational research. Educational Psychology Review. doi:10.1007/s10648-024-09878-x — systematic review of 48 studies.