REVIEW 5 major objections 6 minor 91 references
Streaming, Fast and Slow: Cognitive Load-Aware Streaming for Efficient LLM Serving
T0 review · 5 major / 6 minor · reviewed 2026-08-16 · deepseek-v4-flash
Pith's one-line read Adaptive streaming that slows LLM output for hard text cuts compute by up to 16.79 percent at equal alignment rate.
desk verdict A sensible idea for pacing LLM streams by content difficulty, but the headline savings rest on a reading-speed measurement that may measure tolerance rather than actual reading; worth a serious referee, not worth trusting as gospel. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing machinery is a log-normal model of user reading speeds combined with a tunable speed-allocation rule. Reading speeds within a content group are modeled as log-normal; the alpha-quantile of that distribution gives the slowest stream that still covers fraction alpha of users, and the savings formula compares those quantile speeds against the system's maximum streaming speed. For the live prototype, cognitive load scores (Gunning-Fog or an LLM self-score) are normalized and blended with a uniform allocation by a parameter alpha, so each concurrent segment receives speed proportional to its estimated complexity. The asymmetry that justifies the entire scheme is the claim that streaming faster than the user reads is unused capacity, since users cannot consume tokens quicker than their own reading pace.
What would settle it
Run a direct measurement study where participants read the same passages fully presented rather than streamed, using eye tracking or word-by-word self-paced reading, and compare those actual reading speeds to the PEST-derived comfortable speeds; if the 99th percentile of the PEST speeds (21.20 WPS for the easy passage) exceeds the fastest comprehension-preserving reading speed observed, the paper's savings estimates fail.
Extended reading notes
Core claim
The central claim is that content complexity predicts how fast people can consume streamed text, so a stream that tracks complexity is both cheaper and no worse for users. The authors fit log-normal distributions to PEST-measured comfortable streaming speeds for easy and hard passages, define SRAR as the fraction of users whose reading speed is below the stream rate, and show that streaming at the 99th percentile of each content type instead of the system's maximum rate yields a 63.14 percent compute saving in the two-passage example. For the ten-passage simulation, adaptive allocation achieves any given SRAR with less compute than uniform streaming, with savings growing as the target SRAR rises. The LLM-based cognitive load estimator correlates with measured comfortable speeds at r = 0.955, while the Gunning-Fog index gives r = 0.828, meaning even a nearly free heuristic captures a large share of the benefit.
Load-bearing premise
The entire savings figure rests on treating the PEST-measured comfortable streaming speed as the user's natural reading speed for fully presented text, then extrapolating that to the 99th percentile of a fitted log-normal curve; if those values overestimate real reading speed, the compute savings shrink.
Editorial extensions
If this is right
- At a 95 percent SRAR target, Gunning-Fog allocation needs 10.33 percent less compute than uniform streaming, and the LLM-based allocation needs 16.79 percent less.
- In the two-passage example, streaming at the fitted 99th-percentile speeds instead of the maximum rate reduces compute by 63.14 percent.
- LLM-based cognitive load scores correlate with measured comfortable reading speeds at r = 0.955, and the Gunning-Fog index at r = 0.828, so both are usable signals for pacing.
- Below an average budget of about 6 WPS, adaptive streaming no longer beats uniform streaming, marking a resource floor for the method.
- The tunable parameter alpha blends complexity-proportional allocation with uniform allocation, giving service operators a single knob between user experience and efficiency.
Reading between the lines
- A deployment rule not explored in the paper is to use intention detection to route only reading-intensive interactions (explanations, tutorials, detailed answers) through adaptive pacing, while keeping task-oriented outputs at full speed; the paper's savings figures suggest the gain is concentrated in reading-heavy traffic.
- Because the LLM self-score costs only a few extra tokens per segment, the same prompt could be extended to output a predicted reading-speed quantile directly, removing the need for a separate readability model in the allocator.
- The log-normal fit implies that most savings come from the slowest readers in the tail; a system could cap the stream at a much lower percentile and measure user satisfaction empirically, trading a small amount of comfort for most of the compute gain.
- If GPU power scales roughly linearly with inference speed in the operating range, as the paper's cited evidence suggests, the compute savings would translate directly into energy savings, which the paper does not quantify.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper proposes cognitive-load-aware adaptive streaming for LLM-generated text. It models users' reading speeds for different content types as log-normal distributions, derives a compute-savings formula in Section 3.1, and measures what it calls "comfortable reading speeds" through a PEST procedure in a crowdsourced study (Section 3.2). The reported results are a 63.14% compute reduction in the two-passage example at 99% SRAR and 10.33% (Gunning-Fog) and 16.79% (LLM-based) reductions at 95% SRAR relative to uniform streaming (Table 1). A prototype system estimates cognitive load either from the Gunning-Fog index or from LLM-generated load tags and allocates streaming speeds under a fixed budget via the interpolation rule in Equation 2.
Significance. If the quantitative claims are robust, the paper would make a useful connection between cognitive-load modeling and LLM serving efficiency, with a practical deployment story via DVFS and lightweight estimators. The authors are transparent in acknowledging several modeling assumptions and limitations. However, the headline savings are conditional on an unvalidated equivalence between PEST-derived comfortable streaming speed and natural reading speed, on exclusion rules that enforce the expected complexity-speed monotonicity, and on a linear resource-allocation model. These issues undermine the strength of the empirical support for the central quantitative claims and require substantial revision before the results can be taken at face value.
major comments (5)
- [Section 3.1 and Section 3.2] The model in Section 3.1 defines r as the natural reading speed when the entire passage is fully presented, but the PEST procedure in Section 3.2 measures a self-selected 'comfortable streaming speed' under repeated faster/slower adjustments, with an option to accept the speed as 'the same as my reading speed.' This is a preference or tolerance measure, not a speed verified to support comprehension. The two-choice comprehension check in Section 5.1 is administered after the passage disappears and cannot establish that reading at the selected speed was successful. Because the 63.14% saving in Section 3.2 and all savings in Table 1 are arithmetic on the fitted quantiles of these PEST speeds, the central quantitative claims require either a comprehension-gated validation of the measured speeds or a reframing of the claims as being about comfortable streaming speed rather than natural reading speed.
- [Section 3.2 and Section 5.1] Exclusion rule (b) removes participants whose preferred speed for the more complex passage exceeds their preferred speed for the simpler passage. This directly enforces the monotonic relationship between content complexity and reading speed that the paper aims to demonstrate, and it can inflate the separation between the two fitted distributions and hence the computed savings. The t-test, K-S fit, and all subsequent savings are computed after this selection. The authors should report results with and without the exclusion rule, or justify the rule with an independent criterion rather than the outcome variable itself.
- [Section 3.2] The K-S p-values of 0.35 and 0.39 validate the log-normal shape over the observed range but do not validate the upper tail. The 99th-percentile values of 21.20 WPS and 11.97 WPS that drive the 63.14% saving are sensitive to the fitted tail parameters and are far above the bulk of the data. The paper should provide nonparametric quantile estimates, bootstrap confidence intervals, or a sensitivity analysis for the tail quantiles before presenting the savings as stable results.
- [Section 5.2] The correlation of r = 0.955 for the LLM-based estimator is computed against the median comfortable reading speeds from the same PEST data and the same passages that are then used to define SRAR in Figure 6 and Table 1. This makes the estimator comparison partly circular: the 'ground truth' is the same self-reported speed data used to fit the reading-speed distributions. An independent evaluation should use held-out passages or a separate reading-speed measurement, such as comprehension-gated self-paced reading, to validate the cognitive-load estimators.
- [Section 4.3] Equations (1) and (2) assume that a unit decrease in streaming speed for one request frees exactly a unit increase for another request, so the reported percentage savings depend on a linear resource-allocation model. The authors acknowledge this assumption and cite supporting evidence for approximate linearity in some ranges, but a sensitivity analysis around sublinear transfer would clarify how much of the 10.33% and 16.79% savings is an artifact of the linearity assumption rather than a robust property of adaptive streaming.
minor comments (6)
- [Section 3.2] The sentence '78 is degrees of freedom correspond to 79 effective paired samples' is ungrammatical and should be rephrased.
- [Figure 3] The caption for panel (c) contains a stray '>' symbol at the end that should be removed.
- [Table 1 and Figure 6] The label 'Compute Requirement' is measured in WPS; this conflates streaming speed with computational resource unless the linear allocation assumption holds. Consider labeling it 'average streaming speed' or 'compute under the linear allocation assumption.'
- [Section 3.2] The paired t-test is reported with t and p but not an effect size; adding Cohen's d or a similar measure would strengthen the presentation.
- [References] Reference [29] is a community forum post for the GPT-4o tokens-per-second estimate; if possible, cite an official OpenAI documentation page or a more authoritative benchmark instead.
- [Section 5.1] The phrase 'Each session lasted about 15 minutes to reduce fatigue' is awkward; the intended meaning is presumably that the session duration was chosen to reduce fatigue, and the wording should be revised accordingly.
Circularity Check
No significant circularity: the reported savings are explicit model-based computations from the paper's own fitted distributions, and the main measurement concerns are construct-validity issues, not definitional reductions.
full rationale
Walking the derivation chain: Section 3.1 defines a log-normal model of reading speed and derives the quantile-based saving formula in Eq. (1); Section 3.2 measures “comfortable reading speed” with a PEST procedure, fits log-normal distributions to those speeds, and plugs the fitted 99th percentiles (21.20 and 11.97 WPS) together with an externally sourced s_max (45 WPS) into Eq. (1) to obtain 63.14%. Section 5 repeats this for ten passages, correlates Gunning-Fog and LLM-based cognitive-load scores with the measured comfortable speeds, and simulates average compute at matched SRAR. In none of these steps is the target quantity defined in terms of the input, nor is a fitted parameter renamed as a prediction: the savings are explicitly computed from the fitted distributions rather than claimed as independent out-of-sample measurements, and the cognitive-load estimators are not trained on the speed data, so their correlations are empirical validations rather than circular fits. The substitution of PEST-derived “comfortable reading speed” for the model’s “natural reading speed” is a substantive measurement assumption that the paper itself acknowledges in its limitations and that could be challenged on construct-validity grounds, but it is not an equation-level identity or a self-citation, so it does not meet the bar for circularity. The disclosed exclusion rule (b) may bias the separation between the two passages, but it does not make the savings equivalent to the filter by construction. No load-bearing self-citations or imported uniqueness theorems appear in the argument. Therefore, the paper’s derivation chain is self-contained and the correct finding is no significant circularity.
Assumptions & free parameters
free parameters (5)
- Lognormal parameters for simple bedtime story passage =
mu and sigma not reported; 99th percentile used = 21.20 WPS
- Lognormal parameters for complex economics passage =
mu and sigma not reported; 99th percentile used = 11.97 WPS
- Interpolation hyperparameter alpha in Equation 2 =
0.5
- Maximum streaming speed s_max =
45 WPS
- PEST initial speed and step size =
initial 3-8 WPS; Delta_v starts at 2 WPS
assumptions (6)
- domain assumption Reading speeds within a content group follow log-normal distributions
- ad hoc to paper Comfortable streaming speed measured by PEST equals natural reading speed for fully presented text
- domain assumption Streaming at or above the user's natural reading speed is sufficient, and excess speed provides no benefit
- ad hoc to paper Computational resource can be transferred linearly between streams
- domain assumption LLM self-reported cognitive load tag is a valid proxy for human cognitive load
- domain assumption Adjacent sentences have similar cognitive load
Cite this review
Pith. "Pith review of Streaming, Fast and Slow: Cognitive Load-Aware Streaming for Efficient LLM Serving." pith.science (2026). https://pith.science/paper/YRV7EBKZ
@misc{pith2026250417999,
author = {Pith},
title = {Pith review of: Streaming, Fast and Slow: Cognitive Load-Aware Streaming for Efficient LLM Serving},
year = {2026},
howpublished = {\url{https://pith.science/paper/YRV7EBKZ}},
note = {Machine review of arXiv:2504.17999}
}
read the original abstract
Generative conversational interfaces powered by large language models (LLMs) typically stream output token-by-token at a rate determined by computational budget, often neglecting actual human reading speeds and the cognitive load associated with the content. This mismatch frequently leads to inefficient use of computational resources. For example, in cloud-based services, streaming content faster than users can read appears unnecessary, resulting in wasted computational resources and potential delays for other users, particularly during peak usage periods. To address this issue, we propose an adaptive streaming method that dynamically adjusts the pacing of LLM streaming output in real-time based on inferred cognitive load. Our approach estimates the cognitive load associated with streaming content and strategically slows down the stream during complex or information-rich segments, thereby freeing computational resources for other users. We conducted a statistical analysis and simulation based on a statistical model derived from data collected in a crowdsourced user study across various types of LLM-generated content. Our results show that this adaptive method can effectively reduce computational consumption while largely maintaining streaming speed above user's normal reading speed.
Figures
Reference graph
Works this paper leans on
-
[1]
Dina Acklin and Megan H Papesh. 2017. Modern speed-reading apps do not foster reading comprehension. The American journal of psychology 130, 2 (2017), 183–199
2017
-
[2]
Oswald Barral, Sébastien Lallé, and Cristina Conati. 2020. Understanding the effectiveness of adaptive guidance for narrative visualization: a gaze-based analy- sis. In Proceedings of the 25th international conference on intelligent user interfaces . 1–9
2020
-
[3]
Basili and David H
Victor R. Basili and David H. Hutchens. 1983. An empirical study of a syntactic complexity family. IEEE Transactions on Software Engineering 6 (1983), 664–672
1983
-
[4]
Vance W Berger and YanYan Zhou. 2014. Kolmogorov–smirnov test: Overview. Wiley statsref: Statistics reference online (2014)
2014
-
[5]
Anubhav Bhatti, Prithila Angkan, Behnam Behinaein, Zunayed Mahmud, Dirk Rodenburg, Heather Braund, P James Mclellan, Aaron Ruberto, Geoffery Harri- son, Daryl Wilson, et al. 2024. CLARE: Cognitive Load Assessment in REaltime with Multimodal Data. arXiv preprint arXiv:2404.17098 (2024)
arXiv 2024
-
[6]
Mikołaj Buchwald, Szymo Kupiński, Adam Bykowski, Joanna Marcinkowska, Dawid Ratajczyk, and Marcin Jukiewicz. 2019. Electrodermal activity as a mea- sure of cognitive load: A methodological approach. In 2019 Signal Processing: Algorithms, Architectures, Arrangements, and Applications (SPA). IEEE, 175–179
2019
-
[7]
Jierun Chen, Shiu-hong Kao, Hao He, Weipeng Zhuo, Song Wen, Chul-Ho Lee, and S-H Gary Chan. 2023. Run, don’t walk: chasing higher FLOPS for faster neural networks. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition. 12021–12031
2023
-
[8]
Gautam Chutani. 2024. Streaming LLM Responses: Importance and Implementa- tion. https://gautam75.medium.com/streaming-llm-responses-importance-and- implementation-911b135ef541 Accessed: 2025-03-30
2024
Show all 91 references
-
[9]
Scott Crossley, Aron Heintz, Joon Suh Choi, Jordan Batchelor, Mehrnoush Karimi, and Agnes Malatinszky. 2023. A large-scaled corpus for assessing text readability. Behavior Research Methods 55, 2 (2023), 491–507
2023
-
[10]
Scott A Crossley, Stephen Skalicky, and Mihai Dascalu. 2019. Moving beyond classic readability formulas: New methods and new models. Journal of Research in Reading 42, 3-4 (2019), 541–561
2019
-
[11]
Edgar Dale and Jeanne S Chall. 1948. A formula for predicting readability: Instructions. Educational research bulletin (1948), 37–54
1948
-
[12]
Edgar Dale and Jeanne S Chall. 1949. The concept of readability. Elementary English 26, 1 (1949), 19–26
1949
-
[13]
Elnaz Davoodi and Leila Kosseim. 2016. On the Contribution of Discourse Structure on Text Complexity Assessment. In Proceedings of the 17th Annual Meeting of the Special Interest Group on Discourse and Dialogue . 166–174
2016
-
[14]
Diana DeStefano and Jo-Anne LeFevre. 2007. Cognitive load in hypertext reading: A review. Computers in human behavior 23, 3 (2007), 1616–1641
2007
-
[15]
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019. Bert: Pre-training of deep bidirectional transformers for language understanding. In Proceedings of the 2019 conference of the North American chapter of the association for computational linguistics: human...
2019
-
[16]
Sidney D’Mello, Andrew Olney, Claire Williams, and Patrick Hays. 2012. Gaze tutor: A gaze-reactive intelligent tutoring system.International Journal of human- computer studies 70, 5 (2012), 377–398
2012
-
[17]
Benjamin D Douglas, Patrick J Ewell, and Markus Brauer. 2023. Data qual- ity in online human-subjects research: Comparisons between MTurk, Prolific, CloudResearch, Qualtrics, and SONA. Plos one 18, 3 (2023), e0279720
2023
-
[18]
Budmonde Duinkharjav, Kenneth Chen, Abhishek Tyagi, Jiayi He, Yuhao Zhu, and Qi Sun. 2022. Color-perception-guided display power reduction for virtual reality. ACM Transactions on Graphics (TOG) 41, 6 (2022), 1–16
2022
-
[19]
Mary C Dyson and Mark Haselgrove. 2001. The influence of reading speed and line length on the effectiveness of reading from screen. International Journal of Human-Computer Studies 54, 4 (2001), 585–612
2001
-
[20]
Lisa Jo Elliott, Medina Ljubijanac, and Danielle Wieczorek. 2020. The effect of screen size on reading speed: A comparison of three screens to print. InAdvances in Human Factors in Training, Education, and Learning Sciences: Proceedings of the AHFE 2019 International Conferenc...
2020
-
[21]
Andreas Fink and Aljoscha C Neubauer. 2001. Speed of information processing, psychometric intelligence: And time estimation as an index of cognitive load. Chang Xiao and Brenda Z. Yang Personality and Individual Differences 30, 6 (2001), 1009–1021
2001
-
[22]
Thomas François and Eleni Miltsakaki. 2012. Do NLP and machine learning improve traditional readability formulas?. In Proceedings of the First Workshop on Predicting and Improving Text Readability for target reader populations . 49–57
2012
-
[23]
Gregor Franken, Anja Podlesek, and Klementina Mozina. 2015. Eye-tracking study of reading speed from LCD displays: influence of type style and type size. Journal of Eye Movement Research 8, 1 (2015)
2015
-
[24]
Arthur C Graesser, Shulan Lu, George Tanner Jackson, Heather Hite Mitchell, Mathew Ventura, Andrew Olney, and Max M Louwerse. 2004. AutoTutor: A tutor with dialogue in natural language. Behavior Research Methods, Instruments, & Computers 36 (2004), 180–192
2004
-
[25]
Arthur C Graesser, Danielle S McNamara, Max M Louwerse, and Zhiqiang Cai
-
[26]
Jesse W Grootjen, Philipp Thalhammer, and Thomas Kosch. 2024. Your Eyes on Speed: Using Pupil Dilation to Adaptively Select Speed-Reading Parameters in Virtual Reality. Proceedings of the ACM on Human-Computer Interaction 8, MHCI (2024), 1–17
2024
-
[27]
Jiawei Gu, Xuhui Jiang, Zhichao Shi, Hexiang Tan, Xuehao Zhai, Chengjin Xu, Wei Li, Yinghan Shen, Shengjie Ma, Honghao Liu, et al . 2024. A survey on llm-as-a-judge. arXiv preprint arXiv:2411.15594 (2024)
2024 arXiv
-
[28]
Robert Gunning. 1952. The technique of clear writing. (No Title) (1952)
1952
-
[29]
Duncan Haywood. 2024. Gpt-4o tokens per second comparable to gpt-3.5-turbo. Data and analysis. https://community.openai.com/t/gpt-4o-tokens-per-second- comparable-to-gpt-3-5-turbo-data-and-analysis/768559 Accessed: April 7, 2025
2024
-
[30]
Nora Hollenstein, Jonathan Rotsztejn, Marius Troendle, Andreas Pedroni, Ce Zhang, and Nicolas Langer. 2018. ZuCo, a simultaneous EEG and eye-tracking resource for natural sentence reading. Scientific data 5, 1 (2018), 1–13
2018
-
[31]
Te-Yuan Huang, Ramesh Johari, and Nick McKeown. 2013. Downton abbey without the hiccups: Buffer-based rate adaptation for http video streaming. In Proceedings of the 2013 ACM SIGCOMM workshop on Future human-centric multimedia networking. 9–14
2013
-
[32]
Stephen Hutt, Joseph F Grafsgaard, and Sidney K D’Mello. 2019. Time to scale: Generalizable affect detection for tens of thousands of students across an en- tire school year. In Proceedings of the 2019 CHI conference on human factors in computing systems. 1–14
2019
-
[33]
Brockmole, and Sidney K
Stephen Hutt, Kristina Krasich, James R. Brockmole, and Sidney K. D’Mello. 2021. Breaking out of the lab: Mitigating mind wandering with gaze-based attention- aware technology in classrooms. In Proceedings of the 2021 CHI Conference on Human Factors in Computing Systems . 1–14
2021
-
[34]
Matthew Jörke, Shardul Sapkota, Lyndsea Warkenthien, Niklas Vainio, Paul Schmiedmayer, Emma Brunskill, and James A Landay. 2025. GPTCoach: Towards LLM-Based Physical Activity Coaching. InProceedings of the 2025 CHI Conference on Human Factors in Computing Systems . 1–46
2025
-
[35]
Antony William Joseph and Ramaswamy Murugesh. 2020. Potential eye tracking metrics and indicators to measure cognitive load in human-computer interaction research. J. Sci. Res 64, 1 (2020), 168–175
2020
-
[36]
Daniel Jurafsky and James H. Martin. 2025. Discourse Coherence. In Speech and Language Processing: An Introduction to Natural Language Processing, Com- putational Linguistics, and Speech Recognition with Language Models (3rd ed.). Chapter 24. https://web.stanford.edu/~jurafsky...
2025
-
[37]
Andreas Kosmas Kakolyris, Dimosthenis Masouros, Sotirios Xydis, and Dimitrios Soudris. 2024. Slo-aware gpu dvfs for energy-efficient llm inference serving. IEEE Computer Architecture Letters 23, 2 (2024), 150–153
2024
-
[38]
J Peter Kincaid, Robert P Fishburne Jr, Richard L Rogers, and Brad S Chissom
-
[39]
Thomas Kosch, Jakob Karolus, Johannes Zagermann, Harald Reiterer, Albrecht Schmidt, and Paweł W Woźniak. 2023. A survey on measuring cognitive workload in human-computer interaction. Comput. Surveys 55, 13s (2023), 1–39
2023
-
[40]
Thomas Kosch, Albrecht Schmidt, Simon Thanheiser, and Lewis L Chuang. 2020. One does not simply RSVP: mental workload to select speed reading parameters using electroencephalography. In Proceedings of the 2020 CHI Conference on Human Factors in Computing Systems . 1–13
2020
-
[41]
Naveen Kumar and Jyoti Kumar. 2016. Measurement of cognitive load in HCI systems using EEG power spectrum: an experimental study. Procedia Computer Science 84 (2016), 70–78
2016
-
[42]
Sri Hastuti Kurniawan and Panayiotis Zaphiris. 2001. Reading online or on paper: which is faster? (2001)
2001
-
[43]
Gonzalez, Hao Zhang, and Ion Stoica
Woosuk Kwon, Zhuohan Li, Siyuan Zhuang, Ying Sheng, Lianmin Zheng, Cody Hao Yu, Joseph E. Gonzalez, Hao Zhang, and Ion Stoica. 2023. Efficient Memory Management for Large Language Model Serving with PagedAtten- tion. In Proceedings of the ACM SIGOPS 29th Symposium on Operating...
2023
-
[44]
Sébastien Lallé, Dereck Toker, and Cristina Conati. 2019. Gaze-driven adaptive interventions for magazine-style narrative visualizations. IEEE Transactions on Visualization and Computer Graphics 27, 6 (2019), 2941–2952
2019
-
[45]
Sébastien Lallé, Dereck Toker, Cristina Conati, and Giuseppe Carenini. 2015. Prediction of users’ learning curves for adaptation while using an information visualization. In Proceedings of the 20th International Conference on Intelligent User Interfaces. 357–368
2015
-
[46]
Christian Lander, Marco Speicher, Denise Paradowski, Norine Coenen, Sebastian Biewer, and Antonio Krüger. 2015. Collaborative newspaper: Exploring an adaptive scrolling algorithm in a multi-user reading scenario. In Proceedings of the 4th International Symposium on Pervasive D...
2015
-
[47]
Etienne Le Sueur and Gernot Heiser. 2010. Dynamic voltage and frequency scaling: The laws of diminishing returns. In Proceedings of the 2010 international conference on Power aware computing and systems . 1–8
2010
-
[48]
Hanchen Li, Yuhan Liu, Yihua Cheng, Siddhant Ray, Kuntai Du, and Junchen Jiang. 2024. Eloquent: A More Robust Transmission Scheme for LLM Token Streaming. In Proceedings of the 2024 SIGCOMM Workshop on Networks for AI Computing. 34–40
2024
-
[49]
Zikun Li, Zhuofu Chen, Remi Delacourt, Gabriele Oliaro, Zeyu Wang, Qinghan Chen, Shuhuai Lin, April Yang, Zhihao Zhang, Zhuoming Chen, et al . 2025. AdaServe: SLO-Customized LLM Serving with Fine-Grained Speculative Decod- ing. arXiv preprint arXiv:2501.12162 (2025)
2025 arXiv
-
[50]
Neiwen Ling, Guojun Chen, and Lin Zhong. 2024. TimelyLLM: Segmented LLM Serving System for Time-sensitive Robotic Applications. arXiv preprint arXiv:2412.18695 (2024)
2024 arXiv
-
[51]
Fengkai Liu, Tan Jin, and John SY Lee. 2025. Automatic readability assessment for sentences: neural, hybrid and large language models. Language Resources and Evaluation (2025), 1–32
2025
-
[52]
Jiachen Liu, Jae-Won Chung, Zhiyu Wu, Fan Lai, Myungjin Lee, and Mosharaf Chowdhury. 2024. Andes: Defining and enhancing quality-of-experience in llm-based text streaming services. arXiv preprint arXiv:2404.16283 (2024)
2024 arXiv
-
[53]
Paul Joe Maliakel, Shashikant Ilager, and Ivona Brandic. 2025. Investigating Energy Efficiency and Performance Trade-offs in LLM Inference Across Tasks and DVFS Settings. arXiv preprint arXiv:2501.08219 (2025)
2025
-
[54]
Lena Mamykina, Elizabeth Mynatt, and Michael A Terry. 2001. Time aura: Interfaces for pacing. In Proceedings of the SIGCHI conference on Human factors in computing systems. 144–151
2001
-
[55]
Emma Marsden, Sophie Thompson, and Luke Plonsky. 2018. A methodological synthesis of self-paced reading in second language research. Applied Psycholin- guistics 39, 5 (2018), 861–904
2018
-
[56]
Xinxin Mei, Qiang Wang, and Xiaowen Chu. 2017. A survey and measure- ment study of GPU DVFS on energy conservation. Digital Communications and Networks 3, 2 (2017), 89–100
2017
-
[57]
Tadashi Okoshi, Julian Ramos, Hiroki Nozaki, Jin Nakazawa, Anind K Dey, and Hideyuki Tokuda. 2015. Attelia: Reducing user’s cognitive load due to interruptive notifications on smart phones. In 2015 IEEE international conference on pervasive computing and communications (PerCom...
2015
-
[58]
OpenAI. 2023. What are tokens and how to count them. https://help.openai. com/en/articles/4936856-what-are-tokens-and-how-to-count-them Accessed: April 7, 2025
2023
-
[59]
Gustav Öquist and Mikael Goldstein. 2001. Adaptive rapid serial visual presenta- tion. Unpublished Thesis, Uppsala University, Uppsala, Sweden (2001)
2001
-
[60]
Gustav Öquist and Mikael Goldstein. 2003. Towards an improved readability on mobile devices: evaluating adaptive rapid serial visual presentation. Interacting with Computers 15, 4 (2003), 539–558
2003
-
[61]
Gustav Öquist and Kristin Lundin. 2007. Eye movement study of reading text on a mobile phone using paging, scrolling, leading, and RSVP. In Proceedings of the 6th international conference on Mobile and ubiquitous multimedia . 176–183
2007
-
[62]
Nirmal Patel, Pooja Nagpal, Tirth Shah, Aditya Sharma, Shrey Malvi, and Derek Lomas. 2023. Improving mathematics assessment readability: Do large language models help? Journal of Computer Assisted Learning 39, 3 (2023), 804–822
2023
-
[63]
Ildikó Pilán, Elena Volodina, and Richard Johansson. 2014. Rule-based and machine learning approaches for second language sentence-level readability. In Proceedings of the ninth workshop on innovative use of NLP for building educational applications. 174–184
2014
-
[64]
Jan L Plass, Roxana Moreno, and Roland Brünken. 2010. Cognitive load theory. (2010)
2010
-
[65]
Alec Radford, Jeffrey Wu, Rewon Child, David Luan, Dario Amodei, Ilya Sutskever, et al. 2019. Language models are unsupervised multitask learners. OpenAI blog 1, 8 (2019), 9
2019
-
[66]
Pragnya Ramjee, Mehak Chhokar, Bhuvan Sachdeva, Mahendra Meena, Hamid Abdullah, Aditya Vashistha, Ruchit Nagar, and Mohit Jain. 2025. ASHABot: an LLM-powered chatbot to support the informational needs of community health workers. In Proceedings of the 2025 CHI Conference on Hu...
2025
-
[67]
Keith Rayner, Elizabeth R Schotter, Michael EJ Masson, Mary C Potter, and Rebecca Treiman. 2016. So much to read, so little time: How do we read, and can speed reading help? Psychological Science in the Public Interest 17, 1 (2016), 4–34
2016
-
[68]
Reza Rejaie, Mark Handley, and Deborah Estrin. 2000. Layered quality adaptation for Internet video streaming. IEEE Journal on Selected Areas in Communications 18, 12 (2000), 2530–2543. Streaming, Fast and Slow: Cognitive Load-Aware Streaming for Efficient LLM Serving
2000
-
[69]
Antonios Saravanos, Stavros Zervoudakis, Dongnanzi Zheng, Neil Stott, Bohdan Hawryluk, and Donatella Delfino. 2021. The hidden cost of using Amazon Mechanical Turk for research. In HCI International 2021-Late Breaking Papers: Design and User Experience: 23rd HCI International ...
2021
-
[70]
Erik Schils and Pieter de Haan. 1993. Characteristics of sentence length in running text. Literary and linguistic computing 8, 1 (1993), 20–26
1993
-
[71]
Elliot Schumacher, Maxine Eskenazi, Gwen Frishkoff, and Kevyn Collins- Thompson. 2016. Predicting the Relative Difficulty of Single Sentences With and Without Surrounding Context. In Proceedings of the 2016 Conference on Empirical Methods in Natural Language Processing . 1871–1881
2016
-
[72]
Sally E Shaywitz and Bennett A Shaywitz. 2005. Dyslexia (specific reading disability). Biological psychiatry 57, 11 (2005), 1301–1309
2005
-
[73]
Eva Siegenthaler, Laura Schmid, Michael Wyss, and Pascal Wurtz. 2012. LCD vs. E-ink: An Analysis of the Reading Behavior. Journal of Eye Movement Research 5, 3 (2012)
2012
-
[74]
Shubham Singh. 2025. ChatGPT Statistics (March 2025): Number of Users & Queries. DemandSage (28 February 2025). https://www.demandsage.com/ chatgpt-statistics/
2025
-
[75]
Soroosh Solhjoo, Mark C Haigney, Elexis McBee, Jeroen JG van Merrienboer, Lambert Schuwirth, Anthony R Artino Jr, Alexis Battista, Temple A Ratcliffe, Howard D Lee, and Steven J Durning. 2019. Heart rate and heart rate variability correlate with clinical reasoning performance ...
2019
-
[76]
John Sweller. 1988. Cognitive load during problem solving: Effects on learning. Cognitive science 12, 2 (1988), 257–285
1988
-
[77]
Zhenheng Tang, Yuxin Wang, Qiang Wang, and Xiaowen Chu. 2019. The impact of GPU DVFS on the energy and performance of deep learning: An empirical study. In Proceedings of the Tenth ACM International Conference on Future Energy Systems. 315–325
2019
-
[78]
MiM Taylor, C Douglas Creelman, et al . 1967. PEST: Efficient estimates on probability functions. Journal of the acoustical society of America 41, 4 (1967), 782–787
1967
-
[79]
TOI Tech Desk. 2025. ChatGPT down: AI chatbot back after global outage. The Times of India (30 March 2025). https://timesofindia.indiatimes.com/technology/ tech-news/chatgpt-down-ai-chatbot-outage-affecting-ghibli-style-photo- generation-globally/articleshow/119754435.cms
2025
-
[80]
Sowmya Vajjala and Ivana Lučić. 2018. OneStopEnglish corpus: A new corpus for automatic readability assessment and text simplification. InProceedings of the thirteenth workshop on innovative use of NLP for building educational applications . 297–304
2018
-
[81]
Sowmya Vajjala and Detmar Meurers. 2012. On improving the accuracy of readability classification using insights from second language acquisition. In Proceedings of the seventh workshop on building educational applications using NLP. 163–173
2012
-
[82]
VentureBeat. 2024. How enterprises are using open source LLMs: 16 examples. VentureBeat (30 Jan. 2024). https://venturebeat.com/ai/how-enterprises-are- using-open-source-llms-16-examples Accessed July 2025
2024
-
[83]
Shaun Wallace, Zoya Bylinskii, Jonathan Dobres, Bernard Kerr, Sam Berlow, Rick Treitman, Nirmal Kumawat, Kathleen Arpin, Dave B Miller, Jeff Huang, et al. 2022. Towards individuated reading experiences: Different fonts increase reading speed for different individuals. ACM Tran...
2022
-
[84]
Chen Wu. 2025. Test Time Compute: Balancing Throughput, Speed, and Quality in LLM Thinking. https://medium.com/@chenwuperth/test-time-compute- balancing-throughput-speed-and-quality-in-llm-thinking-8428c5cf757c. Medium (26 Mar 2025)
2025
-
[85]
Wei Xu, Chris Callison-Burch, and Courtney Napoles. 2015. Problems in current text simplification research: New data can help. Transactions of the Association for Computational Linguistics 3 (2015), 283–297
2015
-
[86]
Peifeng Yin, Ping Luo, Wang-Chien Lee, and Min Wang. 2013. Silence is also evidence: interpreting dwell time for recommendation from psychological per- spective. In Proceedings of the 19th ACM SIGKDD international conference on Knowledge discovery and data mining . 989–997
2013
-
[87]
Johannes Zagermann, Ulrike Pfeil, and Harald Reiterer. 2016. Measuring cognitive load using eye tracking technology in visual computing. InProceedings of the sixth workshop on beyond time and errors on novel evaluation methods for visualization . 78–85
2016
-
[88]
Xi Zheng, Zhuoyang Li, Xinning Gui, and Yuhan Luo. 2025. Customizing emo- tional support: How do individuals construct and interact with LLM-powered chatbots. In Proceedings of the 2025 CHI Conference on Human Factors in Comput- ing Systems. 1–20
2025
-
[89]
Yihao Zhu, Zhoutong Ye, Yichen Yuan, Wenxuan Tang, Chun Yu, and Yuanchun Shi. 2025. AutoPBL: An LLM-powered Platform to Guide and Support Individual Learners Through Self Project-based Learning. In Proceedings of the 2025 CHI Conference on Human Factors in Computing Systems . 1–26
2025
-
[1975]
Derivation of new readability formulas (automated readability index, fog count and flesch reading ease formula) for navy enlisted personnel. (1975)
1975
-
[2004]
Behavior research methods, instruments, & computers 36, 2 (2004), 193–202
Coh-Metrix: Analysis of text on cohesion and language. Behavior research methods, instruments, & computers 36, 2 (2004), 193–202
2004
Reviewed August 16, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.