{"id":"e3776f0e-96d5-4811-9093-dcac47a7e8b4","arxiv_id":"1908.01699","paper_version":1,"verdict":"REJECT","confidence":"HIGH","novelty_score":4.0,"correctness_risk":"high","formal_verification":"none","parameter_count":1,"one_line_summary":"Thoth is an RSVP speed reader that shows unfamiliar words for 1.5x longer based on readability formulas, but no experimental evidence supports the claimed reading improvements.","lead":"This paper describes Thoth, an open-source speed-reading tool that lengthens the display time of words it deems unfamiliar using readability formulas. The paper claims this improves reading speed and comprehension, but it includes no user study or measured data to support that claim.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The paper's central claim of faster reading with retained comprehension is entirely unsupported: no user study, no baseline comparison, and the only 'results' reported are tool availability.","rationale":"The reader correctly identifies that the fixed 1.5x unfamiliar-word multiplier is arbitrary and untested; that is a real weakness in the design. However, the more load-bearing concern is broader: the paper's strongest claim is a comparative empirical claim about user reading speed and comprehension, and no user study, baseline, or dataset is presented anywhere in the manuscript. The arbitrary multiplier is one symptom of the missing evaluation, but even a perfectly calibrated per-word timing rule would still need a controlled experiment to establish the claimed benefit. In good faith, the paper is best read as a systems/demonstration note describing an open-source RSVP tool with a plausible NLP-based timing heuristic. The tool may be useful and the heuristic is reasonable, but the conclusion's wording ('has enabled users... faster... retaining more context') asserts a measured outcome that the paper never measures. Section 8.3's future-work item is an explicit admission that the needed study was not conducted. Because the reader's verdict is already REJECT and this concern reinforces rather than redirects that verdict, the recommended verdict is UNCHANGED. The practical path to acceptance would be to reframe the paper as a tool description or to add the controlled user study described in the concrete test.","tokens_in":11000,"tokens_out":2642,"duration_ms":28862,"concrete_test":"Run a preregistered within-subject experiment with at least 30 participants and counterbalanced order, comparing (a) Thoth at default settings, (b) a constant-rate RSVP tool such as Spritz matched on mean words-per-minute, and (c) normal page reading, on the same medium- and low-fidelity texts. Measure comprehension with multiple-choice and free-recall questions and measure effective reading time, including any re-reading. Predefine a threshold, for example that Thoth must be at least 10% faster than normal reading while maintaining comprehension within 5% of normal reading, and must beat the matched-rate RSVP baseline on comprehension at equal mean speed. If Thoth does not meet these thresholds, the Section 7 conclusion is unsupported.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim in Section 7 ('Thoth has enabled users to read through medium and low fidelity content faster on average while retaining more context and comprehension than they otherwise would have with similar RSVP tools') is an empirical causal claim about human reading performance. The paper supplies no user data to support it. Section 5, titled 'Results', contains no measurements, only the statements that the tool is freely available and that it is the only RSVP tool using natural language parsing. Section 8.3 explicitly lists 'User Research' as future work, conceding that the required comparison has not been run. Readability-based per-word timing is a plausible design idea, but the implementation depends on unvalidated parameter choices, including the Dale-Chall familiarity decision and the fixed 1.5x display multiplier for unfamiliar words described in Section 8.1. Section 6 also asserts, without evidence, that 'it doesn't seem to make a significant different which dictionary is used.' Tool availability and plausibility do not entail the claimed speed-and-comprehension benefit. The conclusion is therefore unsupported rather than internally contradicted; there is no experimental or quasi-experimental evidence connecting Thoth's mechanism to the stated outcome.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper describes Thoth, an open-source rapid serial visual presentation (RSVP) speed-reading tool that uses natural language processing and multiple readability formulas (Dale-Chall, Flesch-Kincaid, SMOG, Spache, Coleman-Liau, etc.) to estimate word-level familiarity and assign per-word display times. The central claim, stated in the abstract and in Section 7, is that readability-based per-word timing improves reading speed and comprehension relative to conventional RSVP tools. The manuscript describes the tool's architecture, lists related work, and identifies future directions, including tunable parameters and user research. However, Section 5, titled 'Results,' contains no experimental measurements, and Section 8.3 explicitly defers user studies to future work; the effectiveness claim is therefore asserted rather than demonstrated.","tokens_in":11376,"tokens_out":2817,"duration_ms":29290,"significance":"If the central claim were supported, the idea of adapting RSVP presentation rates to lexical familiarity would be a useful contribution to reading technology and human-computer interaction. The paper has some concrete strengths: the tool is open source and publicly hosted, the design integrates several established readability measures, and the mechanism is described clearly enough to be implemented and tested by others. However, the paper's main claim is an empirical causal claim about human reading speed and comprehension, and no user data, baseline comparison, or measurement of any kind is provided. The plausibility of using word familiarity to modulate display time does not substitute for evidence, and as it stands the paper is a system description without validation of its central assertion.","major_comments":[{"comment":"The central claim that Thoth 'has enabled users to read through medium and low fidelity content faster on average while retaining more context and comprehension' is not supported by any experimental evidence. Section 5, titled 'Results,' contains no measurements, participants, or comparisons; it only reports tool availability and uniqueness. Section 8.3 explicitly lists user research as future work, confirming that the required evaluation has not been performed. To support the conclusion, the paper would need a controlled study measuring reading speed and comprehension for Thoth against at least one baseline condition (e.g., conventional RSVP or normal reading); no such study is present.","section":"§5, §7, §8.3"},{"comment":"The fixed 1.5x display-time multiplier for unfamiliar words is a load-bearing parameter of the proposed mechanism, but the paper provides no empirical justification, user study, or cited prior result for this specific value. If this mapping is incorrect, the claimed comprehension benefit does not follow. The paper itself acknowledges that 'it is possible we are losing time by simply scaling the display time of each unfamiliar word by 1.5,' which underscores that this parameter remains unvalidated.","section":"§8.1"},{"comment":"The statement that 'it doesn't seem to make a significant different which dictionary is used' is presented as a finding, but no analysis or data supporting it is given. The sentence also conflates 'significant' as a statistical term with 'significant' as a substantive judgment, and the claim should either be removed or supported with a formal comparison of the dictionaries under consideration.","section":"§6"}],"minor_comments":[{"comment":"The possessive 'its' is repeatedly written as 'it's' (e.g., Abstract 'It's largest insight,' §3 'it's ease of use,' §4 'it's presentation'); these should be corrected.","section":"Abstract, §3, §4"},{"comment":"The phrase 'significant different' should be 'significant difference.'","section":"§6"},{"comment":"The reference list contains irrelevant or unexplained entries (e.g., #13 'What is the amplitude of a wave?' and #22 'Effects of the Seasons and of Bright Light ...') and duplicates (#10 and #27 are the same Dehaene et al. citation; #11 and #35 are the same Deheane book). Several in-text citations do not match the reference list format (e.g., 'Gelzer et. al, 2015' appears as 'Glezer, L., et al.' in the list).","section":"References"},{"comment":"The caption says 'Source: Rayner, K. sagepub.com' but no complete citation for this figure is provided in the reference list.","section":"Figure 1"},{"comment":"The opening sentence 'The results have been clear' is misleading because no results are presented; consider retitling the section to 'System Availability' or similar.","section":"§5"}],"recommendation":"reject","confidential_remarks":"The manuscript reads more like a system demonstration or technical report than a completed research paper. The decisive issue is the complete absence of evaluation for the central effectiveness claim; a user study comparing Thoth with existing RSVP tools and normal reading would be necessary, and that is a substantial addition rather than a local revision. There are also concerns about the informal and partially irrelevant reference list, which may indicate the manuscript is not yet at the standard expected for an archival venue."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Read this only if you want a clear example of a tool paper without a user study. The one genuinely useful thing here is the admission in Section 2 that it breaks no new technical ground, and the Limitations/Future Work sections are unusually candid: the 1.5x multiplier is called a fixed assumption, the dictionary choice is acknowledged as untested, and Section 8.3 concedes no user research has been run. That honesty is real and should be credited.\n\nWhat is actually new: essentially nothing technically. Combining Dale-Chall and other readability formulas to scale per-word display time in an RSVP reader is a plausible idea, and the open-source implementation with a live hosted version is a concrete, reproducible artifact. The paper's Section 5, titled 'Results,' contains no measurements—just the statement that the tool is freely available. The conclusion that Thoth users read faster with better comprehension is an empirical claim about human performance, and the paper provides no data, no participants, no baseline, no comparison. Even the weaker claim that the choice of dictionary doesn't matter is presented as 'it doesn't seem to make a significant different,' with no analysis shown.\n\nThe soft spots are not subtle. The load-bearing claim is unsupported. Related work is thin: adaptive RSVP and prior work on optimal recognition points are mentioned only in passing, and there is no survey of earlier adaptive timing attempts. The references include irrelevant entries (seasonal affective disorder, amplitude of a wave) and the prose has typos. But the paper is not incoherent: the design is described clearly, the mechanism is transparent, and the author explicitly identifies what a proper evaluation would require.\n\nWho is this for? Someone cataloging RSVP tools or looking for a simple example of word-level timing heuristics might find the code useful. As a scientific paper, it is not ready. The right next step is a controlled user study comparing Thoth with fixed-rate RSVP tools and normal reading, measuring speed and comprehension. As is, I would desk reject it; it does not need referee time. If you work on RSVP or adaptive reading interfaces, the idea is worth a footnote, not a citation.","headline":"An honest tool paper whose central speed/comprehension claim is asserted, not demonstrated; the open-source implementation is real, but there is no user study to back any of the results.","tokens_in":11713,"tokens_out":1530,"would_cite":false,"duration_ms":16984,"reading_group":"no","serious_thinker":"yes","would_accept_peer_review":false},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Per-word timing based on word familiarity can make RSVP speed reading faster and more comprehensible.","keywords":["rapid serial visual presentation","RSVP","speed reading","natural language processing","readability formulas","word familiarity","reading comprehension","text presentation"],"falsifier":"Run a controlled experiment where matched readers see the same passages under Thoth, a fixed-rate RSVP reader, and ordinary static text, then take comprehension tests at matched reading speeds; the central claim collapses if Thoth is neither faster at equal comprehension nor better at comprehension at equal speed. A cheaper proxy: eye-tracking would show whether unfamiliar-word labels actually predict longer fixations under RSVP.","tokens_in":10817,"feed_emoji":"📖","tokens_out":6357,"duration_ms":57083,"temperature":0.7,"pith_summary":"The paper argues that rapid serial visual presentation (RSVP) speed-reading tools are held back by uniform pacing: showing every word for the same fixed time ignores the fact that familiar and unfamiliar words take different amounts of processing. It presents Thoth, an open-source RSVP reader that scores each word with readability formulas and extends the display duration of words flagged as unfamiliar by a fixed multiplier. The central claim is that this word-by-word timing lets people read low- and medium-fidelity content faster while keeping more context and comprehension than conventional RSVP tools. Most of the supporting evidence is drawn from reading science; the paper's own user study is listed as future work.","feed_headline":"Unfamiliar words get more screen time in smarter speed reader","feed_subtitle":"Thoth paces RSVP reading word-by-word so hard words last longer, aiming to beat fixed-rate tools.","key_machinery":"The mechanism is a readability-weighted timing rule: a readability formula's familiar-word list marks each token, and the RSVP engine multiplies the default display duration for any word not on the list by a factor of 1.5. The argument for why this should work rests on the visual word form dictionary, the brain's stored picture-like representations of known words, which make familiar words fast to recognize and unfamiliar words slow. What the rule does is convert a whole-text readability score into a per-word scheduling decision.","core_discovery":"On the paper's own terms, the discovery is that RSVP does not have to treat text as a flat sequence of equal units. Thoth combines several readability measures to estimate a text's required grade level, uses one of them to classify each word as familiar or unfamiliar, and assigns display times accordingly; unfamiliar words are shown roughly 1.5 times as long as familiar ones. Because familiar words can be recognized as whole images while unfamiliar words demand extra labor, the tool spends the limited resource of screen time where it is needed. The paper concludes that this approach yields faster average reading and better retention for the medium- and low-fidelity documents that people skim.","pith_inferences":["The strongest test of the paper's logic is a direct A/B comparison of Thoth, a uniform RSVP reader, and static text on the same passages with comprehension checks; the paper does not report such a study, so the central claim is an engineering prediction rather than a measured result.","A graded difficulty signal such as word frequency or surprisal would likely outperform the binary familiar/unfamiliar split, and the 1.5x multiplier could be tuned per user or per text.","The same scheduling principle—give more time to predicted-hard items—generalizes to flashcard decks, subtitles, and captioning, where pacing is currently uniform."],"forward_implications":["If per-word timing works, RSVP tools can be tuned from text statistics alone, without eye tracking or user calibration.","Skimmers of long documents could keep comprehension close to normal while reading faster than current fixed-rate readers.","Readability formulas gain a new role: not just grading whole texts but scheduling individual words.","An open-source implementation means the timing rule can be tested, improved, and extended by other developers.","The same timing logic could be reversed into a 'speed writing' mode that substitutes unfamiliar words with familiar synonyms before display."],"supporting_citations":[{"why":"Establishes that efficient visual encoding frees processing time for comprehension, the premise for varying display time.","marker":"Jackson & McClelland, 1975"},{"why":"Distinguishes faster reading from accurate recall and frames 'effective reading' as both speed and comprehension.","marker":"Dyson et al., 2001"},{"why":"Shows familiar words are recognized holistically as pictures in the visual dictionary, so unfamiliar words need extra processing.","marker":"Glezer et. al, 2015"},{"why":"Describes the visual word form area and explains why the visual system is not optimized for reading.","marker":"Dehaene & Cohen, 2011"},{"why":"Supplies the familiar-word list that classifies words and triggers the longer display time.","marker":"Dale & Chall, 1948"}],"fun_headline_variants":["Speed reading gets smarter: hard words linger longer","NLP-based RSVP paces reading word by word","Thoth adjusts word timing to boost speed reading","Smart speed reader gives tricky words extra time","Natural language parsing tailors RSVP pacing"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that a word marked as unfamiliar truly needs more display time, and that stretching it by 1.5 times is the correct amount; the paper adopts that factor as a fixed assumption, with no measurements behind it.","fun_headline_variants_meta":{"raw":{"variants":["Speed reading gets smarter: hard words linger longer","NLP-based RSVP paces reading word by word","Thoth adjusts word timing to boost speed reading","Smart speed reader gives tricky words extra time","Natural language parsing tailors RSVP pacing"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000588,"raw_usage":{"total_tokens":2627,"prompt_tokens":680,"completion_tokens":1947,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":296,"completion_tokens_details":{"reasoning_tokens":1891}},"tokens_in":296,"tokens_out":1947,"duration_ms":13849,"temperature":1.0,"reasoning_tokens":1891,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-14T15:05:25.184241+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run a controlled experiment where matched readers see the same passages under Thoth, a fixed-rate RSVP reader, and ordinary static text, then take comprehension tests at matched reading speeds; the central claim collapses if Thoth is neither faster at equal comprehension nor better at comprehension at equal speed. A cheaper proxy: eye-tracking would show whether unfamiliar-word labels actually predict longer fixations under RSVP.","supporting_citations":[{"cited_title":"D., McClelland, J","cited_arxiv_id":null,"evidence_quote":"Establishes that efficient visual encoding frees processing time for comprehension, the premise for varying display time."},{"cited_title":"The Influence of Reading Speed and Line Length on the Effectiveness of Reading from Screen","cited_arxiv_id":null,"evidence_quote":"Distinguishes faster reading from accurate recall and frames 'effective reading' as both speed and comprehension."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Shows familiar words are recognized holistically as pictures in the visual dictionary, so unfamiliar words need extra processing."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Describes the visual word form area and explains why the visual system is not optimized for reading."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies the familiar-word list that classifies words and triggers the longer display time."}],"review_version":1}