{"id":"4fffdfe1-0420-4aed-8209-5307395d3eb3","arxiv_id":"2506.02098","paper_version":1,"verdict":"ACCEPT","confidence":"HIGH","novelty_score":6.0,"correctness_risk":"low","formal_verification":"none","parameter_count":0,"one_line_summary":"LibriBrain is a 52-hour single-subject MEG dataset for speech decoding, the largest of its kind, with public code and baselines showing that more training data improves decoding.","lead":"LibriBrain is a new public dataset of over 50 hours of brain recordings (MEG) from one person listening to Sherlock Holmes audiobooks, the largest single-subject MEG speech dataset to date. It comes with code, standard train/test splits, and baseline results for detecting speech, classifying phonemes, and classifying words.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"No significant objection identified: the >50-hour single-subject size claim is supported by Tables 1, 2 and 6 and is robust to acknowledged limitations; forced-alignment label uncertainty is a caveat for benchmarks, not for the central resource claim.","rationale":"Read in good faith, the paper's central claim is a dataset-resource claim: a single-subject MEG corpus of more than 50 hours, about 5× the per-subject depth of the next public within-subject dataset. That claim is internally supported. Table 2 reports 52.32 released hours, and Table 6's total of 53:02:41 across the seven audiobooks is consistent with more than 50 hours even after subtracting held-out sessions. The comparison to Armeni et al. (2022) at 10 hours per subject establishes the 5× factor. The reader's weakest assumption—Gentle forced-alignment labels without gold-standard onset validation—is legitimate and does threaten the benchmark contribution, since phoneme classification windows are built from those boundaries. However, it is not load-bearing for the size claim, and the paper's own limitations section already acknowledges several secondary caveats (single subject, listening paradigm, content homogeneity, preprocessing). The small session-count and words-per-session inconsistencies are presentation errors rather than structural flaws, and they do not change the more-than-50-hours conclusion. For these reasons I do not recommend changing the ACCEPT verdict; adding a sparse hand-validated alignment subset or an independent onset-error estimate would strengthen the benchmark claims and is the natural next check.","tokens_in":27221,"tokens_out":11719,"duration_ms":109196,"concrete_test":"Recompute the total released MEG duration by summing per-session event-file durations from the public HDF5/TSV release; if the sum is below 50 hours the central size claim fails, and if it matches the reported 52.32 hours the claim is verified.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim—largest single-subject MEG speech dataset, over 50 hours, 5× larger than the next comparable dataset—is supported by the paper's own statistics (Table 1: 52 h vs 10 h/subject for Armeni et al.; Table 2: 52.32 released hours; Table 6: 53:02:41 total audiobook hours) and does not depend on phoneme/word label accuracy. I find no load-bearing concern that would overturn this resource claim. The one caveat worth registering is the reader's: event labels come from Gentle forced alignment validated only by a duration monotonicity check (Table 8), not by onset-level accuracy, so the phoneme and word classification baselines in Sections 4.2–4.3 could be affected by systematic boundary errors. There are also minor internal inconsistencies in session counts (95 vs 93 vs 91) and in the reported average words per session, but even the most conservative reading leaves more than 50 hours of released data, so the headline claim stands.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper introduces LibriBrain, a single-subject MEG dataset of over 50 hours of recordings from one participant listening to Sherlock Holmes audiobooks, together with word- and phoneme-level annotations, a Python API, standardized train/validation/test splits, and baseline results for speech detection, phoneme classification, and word classification. The headline claim is that LibriBrain is the largest single-subject MEG dataset for speech decoding, approximately 5 times larger than the next comparable dataset (Armeni et al., 2022) and 50 times larger than most others. The experiments show statistically significant decoding above random baselines and performance that improves with training data volume, and the word classification results are reproduced with permission from a companion preprint.","tokens_in":27553,"tokens_out":4548,"duration_ms":46128,"significance":"If the resource claim holds, LibriBrain is a valuable community asset: it provides unprecedented per-subject MEG depth for naturalistic speech, enabling scaling studies and more reliable non-invasive speech decoding benchmarks. The paper's strengths include public data and code release, a documented preprocessing pipeline, standard splits to prevent leakage, exact permutation tests against baselines, and explicit reporting of computational requirements. The scaling results, while preliminary, provide a useful reference point. The main caveat—that phoneme/word labels rest on forced alignment validated only by a duration-monotonicity check—does not undermine the central dataset-size claim but does affect the strength of the benchmark contributions.","major_comments":[{"comment":"The validation of the forced-alignment annotations consists solely of a duration-monotonicity check: longer linguistic units receive longer durations. Because the phoneme and word classification baselines in Sections 4.2 and 4.3 depend on millisecond-scale onset accuracy, and because the text states that manual correction was applied, the paper should provide quantitative validation of boundary accuracy (for example, agreement between Gentle and manual boundaries on a held-out subset, or a comparison with an independent aligner). Without such validation, the benchmark results cannot be fully interpreted, although this does not affect the central claim about dataset size.","section":"Appendix B.5, Table 8"}],"minor_comments":[{"comment":"The stated average of 5421.28 words per session does not match Table 2's totals: 466,230 words across 93 sessions gives about 5013 words per session. Please clarify the denominator or correct the figure.","section":"Section 3.2"},{"comment":"The session counts 95, 93, and 91 appear in different places. These are reconcilable (95 total recorded, 2 held out for competitions, leaving 93 released with 91 train + 1 validation + 1 test), but the text should state this relationship explicitly to avoid confusion.","section":"Section 3.2, Table 2, Appendix B.6"},{"comment":"The scaling plots for speech detection and phoneme classification use only three training volumes (5, 10, and 50 hours). The 'approximately logarithmic' claim should be qualified as suggestive, or additional intermediate training volumes should be included to support a stronger scaling-law statement.","section":"Figure 3"},{"comment":"Table 2 totals 52.32 released hours while Table 6 lists total audiobook hours as 53:02:41. The difference is presumably the two competition holdout sessions, but the paper should state this explicitly so readers do not treat the two totals as inconsistent.","section":"Table 2, Table 6"},{"comment":"The Macro F1 score is at chance level (0.0253 vs 0.0258). The text explains that this is due to the power-law phoneme distribution, but the table or caption should mark this clearly so the result is not misread as a failure of the model.","section":"Section 4.2, Table 4"},{"comment":"Occasional typos and formatting issues should be corrected: 'cannon' for 'canon' in Appendix B.1, 'V owel' in the Figure 2 caption, 'learnign' in Appendix B.6, and inconsistent use of 'V AD' vs 'VAD'.","section":"Throughout"}],"recommendation":"minor_revision","confidential_remarks":"The dataset-size claim is solid and the paper is a good contribution to the community. The forced-alignment validation issue is the only substantive concern; it can be addressed in revision without changing the central conclusions. The minor arithmetic inconsistencies should also be cleaned up. The paper is otherwise ready for publication."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Colleague, here's the take: LibriBrain is a real resource contribution. 52 hours of single-subject MEG listening to connected speech, five times the Armeni dataset, with public code, standardized splits, and honest baselines. The headline size claim checks out: Tables 1, 2, and 6 are internally consistent enough that even the most conservative reading leaves over 50 hours. If the data quality holds up on inspection, this will be a standard benchmark for non-invasive speech decoding.\n\nWhat's genuinely new is the resource itself. No prior public MEG dataset has this much within-subject depth for speech. The paper does the release properly: HDF5/TSV per session, pip-installable loader, train/val/test splits with validation and test recorded on a separate day to guard against leakage, permutation tests against random baselines, and compute estimates for reproduction. That's all solid, and the field needs this.\n\nSoft spots: first, the forced-alignment label validation is thin. Table 8 only shows that longer linguistic units get longer durations, not that onset boundaries are accurate enough for 39-way phoneme classification. The authors describe extensive manual correction, but they don't quantify residual alignment error. That's a caveat for the benchmark results, not for the central resource claim. Second, the word classification numbers come from a companion preprint by overlapping authors. They are transparent about this, compare against Armeni on equal terms, and reproduce with permission, but the current paper doesn't independently verify those numbers. A referee should treat them as secondary. Third, minor inconsistencies: session counts appear as 95, 93, and 91 in different places, and the license description shifts between public-domain and CC BY-NC. Paper-polish issues, not red flags.\n\nOverall, the central argument holds up. This is a dataset paper, and the dataset is the contribution. The scaling analysis is a known pattern, and the baselines are standard, but that doesn't undercut the resource value. I'd send this to review. The weaknesses are minor and fixable, and the dataset will be cited widely. Who benefits: anyone building or benchmarking non-invasive speech decoders. A serious referee will find the resource sound and the write-up mostly clean.","headline":"A solid, genuinely useful dataset release; the 50-hour claim holds up, and the forced-alignment label validation is the one real caveat.","tokens_in":27956,"tokens_out":1800,"would_cite":true,"duration_ms":16963,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"A new open dataset records over 50 hours of MEG from one person listening to Sherlock Holmes audiobooks, making it the largest single-subject MEG speech corpus to date and providing reproducible benchmarks for speech detection, phoneme…","keywords":["MEG","speech decoding","single-subject dataset","brain-computer interface","phoneme classification","word classification","speech detection","within-subject scaling"],"falsifier":"Take a held-out subset of LibriBrain audio, have a trained phonetician mark phoneme boundaries by hand on a few thousand tokens, and compare those boundaries to the released event files; if the median absolute onset disagreement is comparable to or larger than the mean phoneme duration (about 80 ms) or is systematically different across phoneme classes, then the phoneme classification benchmark is measuring alignment artifacts rather than neural speech decoding.","tokens_in":27060,"feed_emoji":"🧠","tokens_out":6141,"duration_ms":61220,"temperature":0.7,"pith_summary":"This paper introduces LibriBrain, an open magnetoencephalography dataset of over 50 hours of recordings from a single participant listening to audiobooks, and claims it is the largest single-subject MEG dataset for speech decoding to date—about five times larger than the next comparable resource and tens of times larger than most. The authors' central purpose is to give machine-learning researchers a deep, within-subject resource that supports scaling studies in non-invasive speech decoding, together with standard train/validation/test splits, a Python data loader, and reproducible baselines. The baselines show that decoders trained on the data beat chance on speech detection, phoneme classification, and word classification, and that performance grows roughly logarithmically as training data increases. If the central claim is right, the field gains a public benchmark that removes a major bottleneck: most non-invasive speech datasets give only one to two hours per subject, which is too shallow to train strong decoders.","feed_headline":"52 hours of one brain's MEG data now open for speech decoding","feed_subtitle":"Largest single-subject MEG speech corpus, 5x the next, with benchmarks for detection, phoneme, and word decoding.","key_machinery":"The carrying object is the dataset's within-subject depth: 95 recording sessions from one participant, the same narrator and genre throughout, 306 sensor channels downsampled to 250 Hz with 4 ms samples, and event files giving word and phoneme onsets produced by an automatic forced aligner followed by extensive manual correction. This depth lets a single model train on 51.57 hours of data without cross-subject variability, and the fixed splits—with validation and test sessions recorded on a separate day from all training data—make the benchmark resistant to session-level information leakage. A second mechanism is phoneme averaging: the paper shows that averaging repeated MEG windows for the same phoneme before classification raises balanced accuracy above 60% at 100 repetitions, which simulates a realistic brain-computer interface usage pattern.","core_discovery":"On the paper's own terms, the central discovery is that a single participant can be recorded for 52.32 hours of naturalistic listening and that the resulting brain responses support supervised decoding of speech presence, phoneme identity, and word identity, with performance improving roughly logarithmically as training data increases. The paper reports 306-channel MEG recordings from one native English speaker listening to seven Sherlock Holmes audiobooks read by a single narrator, yielding 466,230 word events and over 1.5 million phoneme events aligned to the audio. On the word classification task, the replicated state-of-the-art model reaches top-10 balanced accuracy of 0.3621 on LibriBrain versus 0.3261 on the next-largest within-subject dataset at matched data volume, and the scaling curves show LibriBrain continuing to improve where the comparison dataset plateaus. The dataset, splits, loaders, and baselines are released so that these results can be reproduced and extended.","pith_inferences":["If the logarithmic scaling law continues past 50 hours, collecting more hours from the same participant should yield further gains; a natural extension is to fit the scaling curve and predict performance at 100 hours before collecting them.","Because the data come from one narrator and one genre, the paper does not establish how models trained on LibriBrain transfer to other voices or speaking styles; a testable extension is fine-tuning on a second narrator and measuring the drop.","The manual correction effort suggests label quality above typical forced alignment, but the paper never quantifies onset error; an independent hand-labelled subset would settle whether phoneme baselines reflect neural signals or alignment noise.","The separate-day holdout design means the benchmark includes day-to-day variability; this makes the baselines conservative, and models that explicitly handle session effects might do better than the reported numbers."],"forward_implications":["Speech detection, phoneme classification, and word classification all beat chance on the held-out sessions, so the dataset demonstrably carries decodable speech information at the single-trial level.","Decoding performance rises roughly logarithmically as training data grows from 5 to 50 hours, so further within-subject scaling is the most direct route to better non-invasive decoders.","Averaging repeated MEG windows of the same phoneme pushes balanced accuracy above 60% at 100 repetitions, which simulates a practical BCI setting where users repeat inputs.","The fixed train/validation/test splits, with validation and test recorded on a separate day, give the community a leakage-resistant standard for comparing methods.","Word classification on LibriBrain outperforms the next-largest dataset at matched training volumes, supporting the claim that depth, not breadth, drives decoding progress."],"supporting_citations":[{"why":"Provides the previous largest within-subject MEG speech dataset (about 10 hours per subject) used as the comparison for the 5x depth claim and as the comparison dataset for word classification.","marker":"Armeni et al. (2022)"},{"why":"Source of the word classification method and reproduced results; their model trained on LibriBrain outperforms the same model trained on the Armeni dataset.","marker":"Jayalath et al. (2025a)"},{"why":"Established that deep single-subject data can outperform broader multi-subject training at matched hours and supplied the word-embedding and transformer classification approach that is replicated.","marker":"d’Ascoli et al. (2024)"},{"why":"Provided the brain encoder and spatial attention architecture used in the word classification pipeline.","marker":"Défossez et al. (2023)"},{"why":"Supplied the forced aligner used to create the phoneme and word event files from the audiobook audio.","marker":"Ochshorn and Hawkins (2015)"},{"why":"Provided the SEANet architecture from which the phoneme classification and speech detection baseline models are adapted.","marker":"Tagliasacchi et al. (2020)"}],"fun_headline_variants":["Largest single-subject MEG speech dataset: 52 hours, open source","One brain, 52 hours: massive MEG speech dataset released","52-hour MEG speech corpus from one listener now public","Biggest within-subject MEG speech data: 52 hours, 5x next","Scaling brain data: 52 hours of MEG speech decoding released"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The phoneme and word labels come from an automatic forced aligner corrected by hand, and the paper only checks that longer linguistic units receive longer duration annotations, not that the onset times are accurate enough for 39-way phoneme classification at the millisecond scale.","fun_headline_variants_meta":{"raw":{"variants":["Largest single-subject MEG speech dataset: 52 hours, open source","One brain, 52 hours: massive MEG speech dataset released","52-hour MEG speech corpus from one listener now public","Biggest within-subject MEG speech data: 52 hours, 5x next","Scaling brain data: 52 hours of MEG speech decoding released"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000939,"raw_usage":{"total_tokens":4012,"prompt_tokens":940,"completion_tokens":3072,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":556,"completion_tokens_details":{"reasoning_tokens":2974}},"tokens_in":556,"tokens_out":3072,"duration_ms":21162,"temperature":1.0,"reasoning_tokens":2974,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-07T11:29:47.037043+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Take a held-out subset of LibriBrain audio, have a trained phonetician mark phoneme boundaries by hand on a few thousand tokens, and compare those boundaries to the released event files; if the median absolute onset disagreement is comparable to or larger than the mean phoneme duration (about 80 ms) or is systematically different across phoneme classes, then the phoneme classification benchmark is measuring alignment artifacts rather than neural speech decoding.","supporting_citations":[],"review_version":1}