{"id":"06b79db3-ddbc-44a6-9058-8d786962b289","arxiv_id":"2603.27043","paper_version":2,"verdict":"ACCEPT","confidence":"LOW","novelty_score":7.0,"correctness_risk":"low","formal_verification":"none","parameter_count":0,"one_line_summary":"MELI is a new 29.8-hour open-source speech corpus from 51 Mandarin-English bilinguals including full transcriptions, word/phone alignments, metadata, and attitude data for quantitative and qualitative bilingual research.","lead":"The paper introduces the MELI Corpus, a new open-source collection of 29.8 hours of audio from 51 Mandarin-English bilingual speakers with matched sessions in both languages and both read and spontaneous styles. Researchers in linguistics and speech technology might read it because the dataset links acoustics to language attitudes and supports cross-language comparisons.","discovery_kind":"new_application","skeptic_critique":{"model":"grok-4.3","headline":"No significant objection identified","rationale":"Reader's weakest assumption targets sample representativeness, which affects generalizability but is not load-bearing for whether the design itself supports the comparisons and linkages described. The paper's claim is internal to the corpus structure and is supported by the listed components.","tokens_in":1706,"tokens_out":242,"duration_ms":28708,"concrete_test":"Confirm in the full methods section that each speaker's Mandarin and English sessions share identical recording hardware, room, and gain settings; if any session-pair differs, recompute a sample acoustic feature (e.g., mean F0) on matched read sentences to quantify the artifact size.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim is that the corpus design (matched Mandarin/English sessions, read + interview styles, transcriptions/alignments, and attitude metadata) enables within-/cross-speaker and within-/cross-language acoustic comparisons plus attitude linkages. The abstract directly enumerates these features without internal contradictions or hidden assumptions that would prevent the claimed functionality. Representativeness of the 51 speakers is not required for the design to support the listed analyses.","agreement_with_reader":"disagree"},"referee_report":{"model":"grok-4.3","summary":"The manuscript introduces the MELI Corpus, an open-source resource of 29.8 hours of speech from 51 Mandarin-English bilingual speakers. It features matched sessions in Mandarin and English, each combining read sentences and spontaneous interviews on language varieties, standardness, and learning experiences. Audio is recorded at 44.1 kHz (16-bit, stereo), with full transcriptions, word- and phone-level force alignments, anonymization, and metadata on speakers' language attitudes. Descriptive statistics cover component durations (~14.7 h Mandarin, ~15.1 h English), token/type counts, and code-switching patterns (frequent in Mandarin sessions, more limited in English). The design is presented as enabling within-/cross-speaker and within-/cross-language acoustic comparisons as well as linkages between acoustics and stated attitudes for quantitative and qualitative analyses. The corpus will be released under CC BY-NC 4.0 with transcriptions, alignments, metadata, and documentation.","tokens_in":1742,"tokens_out":483,"duration_ms":49518,"significance":"If released as described, the corpus fills a notable gap in publicly available Mandarin-English bilingual speech resources by providing matched-language sessions and attitude metadata. This structure directly supports controlled acoustic comparisons and sociolinguistic investigations that are difficult with existing unbalanced or single-language corpora. The open licensing and inclusion of alignments and scans of labelled maps strengthen its potential for reuse in phonetics, code-switching studies, and attitude-acoustic correlation research.","major_comments":[],"minor_comments":[{"comment":"Abstract: the mean session durations (17.3 min Mandarin, 17.8 min English) are reported without clarifying whether these are per-speaker averages or totals; adding this detail would improve clarity of the descriptive statistics.","section":"Abstract"},{"comment":"The manuscript would benefit from an explicit comparison table or paragraph situating MELI against existing bilingual corpora (e.g., in terms of language pair balance, style matching, and attitude metadata) to highlight its distinctive contributions.","section":null},{"comment":"Ensure the full release includes the promised scans of labelled maps and documentation files; a brief appendix listing all released components with file names and formats would aid users.","section":null}],"recommendation":"accept","confidential_remarks":null},"author_rebuttal":{"model":"grok-4.3","summary":"We thank the referee for their positive review of the MELI Corpus manuscript and for recommending acceptance. The assessment accurately captures the corpus design, its matched-language sessions, attitude metadata, and potential for acoustic and sociolinguistic analyses.","responses":[],"tokens_in":1319,"tokens_out":64,"duration_ms":27006,"standing_objections":[]},"desk_editor":{"model":"grok-4.3","letter":"This paper is basically announcing a new corpus for Mandarin-English bilingual speech that has some practical features for comparison work. What they've done is record 51 bilingual speakers in both languages, with sessions that match up in content style: read sentences plus interviews about language attitudes and experiences. That's roughly 15 hours per language. Everything is transcribed and force-aligned, and they've attached metadata on the speakers' attitudes. They point out that code-switching happens more in the Mandarin interviews. The whole thing is going out open source with a non-commercial license, including the alignments and some map scans. The new part is having all those elements together in one place for the same speakers. It lets you look at acoustic differences within a person across languages or styles, and tie that to what they say about language attitudes. The stats they give are straightforward and show the scale. One soft spot is the lack of detail on recruitment and exact processing steps, since we're looking at the abstract. That means we don't know much about possible biases in who the speakers are or how reliable the alignments turned out. But the design itself doesn't have obvious flaws that would stop the intended uses. This kind of paper is aimed at researchers who study bilingual phonetics or sociolinguistics and need ready data for those cross-language questions. If you're doing work that requires controlled comparisons between Mandarin and English, this could save some time. It should go to a serious referee so the collection methods get checked. I'd recommend putting it through peer review.","headline":"A new matched bilingual corpus with attitude metadata that supports cross-language acoustic comparisons.","tokens_in":2221,"tokens_out":357,"would_cite":true,"duration_ms":51466,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":{"model":"grok-4.3","evidence":[],"headline":"MELI corpus introduces bilingual speech data with no RS-shaped machinery","alignment":"orthogonal","rationale":"Paper describes corpus design, recruitment, transcription/alignment pipeline, and sociophonetic metadata for Mandarin-English interviews. Central claims concern within-/cross-speaker acoustic comparison and attitude linkage. RS framework (reality_from_one_distinction, J-cost uniqueness, 8-tick periodicity, AlexanderDuality for D=3, phi-ladder constants) has no opinion on linguistic corpora or speech data collection.","tokens_in":49136,"confidence":"high","tokens_out":124,"duration_ms":11356,"cache_read_input_tokens":38528,"cache_creation_input_tokens":0},"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"grok-4.3","headline":"The MELI Corpus provides matched Mandarin and English speech from 51 bilingual speakers to enable acoustic comparisons linked to language attitudes.","keywords":["Mandarin-English bilinguals","speech corpus","language attitudes","acoustic comparison","code-switching","spontaneous interviews","read speech"],"falsifier":"If the speakers' acoustic data shows no reliable links to their stated attitudes or if the sample is found to be skewed by recruitment methods, the corpus's utility for the intended comparisons would be reduced.","tokens_in":2580,"feed_emoji":"🎙️","tokens_out":600,"duration_ms":58172,"temperature":0.7,"pith_summary":"The paper introduces the MELI Corpus, an open-source collection of 29.8 hours of speech from 51 Mandarin-English bilingual speakers. It features matched sessions in both languages with read sentences and spontaneous interviews about language varieties, standardness, and learning. The design allows for within- and cross-speaker, within- and cross-language acoustic comparisons and connects these to speakers' stated attitudes. This matters for researchers who want to study bilingual speech production alongside personal views on language. It supports quantitative acoustic work as well as qualitative analysis of attitudes in a single resource.","feed_headline":"New corpus connects bilingual speech acoustics to attitudes","feed_subtitle":"MELI Corpus offers 29.8 hours of matched Mandarin and English recordings from 51 speakers for cross-language analysis.","key_machinery":"The MELI Corpus, which integrates matched bilingual recordings with content on language attitudes and provides transcriptions and alignments.","core_discovery":"The authors present the MELI Corpus as a resource of 29.8 hours of speech from 51 bilingual speakers, with ~14.7 hours in Mandarin and ~15.1 hours in English. Each speaker completed read sentence tasks and spontaneous interviews in both languages. All audio is recorded at 44.1 kHz stereo, transcribed, force-aligned at word and phone levels, and anonymized. The corpus is designed to support acoustic comparisons across speakers and languages while linking those measurements to the speakers' expressed language attitudes.","pith_inferences":["The corpus could reveal how language attitudes influence phonetic variation in bilinguals.","It might be used to improve speech technology for code-switched Mandarin-English.","Future studies could test if attitudes predict specific pronunciation traits.","Connections to other bilingual corpora could be explored for cross-linguistic patterns."],"forward_implications":["Acoustic features can be compared for the same speaker in Mandarin versus English.","Code-switching patterns can be examined in relation to attitudes.","Both read and spontaneous styles are available for style-based comparisons.","The data is released with metadata and map scans for further study.","Quantitative and qualitative methods can be combined using the same speakers."],"fun_headline_variants":["MELI Corpus: 29.8 hours of Mandarin-English bilingual speech","MELI connects bilingual acoustics to language attitudes","51 speakers provide matched Mandarin and English interview data","Bilingual corpus enables cross-language acoustic comparisons"],"cache_read_input_tokens":2112,"weakest_assumption_plain":"The 51 speakers and their sessions represent typical Mandarin-English bilingual speech and attitudes without significant biases from recruitment or self-reporting.","fun_headline_variants_meta":{"raw":{"variants":["MELI Corpus: 29.8 hours of Mandarin-English bilingual speech","MELI connects bilingual acoustics to language attitudes","51 speakers provide matched Mandarin and English interview data","Bilingual corpus enables cross-language acoustic comparisons"]},"model":"grok-4.3","cost_usd":0.007589,"raw_usage":{"total_tokens":3400,"prompt_tokens":675,"num_sources_used":0,"completion_tokens":60,"cost_in_usd_ticks":75890500,"prompt_tokens_details":{"text_tokens":675,"audio_tokens":0,"image_tokens":0,"cached_tokens":64},"completion_tokens_details":{"audio_tokens":0,"reasoning_tokens":2665,"accepted_prediction_tokens":0,"rejected_prediction_tokens":0}},"tokens_in":675,"tokens_out":60,"duration_ms":38423,"temperature":1.0,"reasoning_tokens":2665,"cache_read_input_tokens":64,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-05-19T17:12:39.577800+00:00","model_set":{"reader":"grok-4.3"},"falsifier":"If the speakers' acoustic data shows no reliable links to their stated attitudes or if the sample is found to be skewed by recruitment methods, the corpus's utility for the intended comparisons would be reduced.","supporting_citations":[],"review_version":1}