{"id":"0163a20d-da96-49d2-a967-9a8b8d484086","arxiv_id":"2412.16771","paper_version":1,"verdict":"REJECT","confidence":"MODERATE","novelty_score":4.0,"correctness_risk":"high","formal_verification":"none","parameter_count":3,"one_line_summary":"SilVar fuses speech, image, and text via CLIP, Whisper, and LLaMA to perform reasoning visual question answering and object localization, but its claimed state-of-the-art results are not supported by its own tables.","lead":"This paper introduces SilVar, an open-source style multimodal model that combines CLIP, Whisper, and LLaMA so users can ask questions about images by speaking. It also contributes a new speech-instruction dataset for object localization with GPT-4 generated reasoning.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The paper's central 'SOTA' claim is contradicted by its own Tables 4 and 5: SilVar trails LLaVA-1.5-13B (31.8 vs 36.4) on MMMU and LLaVA-13B (63.21 vs 90.92) on ScienceQA, so the headline result fails on the reported numbers.","rationale":"The reader's verdict is REJECT, and our stress-test supports it. We identify the single most load-bearing concern as the internal contradiction between the claimed SOTA status and the paper's own Tables 4 and 5. This is a correctness risk independent of the data-leakage question: even if the MMMU validation scores were perfectly clean, SilVar's 31.8 still trails LLaVA-1.5-13B's 36.4, and its ScienceQA 63.21 trails LLaVA-13B's 90.92. The reader's weakest_assumption focused on potential MMMU validation leakage (Section 3 preprocessing of 775 validation samples), which is a real but secondary issue; excluding those samples would only lower the reported number, not restore the SOTA claim. The paper also lacks any comparison against other speech-driven multimodal models on these benchmarks, so no 'SOTA among speech models' interpretation is supported. The architecture and dataset may have merit, but the headline empirical claim is not merely overstated; it is contradicted by the reported numbers. Thus the reject verdict stands unchanged.","tokens_in":14565,"tokens_out":4110,"duration_ms":31844,"concrete_test":"Perform a direct arithmetic comparison of SilVar's reported scores against every competitor row in Tables 4 and 5. Since LLaVA-1.5-13B's 36.4 exceeds SilVar Text's 31.8 on MMMU, and LLaVA-13B's 90.92 exceeds SilVar Speech's 63.21 on ScienceQA, the SOTA statement is refuted under the paper's own reported numbers.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The abstract and conclusion assert that SilVar 'achieves SOTA performance on the MMMU and ScienceQA benchmarks.' This is the paper's central claim, and it is internally contradicted by the comparison tables. Table 4 shows SilVar Text at 31.8 on MMMU validation, below LLaVA-1.5-13B (36.4) and Qwen-VL-7B-Chat (35.9); SilVar Speech scores 30.2/30.4, lower still. Table 5 shows SilVar Speech at 63.21 and end-to-end at 63.45 on ScienceQA, far below LLaVA-13B (90.92), LaVIN-13B (90.83), Chat-UniVi (88.78), and LLaMA-Adapter (85.19). The authors offer no restricted comparison class such as 'speech-driven open-source models,' and Table 5 contains no speech-input baselines at all. Section 5.2 justifies validation-only evaluation by noting that SOTA models achieve similar scores on test and validation, but that does not support a SOTA claim. Therefore, the central claim is false as stated, independent of any additional concern about the 775 MMMU validation samples used in preprocessing (Section 3) potentially overlapping with the reported evaluation.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"SilVar is an end-to-end speech-driven visual language model built on CLIP, Whisper, and LLaMA 3.1-8B, trained in two stages on ScienceQA, MMMU, LISA, and a newly introduced SilVar dataset for speech-based reasoning and object localization. The paper claims state-of-the-art performance on MMMU and ScienceQA while also investigating conversational, simple, and complex reasoning instructions in both speech and text modalities. The model architecture, training pipeline, dataset generation procedure, and a small ablation of audio adapters are described.","tokens_in":14824,"tokens_out":3626,"duration_ms":29360,"significance":"Speech-driven reasoning for visual question answering is a timely and useful direction, and the release of a new dataset with human-verified bounding boxes and speech instructions is a positive contribution. The two-stage training pipeline and the comparison of linear, MLP, and Transformer audio adapters are also useful engineering results. However, the central claim of state-of-the-art performance is contradicted by the paper's own tables, and the MMMU evaluation is compromised by the likely overlap between validation samples used in training and the reported validation evaluation. The dataset and code release, if made accessible with a proper URL, would be a valuable resource for the community even if the benchmark claims are revised downward.","major_comments":[{"comment":"The abstract and conclusion claim that SilVar 'achieves SOTA performance on the MMMU and ScienceQA benchmarks,' but the paper's own tables contradict this. On MMMU validation, SilVar Text scores 31.8, below LLaVA-1.5-13B (36.4) and Qwen-VL-7B-Chat (35.9). On ScienceQA, SilVar Speech scores 63.21, far below LLaVA-13B (90.92), LaVIN-13B (90.83), and Chat-UniVi (88.78). The paper does not define a restricted comparison class such as 'speech-driven open-source models,' and Table 5 contains no speech-input baselines at all. The SOTA claim is therefore not supported by the reported results.","section":"§5.2, Table 4 and §5.3, Table 5"},{"comment":"The MMMU evaluation is likely in-distribution. The paper reports curating a subset that includes 775 samples from the MMMU validation set for preprocessing, Table 2 lists MMMU as used in both Stage 1 and Stage 2 training, and Section 5.2 evaluates only on the MMMU validation set. The paper never states that the 775 preprocessed validation samples were excluded from training. If they were not excluded, the reported validation score of 31.8 is an in-distribution accuracy, not a legitimate benchmark result. The authors must clarify the exact data split and, ideally, evaluate on the official MMMU test set.","section":"§3, Table 1, Table 2, §5.2"},{"comment":"The justification for evaluating only on the validation set—'SOTA models achieve similar performance on both test and validation datasets'—is not a substitute for reporting test-set numbers and does not support a SOTA claim. The official MMMU test set contains 10,500 samples, and the validation set is not a held-out benchmark if any part of it contributed to training. The paper should report official test-set accuracy or clearly and verifiably document that the validation samples used in preprocessing were fully excluded from all training stages.","section":"§5.2"},{"comment":"The localization accuracy at IoU = 0.5 is between 21.11% and 26.32% across all instruction types. Since object localization is one of the paper's two core contributions, these numbers need to be placed in context with baselines or an explicit discussion of what accuracy level constitutes a meaningful capability. Without any comparison for the localization task, the reported numbers do not substantiate the claim that SilVar can perform reasoning-based object localization effectively.","section":"§5.1, Table 3"}],"minor_comments":[{"comment":"'ScieneQA' is a typo that should be 'ScienceQA'.","section":"§5.3"},{"comment":"Cross-references are often given as bare numbers (e.g., 'as shown in 2', 'as shown in 1') instead of 'Figure 2' and 'Figure 1'; the captions should also number all figures consistently.","section":"Throughout"},{"comment":"The 'Test' column heading is misleading because SilVar rows report no test values; consider renaming the column to 'Val/Test' or clarifying which models have official test scores.","section":"Table 4"},{"comment":"The abstract and conclusion state 'Our code and dataset are available here,' but no URL, repository identifier, or anonymous link is provided in the manuscript.","section":"Abstract and §7"},{"comment":"The caption refers to 'blue' and 'cyan' highlights, which are not reproducible in monochrome print; please use boldface or symbols instead.","section":"Table 3"}],"recommendation":"reject","confidential_remarks":"The paper is a systems/application paper whose headline contribution is a SOTA benchmark claim. That claim is internally contradicted by Tables 4 and 5, and the MMMU evaluation appears to use validation samples that overlap with training data. These are load-bearing flaws that cannot be fixed by minor rewriting; they require new experiments and a substantial reframing of the claims. The dataset and the speech-instruction training pipeline are potentially useful, so a future revision that corrects the evaluation and drops the SOTA framing could be considered, but the current manuscript does not meet the bar for publication."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Quick read of SilVar. The useful core is the dataset: a small set of 998 COCO images with GPT-4-generated reasoning questions at three levels, human-verified, with manually labeled bounding boxes, converted to speech with 50+ voices. That is a real resource for anyone working on spoken instructions for VQA and grounding. The system itself is a straightforward LLaVA-style assembly of CLIP, Whisper, and LLaMA-3.1-8B with learned projectors and a two-stage training pipeline. Nothing wrong with that, but it is not a new method.\n\nThe problem is the headline. The abstract and conclusion claim SOTA on MMMU and ScienceQA. The paper's own tables refute that. On MMMU val, SilVar text gets 31.8 vs LLaVA-1.5-13B's 36.4; on ScienceQA, SilVar speech gets 63.21 vs LLaVA-13B's 90.92. The authors never define a restricted class like 'among open-source speech-driven models' that would make those numbers competitive, and they include no speech-input baselines on ScienceQA.\n\nThe MMMU numbers are also suspect for a second reason. Section 3 says they curated 775 samples from the MMMU validation set for their pipeline, and Table 2 shows MMMU used in both training stages. They then report accuracy only on the MMMU validation split. That is a leakage risk the paper never addresses. They hand-wave that test and val scores are similar for SOTA models, but that does not justify skipping the test set, and it does not establish anything about their own model.\n\nSo the central empirical claim fails on the reported numbers, and the evaluation protocol is weak. That said, the dataset and the comparison of reasoning levels for speech input could be useful to the community if the code and data are actually released and the evaluation is redone properly. A revision that re-runs on official test splits, rules out overlap, adds an ASR+text-VLM baseline, and drops the SOTA wording would make this a reasonable workshop-quality or maybe even conference-short-paper contribution. As is, it is not publishable.\n\nMy recommendation: send it to review if you want a referee to verify the dataset and push for honest evaluation; it's a borderline desk-reject, but the dataset gives it enough substance to warrant referee time. I would not cite it in its current form.","headline":"A genuinely useful speech-reasoning dataset wrapped in a paper whose SOTA claim is refuted by its own tables.","tokens_in":15408,"tokens_out":2847,"would_cite":false,"duration_ms":23144,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"SilVar is an open-source multimodal model that lets users ask an image questions by voice and get reasoned answers with bounding boxes, at a measurable accuracy cost relative to text.","keywords":["speech-driven visual question answering","object localization","multimodal reasoning","Whisper audio encoder","CLIP vision encoder","LLaMA 3.1","speech instruction tuning","SilVar-Bench dataset"],"falsifier":"Inspect the released data splits and training scripts: if any of the 775 preprocessed MMMU validation samples appear in Stage 1 or Stage 2 training, the reported 31.8 validation score is in-distribution and cannot be compared with published MMMU results; a clean check is to run the released model on the official MMMU test set and compare it with the same baselines.","tokens_in":14300,"feed_emoji":"🗣️","tokens_out":8792,"duration_ms":70857,"temperature":0.7,"pith_summary":"SilVar is an open-source multimodal model that takes a spoken or typed question together with an image and produces a reasoned text answer plus bounding boxes for the object in question. The paper is trying to establish that speech can be a first-class instruction modality for visual question answering and object localization, rather than just a front-end transcription step, and it contributes a two-stage training pipeline plus a new benchmark to make that case. The pipeline aligns a Whisper audio encoder, a CLIP visual encoder, and the LLaMA 3.1-8B language model, and the new dataset supplies 998 COCO images with GPT-4-generated conversational, simple-reasoning, and complex-reasoning speech questions. The abstract states that SilVar achieves state-of-the-art results on MMMU and ScienceQA, while the paper's own tables show a more modest outcome: SilVar is competitive with strong text-based models but trails them on both benchmarks. Read in good faith, the contribution is a working open-source speech-driven reasoning pipeline and evidence that spoken instructions cost a measurable but not prohibitive amount of accuracy.","feed_headline":"Voice questions drive a visual model to reason and localize","feed_subtitle":"Open-source SilVar pairs Whisper, CLIP and LLaMA 3.1 so users can ask images questions by voice.","key_machinery":"The carrying mechanism is a two-stage speech-instruction tuning pipeline built from three open-source components: Whisper (an audio encoder), CLIP (a visual encoder), and LLaMA 3.1-8B (the language backbone). Stage 1 trains the audio side to align speech with text in the reasoning domain using science and multimodal-reasoning datasets; Stage 2 fine-tunes the whole model to answer from direct audio input on those datasets plus the new SilVar-Bench. The adapter design is intentionally simple, with linear or MLP projections for audio and two linear layers with GELU for vision, and the ablation shows that larger Transformer audio adapters buy almost nothing, implying that the Whisper encoder's final layer already carries the needed structure. The new dataset supplies the speech-instruction signal: 998 COCO images, three reasoning question types per object, human-verified answers, manual bounding boxes, and synthetic speech from over fifty voices.","core_discovery":"The central claim, on the paper's own terms, is that a single open-source model can align audio, image, and language features well enough to answer complex reasoning questions and localize objects from spoken instructions. SilVar does this by encoding speech with Whisper into 768-dimensional features, projecting them to the LLaMA input space, and concatenating them with CLIP visual tokens before the language model generates text and bounding boxes. The paper's experiments indicate that complex-reasoning speech prompts outperform simple and conversational speech prompts on its own benchmark, and that end-to-end training with speech slightly improves the MMMU and ScienceQA scores over the pipeline variant. The same tables show that text instructions consistently score higher than speech instructions, so the defensible finding is that speech-driven reasoning works at a measurable accuracy cost, not that it surpasses text-based state of the art.","pith_inferences":["Editorial: the abstract's 'state-of-the-art' phrasing is stronger than the tables support; the reproducible contribution is the first open speech-driven reasoning pipeline and benchmark, not a new accuracy record.","Editorial: the speech-to-text gap is likely dominated by the tiny 39-million-parameter Whisper encoder; retraining with a larger ASR encoder and measuring the same benchmarks would test this directly.","Editorial: the MMMU comparison is only meaningful if the 775 preprocessed validation samples were excluded from both training stages, which the paper does not state; releasing the exact split would settle the question.","Editorial: because the benchmark questions are GPT-4-generated from captions, some questions may leak the target object through wording; a no-image human baseline would quantify how much of the 'reasoning' is linguistic rather than visual."],"forward_implications":["Voice becomes a usable input modality for open-source visual reasoning, not just a speech-to-text front end, so applications can be built without proprietary speech models.","Spoken complex-reasoning prompts produce better grounded answers than simple or conversational prompts, so prompt-design techniques transfer to the audio channel.","A model trained on SilVar-Bench can output bounding boxes from speech, opening a route to hands-free assistive tools such as scene description and navigation aids.","Because the paper's own measurements show text instructions are more accurate, deployed speech-driven systems should keep a text fallback for high-stakes questions."],"supporting_citations":[{"why":"Supplies the CLIP ViT-B/32 visual encoder whose tokens are projected for the language model.","marker":"[43]"},{"why":"Supplies the Whisper audio encoder that turns spoken instructions into 768-dimensional features.","marker":"[44]"},{"why":"Provides the LLaMA 3.1-8B language backbone that generates the reasoned text and bounding boxes.","marker":"[17]"},{"why":"Provides the GPT-assisted instruction-data recipe that the SilVar-Bench generation is modeled on.","marker":"[33]"},{"why":"The MMMU benchmark used for speech-rendered training and for the validation-score comparison.","marker":"[66]"},{"why":"The ScienceQA benchmark used for speech-rendered training and evaluation.","marker":"[36]"},{"why":"The LISA dataset that supplies the reasoning-based object localization training signal.","marker":"[27]"},{"why":"The MiniGPT-v2-style two-layer GELU visual adapter design that SilVar adopts for image projection.","marker":"[71]"},{"why":"The COCO 2014 dataset from which the 998 SilVar-Bench images are drawn.","marker":"[31]"},{"why":"GPT-4, the model that generated the questions and detailed reasoning answers for SilVar-Bench.","marker":"[41]"}],"fun_headline_variants":["Speech prompts let a multimodal model reason and locate objects, with a text-accuracy gap","Voice-driven visual reasoning: SilVar pairs Whisper, CLIP, LLaMA, but trails text inputs","For spoken questions, SilVar reasons and localizes—just less accurately than typed ones","Multimodal model answers spoken visual queries; text still wins on accuracy","SilVar: speech-to-vision reasoning works, yet lags behind text-based instructions"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The whole benchmark comparison rests on the assumption that the 775 MMMU validation samples used in the speech pipeline were not included in either training stage, but the paper never states that exclusion.","fun_headline_variants_meta":{"raw":{"variants":["Speech prompts let a multimodal model reason and locate objects, with a text-accuracy gap","Voice-driven visual reasoning: SilVar pairs Whisper, CLIP, LLaMA, but trails text inputs","For spoken questions, SilVar reasons and localizes—just less accurately than typed ones","Multimodal model answers spoken visual queries; text still wins on accuracy","SilVar: speech-to-vision reasoning works, yet lags behind text-based instructions"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000767,"raw_usage":{"total_tokens":3400,"prompt_tokens":947,"completion_tokens":2453,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":563,"completion_tokens_details":{"reasoning_tokens":2340}},"tokens_in":563,"tokens_out":2453,"duration_ms":14992,"temperature":1.0,"reasoning_tokens":2340,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-11T10:15:02.019587+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Inspect the released data splits and training scripts: if any of the 775 preprocessed MMMU validation samples appear in Stage 1 or Stage 2 training, the reported 31.8 validation score is in-distribution and cannot be compared with published MMMU results; a clean check is to run the released model on the official MMMU test set and compare it with the same baselines.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"GPT-4, the model that generated the questions and detailed reasoning answers for SilVar-Bench."},{"cited_title":"Ro- bust speech recognition via large-scale weak supervi- sion","cited_arxiv_id":null,"evidence_quote":"Supplies the Whisper audio encoder that turns spoken instructions into 768-dimensional features."},{"cited_title":"Visual instruction tuning","cited_arxiv_id":null,"evidence_quote":"Provides the GPT-assisted instruction-data recipe that the SilVar-Bench generation is modeled on."},{"cited_title":"Learn to explain: Mul- timodal reasoning via thought chains for science ques- tion answering","cited_arxiv_id":null,"evidence_quote":"The ScienceQA benchmark used for speech-rendered training and evaluation."},{"cited_title":"Lisa: Reasoning segmentation via large language model","cited_arxiv_id":null,"evidence_quote":"The LISA dataset that supplies the reasoning-based object localization training signal."},{"cited_title":"Microsoft coco: Com- mon objects in context","cited_arxiv_id":null,"evidence_quote":"The COCO 2014 dataset from which the 998 SilVar-Bench images are drawn."}],"review_version":1}