{"id":"46c0b289-509c-46fb-b071-ebc95af517e9","arxiv_id":"2509.12263","paper_version":3,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":4,"one_line_summary":"Large multimodal models do far worse when collision videos violate familiar physics, and their small gains come from text exemplars, not the videos.","lead":"A new benchmark tests whether AI models that watch videos can learn unusual physics rules from a few examples. The results show 13 multimodal models mostly copy text hints in the examples and largely fail to use the videos.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The Fig. 4 regular–irregular accuracy gap is not calibrated against human/oracle performance, and several irregular scenarios are compared against LMC(Reg) despite different question types and visual properties; the headline 'weak inductive reasoning' may partly reflect task difficulty.","rationale":"The reader's weakest assumption already identifies the uncalibrated regular–irregular gap as the principal threat to the central claim. My stress-test agrees and sharpens the concern: the comparison is not merely uncalibrated, but also unmatched in several scenarios, since AMC, Red-LMC, Red-Pass, and CC are scored against LMC(Reg) with different question types and visual content. This is a concrete correctness risk for Finding 2 and the headline conclusion, not just a missing baseline. The proposed human/oracle calibration directly tests whether the gap reflects deficient inductive reasoning or task difficulty. I am not raising a secondary language-bias objection as the primary issue because the central novelty and the paper's headline discovery depend on the gap metric; the video-only comparison in Sec. 4.5 is important but secondary. The reader's conditional verdict remains appropriate: the concern is addressable, and a matched or calibrated evaluation would either confirm or refute the claim.","tokens_in":34136,"tokens_out":6536,"duration_ms":85384,"concrete_test":"Run a human-participant study (n >= 20 per scenario) on the exact InPhyRe items from Sec. 4.4, using the same three exemplars and the same regular/irregular pairs, and compute the same accuracy-gap metric. If the human mean gap is statistically indistinguishable from the LMM average, the metric does not specifically measure LMM induction failure; if humans remain near ceiling on irregular scenarios while LMMs drop, the concern is resolved. As a cheaper companion check, construct a matched regular version of AMC that uses the same rotation question and answer options but preserves angular momentum, and verify whether the AMC-relevant gap is reproduced; if it is not, AMC should not be benchmarked against LMC(Reg).","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central metric in Sec. 4.4 equates a negative difference between irregular 3-shot accuracy and the best regular-scenario accuracy with poor inductive physical reasoning. That interpretation requires the regular and irregular versions to be matched in every respect except the violated law, and requires the intended law to be uniquely inferable from three exemplars. Neither condition is established. In particular, AMC, Red-LMC, Red-Pass, and CC are all compared to LMC(Reg), even though their questions concern rotation, color-conditional motion, passing-through, and object permanence rather than the same velocity-change question. A negative gap can therefore reflect task- or question-specific difficulty, not an induction deficit. The paper also asserts in Sec. 1 that humans would \"easily adapt\" from demonstrations, but no human or oracle baseline is provided, so the absolute irregular accuracies in Table 3 have no calibration against what a competent inductive reasoner should achieve. Some irregular accuracies are near ceiling (e.g., InternVL3-8B at 94–100% on LMC, Wall, AMC), so the aggregate negative averages are not a uniform failure signature. Without calibration, the benchmark may still be useful as a relative comparator, but the paper's headline discovery—that LMMs have weak inductive physical reasoning—rests on an unvalidated metric.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper introduces InPhyRe, a synthetic video question-answering benchmark designed to measure inductive physical reasoning in large multimodal models (LMMs). Scenarios depict collision events that either follow or violate universal physical laws such as momentum conservation. Models are evaluated zero-shot and few-shot, with exemplars containing either videos plus question-answer pairs or videos only. The main findings are: (1) LMMs have limited and poorly applied parametric knowledge of physical laws; (2) LMMs show weak inductive physical reasoning when exemplars violate the laws encoded in their parameters; and (3) the observed inductive behavior is driven primarily by language, with visual inputs playing little role. The headline metric is the accuracy gap between irregular and regular scenarios, with negative gaps interpreted as poor inductive physical reasoning.","tokens_in":34406,"tokens_out":4576,"duration_ms":50624,"significance":"If properly validated, InPhyRe would be a valuable benchmark for a question that is both scientifically interesting and practically important: whether LMMs can adapt their physical reasoning to novel or counter-physical environments from a few demonstrations. The paper evaluates a diverse cohort of 13 LMMs, uses a reproducible synthetic generation pipeline, and reports clear, structured accuracy tables. The finding that models often fail when exemplars contradict parametric knowledge, and that performance collapses in video-only settings, is suggestive and worth community attention. However, the strength of the conclusions depends on the validity of the regular–irregular gap as a measure of inductive ability, and on the absence of human/oracle calibration. Those issues are currently unresolved.","major_comments":[{"comment":"The central metric compares every irregular scenario against the best regular scenario accuracy, but for AMC, Red-LMC, Red-Pass, and CC, the 'corresponding regular' is LMC(Reg), despite these scenarios differing in question type (rotation, color-conditional motion, passing-through, shape/object permanence) and visual properties. A negative gap therefore conflates induction failure with task-specific difficulty. To support the claim that 'a negative value indicates poor inductive physical reasoning,' the paper needs per-scenario matched regular controls or some other calibration that controls for task difficulty.","section":"§4.4, Fig. 4"},{"comment":"There is no human baseline or oracle calibration for irregular scenarios. The paper asserts in §1 that humans would 'easily adapt' from demonstrations, but this is not tested. Meanwhile, several irregular accuracies are near ceiling (e.g., InternVL3-8B at 94–100% on LMC, Wall, AMC), so the aggregate negative average is not a uniform failure signature. Without a competent-reasoner reference, the absolute irregular accuracies cannot be interpreted as showing weak inductive physical reasoning.","section":"§4.4, Table 3"},{"comment":"In the video-only condition, exemplars include videos plus randomly chosen option letters. In the video-text condition, exemplars include correct question-answer pairs. The comparison therefore varies not only the presence of textual information but also whether the labels are informative. A model may perform worse in the video-only condition because the random labels provide no usable signal, not necessarily because it is visually incapable. A control that keeps labels informative but removes the question text, or another design that separates label informativeness from modality, is needed to support Finding 3.","section":"§4.5, Fig. 5"},{"comment":"The reported accuracy differences lack confidence intervals or repeated-seed variability. Many entries are small (e.g., -0.20, +0.35, +0.30), and without uncertainty quantification it is hard to distinguish genuine effects from noise. Since each scenario contains around 2000 samples and 13 models are evaluated, paired bootstrap or stratified sampling would be straightforward and should be reported.","section":"§4.4, Fig. 4"}],"minor_comments":[{"comment":"The 'zero-shot' setting includes three assistant messages with random option labels. This is a reasonable formatting control, but calling it 'zero-shot' is potentially confusing. Clarify in the main text that random options are used only to enforce a constrained answer format and do not provide task information.","section":"§B.2"},{"comment":"The sentence 'Almost all LMMs showed significant deterioration in performance' uses 'significant' in a non-statistical sense. Recommend replacing with 'substantial' or adding a statistical test.","section":"§4.4 conclusion"},{"comment":"The claim that InPhyRe is 'the first visual question answering benchmark to measure inductive physical reasoning in LMMs' should be softened to 'to our knowledge,' as the paper does not exhaustively survey all recent benchmarks.","section":"§3"},{"comment":"The heatmap labels and the additional 'Average over LMMs' and 'Average over scenarios' rows are visually dense and difficult to read. Consider a cleaner formatting or a separate table for the averages.","section":"Fig. 4 and Fig. 5"},{"comment":"The heading 'AMC (regular)' in E.3 appears to be a misnomer, since AMC is an irregular scenario; the subsection describes open-ended outputs for AMC. Rename for clarity.","section":"§E.3"}],"recommendation":"major_revision","confidential_remarks":"The benchmark could become a useful community resource, and the multi-model evaluation is a strength. However, the headline discovery — weak inductive physical reasoning — rests on the regular–irregular gap, which is not yet adequately controlled for task difficulty. The authors should either construct matched regular versions for every irregular scenario or provide human/oracle calibration. Without that, the central claim is not ready for publication in a top venue. The language-bias finding also needs a cleaner experimental contrast."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Quick take: InPhyRe is a genuine contribution—the first visual benchmark for inductive physical reasoning—and the language-bias finding is the strongest result. The paper's other headline, 'weak inductive reasoning,' is less solid because the regular–irregular accuracy gap is not calibrated against any competent baseline.\n\nWhat's new: the benchmark design (synthetic collisions violating momentum/object permanence, with few-shot exemplars) and the 13-model evaluation. The ablations are thorough: visual-only vs visual-text, number of exemplars, NNER retrieval, quantization, plus detailed qualitative analysis of model outputs. That qualitative work gives real texture—models can recite momentum conservation but fail to apply it.\n\nThe soft spots: the central metric in Fig. 4 equates a negative gap between irregular 3-shot accuracy and the best regular-scenario accuracy with poor inductive reasoning. That requires regular and irregular versions to be matched except for the violated law, and the intended law to be uniquely inferable from three exemplars. Neither is established. Several irregular scenarios (AMC, Red-LMC, Red-Pass, CC) are compared against LMC(Reg) even though their questions target different physical quantities. So the aggregate drop could partly reflect task difficulty, not an induction deficit. There's also no human or oracle baseline, so we don't know what a competent inductive reasoner would score. Some irregular accuracies are near ceiling, which the aggregate hides.\n\nThe language-bias finding is on firmer ground: comparing video-only to video-text exemplars shows a large, consistent drop, and that's a within-scenario comparison. I'd trust that as an ordinal claim, though the absolute numbers still lack uncertainty estimates.\n\nAlso: the abstract says 'proprietary' models but the cohort is all open-weight; and there's no dataset/code link, which hurts reproducibility.\n\nVerdict: this deserves a serious referee. The benchmark is useful as a relative evaluation tool even if the 'weak inductive reasoning' interpretation needs calibration. A revision with a human/oracle baseline, matched regular-regular comparisons, repeated runs with CIs, and a public release would substantially strengthen it.","headline":"A solid new benchmark with a robust language-bias finding, but the 'weak inductive reasoning' claim needs calibration against human/oracle baselines.","tokens_in":34908,"tokens_out":2306,"would_cite":true,"duration_ms":25726,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Large multimodal models fail at inductive physical reasoning: a new benchmark shows they cannot infer physics that contradicts what they learned in training, and what little reasoning they do is driven by language, not vision.","keywords":["inductive physical reasoning","large multimodal models","visual question answering","benchmark","physical reasoning","language bias","momentum conservation","in-context learning"],"falsifier":"If human participants, given the same images and three exemplars, showed a similar or larger accuracy drop on the irregular scenarios, the gap would reflect task ambiguity or difficulty rather than a model-specific deficit in inductive physical reasoning.","tokens_in":33984,"feed_emoji":"📉","tokens_out":4567,"duration_ms":44381,"temperature":0.7,"pith_summary":"The paper introduces InPhyRe, a visual-question-answering benchmark that tests whether large multimodal models (LMMs) can infer physical laws from a few demonstration videos, including collisions that violate universal laws such as momentum conservation. Across 13 open and proprietary models, it finds that LMMs use demonstrations only when those demonstrations confirm what the model already knows; when the demonstrated physics contradicts their parametric knowledge, few-shot accuracy drops sharply. The paper also shows that the limited inductive physical reasoning that does occur is largely language-driven—when exemplars contain only videos, accuracy falls to near chance for several models. The authors conclude that LMMs treat physical laws as fixed rules rather than transferable mathematical models, casting doubt on their trustworthiness in safety-critical applications.","feed_headline":"LMMs fail at inductive physics from videos","feed_subtitle":"When demos contradict known physics, 13 LMMs drop sharply—and text, not video, drives their reasoning.","key_machinery":"The central object is InPhyRe, a synthetic benchmark of collision videos whose trajectories are generated by manually overriding a PyBullet simulation at the moment of collision so that they violate laws such as momentum conservation. Scenarios are grouped into irregular (law-violating) and regular (law-abiding) counterparts, and the key metric is the difference between few-shot accuracy in the irregular scenario and the model's best regular-scenario accuracy; a negative value is interpreted as weak inductive physical reasoning. A second experimental manipulation—exemplars with both video and question-answer pairs versus exemplars with video only—isolates the contribution of language and exp","core_discovery":"InPhyRe is the first visual question-answering benchmark purpose-built to measure inductive physical reasoning in LMMs by confronting them with collision videos that violate real-world physical laws, generated by intervening in a physics simulator. The paper reports three findings from 13 models: (1) LMMs can recite momentum and energy conservation but apply these laws inconsistently even in regular scenarios; (2) when exemplar videos follow laws unseen in training, almost all models show a substantial accuracy drop relative to regular scenarios, indicating weak inductive physical reasoning; and (3) removing question-answer pairs from the exemplars (video-only) cuts accuracy dramatically, sh","pith_inferences":["The authors do not draw this conclusion, but the irregular-versus-regular gap likely conflates inductive reasoning with task difficulty: if the irregular versions are harder for reasons other than the violated law, part of the drop would appear even in a perfectly inductive agent.","A testable extension the paper leaves implicit: adding human participants to the same scenarios would calibrate the gap; if humans show a comparable drop, the metric is not measuring model-specific inductive ability.","One could also vary exemplar count beyond three and provide explicit textual statements of the law to separate failure of visual perception, rule induction, and rule application.","The observation that larger models show larger language bias suggests that scaling up models may worsen, not fix, the reliance on text, which runs counter to the usual assumption that larger models are more robust reasoners."],"forward_implications":["In safety-critical settings where novel physics can occur, an LMM cannot be assumed to adapt from demonstrations; its predictions will default to parametric knowledge.","Exemplars help LMMs only when they align with the physical laws already encoded in the model's parameters; conflicting demonstrations are not incorporated.","The language-bias result implies that standard visual-question-answering accuracy can overstate a model's visual understanding; multimodal evaluation should separate textual and visual contributions.","Instruction tuning as currently practiced does not address this gap; the authors suggest simulation-based feedback signals, similar to reinforcement learning from human feedback, as a direction.","The same benchmark methodology—impossible or law-violating scenarios—can be applied to other branches of physics beyond mechanics."],"fun_headline_variants":["InPhyRe: LMMs struggle when physics rules are novel","LMMs can't induce physics from video, new benchmark shows","LMMs ignore visuals and rely on text for physics reasoning","Video-only demos expose LMMs' weak inductive physics","LMMs fail inductive reasoning when laws are unseen"],"cache_read_input_tokens":2304,"weakest_assumption_plain":"The paper assumes that the accuracy gap between irregular and regular scenarios measures inductive physical reasoning—that is, that the regular and irregular versions are matched in all respects except the violated law and that the intended law is uniquely inferable from three exemplars.","fun_headline_variants_meta":{"raw":{"variants":["InPhyRe: LMMs struggle when physics rules are novel","LMMs can't induce physics from video, new benchmark shows","LMMs ignore visuals and rely on text for physics reasoning","Video-only demos expose LMMs' weak inductive physics","LMMs fail inductive reasoning when laws are unseen"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.001036,"raw_usage":{"total_tokens":4219,"prompt_tokens":786,"completion_tokens":3433,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":530,"completion_tokens_details":{"reasoning_tokens":3346}},"tokens_in":530,"tokens_out":3433,"duration_ms":22203,"temperature":1.0,"reasoning_tokens":3346,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-04T17:41:52.966797+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"If human participants, given the same images and three exemplars, showed a similar or larger accuracy drop on the irregular scenarios, the gap would reflect task ambiguity or difficulty rather than a model-specific deficit in inductive physical reasoning.","supporting_citations":[],"review_version":1}