{"id":"bb692aec-ba21-4a44-bde3-fd67afe2e87f","arxiv_id":"2506.14212","paper_version":1,"verdict":"REJECT","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":2,"one_line_summary":"A neurosymbolic pipeline combining LLM parsing, audio classification, and Bayesian reasoning achieves r=0.78 correlation with human judgments on a new hidden-object guessing task.","lead":"This paper presents a neurosymbolic model that parses video, audio, and language cues into structured descriptions and then uses a Bayesian-style inference to guess which hidden objects are in which boxes. In a human experiment, the model's guesses correlated with human judgments more strongly than unimodal ablations or a large vision-language model.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Post hoc exclusion of 5 of 25 stimuli based on human agreement, with no confidence intervals or significance tests on the correlations, leaves the headline r=0.78 versus 0.55/0.52/0.31 unverified.","rationale":"I considered the reader's identified weakest assumption (using the audio posterior as a likelihood and assuming conditional independence). That is a genuine mathematical flaw and could undermine the claim that the model performs principled Bayesian inference. However, the most load-bearing concern for the central claim (r=0.78 vs baselines) is the post hoc exclusion of 5 stimuli based on human agreement. This exclusion is explicitly documented in the paper, so it is not an artifact. The paper gives no analysis on the full sample and no statistical comparison of correlations. Because the central empirical claim rests entirely on these correlations, the missing robustness checks make the claim unverified. My proposed test would settle whether the headline result survives without exclusions. The reader's verdict of REJECT is appropriate; my concern reinforces it. Therefore I do not change the verdict.","tokens_in":594,"tokens_out":2696,"duration_ms":141335,"concrete_test":"Recompute the model and baseline correlations on the full 25-stimulus dataset with no exclusions, and report 95% confidence intervals for each r and for the differences (full model minus each baseline) using a bootstrap or Fisher z-transform. If the full model's correlation is not significantly higher than the best baseline or drops below 0.7, the central claim fails.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper's headline claim is that the full neurosymbolic model correlates strongly with human judgment (r=0.78) and outperforms unimodal and VLM baselines (r=0.55, 0.52, 0.31). The Experiment section states: 'We excluded 5 stimuli in the experiment due to low agreement among human participants (split-half correlation less than 0.8).' This is a post hoc exclusion of 20% of the stimuli based on the outcome variable (human consistency). It is unclear whether this exclusion inflates the full model's correlation more than the baselines'. The Results section reports only point estimates and no confidence intervals or significance tests for the difference between correlations. Without the full 25-stimulus analysis or a rigorous comparison, the claimed advantage could be an artifact of selective reporting. This is the load-bearing fragility: the central empirical claim is not statistically established. The reader's concern about substituting the audio posterior for the likelihood is real, but it affects the model's interpretation as 'Bayesian' rather than directly threatening the correlation result; the exclusion rule is more directly load-bearing for the claimed r values.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper presents a neurosymbolic model for reasoning about hidden objects in boxes from multimodal cues. The model parses language, vision, and audio into structured representations using LLMs, VLMs, and CLAP, then performs Bayesian inference over all possible object-placement hypotheses. The authors introduce the 'What's in the Box?' task, collect human judgments from 54 participants on 25 stimuli (with 5 excluded), and report that their full model correlates with human judgments at r=0.78, outperforming audio-only (0.55), vision-only (0.52), and Gemini 2.0 Flash (0.31) baselines. They argue that structured Bayesian integration of neural perceptual outputs yields more human-like inference than monolithic multimodal models.","tokens_in":7878,"tokens_out":6241,"duration_ms":62616,"significance":"If the results hold, the paper would provide a useful demonstration that neurosymbolic inference over neural perceptual outputs can capture human multimodal reasoning in an open-ended task, and the WiTB paradigm could be a valuable benchmark. Strengths include a clear architectural decomposition, no fitting of the model to human judgments (so the correlation is not circular in the standard sense), a novel stimulus set, and qualitative analyses illustrating the complementary roles of audio and visual cues. The main weaknesses are statistical: the headline correlation advantage rests on post hoc stimulus exclusion and point estimates without confidence intervals, and the Bayesian derivation contains an unjustified substitution of a classifier posterior for a likelihood.","major_comments":[{"comment":"Equation (1) replaces the audio likelihood P(A|H) with the posterior P(H|A), which is only valid up to a constant if the prior over hypotheses is uniform and the observation A is fixed; the subsequent factorization in Eq. (2) is more problematic. The notation conflates the hypothesis index i with the box index: P(H_i|A_i) in the product is not well-defined because H_i is a full placement hypothesis while A_i is the audio of a single box. Moreover, P(H_i^n|A_i) is computed as the product of per-object classifier posteriors ∏_{o∈H_i^n} P(o|A_i), which assumes that the presence of objects within a box is conditionally independent given the box audio and penalizes boxes containing more objects. This is a load-bearing heuristic for the model's posterior, and the paper provides no generative justification or validation. The authors should either derive a proper likelihood P(A_i|contents) or demonstrate that this approximation does not drive the reported correlation.","section":"Generating and evaluating hypotheses, Eq. (1)-(2)"},{"comment":"The exclusion of 5 of 25 stimuli based on a split-half correlation of human judgments below 0.8 is post hoc selection on the outcome variable. No results for the full stimulus set or a sensitivity analysis are reported, so the headline r=0.78 versus the baseline correlations could be an artifact of favorable stimulus selection. In addition, the Results section reports only point estimates; there are no confidence intervals for the correlations and no significance tests for the differences between the full model and the unimodal or VLM baselines, despite the Figure 3 caption claiming a 'significantly better fit.' The authors should report the full-data correlations, bootstrap or Fisher-transformed CIs, and a formal comparison of dependent correlations (e.g., Steiger's test or a permutation test).","section":"Experiment and Results"},{"comment":"The CLAP model's posterior P(o|A_i) is used as the audio evidence term without calibration or validation. The audio of a shaken box is a joint acoustic mixture of all objects inside, so the per-object probabilities output by CLAP for candidate labels are not the same as the likelihood of observing that audio given a particular set of contents. The error analysis notes that CLAP misses nuanced sounds in mixtures, but the model's scores are treated as probabilities in the Bayesian update. This is closely related to Major Comment 1, and it should be addressed either by reframing the model as a heuristic scoring function rather than a Bayesian model, or by providing evidence that the uncalibrated posteriors behave like likelihoods in this task.","section":"Audio component; Discussion, Error Analysis"}],"minor_comments":[{"comment":"There are several typographical errors: 'Keywords:perception' is missing a space, 'Ernst, M. ) (2007)' has a stray parenthesis, and the Alais & Burr reference has a doubled closing parenthesis.","section":"Abstract and References"},{"comment":"The sentence 'where the of scale for each item automatically sums to 100' appears to be missing a word; it should likely read 'where the value of the scale for each item automatically sums to 100.'","section":"Experiment"},{"comment":"The caption states 'Error bars show standard error and CI indicates 95% confidence interval,' but the figure does not clearly distinguish which visual elements correspond to standard error versus confidence intervals; please clarify the plotting convention.","section":"Figure 3"},{"comment":"The paper does not state whether stimuli, human data, or code will be made available; for reproducibility, please add a data/code availability statement.","section":"General"},{"comment":"The notation in Eq. (1) is inconsistent: H_i denotes both a hypothesis and a box index; please use distinct subscripts for boxes (e.g., n) and hypotheses (e.g., i) throughout.","section":"Computational Model, Eq. (1)"}],"recommendation":"major_revision","confidential_remarks":"The paper proposes a genuinely interesting benchmark and a clear neurosymbolic architecture, and the fact that the model is not fitted to the human data is a real strength. However, the statistical foundation of the headline claim is currently too weak: the post hoc exclusion of 20% of the stimuli and the absence of any confidence intervals or significance tests mean the reported r-values cannot be taken at face value. The derivation issue in Eq. (1)-(2) is fixable by reframing the model as a heuristic scoring procedure or by providing a proper likelihood treatment, and the reanalysis of the full stimulus set is straightforward if the data are retained. I therefore recommend major revision rather than rejection, provided the authors can supply the full-data analysis and the requested statistical tests."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"First, the useful thing: WiTB is a clean new benchmark, and the pipeline—LLM-generated object priors, CLAP audio classification, Gemini visual parsing, and a combinatoric hypothesis evaluation—is a reasonable engineering recipe for a hard multimodal reasoning task. The qualitative examples are genuinely illustrative, and the comparison against a monolithic VLM baseline is the right kind of baseline to include. Credit where due: the authors are upfront that the audio likelihood is replaced by the classifier's posterior.\n\nNow the problems. The inference equation is not Bayesian as written. Equation (1) substitutes P(H|A) for P(A|H), and the paper explicitly admits this 'manipulation.' That substitution is not a harmless approximation; it changes the meaning of the update. The product of a likelihood from vision and a posterior from audio is a heuristic score, not a posterior. Conditional independence of audio and visual observations is also asserted without support. These are legitimate concerns, but note that the correlation with human judgment doesn't depend on the update being Bayesian in the strict sense—the model could still be predictive. So I'd treat this as a framing/justification flaw, not necessarily the thing that makes the empirical claim false.\n\nMore load-bearing is the stimulus exclusion. The authors built 25 stimuli and then dropped 5 because human split-half agreement was below 0.8. That is post hoc filtering on the outcome variable, and with 20 stimuli left the difference between r=0.78 and r=0.55/0.52/0.31 is reported without confidence intervals, significance tests, or a full-data analysis. The figure caption mentions 95% CIs, but the text gives no numbers. This selective reporting directly undermines the headline comparison.\n\nAlso, no code or data is released, which makes it hard to check the parsing, the priors, or the correlations.\n\nWho is this for? People interested in neurosymbolic reasoning, multimodal perception, and new benchmarks for human-like inference. The WiTB task itself is a useful addition to the toolkit. The model as described is not a validated Bayesian account, and the empirical advantage is not statistically supported as presented. But those are fixable with a revision: reframe the update as a heuristic or derive a proper likelihood, justify or remove the exclusion rule, add significance testing, and release the stimuli.\n\nRecommendation: send to peer review. The benchmark deserves referee time, and the flaws, while significant, are addressable. A serious referee should push for the above rather than a desk reject.","headline":"Useful new benchmark, but the Bayesian claim is heuristic and the headline correlations are statistically ungrounded.","tokens_in":8387,"tokens_out":2730,"would_cite":true,"duration_ms":29339,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"A neurosymbolic model parses audio, vision, and language with neural networks and combines them with a Bayesian engine, judging which unseen objects are in which box with $r=0.78$ agreement with humans, while unimodal and vision-language…","keywords":["neurosymbolic model","multimodal reasoning","Bayesian inference","audio-visual integration","hidden object reasoning","What's in the Box","human judgment correlation","vision-language model"],"falsifier":"A controlled replication could record each object-box configuration several times, estimate the true audio likelihood $P(A_i|H_i)$ empirically, and re-run the model with that likelihood in place of the classifier posterior; if placement rankings shift in cases where human judgments stay stable, the reported correlation depends on the substitution rather than on Bayesian integration itself.","tokens_in":7428,"feed_emoji":"📦","tokens_out":10480,"duration_ms":95913,"temperature":0.7,"pith_summary":"This paper aims to establish that human-like reasoning about objects hidden from view can be produced by a neurosymbolic pipeline: neural networks convert sound, video, and a list of object names into structured descriptions, and a Bayesian engine then scores every way the objects could be placed inside the boxes. On a new 'What's in the Box?' task, the full model's graded judgments correlate with human judgments at $r=0.78$, while audio-only and vision-only versions fall to $r=0.55$ and $r=0.52$, and a large vision-language model baseline drops to $r=0.31$. The result matters because it suggests that monolithic neural models, however large, do not currently capture the joint, constraint-based inference humans perform across senses, whereas even modest neural parsers followed by structured probabilistic reasoning can. The paper also argues that this integration is not a simple average of the modalities: the joint posterior over complete placements behaves differently from either unimodal judgment.","feed_headline":"Sound plus sight plus Bayes matches human box-guessing","feed_subtitle":"Full model matches human judgments at r=0.78; unimodal and large-model baselines lag at 0.55, 0.52, 0.31.","key_machinery":"The load-bearing mechanism is the Bayesian hypothesis evaluator. It defines the state space as every ordered partition of the $N$ object names into $K$ boxes, $|H|=K!S(N,K)$, assigns a uniform prior, and writes the joint likelihood as $P(O,A|H)\\propto \\prod_i P(O_i|H_i)\\,P(H_i|A_i)$, treating audio and visual evidence as conditionally independent. The visual factor $P(O_i|H_i)$ is computed by rejection sampling: drawing object and box dimensions from normal distributions and checking fit 1000 times. The audio factor is obtained by substituting the audio classifier's posterior $P(H_i|A_i)$ for the unavailable likelihood $P(A_i|H_i)$, a shortcut the paper acknowledges is not derived from a generative audio model. The posterior over hypotheses is then marginalized to produce per-object, per-box probability ratings.","core_discovery":"Using the 'What's in the Box?' game, the paper claims that combining neural perception with explicit Bayesian inference produces human-like graded beliefs about occluded objects. Given $N$ objects and $K$ boxes, the model enumerates all $K!S(N,K)$ placement hypotheses, assumes a uniform prior, and updates each hypothesis with $P(O,A|H)$, factored as a product over boxes of a visual fit term and an audio term. The visual fit comes from rejection-sampling object and box dimensions under normal distributions to check fit 1000 times; the audio term is obtained by querying the audio classifier with the sound of each box and candidate object labels. The authors report that this full model matches human probability ratings with $r=0.78$, that removing either modality lowers the correlation to roughly $0.5$, and that the vision-language baseline given the same video and instructions reaches only $r=0.31$. Qualitative examples show the full model resolving cases where vision and audio individually disagree, such as a yoga mat constrained by box size plus a laptop identified by collision sounds, and coins located by jingling.","pith_inferences":["One step beyond the paper: if the classifier-posterior-for-likelihood substitution is as fragile as it sounds, the model should degrade sharply when the audio classifier is miscalibrated or when two candidate objects make similar sounds; a calibration study on near-confusable object sets would show where the shortcut holds.","The conditional-independence approximation will likely break in richer scenes where one physical shake generates both the sound and the visible box motion; adding learned reliability weights per modality, which the paper lists as a limitation, would turn this approximation into a testable model of adaptive cue weighting.","The task itself could serve as a compact behavioral probe for physical reasoning in vision-language models: varying box-size constraints and object counts changes the difficulty of second-order reasoning such as two objects sharing a box, making it easy to test whether future foundation models close the gap.","In robotics, the same pipeline could maintain a belief distribution over sealed container contents from a single shake, a use the paper mentions only as a future direction."],"forward_implications":["Audio alone and vision alone in the same architecture both lose about a quarter of the correlation with human judgments, so the integration step, not any single parser, is doing the work.","A monolithic vision-language model given identical video and instructions trails the neurosymbolic model by a large margin, implying that large pretrained multimodal models are not yet a substitute for structured joint inference on ambiguous physical scenes.","The full model's judgments are not a weighted average of unimodal outputs, since joint inference over whole placements produces different marginals than either modality alone.","The hypothesis space grows combinatorially as $K!S(N,K)$, so applying the approach to more boxes or more objects will require approximate inference over placements."],"supporting_citations":[{"why":"Supplies CLAP, the audio classifier whose posterior P(o|A_i) fills the audio likelihood slot in the Bayesian update.","marker":"Elizalde et al., 2022"},{"why":"Supplies the vision parser for box dimensions and the monolithic vision-language baseline whose low correlation is the key comparison.","marker":"Gemini Team et al., 2024"},{"why":"Supplies the particle-filtering perspective used to frame hypothesis generation and sequential belief update.","marker":"Wills & Schön, 2023"},{"why":"Precedent for the two-module neurosymbolic architecture of neural parsing followed by a symbolic probabilistic engine.","marker":"Hsu et al., 2023"},{"why":"Precedent that human audiovisual perception integrates cues near-optimally with Bayes' rule, motivating the integration target.","marker":"Alais & Burr, 2004"},{"why":"Precedent for causal inference in multisensory perception, supporting the conditional-independence treatment of the cues.","marker":"Körding et al., 2007"}],"fun_headline_variants":["Neurosymbolic model nails human box-guessing","Bayes + neural nets beat large models at box guessing","Combining audio, vision, Bayes yields human-level guesses","Multimodal Bayesian model guesses like a human","What's in the box? Model reasoning matches humans"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The inference engine's output is only as valid as the shortcut that replaces the audio likelihood $P(A_i|H_i)$ with the audio classifier's posterior $P(H_i|A_i)$, combined with the assumption that sound and visual fit are conditionally independent.","fun_headline_variants_meta":{"raw":{"variants":["Neurosymbolic model nails human box-guessing","Bayes + neural nets beat large models at box guessing","Combining audio, vision, Bayes yields human-level guesses","Multimodal Bayesian model guesses like a human","What's in the box? Model reasoning matches humans"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000979,"raw_usage":{"total_tokens":4156,"prompt_tokens":941,"completion_tokens":3215,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":557,"completion_tokens_details":{"reasoning_tokens":3138}},"tokens_in":557,"tokens_out":3215,"duration_ms":22854,"temperature":1.0,"reasoning_tokens":3138,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-07T00:18:19.594898+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"A controlled replication could record each object-box configuration several times, estimate the true audio likelihood $P(A_i|H_i)$ empirically, and re-run the model with that likelihood in place of the classifier posterior; if placement rankings shift in cases where human judgments stay stable, the reported correlation depends on the substitution rather than on Bayesian integration itself.","supporting_citations":[{"cited_title":"A., & Wang, H","cited_arxiv_id":null,"evidence_quote":"Supplies CLAP, the audio classifier whose posterior P(o|A_i) fills the audio likelihood slot in the Bayesian update."},{"cited_title":"G., & Sch¨on, T","cited_arxiv_id":null,"evidence_quote":"Supplies the particle-filtering perspective used to frame hypothesis generation and sequential belief update."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Precedent for the two-module neurosymbolic architecture of neural parsing followed by a symbolic probabilistic engine."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Precedent that human audiovisual perception integrates cues near-optimally with Bayes' rule, motivating the integration target."}],"review_version":1}