{"id":"601a96eb-488a-4aad-a08c-77232a5e584f","arxiv_id":"2508.00726","paper_version":1,"verdict":"UNVERDICTED","confidence":"LOW","novelty_score":6.0,"correctness_risk":"unknown","formal_verification":"none","parameter_count":0,"one_line_summary":"MIHBench is a new three-task benchmark for multi-image object hallucination in multimodal LLMs, along with an attention-balancing mitigation that lowers error rates.","lead":"This paper introduces MIHBench, a benchmark for measuring when multimodal AI models hallucinate objects across multiple images. It also reports patterns in those errors and proposes an attention-balancing method that reduces them.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The claimed single-image-to-multi-image hallucination correlation may be an artifact of shared benchmark construction, not a discovered behavioral law.","rationale":"The reader's weakest assumption is that MIHBench's design choices generalize. I narrow this to a specific testable implication: the correlation between single- and multi-image hallucination rates may be an in-construction correlation. This matters because the paper's central claim of being the first systematic study rests on discovering factors, not creating them. If the correlation is an artifact, the conclusion that single-image hallucination tendencies predict multi-image behavior — a conclusion that directly motivates the Dynamic Attention Balancing mechanism — would be unsupported. The proposed concrete test isolates construction effects by breaking shared image pools, object categories, and prompt templates. A passing test would restore confidence; a failing test would require reinterpreting the results. Since the full text is unavailable, this does not change the reader's UNVERDICTED status; it sharpens what evidence is needed.","tokens_in":822,"tokens_out":3754,"duration_ms":43753,"concrete_test":"Construct a control variant of MIHBench in which the multi-image trials are assembled from an image pool disjoint from the single-image hallucination test, using object categories not shared with that test and a prompt template that does not reuse the single-image stimulus wording. On this control, recompute the correlation between per-model single-image hallucination rate (measured on a standard single-image benchmark such as CHAIR or POPE) and multi-image hallucination rate. If the correlation drops to near zero (e.g., Spearman rho < 0.3), the published 'strong correlation' is a construction artifact. If it remains substantial (>0.6), the claim is robust. Additionally, run a logistic regression controlling for per-image task difficulty (e.g., object size, category frequency, ambiguity) to test whether 'number of images' retains a significant coefficient once difficulty is controlled.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The most load-bearing empirical finding is the reported 'strong correlation between single-image hallucination tendencies and those observed in multi-image contexts' (Abstract). This correlation is used as support for the claim that single-image hallucination behavior transfers to multi-image settings and motivates a mitigation that rebalances attention across images. However, the abstract does not state how multi-image test items are assembled. A plausible construction is to compose multi-image trials from the same image pools, object annotations, and prompt templates as the single-image tests, possibly by selecting images that already trigger single-image hallucinations. Under that construction, the correlation is partly a logical consequence of the benchmark's shared priors: models that are prone to hallucinate certain object categories on single images will naturally hallucinate those same categories when those images appear in multi-image sets, regardless of any distinct multi-image mechanism. The same worry applies to the 'progressive relationship between the number of image inputs and hallucination likelihood' — adding images increases memory load and task difficulty, yet the abstract does not control for per-image difficulty. Because the paper is abstract-only, this is a risk rather than a demonstrated flaw, but it is the single most load-bearing concern: if the correlation is an artifact, the central empirical contribution and the motivation for Dynamic Attention Balancing are substantially weakened.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"This paper introduces MIHBench, a benchmark for evaluating object-related hallucinations in multimodal large language models (MLLMs) when processing multiple images. The abstract describes three core tasks: multi-image object existence hallucination, multi-image object count hallucination, and object identity consistency hallucination. It further claims to identify key factors associated with multi-image hallucinations, including a progressive relationship with the number of images, a strong correlation between single-image and multi-image hallucination tendencies, and the influence of same-object image ratios and negative sample positions. The paper then proposes a Dynamic Attention Balancing mechanism and claims that experiments across multiple state-of-the-art MLLMs demonstrate that it reduces hallucination occurrences. However, the abstract provides no quantitative results, no benchmark construction details, no error bars, and no statistical tests, making the central claims unverifiable from the submitted material.","tokens_in":1018,"tokens_out":3672,"duration_ms":43491,"significance":"If the claims hold, MIHBench would address an important gap in hallucination evaluation for multi-image scenarios, and the proposed Dynamic Attention Balancing mechanism could provide a practical mitigation. The task taxonomy (existence, count, identity consistency) is a reasonable decomposition of object-related hallucinations, and the focus on multi-image settings is timely. The abstract also suggests a transferability hypothesis between single-image and multi-image hallucination behavior, which would be useful if established. However, the lack of any supporting evidence in the abstract means that the significance is conditional and cannot be assessed. The paper does not demonstrate the existence of the benchmark, the validity of the measurements, or the effectiveness of the proposed method in this submission.","major_comments":[{"comment":"The abstract asserts a strong correlation between single-image hallucination tendencies and multi-image contexts without describing how the multi-image test items are assembled. If multi-image trials are constructed from the same image pools, object annotations, and prompt templates as the single-image trials, then the correlation may partly reflect shared benchmark priors (e.g., object categories that already trigger single-image hallucinations) rather than a distinct multi-image mechanism. The paper must specify the construction protocol, and ideally demonstrate that the correlation holds when multi-image trials are generated independently from single-image benchmarks or when controlling for per-image difficulty and object-category frequency. This point is load-bearing because the correlation motivates the cross-setting transferability assumption behind Dynamic Attention Balancing.","section":"Abstract, strong correlation sentence"},{"comment":"The claimed progressive relationship between the number of image inputs and hallucination likelihood is confounded with task difficulty: adding images increases memory load, the number of objects to reason about, and the complexity of cross-image comparisons. Without an experimental design that varies image count while holding image content and question type constant, or statistically controlling for per-image difficulty, the monotonic trend is not established. The abstract should report the design or explicitly acknowledge this confound, and the full paper must include such controls to support the causal claim.","section":"Abstract, progressive relationship sentence"},{"comment":"The abstract provides no quantitative outcomes, no model names, no per-task results, and no statistical significance tests for the Dynamic Attention Balancing mechanism. A claim that the method 'effectively reduces hallucination occurrences' requires, at minimum, hallucination rates before and after application, error bars, and ideally a comparison with existing mitigation baselines. The abstract should include a concise summary of the main quantitative results to allow preliminary assessment of the effectiveness claim.","section":"Abstract, experiments demonstrate sentence"}],"minor_comments":[{"comment":"The abstract claims 'the first systematic study' of multi-image hallucinations but cites no prior work. Please clarify the novelty relative to existing hallucination benchmarks and multi-image reasoning evaluations, or soften the claim.","section":"Abstract, 'first systematic study'"},{"comment":"The mechanism name 'Dynamic Attention Balancing' is not defined. A one-sentence description of how inter-image attention distributions are adjusted while preserving overall visual attention would improve the readability of the abstract.","section":"Abstract, Dynamic Attention Balancing"},{"comment":"The abstract does not mention the size of the benchmark, the number of models evaluated, or the evaluation protocol. Adding these details would help place the contribution in context.","section":"General"},{"comment":"The abstract would benefit from a sentence acknowledging limitations, such as the potential restricted generalizability of the benchmark's construction choices to real-world multi-image tasks.","section":"General"}],"recommendation":"major_revision","confidential_remarks":"The submission is abstract-only, so a meaningful technical assessment is impossible. The editor should require the full manuscript before further consideration. The novelty claim 'first systematic study' should be checked against recent work on multi-image hallucination and multimodal benchmarks, as such claims are difficult to support without a thorough literature review. The major concerns (benchmark construction, causal controls, and quantitative evidence) are addressable if the full paper contains the missing details, but as submitted the abstract does not meet the evidentiary bar."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Short version: this looks like a genuine first attempt at a real gap, but the abstract alone doesn't let you check the load-bearing correlation. I'd send it to review, and press for benchmark construction details and code.\n\nWhat's new: multi-image hallucination is indeed untouched compared to single-image benchmarks. The three tasks — existence, count, identity consistency — are natural categories. The Dynamic Attention Balancing idea is simple and plausible: keep total visual attention fixed but rebalance between images. If the full paper ships code and data, that is real evidence.\n\nWhere it's soft: the whole thing rests on an abstract. No construction details, no numbers, no ablations. The strongest empirical claim, the correlation between single- and multi-image hallucination tendencies, could be an artifact of building multi-image trials from the same image pools and prompts as the single-image tests. That's a risk, not a demonstrated flaw, but it should be addressed head-on. Same for the 'progressive relationship' with number of images — that might just be memory load, not a multi-image-specific mechanism. The mitigation method is evaluated on the authors' own benchmark; that is standard, but the reviewer should check whether benchmark design was tuned to the method.\n\nBottom line: the paper is for the MLLM evaluation community. It deserves a serious referee, but the referee needs full benchmark documentation, code, and ideally an independent task construction or at least a clear explanation of how multi-image items are assembled. If those do not materialize, the headline correlation and the effectiveness claim cannot be believed.\n\nRecommendation: accept for peer review, with a strong request for reproducibility artifacts and direct addressing of the shared-construction concern.","headline":"A plausible first benchmark for multi-image hallucination, but the abstract-only presentation leaves the load-bearing correlation and the method's effectiveness unverified.","tokens_in":1536,"tokens_out":1821,"would_cite":false,"duration_ms":21216,"reading_group":"maybe","serious_thinker":"unclear","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Multi-image hallucination is a distinct failure mode with measurable drivers, and an attention-rebalancing mechanism reduces it.","keywords":["multi-image hallucination","multimodal large language models","benchmark","object existence","object count","identity consistency","dynamic attention balancing","hallucination mitigation"],"falsifier":"A reader could test whether the reported factors hold up on a benchmark constructed with a different sampling strategy: for instance, if randomizing the position of negative samples eliminates the effect on identity consistency hallucination, or if the correlation between single-image and multi-image tendencies disappears when the task distribution changes, the causal role claimed for those factors would be weakened. Similarly, if Dynamic Attention Balancing shows no gain on a held-out multi-image task suite built from naturally occurring image sets, the generalizability claim would be falsified.","tokens_in":646,"feed_emoji":"🖼️","tokens_out":3137,"duration_ms":34446,"temperature":0.7,"pith_summary":"The paper claims that hallucination in multimodal large language models is not just a single-image problem: when models must reason across several images, they systematically confuse whether objects exist, how many there are, and whether two views show the same object. To make this measurable, it introduces MIHBench, a benchmark with three tasks targeting existence, count, and cross-view identity consistency. Evaluating current models, the paper reports that hallucination likelihood rises with the number of input images, tracks each model's single-image hallucination rate, and is shaped by how many images show the same object and where a negative sample sits in the sequence. It then proposes Dynamic Attention Balancing, which redistributes attention between images without shrinking the overall visual attention budget, and reports that this reduces hallucinations and improves reasoning stability across several state-of-the-art MLLMs.","feed_headline":"Benchmark tracks AI hallucinations across multiple images","feed_subtitle":"Three tasks measure existence, count, and identity errors; an attention fix cuts them.","key_machinery":"The load-bearing pieces are the MIHBench task suite and the Dynamic Attention Balancing mechanism. MIHBench specifies three tasks with controlled negative samples and same-object image ratios; these design choices are what make hallucination measurable and what the reported correlations depend on. Dynamic Attention Balancing is the proposed intervention: it reweights attention between the input images (inter-image attention) while keeping the total visual attention share constant, so the model still attends to all images but distributes its cross-image focus differently.","core_discovery":"On the paper's own terms, the central discovery is that multi-image object hallucination is a distinct, structured failure mode with its own drivers, and that it can be mitigated by adjusting inter-image attention. The three MIHBench tasks operationalize hallucination as (1) claiming an object exists when it does not, (2) giving an incorrect count of objects across images, and (3) failing to keep object identity consistent across views. The evaluation identifies a progressive relationship between image count and hallucination likelihood, a strong correlation between single-image and multi-image hallucination tendencies, and the influence of same-object image ratio and negative-sample position on identity consistency errors. The proposed Dynamic Attention Balancing mechanism adjusts inter-image attention distributions while preserving the overall visual attention proportion, and experiments across multiple state-of-the-art MLLMs show reduced hallucination occurrence and enhanced semantic integration and reasoning stability.","pith_inferences":["The paper does not claim, but its results imply, that single-image hallucination benchmarks could serve as a cheap screening tool before full multi-image evaluation.","A testable extension of the positional finding would be to check whether simply reordering the same set of images changes a model's answer accuracy in real multi-image applications.","Because the benchmark controls negative sample placement, it cannot yet tell us how hallucinations behave in open-ended multi-image conversations where no explicit negative is present; extending the tasks to unconstrained prompts would test that boundary.","If Dynamic Attention Balancing preserves overall visual attention proportion, it may combine with other attention-based mitigation methods, though the paper does not test such combinations."],"forward_implications":["Hallucination research should treat multi-image settings as separate from single-image settings, because error rates scale with image count and with single-image tendencies.","Benchmark designers can use the three MIHBench tasks as a standard way to measure object existence, count, and identity consistency errors.","MLLM developers can apply Dynamic Attention Balancing as a lightweight intervention to reduce multi-image hallucinations.","Identity consistency hallucination can be controlled during evaluation by arranging same-object image ratios and negative-sample positions.","The reported correlation between single-image and multi-image tendencies suggests that single-image hallucination benchmarks may help predict which models will struggle with multiple images."],"supporting_citations":[],"fun_headline_variants":["Multi-image AI hallucination gets a benchmark and a mitigation","Attention tweak reduces AI hallucination in multi-image tasks","Study: image count correlates with AI hallucination likelihood","MIHBench probes object existence, count, and identity across images","New benchmark and method tackle multi-image AI hallucinations"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The results depend on the assumption that the benchmark's task definitions, negative samples, and same-object image ratios reflect real multi-image hallucination behavior rather than quirks of how the benchmark was built.","fun_headline_variants_meta":{"raw":{"variants":["Multi-image AI hallucination gets a benchmark and a mitigation","Attention tweak reduces AI hallucination in multi-image tasks","Study: image count correlates with AI hallucination likelihood","MIHBench probes object existence, count, and identity across images","New benchmark and method tackle multi-image AI hallucinations"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000388,"raw_usage":{"total_tokens":2047,"prompt_tokens":949,"completion_tokens":1098,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":565,"completion_tokens_details":{"reasoning_tokens":1018}},"tokens_in":565,"tokens_out":1098,"duration_ms":12323,"temperature":1.0,"reasoning_tokens":1018,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-06T05:57:10.762517+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"A reader could test whether the reported factors hold up on a benchmark constructed with a different sampling strategy: for instance, if randomizing the position of negative samples eliminates the effect on identity consistency hallucination, or if the correlation between single-image and multi-image tendencies disappears when the task distribution changes, the causal role claimed for those factors would be weakened. Similarly, if Dynamic Attention Balancing shows no gain on a held-out multi-image task suite built from naturally occurring image sets, the generalizability claim would be falsified.","supporting_citations":[],"review_version":1}