{"record_type":"pith_number_record","schema_url":"https://pith.science/schemas/pith-number/v1.json","pith_number":"pith:2024:Q45DGSPFMM4MN4QQLYKP72QEW6","short_pith_number":"pith:Q45DGSPF","schema_version":"1.0","canonical_sha256":"873a3349e56338c6f2105e14ffea04b7a93cf2a85d79781ecc2c6f664db07f81","source":{"kind":"arxiv","id":"2411.02937","version":5},"attestation_state":"computed","paper":{"title":"Benchmarking Multimodal Retrieval Augmented Generation with Dynamic VQA Dataset and Self-adaptive Planning Agent","license":"http://arxiv.org/licenses/nonexclusive-distrib/1.0/","headline":"","cross_cats":[],"primary_cat":"cs.CL","authors_text":"Fei Huang, Hai-Tao Zheng, Hui Wang, Jingren Zhou, Philip S. Yu, Xinran Zheng, Xinyu Wang, Yangning Li, Yinghui Li, Yong Jiang, Zhen Zhang","submitted_at":"2024-11-05T09:27:21Z","abstract_excerpt":"Multimodal Retrieval Augmented Generation (mRAG) plays an important role in mitigating the \"hallucination\" issue inherent in multimodal large language models (MLLMs). Although promising, existing heuristic mRAGs typically predefined fixed retrieval processes, which causes two issues: (1) Non-adaptive Retrieval Queries. (2) Overloaded Retrieval Queries. However, these flaws cannot be adequately reflected by current knowledge-seeking visual question answering (VQA) datasets, since the most required knowledge can be readily obtained with a standard two-step retrieval. To bridge the dataset gap, w"},"verification_status":{"content_addressed":true,"pith_receipt":true,"author_attested":false,"weak_author_claims":0,"strong_author_claims":0,"externally_anchored":false,"storage_verified":false,"citation_signatures":0,"replication_records":0,"graph_snapshot":true,"references_resolved":false,"formal_links_present":false},"canonical_record":{"source":{"id":"2411.02937","kind":"arxiv","version":5},"metadata":{"license":"http://arxiv.org/licenses/nonexclusive-distrib/1.0/","primary_cat":"cs.CL","submitted_at":"2024-11-05T09:27:21Z","cross_cats_sorted":[],"title_canon_sha256":"2a11fb552a5af482f2f2c6e30ca371d69f641551db271bff841c2e452b4b621e","abstract_canon_sha256":"337d54d82e9c6002fd3c0c5e2e18432edf8c5b51572bd102329dada048fd3554"},"schema_version":"1.0"},"receipt":{"kind":"pith_receipt","key_id":"pith-v1-2026-05","algorithm":"ed25519","signed_at":"2026-07-05T11:09:03.789724Z","signature_b64":"LX9pZBsW1hOzJE5aYaacBZyvlqwjqnLPVplIvxfuKUC7AHs//3ukHhpZlP0Qqu9OcVjQj1dyyUIl1VRQOoNDBw==","signed_message":"canonical_sha256_bytes","builder_version":"pith-number-builder-2026-05-17-v1","receipt_version":"0.3","canonical_sha256":"873a3349e56338c6f2105e14ffea04b7a93cf2a85d79781ecc2c6f664db07f81","last_reissued_at":"2026-07-05T11:09:03.789261Z","signature_status":"signed_v1","first_computed_at":"2026-07-05T11:09:03.789261Z","public_key_fingerprint":"8d4b5ee74e4693bcd1df2446408b0d54"},"graph_snapshot":{"paper":{"title":"Benchmarking Multimodal Retrieval Augmented Generation with Dynamic VQA Dataset and Self-adaptive Planning Agent","license":"http://arxiv.org/licenses/nonexclusive-distrib/1.0/","headline":"","cross_cats":[],"primary_cat":"cs.CL","authors_text":"Fei Huang, Hai-Tao Zheng, Hui Wang, Jingren Zhou, Philip S. Yu, Xinran Zheng, Xinyu Wang, Yangning Li, Yinghui Li, Yong Jiang, Zhen Zhang","submitted_at":"2024-11-05T09:27:21Z","abstract_excerpt":"Multimodal Retrieval Augmented Generation (mRAG) plays an important role in mitigating the \"hallucination\" issue inherent in multimodal large language models (MLLMs). Although promising, existing heuristic mRAGs typically predefined fixed retrieval processes, which causes two issues: (1) Non-adaptive Retrieval Queries. (2) Overloaded Retrieval Queries. However, these flaws cannot be adequately reflected by current knowledge-seeking visual question answering (VQA) datasets, since the most required knowledge can be readily obtained with a standard two-step retrieval. To bridge the dataset gap, w"},"claims":{"count":0,"items":[],"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57"},"source":{"id":"2411.02937","kind":"arxiv","version":5},"verdict":{"id":null,"model_set":{},"created_at":null,"strongest_claim":"","one_line_summary":"","pipeline_version":null,"weakest_assumption":"","pith_extraction_headline":""},"integrity":{"clean":true,"summary":{"advisory":0,"critical":0,"by_detector":{},"informational":0},"endpoint":"/pith/2411.02937/integrity.json","findings":[],"available":true,"detectors_run":[],"snapshot_sha256":"c28c3603d3b5d939e8dc4c7e95fa8dfce3d595e45f758748cecf8e644a296938"},"references":{"count":0,"sample":[],"resolved_work":0,"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57","internal_anchors":0},"formal_canon":{"evidence_count":0,"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57"},"author_claims":{"count":0,"strong_count":0,"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57"},"builder_version":"pith-number-builder-2026-05-17-v1"},"aliases":[{"alias_kind":"arxiv","alias_value":"2411.02937","created_at":"2026-07-05T11:09:03.789320+00:00"},{"alias_kind":"arxiv_version","alias_value":"2411.02937v5","created_at":"2026-07-05T11:09:03.789320+00:00"},{"alias_kind":"doi","alias_value":"10.48550/arxiv.2411.02937","created_at":"2026-07-05T11:09:03.789320+00:00"},{"alias_kind":"pith_short_12","alias_value":"Q45DGSPFMM4M","created_at":"2026-07-05T11:09:03.789320+00:00"},{"alias_kind":"pith_short_16","alias_value":"Q45DGSPFMM4MN4QQ","created_at":"2026-07-05T11:09:03.789320+00:00"},{"alias_kind":"pith_short_8","alias_value":"Q45DGSPF","created_at":"2026-07-05T11:09:03.789320+00:00"}],"events":[],"event_summary":{},"paper_claims":[],"inbound_citations":{"count":9,"internal_anchor_count":1,"sample":[{"citing_arxiv_id":"2607.07383","citing_title":"MMAgent-R$^2$: Learning to Rerank and Reject for Agentic mRAG","ref_index":16,"is_internal_anchor":true},{"citing_arxiv_id":"2605.17946","citing_title":"SVFSearch: A Multimodal Knowledge-Intensive Benchmark for Short-Video Frame Search in the Gaming Vertical Domain","ref_index":52,"is_internal_anchor":false},{"citing_arxiv_id":"2605.17946","citing_title":"SVFSearch: A Multimodal Knowledge-Intensive Benchmark for Short-Video Frame Search in the Gaming Vertical Domain","ref_index":52,"is_internal_anchor":false},{"citing_arxiv_id":"2505.22095","citing_title":"Mixture-of-Retrieval Experts for Reasoning-Guided Multimodal Knowledge Exploitation","ref_index":23,"is_internal_anchor":false},{"citing_arxiv_id":"2601.12538","citing_title":"Agentic Reasoning for Large Language Models","ref_index":183,"is_internal_anchor":false},{"citing_arxiv_id":"2504.19678","citing_title":"From LLM Reasoning to Autonomous AI Agents: A Comprehensive Review","ref_index":9,"is_internal_anchor":false},{"citing_arxiv_id":"2511.20857","citing_title":"Evo-Memory: Benchmarking LLM Agent Test-time Learning with Self-Evolving Memory","ref_index":198,"is_internal_anchor":false},{"citing_arxiv_id":"2604.04017","citing_title":"GeoBrowse: A Geolocation Benchmark for Agentic Tool Use with Expert-Annotated Reasoning Traces","ref_index":27,"is_internal_anchor":false},{"citing_arxiv_id":"2604.15736","citing_title":"RefereeBench: Are Video MLLMs Ready to be Multi-Sport Referees","ref_index":29,"is_internal_anchor":false}]},"formal_canon":{"evidence_count":0,"sample":[],"anchors":[]},"links":{"html":"https://pith.science/pith/Q45DGSPFMM4MN4QQLYKP72QEW6","json":"https://pith.science/pith/Q45DGSPFMM4MN4QQLYKP72QEW6.json","graph_json":"https://pith.science/api/pith-number/Q45DGSPFMM4MN4QQLYKP72QEW6/graph.json","events_json":"https://pith.science/api/pith-number/Q45DGSPFMM4MN4QQLYKP72QEW6/events.json","paper":"https://pith.science/paper/Q45DGSPF"},"agent_actions":{"view_html":"https://pith.science/pith/Q45DGSPFMM4MN4QQLYKP72QEW6","download_json":"https://pith.science/pith/Q45DGSPFMM4MN4QQLYKP72QEW6.json","view_paper":"https://pith.science/paper/Q45DGSPF","resolve_alias":"https://pith.science/api/pith-number/resolve?arxiv=2411.02937&json=true","fetch_graph":"https://pith.science/api/pith-number/Q45DGSPFMM4MN4QQLYKP72QEW6/graph.json","fetch_events":"https://pith.science/api/pith-number/Q45DGSPFMM4MN4QQLYKP72QEW6/events.json","actions":{"anchor_timestamp":"https://pith.science/pith/Q45DGSPFMM4MN4QQLYKP72QEW6/action/timestamp_anchor","attest_storage":"https://pith.science/pith/Q45DGSPFMM4MN4QQLYKP72QEW6/action/storage_attestation","attest_author":"https://pith.science/pith/Q45DGSPFMM4MN4QQLYKP72QEW6/action/author_attestation","sign_citation":"https://pith.science/pith/Q45DGSPFMM4MN4QQLYKP72QEW6/action/citation_signature","submit_replication":"https://pith.science/pith/Q45DGSPFMM4MN4QQLYKP72QEW6/action/replication_record"}},"created_at":"2026-07-05T11:09:03.789320+00:00","updated_at":"2026-07-05T11:09:03.789320+00:00"}