Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T05:49:53.545730Z
Paper Citation Record · LEDGER
As of 7 August 2026, this Paper Citation Record lists 100 of 127 outbound references and 3 inbound Pith citation observations for arXiv:2506.07016.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T05:49:53.545730Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-06T22:41:05.001058Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-06-28T19:22:34.453758Z
100 of 127 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation c9c349f6-6edd-46a0-bdcb-665277de79c7 · outbound
MAGNET: A Multi-agent Framework for Finding Audio-Visual Needles by Reasoning over Multi-Video Haystacks Crab: A Unified Audio-Visual Scene Understanding Model with Explicit Cooperation
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation ae26d71d-10cd-4f1b-987c-f9fc48a03f20 · outbound
MAGNET: A Multi-agent Framework for Finding Audio-Visual Needles by Reasoning over Multi-Video Haystacks Meerkat: Audio-visual large language model for grounding in space and time
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6fbb984b-5ef9-49ec-abb0-41193774fd76 · outbound
MAGNET: A Multi-agent Framework for Finding Audio-Visual Needles by Reasoning over Multi-Video Haystacks VITA: Towards Open-Source Interactive Omni Multimodal LLM
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9281e5fc-12a0-44b1-b467-6ee8a1b32bbc · outbound
MAGNET: A Multi-agent Framework for Finding Audio-Visual Needles by Reasoning over Multi-Video Haystacks Unified-io 2: Scaling autoregressive multimodal models with vision language audio and action
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cb2ada9f-3ebe-453f-974a-ce660e33b6ce · outbound
MAGNET: A Multi-agent Framework for Finding Audio-Visual Needles by Reasoning over Multi-Video Haystacks CREMA: Generalizable and Efficient Video-Language Reasoning via Multimodal Modular Fusion
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 07fa958e-1721-4972-a904-1b52d70a42c5 · outbound
MAGNET: A Multi-agent Framework for Finding Audio-Visual Needles by Reasoning over Multi-Video Haystacks Avicuna: Audio-visual llm with interleaver and context-boundary alignment for temporal referential dialogue
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2aa1aa53-cd40-479c-b381-cab4192fa0bd · outbound
MAGNET: A Multi-agent Framework for Finding Audio-Visual Needles by Reasoning over Multi-Video Haystacks MLVU: Benchmarking Multi-task Long Video Understanding
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ffb5b741-b0a6-47a7-9850-a1b6c96facdf · outbound
MAGNET: A Multi-agent Framework for Finding Audio-Visual Needles by Reasoning over Multi-Video Haystacks Sharegpt4video: Improving video understanding and generation with better captions
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a55a95d4-5a08-4bdb-a540-3c8398416327 · outbound
MAGNET: A Multi-agent Framework for Finding Audio-Visual Needles by Reasoning over Multi-Video Haystacks CinePile: A Long Video Question Answering Dataset and Benchmark
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b68cb8c7-9fec-421a-85ad-614890e1b7d6 · outbound
MAGNET: A Multi-agent Framework for Finding Audio-Visual Needles by Reasoning over Multi-Video Haystacks Moviechat: From dense token to sparse memory for long video understanding
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0642fdb3-3293-452d-aace-d241ec4e6a73 · outbound
MAGNET: A Multi-agent Framework for Finding Audio-Visual Needles by Reasoning over Multi-Video Haystacks Dense-localizing audio-visual events in untrimmed videos: A large-scale benchmark and baseline
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 28cd0823-f30f-498c-9861-143eb3a8bc9f · outbound
MAGNET: A Multi-agent Framework for Finding Audio-Visual Needles by Reasoning over Multi-Video Haystacks Vast: A vision-audio-subtitle-text omni-modality foundation model and dataset
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2f204c4f-25e5-4659-bdd6-a4b9f95b604a · outbound
MAGNET: A Multi-agent Framework for Finding Audio-Visual Needles by Reasoning over Multi-Video Haystacks Avqa: A dataset for audio-visual question answering on videos
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e8a4da05-4271-4769-9f0e-d9b74fa231d7 · outbound
MAGNET: A Multi-agent Framework for Finding Audio-Visual Needles by Reasoning over Multi-Video Haystacks Cat: Enhancing multimodal large language model to answer questions in dynamic audio-visual scenarios
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 86e2d76f-14bd-411b-a289-923b64fd4cb2 · outbound
MAGNET: A Multi-agent Framework for Finding Audio-Visual Needles by Reasoning over Multi-Video Haystacks Learning to answer questions in dynamic audio-visual scenarios
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2b52dfc9-6683-4c2b-abb7-c81cab78b27a · outbound
MAGNET: A Multi-agent Framework for Finding Audio-Visual Needles by Reasoning over Multi-Video Haystacks Vggsound: A large-scale audio- visual dataset
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 74a42b6f-062f-4581-8e7b-74b0436b721c · outbound
MAGNET: A Multi-agent Framework for Finding Audio-Visual Needles by Reasoning over Multi-Video Haystacks SAVEn-Vid: Synergistic Audio-Visual Integration for Enhanced Understanding in Long Video Context
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 9743a3e4-b48c-45de-a516-1a9cbe36c5de · outbound
MAGNET: A Multi-agent Framework for Finding Audio-Visual Needles by Reasoning over Multi-Video Haystacks Imagebind: One embedding space to bind them all
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fd4a3517-5586-4f10-a753-2a6c64bfd37b · outbound
MAGNET: A Multi-agent Framework for Finding Audio-Visual Needles by Reasoning over Multi-Video Haystacks Gemini: Google’s multimodal ai model
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 14d0b676-6f3f-4568-95a6-a122675a9208 · outbound
MAGNET: A Multi-agent Framework for Finding Audio-Visual Needles by Reasoning over Multi-Video Haystacks Qwen2.5 Technical Report
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d21a88a6-d50d-4291-8c0a-83ba536d916d · outbound
MAGNET: A Multi-agent Framework for Finding Audio-Visual Needles by Reasoning over Multi-Video Haystacks GPT-4o System Card
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 43ef4468-98cc-4c9e-9be2-577564468c64 · outbound
MAGNET: A Multi-agent Framework for Finding Audio-Visual Needles by Reasoning over Multi-Video Haystacks Towards General Text Embeddings with Multi-stage Contrastive Learning
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0cc2a8d5-55cd-4c78-9013-06881199978f · outbound
MAGNET: A Multi-agent Framework for Finding Audio-Visual Needles by Reasoning over Multi-Video Haystacks Variants of the hungarian method for assignment problems
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c9527294-ac5e-4fa7-bc35-06407f19f055 · outbound
MAGNET: A Multi-agent Framework for Finding Audio-Visual Needles by Reasoning over Multi-Video Haystacks Video-rag: Visually-aligned retrieval-augmented long video comprehension
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 75b41b82-8f33-443b-a107-50688a16e005 · outbound
MAGNET: A Multi-agent Framework for Finding Audio-Visual Needles by Reasoning over Multi-Video Haystacks video-SALMONN: Speech-Enhanced Audio-Visual Large Language Models
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 861d624e-75d6-47de-be27-79ed826e7e49 · outbound
MAGNET: A Multi-agent Framework for Finding Audio-Visual Needles by Reasoning over Multi-Video Haystacks High-fidelity audio compression with improved rvqgan
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4faf8e1e-e314-4294-9137-d87b8caf369f · outbound
MAGNET: A Multi-agent Framework for Finding Audio-Visual Needles by Reasoning over Multi-Video Haystacks Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 521efbb0-9174-4f8f-9d04-0aca524bec44 · outbound
MAGNET: A Multi-agent Framework for Finding Audio-Visual Needles by Reasoning over Multi-Video Haystacks Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 64c22047-d71b-4f29-9eed-692710e1357a · outbound
MAGNET: A Multi-agent Framework for Finding Audio-Visual Needles by Reasoning over Multi-Video Haystacks Mvbench: A comprehensive multi-modal video understanding benchmark
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1e9f3a76-d7a7-4118-adb0-2a0b0df3377f · outbound
MAGNET: A Multi-agent Framework for Finding Audio-Visual Needles by Reasoning over Multi-Video Haystacks Video Question Answering: Datasets, Algorithms and Challenges
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d005edee-7e59-453d-a80b-70df346af50f · outbound
MAGNET: A Multi-agent Framework for Finding Audio-Visual Needles by Reasoning over Multi-Video Haystacks LVBench: An Extreme Long Video Understanding Benchmark
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 44fd02b2-02c6-4094-a565-5d80afc81ff9 · outbound
MAGNET: A Multi-agent Framework for Finding Audio-Visual Needles by Reasoning over Multi-Video Haystacks Movieqa: Understanding stories in movies through question-answering
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1e25c88d-2c88-4a7a-9074-80b411e855b5 · outbound
MAGNET: A Multi-agent Framework for Finding Audio-Visual Needles by Reasoning over Multi-Video Haystacks Are we asking the right questions in movieqa? In Proceedings of the IEEE/CVF International Conference on Computer Vision Workshops, pages 0–0, 2019
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e13897b5-9982-4b69-b512-dcf1b32f27c5 · outbound
MAGNET: A Multi-agent Framework for Finding Audio-Visual Needles by Reasoning over Multi-Video Haystacks Activitynet-qa: A dataset for understanding complex web videos via question answering
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 04042d7a-4529-42c6-add3-d490954fb018 · outbound
MAGNET: A Multi-agent Framework for Finding Audio-Visual Needles by Reasoning over Multi-Video Haystacks How2: A Large-scale Dataset for Multimodal Language Understanding
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dd516ceb-d6e4-4528-8a6e-b47a87b04001 · outbound
MAGNET: A Multi-agent Framework for Finding Audio-Visual Needles by Reasoning over Multi-Video Haystacks Next-qa: Next phase of question-answering to explaining temporal actions
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4de816fd-bc2a-4bcb-90f1-f11d582fca64 · outbound
MAGNET: A Multi-agent Framework for Finding Audio-Visual Needles by Reasoning over Multi-Video Haystacks Perception test: A diagnostic benchmark for multimodal video models
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 28fd6f61-85cf-4000-9df3-5bdea4755f80 · outbound
MAGNET: A Multi-agent Framework for Finding Audio-Visual Needles by Reasoning over Multi-Video Haystacks STAR: A Benchmark for Situated Reasoning in Real-World Videos
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c14e6391-f72e-47a3-9cbe-5663d75b4092 · outbound
MAGNET: A Multi-agent Framework for Finding Audio-Visual Needles by Reasoning over Multi-Video Haystacks Agqa: A benchmark for compositional spatio-temporal reasoning
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dfa4da0a-2546-47a2-9d07-f89a2ae3df00 · outbound
MAGNET: A Multi-agent Framework for Finding Audio-Visual Needles by Reasoning over Multi-Video Haystacks Egoschema: A diagnostic benchmark for very long-form video language understanding
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4129b2ae-f882-4af7-9895-bbe0823ac0cb · outbound
MAGNET: A Multi-agent Framework for Finding Audio-Visual Needles by Reasoning over Multi-Video Haystacks Longvideobench: A benchmark for long-context interleaved video-language understanding
Reference 46
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c8204057-19fa-4593-b216-cc2fbaa405a1 · outbound
MAGNET: A Multi-agent Framework for Finding Audio-Visual Needles by Reasoning over Multi-Video Haystacks Just ask: Learning to answer questions from millions of narrated videos
Reference 47
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b3a54a4b-e9d3-4d54-a0b1-7cbe0b1e541f · outbound
MAGNET: A Multi-agent Framework for Finding Audio-Visual Needles by Reasoning over Multi-Video Haystacks InstructionBench: An Instructional Video Understanding Benchmark
Reference 48
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cb94633a-c344-46d0-a195-ee1cbbcc352c · outbound
MAGNET: A Multi-agent Framework for Finding Audio-Visual Needles by Reasoning over Multi-Video Haystacks HD-EPIC: A Highly-Detailed Egocentric Video Dataset
Reference 49
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7fa9c4d4-f755-44ca-9822-11dfc0246416 · outbound
MAGNET: A Multi-agent Framework for Finding Audio-Visual Needles by Reasoning over Multi-Video Haystacks Ego4d: Around the world in 3,000 hours of egocentric video
Reference 50
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 812bdbe8-357f-4318-92ae-ef9fe1546578 · outbound
MAGNET: A Multi-agent Framework for Finding Audio-Visual Needles by Reasoning over Multi-Video Haystacks Ego-exo4d: Understand- ing skilled human activity from first-and third-person perspectives
Reference 51
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4a9a7a8e-fa8e-465d-a7c3-0011f7212f13 · outbound
MAGNET: A Multi-agent Framework for Finding Audio-Visual Needles by Reasoning over Multi-Video Haystacks Scaling egocentric vision: The epic-kitchens dataset
Reference 52
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5c323aad-84cb-4215-9c20-e0af2b7a6b7e · outbound
MAGNET: A Multi-agent Framework for Finding Audio-Visual Needles by Reasoning over Multi-Video Haystacks Vicuna: An open-source chatbot impressing gpt-4 with 90%* chatgpt quality
Reference 53
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f5324068-7c28-4598-a0ac-5e7f23433fe8 · outbound
MAGNET: A Multi-agent Framework for Finding Audio-Visual Needles by Reasoning over Multi-Video Haystacks LLaMA: Open and Efficient Foundation Language Models
Reference 54
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4714920f-c24b-4c1a-a7e2-0299bb14b562 · outbound
MAGNET: A Multi-agent Framework for Finding Audio-Visual Needles by Reasoning over Multi-Video Haystacks Mistral 7B
Reference 55
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 20fd0780-27b5-4a90-872a-f5ccf4089a5d · outbound
MAGNET: A Multi-agent Framework for Finding Audio-Visual Needles by Reasoning over Multi-Video Haystacks VideoChat: Chat-Centric Video Understanding
Reference 56
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 78fa4c01-ccd7-4fb7-b01c-789b8ccbdd51 · outbound
MAGNET: A Multi-agent Framework for Finding Audio-Visual Needles by Reasoning over Multi-Video Haystacks Video-LLaMA: An Instruction-tuned Audio-Visual Language Model for Video Understanding
Reference 57
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3e0fbcc7-cc27-4663-a7cd-b84c9c0db2f4 · outbound
MAGNET: A Multi-agent Framework for Finding Audio-Visual Needles by Reasoning over Multi-Video Haystacks LLaVA-OneVision: Easy Visual Task Transfer
Reference 58
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 01eb85a9-d3e8-4f7e-9a2d-ffdb046686dc · outbound
MAGNET: A Multi-agent Framework for Finding Audio-Visual Needles by Reasoning over Multi-Video Haystacks World Model on Million-Length Video And Language With Blockwise RingAttention
Reference 59
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 372c0c6c-ddb6-46d8-8879-8a2903415c3f · outbound
MAGNET: A Multi-agent Framework for Finding Audio-Visual Needles by Reasoning over Multi-Video Haystacks Long Context Transfer from Language to Vision
Reference 60
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ea34d88d-cd06-4fee-838a-1704492c1e8c · outbound
MAGNET: A Multi-agent Framework for Finding Audio-Visual Needles by Reasoning over Multi-Video Haystacks LongVILA: Scaling Long-Context Visual Language Models for Long Videos
Reference 61
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 785eeebb-0c74-458b-91e0-3df35fdf5fc7 · outbound
MAGNET: A Multi-agent Framework for Finding Audio-Visual Needles by Reasoning over Multi-Video Haystacks Longllava: Scaling multi-modal llms to 1000 images efficiently via hybrid architecture
Reference 62
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dc8fecce-635a-4350-823e-a604ad8268a5 · outbound
MAGNET: A Multi-agent Framework for Finding Audio-Visual Needles by Reasoning over Multi-Video Haystacks LLaMA-VID: An Image is Worth 2 Tokens in Large Language Models
Reference 63
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5c7ac51c-e03f-4c20-be75-737f1f5e4b4e · outbound
MAGNET: A Multi-agent Framework for Finding Audio-Visual Needles by Reasoning over Multi-Video Haystacks Flash-VStream: Memory-Based Real-Time Understanding for Long Video Streams
Reference 64
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1eca1a7e-e2ae-4794-9fb8-a2cf4f1428ac · outbound
MAGNET: A Multi-agent Framework for Finding Audio-Visual Needles by Reasoning over Multi-Video Haystacks VideoAgent: A Memory-augmented Multimodal Agent for Video Understanding
Reference 65
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c0c07cc0-f991-4305-80a8-37b2a2c996c5 · outbound
MAGNET: A Multi-agent Framework for Finding Audio-Visual Needles by Reasoning over Multi-Video Haystacks VideoAgent: Long-form Video Understanding with Large Language Model as Agent
Reference 66
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 28f45b09-2bf9-437a-841d-7b2e7489b3af · outbound
MAGNET: A Multi-agent Framework for Finding Audio-Visual Needles by Reasoning over Multi-Video Haystacks VideoTree: Adaptive Tree-based Video Representation for LLM Reasoning on Long Videos
Reference 67
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 307dc06e-81df-428d-b77f-398c091a9c6c · outbound
MAGNET: A Multi-agent Framework for Finding Audio-Visual Needles by Reasoning over Multi-Video Haystacks Retrieval-augmented generation for knowledge-intensive nlp tasks
Reference 68
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b84ea8fc-79e5-4af5-89b9-ee20d9c47a53 · outbound
MAGNET: A Multi-agent Framework for Finding Audio-Visual Needles by Reasoning over Multi-Video Haystacks Active retrieval augmented generation
Reference 69
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8afc8e91-0f1c-47c2-8a98-f9cd1ff1daed · outbound
MAGNET: A Multi-agent Framework for Finding Audio-Visual Needles by Reasoning over Multi-Video Haystacks Sentence-level prompts benefit composed image retrieval
Reference 70
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7e026fbc-b4f4-4fb3-806c-4d987713efd9 · outbound
MAGNET: A Multi-agent Framework for Finding Audio-Visual Needles by Reasoning over Multi-Video Haystacks VQA4CIR: Boosting Composed Image Retrieval with Visual Question Answering
Reference 71
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fd6652c5-5c9d-4bb3-9842-a4c75f95a24c · outbound
MAGNET: A Multi-agent Framework for Finding Audio-Visual Needles by Reasoning over Multi-Video Haystacks Searching for Best Practices in Retrieval-Augmented Generation
Reference 72
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 211330bc-be32-4ab6-8105-170281393d0f · outbound
MAGNET: A Multi-agent Framework for Finding Audio-Visual Needles by Reasoning over Multi-Video Haystacks Retrieval-Augmented Generation for Natural Language Processing: A Survey
Reference 73
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3de5b9a1-3313-48b8-a56f-8d2929889cf1 · outbound
MAGNET: A Multi-agent Framework for Finding Audio-Visual Needles by Reasoning over Multi-Video Haystacks A Survey on Retrieval-Augmented Text Generation for Large Language Models
Reference 74
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cea28d7a-102d-4cb3-856a-122697f58eb1 · outbound
MAGNET: A Multi-agent Framework for Finding Audio-Visual Needles by Reasoning over Multi-Video Haystacks Retrieval-Augmented Generation for AI-Generated Content: A Survey
Reference 75
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bf8931c0-9ddf-442f-9271-4b1c7f1279ff · outbound
MAGNET: A Multi-agent Framework for Finding Audio-Visual Needles by Reasoning over Multi-Video Haystacks Self-RAG: Learning to Retrieve, Generate, and Critique through Self-Reflection
Reference 76
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 82e24110-d08c-4dbe-9695-4b2b63acb26f · outbound
MAGNET: A Multi-agent Framework for Finding Audio-Visual Needles by Reasoning over Multi-Video Haystacks Realm: Retrieval- augmented language model pre-training
Reference 77
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a1e3442f-a6d3-46a9-8103-6820d27ca7e6 · outbound
MAGNET: A Multi-agent Framework for Finding Audio-Visual Needles by Reasoning over Multi-Video Haystacks SAIL: Search-Augmented Instruction Learning
Reference 78
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f6995908-3bb9-49a2-8ee2-fe0fc8b4e038 · outbound
MAGNET: A Multi-agent Framework for Finding Audio-Visual Needles by Reasoning over Multi-Video Haystacks Dense passage retrieval for open-domain question answering
Reference 79
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9d1c0ee8-f18d-4add-88fc-4f40f10f5aea · outbound
MAGNET: A Multi-agent Framework for Finding Audio-Visual Needles by Reasoning over Multi-Video Haystacks Document haystacks: Vision-language reasoning over piles of 1000+ documents
Reference 80
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ecef7520-2cff-45d5-9b6d-d3d4fb5e731a · outbound
MAGNET: A Multi-agent Framework for Finding Audio-Visual Needles by Reasoning over Multi-Video Haystacks Unresolved cited work
Reference 81
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 03f03589-0a7e-44de-a1b4-023e511b63e0 · outbound
MAGNET: A Multi-agent Framework for Finding Audio-Visual Needles by Reasoning over Multi-Video Haystacks Visual Haystacks: A Vision-Centric Needle-In-A-Haystack Benchmark
Reference 82
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b601c34f-c801-408f-b65c-b77ebb3da369 · outbound
MAGNET: A Multi-agent Framework for Finding Audio-Visual Needles by Reasoning over Multi-Video Haystacks VisRAG: Vision-based Retrieval-augmented Generation on Multi-modality Documents
Reference 83
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 326e72fd-02cb-4c9e-924a-f48db5577fca · outbound
MAGNET: A Multi-agent Framework for Finding Audio-Visual Needles by Reasoning over Multi-Video Haystacks Visual-RAG: Benchmarking Text-to-Image Retrieval Augmented Generation for Visual Knowledge Intensive Queries
Reference 84
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ffd5fd2a-22ec-48b0-bdf5-303b0fafa39b · outbound
MAGNET: A Multi-agent Framework for Finding Audio-Visual Needles by Reasoning over Multi-Video Haystacks Vdocrag: Retrieval-augmented generation over visually-rich documents
Reference 85
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cc435739-a5a2-4dbf-bfbf-127ac913f76d · outbound
MAGNET: A Multi-agent Framework for Finding Audio-Visual Needles by Reasoning over Multi-Video Haystacks Retrieval Augmented Visual Question Answering with Outside Knowledge
Reference 86
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 13bf5ae5-ae74-47dc-b20a-7252476bba83 · outbound
MAGNET: A Multi-agent Framework for Finding Audio-Visual Needles by Reasoning over Multi-Video Haystacks Fact-Aware Multimodal Retrieval Augmentation for Accurate Medical Radiology Report Generation
Reference 87
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3a3c1727-6ef2-48e4-88fb-1affe423f913 · outbound
MAGNET: A Multi-agent Framework for Finding Audio-Visual Needles by Reasoning over Multi-Video Haystacks Rule: Reliable multimodal rag for factuality in medical vision language models
Reference 88
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fe7e62fc-4e71-4d67-8fa1-ece2db426d5e · outbound
MAGNET: A Multi-agent Framework for Finding Audio-Visual Needles by Reasoning over Multi-Video Haystacks Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution
Reference 89
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 27713dc0-bfae-43f4-8ab4-52514fc9136e · outbound
MAGNET: A Multi-agent Framework for Finding Audio-Visual Needles by Reasoning over Multi-Video Haystacks Gpt-4o: Enhanced multimodal language model
Reference 90
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 740f4cb2-fa07-4a92-9683-d5d6d5972ec0 · outbound
MAGNET: A Multi-agent Framework for Finding Audio-Visual Needles by Reasoning over Multi-Video Haystacks Minigpt-4: Enhancing vision-language understanding with advanced large language models
Reference 91
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8129eeaa-5041-402e-90f5-2c55f99f7835 · outbound
MAGNET: A Multi-agent Framework for Finding Audio-Visual Needles by Reasoning over Multi-Video Haystacks Unified-IO: A Unified Model for Vision, Language, and Multi-Modal Tasks
Reference 92
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e9e009aa-e610-432d-a75e-7b9b36171587 · outbound
MAGNET: A Multi-agent Framework for Finding Audio-Visual Needles by Reasoning over Multi-Video Haystacks MMMU: A Massive Multi-discipline Multimodal Understanding and Reasoning Benchmark for Expert AGI
Reference 93
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f397db9d-cebe-4215-ac62-ec113453365f · outbound
MAGNET: A Multi-agent Framework for Finding Audio-Visual Needles by Reasoning over Multi-Video Haystacks Making the V in VQA matter: Elevating the role of image understanding in Visual Question Answering
Reference 94
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fd642cb6-bd08-4575-a76d-a41451365bf0 · outbound
MAGNET: A Multi-agent Framework for Finding Audio-Visual Needles by Reasoning over Multi-Video Haystacks MathVerse: Does Your Multi-modal LLM Truly See the Diagrams in Visual Math Problems?
Reference 95
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 10e2c647-8096-410e-bede-17e3d463131c · outbound
MAGNET: A Multi-agent Framework for Finding Audio-Visual Needles by Reasoning over Multi-Video Haystacks MathVista: Evaluating Mathematical Reasoning of Foundation Models in Visual Contexts
Reference 96
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8712754c-c291-4b51-9ca2-36203477e74a · outbound
MAGNET: A Multi-agent Framework for Finding Audio-Visual Needles by Reasoning over Multi-Video Haystacks VideoRAG: Retrieval-Augmented Generation with Extreme Long-Context Videos
Reference 97
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 94a4611f-a6a2-4342-a19f-f5cbfc809bc6 · outbound
MAGNET: A Multi-agent Framework for Finding Audio-Visual Needles by Reasoning over Multi-Video Haystacks V-desirr: Very fast deep embedded single image reflection removal
Reference 98
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cb64ee82-bb51-4f1b-a730-c2ccfc1ca690 · outbound
MAGNET: A Multi-agent Framework for Finding Audio-Visual Needles by Reasoning over Multi-Video Haystacks Measured albedo in the wild: Filling the gap in intrinsics evaluation
Reference 99
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ccf0e7a5-89b1-422e-b981-79f3699f5364 · outbound
MAGNET: A Multi-agent Framework for Finding Audio-Visual Needles by Reasoning over Multi-Video Haystacks Adverb: Visually guided audio dereverberation
Reference 100
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 24a613e3-e1f0-43fe-89fb-6bf65e1a63a2 · outbound
MAGNET: A Multi-agent Framework for Finding Audio-Visual Needles by Reasoning over Multi-Video Haystacks Melfusion: Synthesizing music from image and language cues using diffusion models
Reference 101
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ece97e52-8192-434d-8c9a-7279ccfaf2c2 · outbound
MAGNET: A Multi-agent Framework for Finding Audio-Visual Needles by Reasoning over Multi-Video Haystacks Foleygen: Visually-guided audio generation
Reference 102
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e61e0399-ea2b-4839-8675-591d47f90ef7 · outbound
MAGNET: A Multi-agent Framework for Finding Audio-Visual Needles by Reasoning over Multi-Video Haystacks Codi-2: In-context interleaved and interactive any-to-any generation
Reference 103
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 95e292b2-e55e-489f-8221-178eeb0a9785 · outbound
MAGNET: A Multi-agent Framework for Finding Audio-Visual Needles by Reasoning over Multi-Video Haystacks Listen to the pixels
Reference 104
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 8df9726a-60c8-4102-83a5-372640adcbd6 · outbound
MAGNET: A Multi-agent Framework for Finding Audio-Visual Needles by Reasoning over Multi-Video Haystacks Audvisum: Self- supervised deep reinforcement learning for diverse audio-visual summary generation
Reference 105
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation b7e657f6-9660-4ad7-a8ed-513e02acdef8 · inbound
EgoAdapt: Adaptive Multisensory Distillation and Policy Learning for Efficient Egocentric Perception MAGNET: A Multi-agent Framework for Finding Audio-Visual Needles by Reasoning over Multi-Video Haystacks
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dbe2db90-2cb4-4672-b286-ff4a14a70b56 · inbound
EgoSound: Benchmarking Sound Understanding in Egocentric Videos MAGNET: A Multi-agent Framework for Finding Audio-Visual Needles by Reasoning over Multi-Video Haystacks
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation ea59bb54-f6e6-48c1-b475-63155e155898 · inbound
Through the PRISM: Principle-Aware, Interpretable, and Multi-Scale Evaluation of Visual Designs MAGNET: A Multi-agent Framework for Finding Audio-Visual Needles by Reasoning over Multi-Video Haystacks
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.