Pith. sign in

Paper Citation Record · LEDGER

MAGNET: A Multi-agent Framework for Finding Audio-Visual Needles by Reasoning over Multi-Video Haystacks

As of 7 August 2026, this Paper Citation Record lists 100 of 127 outbound references and 3 inbound Pith citation observations for arXiv:2506.07016.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.07016 v2

Coverage vector

measured 100 of 127 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T05:49:53.545730Z

measured 103 of 103 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 3 of 3 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-06T22:41:05.001058Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-06-28T19:22:34.453758Z

Reference resolution

100 of 127 outbound references displayed

  • verified exact2
  • verified fuzzy2
  • unresolved96
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation c9c349f6-6edd-46a0-bdcb-665277de79c7 · outbound

This paper cites Crab: A Unified Audio-Visual Scene Understanding Model with Explicit Cooperation.

MAGNET: A Multi-agent Framework for Finding Audio-Visual Needles by Reasoning over Multi-Video Haystacks Crab: A Unified Audio-Visual Scene Understanding Model with Explicit Cooperation

Reference 1

Resolution
verified exact
local_arxiv, observed 2026-08-07T05:49:54.269852Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:49:53.235638Z digest=sha256:232b3deddfedc1cbc868c44bef802180f2f5804a24ce7a293353b2c810507621

Observation ae26d71d-10cd-4f1b-987c-f9fc48a03f20 · outbound

This paper cites Meerkat: Audio-visual large language model for grounding in space and time.

MAGNET: A Multi-agent Framework for Finding Audio-Visual Needles by Reasoning over Multi-Video Haystacks Meerkat: Audio-visual large language model for grounding in space and time

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T05:49:53.239537Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:49:53.239537Z digest=sha256:d4fddafc1ef3f33adfa662ab23665bc016238a15308366185fd182c84d100411

Observation 6fbb984b-5ef9-49ec-abb0-41193774fd76 · outbound

This paper cites VITA: Towards Open-Source Interactive Omni Multimodal LLM.

MAGNET: A Multi-agent Framework for Finding Audio-Visual Needles by Reasoning over Multi-Video Haystacks VITA: Towards Open-Source Interactive Omni Multimodal LLM

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T05:49:53.242431Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:49:53.242431Z digest=sha256:1114c5177c320c7debb20434d33802acee372e327500b28629b0be010a4939bf

Observation 9281e5fc-12a0-44b1-b467-6ee8a1b32bbc · outbound

This paper cites Unified-io 2: Scaling autoregressive multimodal models with vision language audio and action.

MAGNET: A Multi-agent Framework for Finding Audio-Visual Needles by Reasoning over Multi-Video Haystacks Unified-io 2: Scaling autoregressive multimodal models with vision language audio and action

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T05:49:53.245607Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:49:53.245607Z digest=sha256:c1eaf10d022333c0a3871267605bf4bc371af9b52e54e5636fd1b6d312e70d67

Observation cb2ada9f-3ebe-453f-974a-ce660e33b6ce · outbound

This paper cites CREMA: Generalizable and Efficient Video-Language Reasoning via Multimodal Modular Fusion.

MAGNET: A Multi-agent Framework for Finding Audio-Visual Needles by Reasoning over Multi-Video Haystacks CREMA: Generalizable and Efficient Video-Language Reasoning via Multimodal Modular Fusion

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T05:49:53.248569Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:49:53.248569Z digest=sha256:75b210ae7a321ca5aa00f477bbe8bf204ca997cb533ed0f0a9030b93a67a9f71

Observation 07fa958e-1721-4972-a904-1b52d70a42c5 · outbound

This paper cites Avicuna: Audio-visual llm with interleaver and context-boundary alignment for temporal referential dialogue.

MAGNET: A Multi-agent Framework for Finding Audio-Visual Needles by Reasoning over Multi-Video Haystacks Avicuna: Audio-visual llm with interleaver and context-boundary alignment for temporal referential dialogue

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T05:49:53.251727Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:49:53.251727Z digest=sha256:c4dec5679fdc95c7a564ef9fd0e452ac596a7f0905279e8ca343544a539ed936

Observation 2aa1aa53-cd40-479c-b381-cab4192fa0bd · outbound

This paper cites MLVU: Benchmarking Multi-task Long Video Understanding.

MAGNET: A Multi-agent Framework for Finding Audio-Visual Needles by Reasoning over Multi-Video Haystacks MLVU: Benchmarking Multi-task Long Video Understanding

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T05:49:53.267029Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:49:53.267029Z digest=sha256:48d9651f4435f593916469488accb5918c2f92c8a9e0cdb56b2f860da6668542

Observation ffb5b741-b0a6-47a7-9850-a1b6c96facdf · outbound

This paper cites Sharegpt4video: Improving video understanding and generation with better captions.

MAGNET: A Multi-agent Framework for Finding Audio-Visual Needles by Reasoning over Multi-Video Haystacks Sharegpt4video: Improving video understanding and generation with better captions

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T05:49:53.270357Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:49:53.270357Z digest=sha256:62cec1c8c4fe05b3ec557c99ce95440cc9a3473b9ef3aa2e910eb285c6b44318

Observation a55a95d4-5a08-4bdb-a540-3c8398416327 · outbound

This paper cites CinePile: A Long Video Question Answering Dataset and Benchmark.

MAGNET: A Multi-agent Framework for Finding Audio-Visual Needles by Reasoning over Multi-Video Haystacks CinePile: A Long Video Question Answering Dataset and Benchmark

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T05:49:53.273176Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:49:53.273176Z digest=sha256:e4606086abc6e104156e97f0bc486efa79338c9a68231b6cf837703c883edbff

Observation b68cb8c7-9fec-421a-85ad-614890e1b7d6 · outbound

This paper cites Moviechat: From dense token to sparse memory for long video understanding.

MAGNET: A Multi-agent Framework for Finding Audio-Visual Needles by Reasoning over Multi-Video Haystacks Moviechat: From dense token to sparse memory for long video understanding

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T05:49:53.279518Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:49:53.279518Z digest=sha256:740a07894acf96f574fcb74dcecc36f212dce7e8220bb3d47e47fb810ac77c67

Observation 0642fdb3-3293-452d-aace-d241ec4e6a73 · outbound

This paper cites Dense-localizing audio-visual events in untrimmed videos: A large-scale benchmark and baseline.

MAGNET: A Multi-agent Framework for Finding Audio-Visual Needles by Reasoning over Multi-Video Haystacks Dense-localizing audio-visual events in untrimmed videos: A large-scale benchmark and baseline

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T05:49:53.281960Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:49:53.281960Z digest=sha256:739c2a8e96c11462af45ef2737db0bf07cbdf568d01c487dd49d0b2f33840f2e

Observation 28cd0823-f30f-498c-9861-143eb3a8bc9f · outbound

This paper cites Vast: A vision-audio-subtitle-text omni-modality foundation model and dataset.

MAGNET: A Multi-agent Framework for Finding Audio-Visual Needles by Reasoning over Multi-Video Haystacks Vast: A vision-audio-subtitle-text omni-modality foundation model and dataset

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T05:49:53.284473Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:49:53.284473Z digest=sha256:0d60cc08901ccc9665096243c803c71d4698e5dc6ea7ebe40330995371b36864

Observation 2f204c4f-25e5-4659-bdd6-a4b9f95b604a · outbound

This paper cites Avqa: A dataset for audio-visual question answering on videos.

MAGNET: A Multi-agent Framework for Finding Audio-Visual Needles by Reasoning over Multi-Video Haystacks Avqa: A dataset for audio-visual question answering on videos

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T05:49:53.287137Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:49:53.287137Z digest=sha256:4a334ff2dacca8d9f7ddd075265731f9ceeea80cfda7ea9e33792b462bf4e9e3

Observation e8a4da05-4271-4769-9f0e-d9b74fa231d7 · outbound

This paper cites Cat: Enhancing multimodal large language model to answer questions in dynamic audio-visual scenarios.

MAGNET: A Multi-agent Framework for Finding Audio-Visual Needles by Reasoning over Multi-Video Haystacks Cat: Enhancing multimodal large language model to answer questions in dynamic audio-visual scenarios

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T05:49:53.290455Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:49:53.290455Z digest=sha256:ab12268e9405eaf395d81ed8ac909507aa462d988cc3877d3877d1523fb77272

Observation 86e2d76f-14bd-411b-a289-923b64fd4cb2 · outbound

This paper cites Learning to answer questions in dynamic audio-visual scenarios.

MAGNET: A Multi-agent Framework for Finding Audio-Visual Needles by Reasoning over Multi-Video Haystacks Learning to answer questions in dynamic audio-visual scenarios

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T05:49:53.293237Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:49:53.293237Z digest=sha256:6c0e554af474146553c061d660bef11875abf4e9cea7fcc9843abc00f19c6ddb

Observation 2b52dfc9-6683-4c2b-abb7-c81cab78b27a · outbound

This paper cites Vggsound: A large-scale audio- visual dataset.

MAGNET: A Multi-agent Framework for Finding Audio-Visual Needles by Reasoning over Multi-Video Haystacks Vggsound: A large-scale audio- visual dataset

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T05:49:53.296583Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:49:53.296583Z digest=sha256:76ae5ef9f70ef25f962db4f8491d682d4bd3710a1a4357251dd020b80b73cad2

Observation 74a42b6f-062f-4581-8e7b-74b0436b721c · outbound

This paper cites SAVEn-Vid: Synergistic Audio-Visual Integration for Enhanced Understanding in Long Video Context.

MAGNET: A Multi-agent Framework for Finding Audio-Visual Needles by Reasoning over Multi-Video Haystacks SAVEn-Vid: Synergistic Audio-Visual Integration for Enhanced Understanding in Long Video Context

Reference 22

Resolution
verified exact
local_arxiv, observed 2026-08-07T05:49:54.215338Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:49:53.299116Z digest=sha256:7a547e5431c5947234493d81a560b0a6e3227fa009d4e19b0ca5cc6e2bc2f7f6

Observation 9743a3e4-b48c-45de-a516-1a9cbe36c5de · outbound

This paper cites Imagebind: One embedding space to bind them all.

MAGNET: A Multi-agent Framework for Finding Audio-Visual Needles by Reasoning over Multi-Video Haystacks Imagebind: One embedding space to bind them all

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T05:49:53.302017Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:49:53.302017Z digest=sha256:4bbe58abe4a5b4c15e9b48c8fcf384f84db1f2f4cfb2446028235a1ef03ba657

Observation fd4a3517-5586-4f10-a753-2a6c64bfd37b · outbound

This paper cites Gemini: Google’s multimodal ai model.

MAGNET: A Multi-agent Framework for Finding Audio-Visual Needles by Reasoning over Multi-Video Haystacks Gemini: Google’s multimodal ai model

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T05:49:53.305234Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:49:53.305234Z digest=sha256:814ab4f92d06560363f71c4702353df9974af1dffc5b349e79a62ec8dbf4548a

Observation 14d0b676-6f3f-4568-95a6-a122675a9208 · outbound

This paper cites Qwen2.5 Technical Report.

MAGNET: A Multi-agent Framework for Finding Audio-Visual Needles by Reasoning over Multi-Video Haystacks Qwen2.5 Technical Report

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T05:49:53.307666Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:49:53.307666Z digest=sha256:093eb68e1e198636c2774ed4f0f044266a8c566b9c7a31189b90ba2bbad0bc97

Observation d21a88a6-d50d-4291-8c0a-83ba536d916d · outbound

This paper cites GPT-4o System Card.

MAGNET: A Multi-agent Framework for Finding Audio-Visual Needles by Reasoning over Multi-Video Haystacks GPT-4o System Card

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T05:49:53.310525Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:49:53.310525Z digest=sha256:43387ac3e247b3723326ed2f0d7e3775b77be4830b60868cb3f91d958bacbdf9

Observation 43ef4468-98cc-4c9e-9be2-577564468c64 · outbound

This paper cites Towards General Text Embeddings with Multi-stage Contrastive Learning.

MAGNET: A Multi-agent Framework for Finding Audio-Visual Needles by Reasoning over Multi-Video Haystacks Towards General Text Embeddings with Multi-stage Contrastive Learning

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T05:49:53.313653Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:49:53.313653Z digest=sha256:e20e9cde55c556121297a9f5b210decaf5e50fbbc7fe896576713fe05d1524c3

Observation 0cc2a8d5-55cd-4c78-9013-06881199978f · outbound

This paper cites Variants of the hungarian method for assignment problems.

MAGNET: A Multi-agent Framework for Finding Audio-Visual Needles by Reasoning over Multi-Video Haystacks Variants of the hungarian method for assignment problems

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T05:49:53.316432Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:49:53.316432Z digest=sha256:8960e322d595861b214ef2ea72bfa59c3641a53101f208800b09f4e8f80244d5

Observation c9527294-ac5e-4fa7-bc35-06407f19f055 · outbound

This paper cites Video-rag: Visually-aligned retrieval-augmented long video comprehension.

MAGNET: A Multi-agent Framework for Finding Audio-Visual Needles by Reasoning over Multi-Video Haystacks Video-rag: Visually-aligned retrieval-augmented long video comprehension

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T05:49:53.319165Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:49:53.319165Z digest=sha256:10ce0d79275c6d28d924b069cae3fde16351da09ef7455f891ea1caf306538b3

Observation 75b41b82-8f33-443b-a107-50688a16e005 · outbound

This paper cites video-SALMONN: Speech-Enhanced Audio-Visual Large Language Models.

MAGNET: A Multi-agent Framework for Finding Audio-Visual Needles by Reasoning over Multi-Video Haystacks video-SALMONN: Speech-Enhanced Audio-Visual Large Language Models

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T05:49:53.321734Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:49:53.321734Z digest=sha256:0cbb1dff8c2609a1b6e13493f59f6f6664c861d16c986bbb3b1c19ebe913b811

Observation 861d624e-75d6-47de-be27-79ed826e7e49 · outbound

This paper cites High-fidelity audio compression with improved rvqgan.

MAGNET: A Multi-agent Framework for Finding Audio-Visual Needles by Reasoning over Multi-Video Haystacks High-fidelity audio compression with improved rvqgan

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T05:49:53.324462Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:49:53.324462Z digest=sha256:e48a83bdba62373f0544b535888fc17cc84ddd6facce084de207303055ac9bd0

Observation 4faf8e1e-e314-4294-9137-d87b8caf369f · outbound

This paper cites Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context.

MAGNET: A Multi-agent Framework for Finding Audio-Visual Needles by Reasoning over Multi-Video Haystacks Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T05:49:53.326881Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:49:53.326881Z digest=sha256:0777ed424232a098fc97b3954cc00d326ecfdd7944e1a3d14106738f82f9d9d5

Observation 521efbb0-9174-4f8f-9d04-0aca524bec44 · outbound

This paper cites Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis.

MAGNET: A Multi-agent Framework for Finding Audio-Visual Needles by Reasoning over Multi-Video Haystacks Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T05:49:53.329657Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:49:53.329657Z digest=sha256:8bf281c549b3e7e3d0ac003056bec28e9192c915f76310920245219edb7398e4

Observation 64c22047-d71b-4f29-9eed-692710e1357a · outbound

This paper cites Mvbench: A comprehensive multi-modal video understanding benchmark.

MAGNET: A Multi-agent Framework for Finding Audio-Visual Needles by Reasoning over Multi-Video Haystacks Mvbench: A comprehensive multi-modal video understanding benchmark

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T05:49:53.332855Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:49:53.332855Z digest=sha256:3f9b7039744aa23d5f062bfa7169fee4353f3f10ee220d7ba1ccdf61bd3a4b84

Observation 1e9f3a76-d7a7-4118-adb0-2a0b0df3377f · outbound

This paper cites Video Question Answering: Datasets, Algorithms and Challenges.

MAGNET: A Multi-agent Framework for Finding Audio-Visual Needles by Reasoning over Multi-Video Haystacks Video Question Answering: Datasets, Algorithms and Challenges

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T05:49:53.335527Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:49:53.335527Z digest=sha256:3204d36713c28c4098fca68fefddc597a56279d56f3156d05e627f5c09c9af29

Observation d005edee-7e59-453d-a80b-70df346af50f · outbound

This paper cites LVBench: An Extreme Long Video Understanding Benchmark.

MAGNET: A Multi-agent Framework for Finding Audio-Visual Needles by Reasoning over Multi-Video Haystacks LVBench: An Extreme Long Video Understanding Benchmark

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T05:49:53.338244Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:49:53.338244Z digest=sha256:b835e9f4aa884bae54333967bf2c386994bfe13e373d40b8ecf2ee62e3770e7b

Observation 44fd02b2-02c6-4094-a565-5d80afc81ff9 · outbound

This paper cites Movieqa: Understanding stories in movies through question-answering.

MAGNET: A Multi-agent Framework for Finding Audio-Visual Needles by Reasoning over Multi-Video Haystacks Movieqa: Understanding stories in movies through question-answering

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-07T05:49:53.341351Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:49:53.341351Z digest=sha256:e517d365507b79d6fe37f96f9978d708b51936bd2c8001b84343e3b6bc75ccdd

Observation 1e25c88d-2c88-4a7a-9074-80b411e855b5 · outbound

This paper cites Are we asking the right questions in movieqa? In Proceedings of the IEEE/CVF International Conference on Computer Vision Workshops, pages 0–0, 2019.

MAGNET: A Multi-agent Framework for Finding Audio-Visual Needles by Reasoning over Multi-Video Haystacks Are we asking the right questions in movieqa? In Proceedings of the IEEE/CVF International Conference on Computer Vision Workshops, pages 0–0, 2019

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-07T05:49:53.344344Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:49:53.344344Z digest=sha256:d579328930b98e9205dad7bea3b18680fb225754aa5c5e50a75c3805139acb56

Observation e13897b5-9982-4b69-b512-dcf1b32f27c5 · outbound

This paper cites Activitynet-qa: A dataset for understanding complex web videos via question answering.

MAGNET: A Multi-agent Framework for Finding Audio-Visual Needles by Reasoning over Multi-Video Haystacks Activitynet-qa: A dataset for understanding complex web videos via question answering

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-07T05:49:53.347140Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:49:53.347140Z digest=sha256:aaa7bec6a7c800c692e2a458570e0a30814ce5ed223af59a768dc6d558093892

Observation 04042d7a-4529-42c6-add3-d490954fb018 · outbound

This paper cites How2: A Large-scale Dataset for Multimodal Language Understanding.

MAGNET: A Multi-agent Framework for Finding Audio-Visual Needles by Reasoning over Multi-Video Haystacks How2: A Large-scale Dataset for Multimodal Language Understanding

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-07T05:49:53.349615Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:49:53.349615Z digest=sha256:8bbb478894a5a624b7e2b3f7cb138298128f99e4d8dba555d9ee2a54aa9548a5

Observation dd516ceb-d6e4-4528-8a6e-b47a87b04001 · outbound

This paper cites Next-qa: Next phase of question-answering to explaining temporal actions.

MAGNET: A Multi-agent Framework for Finding Audio-Visual Needles by Reasoning over Multi-Video Haystacks Next-qa: Next phase of question-answering to explaining temporal actions

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-07T05:49:53.352727Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:49:53.352727Z digest=sha256:a1d33fd459eb1ef5da9bcc6475fd0762a240adad91a477c78396553543e47aa6

Observation 4de816fd-bc2a-4bcb-90f1-f11d582fca64 · outbound

This paper cites Perception test: A diagnostic benchmark for multimodal video models.

MAGNET: A Multi-agent Framework for Finding Audio-Visual Needles by Reasoning over Multi-Video Haystacks Perception test: A diagnostic benchmark for multimodal video models

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-07T05:49:53.354968Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:49:53.354968Z digest=sha256:05df6383386ed664068200c106853892d18da6091e2527a727c41d2f5c75a70f

Observation 28fd6f61-85cf-4000-9df3-5bdea4755f80 · outbound

This paper cites STAR: A Benchmark for Situated Reasoning in Real-World Videos.

MAGNET: A Multi-agent Framework for Finding Audio-Visual Needles by Reasoning over Multi-Video Haystacks STAR: A Benchmark for Situated Reasoning in Real-World Videos

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-07T05:49:53.357596Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:49:53.357596Z digest=sha256:059cf74a0bda302b0cc821a7aace7c35ca066e03f37a1ccdebc8eb8275047442

Observation c14e6391-f72e-47a3-9cbe-5663d75b4092 · outbound

This paper cites Agqa: A benchmark for compositional spatio-temporal reasoning.

MAGNET: A Multi-agent Framework for Finding Audio-Visual Needles by Reasoning over Multi-Video Haystacks Agqa: A benchmark for compositional spatio-temporal reasoning

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-07T05:49:53.360494Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:49:53.360494Z digest=sha256:6f9eefe5b1cbc676fa3d7bd5a4a585ca23f8a7fd32c8aa5cce01a569e69f06a2

Observation dfa4da0a-2546-47a2-9d07-f89a2ae3df00 · outbound

This paper cites Egoschema: A diagnostic benchmark for very long-form video language understanding.

MAGNET: A Multi-agent Framework for Finding Audio-Visual Needles by Reasoning over Multi-Video Haystacks Egoschema: A diagnostic benchmark for very long-form video language understanding

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-07T05:49:53.363689Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:49:53.363689Z digest=sha256:f6e209d6468eaa443f9f856822205ec0d8d4086ee7f42ad0adaf174822344cd4

Observation 4129b2ae-f882-4af7-9895-bbe0823ac0cb · outbound

This paper cites Longvideobench: A benchmark for long-context interleaved video-language understanding.

MAGNET: A Multi-agent Framework for Finding Audio-Visual Needles by Reasoning over Multi-Video Haystacks Longvideobench: A benchmark for long-context interleaved video-language understanding

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-07T05:49:53.366642Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:49:53.366642Z digest=sha256:a988d70dd81f57da7e39208a841fbb94ded588d50d00eb0f8c4eea7015c325fa

Observation c8204057-19fa-4593-b216-cc2fbaa405a1 · outbound

This paper cites Just ask: Learning to answer questions from millions of narrated videos.

MAGNET: A Multi-agent Framework for Finding Audio-Visual Needles by Reasoning over Multi-Video Haystacks Just ask: Learning to answer questions from millions of narrated videos

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-07T05:49:53.369194Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:49:53.369194Z digest=sha256:037283c1ea8132a9016ff28e3574494e1bad07c418ca3873c5c456d2a282f696

Observation b3a54a4b-e9d3-4d54-a0b1-7cbe0b1e541f · outbound

This paper cites InstructionBench: An Instructional Video Understanding Benchmark.

MAGNET: A Multi-agent Framework for Finding Audio-Visual Needles by Reasoning over Multi-Video Haystacks InstructionBench: An Instructional Video Understanding Benchmark

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-07T05:49:53.372532Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:49:53.372532Z digest=sha256:2172b2fc55c1c74ee28446dc287c95c3a40874ca7625f57179601e5705bef294

Observation cb94633a-c344-46d0-a195-ee1cbbcc352c · outbound

This paper cites HD-EPIC: A Highly-Detailed Egocentric Video Dataset.

MAGNET: A Multi-agent Framework for Finding Audio-Visual Needles by Reasoning over Multi-Video Haystacks HD-EPIC: A Highly-Detailed Egocentric Video Dataset

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-07T05:49:53.375278Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:49:53.375278Z digest=sha256:110b0a181b7d79905902010627ad27a8e02e06cd0449c69bdb084b6ed121759d

Observation 7fa9c4d4-f755-44ca-9822-11dfc0246416 · outbound

This paper cites Ego4d: Around the world in 3,000 hours of egocentric video.

MAGNET: A Multi-agent Framework for Finding Audio-Visual Needles by Reasoning over Multi-Video Haystacks Ego4d: Around the world in 3,000 hours of egocentric video

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-07T05:49:53.378573Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:49:53.378573Z digest=sha256:81259628fdf644e3eae64e260d70e5eb93f18ea019b7929e15caa857ce635521

Observation 812bdbe8-357f-4318-92ae-ef9fe1546578 · outbound

This paper cites Ego-exo4d: Understand- ing skilled human activity from first-and third-person perspectives.

MAGNET: A Multi-agent Framework for Finding Audio-Visual Needles by Reasoning over Multi-Video Haystacks Ego-exo4d: Understand- ing skilled human activity from first-and third-person perspectives

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-07T05:49:53.381710Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:49:53.381710Z digest=sha256:d68926fa0cc3be28c82405fc5a77a13f13c5030a35fe0f8d171803a4624b4b4d

Observation 4a9a7a8e-fa8e-465d-a7c3-0011f7212f13 · outbound

This paper cites Scaling egocentric vision: The epic-kitchens dataset.

MAGNET: A Multi-agent Framework for Finding Audio-Visual Needles by Reasoning over Multi-Video Haystacks Scaling egocentric vision: The epic-kitchens dataset

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-07T05:49:53.384201Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:49:53.384201Z digest=sha256:c8c8960b54180a159986cf867ce10a6116b5b9120b38ee619d8aa172e424a0b8

Observation 5c323aad-84cb-4215-9c20-e0af2b7a6b7e · outbound

This paper cites Vicuna: An open-source chatbot impressing gpt-4 with 90%* chatgpt quality.

MAGNET: A Multi-agent Framework for Finding Audio-Visual Needles by Reasoning over Multi-Video Haystacks Vicuna: An open-source chatbot impressing gpt-4 with 90%* chatgpt quality

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-07T05:49:53.387583Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:49:53.387583Z digest=sha256:7f33741fd1ec44c1cae5a12bb62deebbbfd509cb07b6b2e74db14c5fce16e311

Observation f5324068-7c28-4598-a0ac-5e7f23433fe8 · outbound

This paper cites LLaMA: Open and Efficient Foundation Language Models.

MAGNET: A Multi-agent Framework for Finding Audio-Visual Needles by Reasoning over Multi-Video Haystacks LLaMA: Open and Efficient Foundation Language Models

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-07T05:49:53.390941Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:49:53.390941Z digest=sha256:797c982f2aa4dde08842e57bcd3d1eb8d147322b04223fb71519d824f538ea25

Observation 4714920f-c24b-4c1a-a7e2-0299bb14b562 · outbound

This paper cites Mistral 7B.

MAGNET: A Multi-agent Framework for Finding Audio-Visual Needles by Reasoning over Multi-Video Haystacks Mistral 7B

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-07T05:49:53.393475Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:49:53.393475Z digest=sha256:c440977f92c8787f7e39066562843beb76b48ce1cd8395c82dbafa8dc72cc949

Observation 20fd0780-27b5-4a90-872a-f5ccf4089a5d · outbound

This paper cites VideoChat: Chat-Centric Video Understanding.

MAGNET: A Multi-agent Framework for Finding Audio-Visual Needles by Reasoning over Multi-Video Haystacks VideoChat: Chat-Centric Video Understanding

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-07T05:49:53.397026Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:49:53.397026Z digest=sha256:caf98fa1c1bec3281e3362971c19a6944f60b2cda799aa73cf0a6eef9948cb4d

Observation 78fa4c01-ccd7-4fb7-b01c-789b8ccbdd51 · outbound

This paper cites Video-LLaMA: An Instruction-tuned Audio-Visual Language Model for Video Understanding.

MAGNET: A Multi-agent Framework for Finding Audio-Visual Needles by Reasoning over Multi-Video Haystacks Video-LLaMA: An Instruction-tuned Audio-Visual Language Model for Video Understanding

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-07T05:49:53.399967Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:49:53.399967Z digest=sha256:da65355c550e4c3d5ee935b6c8960c0567696e3fb21abee32f0225748b0a1422

Observation 3e0fbcc7-cc27-4663-a7cd-b84c9c0db2f4 · outbound

This paper cites LLaVA-OneVision: Easy Visual Task Transfer.

MAGNET: A Multi-agent Framework for Finding Audio-Visual Needles by Reasoning over Multi-Video Haystacks LLaVA-OneVision: Easy Visual Task Transfer

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-07T05:49:53.403166Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:49:53.403166Z digest=sha256:2d0368e03d3e70e230ca5461f4e1ac55685f642286bac87e0ac82791dbf41cc6

Observation 01eb85a9-d3e8-4f7e-9a2d-ffdb046686dc · outbound

This paper cites World Model on Million-Length Video And Language With Blockwise RingAttention.

MAGNET: A Multi-agent Framework for Finding Audio-Visual Needles by Reasoning over Multi-Video Haystacks World Model on Million-Length Video And Language With Blockwise RingAttention

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-07T05:49:53.405853Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:49:53.405853Z digest=sha256:79368b215197ff542a84466279c5edfe3b8fabe869f9bcc5c7d4a1ff8f363926

Observation 372c0c6c-ddb6-46d8-8879-8a2903415c3f · outbound

This paper cites Long Context Transfer from Language to Vision.

MAGNET: A Multi-agent Framework for Finding Audio-Visual Needles by Reasoning over Multi-Video Haystacks Long Context Transfer from Language to Vision

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-07T05:49:53.408552Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:49:53.408552Z digest=sha256:d920340613ca04a15c11c69963864bdb6f9fa1c2780a7254dae20fcad2a3be3c

Observation ea34d88d-cd06-4fee-838a-1704492c1e8c · outbound

This paper cites LongVILA: Scaling Long-Context Visual Language Models for Long Videos.

MAGNET: A Multi-agent Framework for Finding Audio-Visual Needles by Reasoning over Multi-Video Haystacks LongVILA: Scaling Long-Context Visual Language Models for Long Videos

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-07T05:49:53.411481Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:49:53.411481Z digest=sha256:bb33e15a1969fe0487140d6e8b4ceb4bb2ed42541a90f4b2780f8e2ebab4da5e

Observation 785eeebb-0c74-458b-91e0-3df35fdf5fc7 · outbound

This paper cites Longllava: Scaling multi-modal llms to 1000 images efficiently via hybrid architecture.

MAGNET: A Multi-agent Framework for Finding Audio-Visual Needles by Reasoning over Multi-Video Haystacks Longllava: Scaling multi-modal llms to 1000 images efficiently via hybrid architecture

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-07T05:49:53.414273Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:49:53.414273Z digest=sha256:bb89c3bcfa7b2461563541d1c4aa448a70dfec2f449ff2d791526f6169e4aa5f

Observation dc8fecce-635a-4350-823e-a604ad8268a5 · outbound

This paper cites LLaMA-VID: An Image is Worth 2 Tokens in Large Language Models.

MAGNET: A Multi-agent Framework for Finding Audio-Visual Needles by Reasoning over Multi-Video Haystacks LLaMA-VID: An Image is Worth 2 Tokens in Large Language Models

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-07T05:49:53.416913Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:49:53.416913Z digest=sha256:46ca3ef6b87be7731b298d441b06370df71a62f9ac58dc54484796f69ceef3b7

Observation 5c7ac51c-e03f-4c20-be75-737f1f5e4b4e · outbound

This paper cites Flash-VStream: Memory-Based Real-Time Understanding for Long Video Streams.

MAGNET: A Multi-agent Framework for Finding Audio-Visual Needles by Reasoning over Multi-Video Haystacks Flash-VStream: Memory-Based Real-Time Understanding for Long Video Streams

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-07T05:49:53.419480Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:49:53.419480Z digest=sha256:cd9250ada6b6c604c76d9d69d5242e44f05e6bb30696c950ec2bd72b0756d772

Observation 1eca1a7e-e2ae-4794-9fb8-a2cf4f1428ac · outbound

This paper cites VideoAgent: A Memory-augmented Multimodal Agent for Video Understanding.

MAGNET: A Multi-agent Framework for Finding Audio-Visual Needles by Reasoning over Multi-Video Haystacks VideoAgent: A Memory-augmented Multimodal Agent for Video Understanding

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-07T05:49:53.422447Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:49:53.422447Z digest=sha256:0ba5e7c9562b7604224cd7a1b79f4ff07f7dcf5508e774279a5b155d35022e70

Observation c0c07cc0-f991-4305-80a8-37b2a2c996c5 · outbound

This paper cites VideoAgent: Long-form Video Understanding with Large Language Model as Agent.

MAGNET: A Multi-agent Framework for Finding Audio-Visual Needles by Reasoning over Multi-Video Haystacks VideoAgent: Long-form Video Understanding with Large Language Model as Agent

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-07T05:49:53.425309Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:49:53.425309Z digest=sha256:eec69ef149e15819c783e0b1b4139eececbeb4d8e02594b979314de533c8238f

Observation 28f45b09-2bf9-437a-841d-7b2e7489b3af · outbound

This paper cites VideoTree: Adaptive Tree-based Video Representation for LLM Reasoning on Long Videos.

MAGNET: A Multi-agent Framework for Finding Audio-Visual Needles by Reasoning over Multi-Video Haystacks VideoTree: Adaptive Tree-based Video Representation for LLM Reasoning on Long Videos

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-07T05:49:53.428376Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:49:53.428376Z digest=sha256:8d3d8c78dbf86bac1eec7028c81e6c0fe3e996816b79d7ac21331230b04b2e79

Observation 307dc06e-81df-428d-b77f-398c091a9c6c · outbound

This paper cites Retrieval-augmented generation for knowledge-intensive nlp tasks.

MAGNET: A Multi-agent Framework for Finding Audio-Visual Needles by Reasoning over Multi-Video Haystacks Retrieval-augmented generation for knowledge-intensive nlp tasks

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-07T05:49:53.431459Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:49:53.431459Z digest=sha256:f0b342b8af67461f351279cbab3e1a39911dae592fb5abc75893a40362d42cc2

Observation b84ea8fc-79e5-4af5-89b9-ee20d9c47a53 · outbound

This paper cites Active retrieval augmented generation.

MAGNET: A Multi-agent Framework for Finding Audio-Visual Needles by Reasoning over Multi-Video Haystacks Active retrieval augmented generation

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-07T05:49:53.434530Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:49:53.434530Z digest=sha256:e1bd241ff9665956e92c03ed4d7782b34a8e2329323a15b159037efa6580d0f7

Observation 8afc8e91-0f1c-47c2-8a98-f9cd1ff1daed · outbound

This paper cites Sentence-level prompts benefit composed image retrieval.

MAGNET: A Multi-agent Framework for Finding Audio-Visual Needles by Reasoning over Multi-Video Haystacks Sentence-level prompts benefit composed image retrieval

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-07T05:49:53.437517Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:49:53.437517Z digest=sha256:bec8234759d4b29af72e58780b1c44857d634f240d1c77c83ffc7c3b2be3f179

Observation 7e026fbc-b4f4-4fb3-806c-4d987713efd9 · outbound

This paper cites VQA4CIR: Boosting Composed Image Retrieval with Visual Question Answering.

MAGNET: A Multi-agent Framework for Finding Audio-Visual Needles by Reasoning over Multi-Video Haystacks VQA4CIR: Boosting Composed Image Retrieval with Visual Question Answering

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-07T05:49:53.440763Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:49:53.440763Z digest=sha256:ee3865cdda6f8dd894dbf32bb07fa0f61696b2a5d1a3b0b314439163a7a358db

Observation fd6652c5-5c9d-4bb3-9842-a4c75f95a24c · outbound

This paper cites Searching for Best Practices in Retrieval-Augmented Generation.

MAGNET: A Multi-agent Framework for Finding Audio-Visual Needles by Reasoning over Multi-Video Haystacks Searching for Best Practices in Retrieval-Augmented Generation

Reference 72

Resolution
unresolved
no resolver link, observed 2026-08-07T05:49:53.443996Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:49:53.443996Z digest=sha256:e3f15e4d15572dea03077cbf313211a25c515842c1ebc6440a922a30d842b966

Observation 211330bc-be32-4ab6-8105-170281393d0f · outbound

This paper cites Retrieval-Augmented Generation for Natural Language Processing: A Survey.

MAGNET: A Multi-agent Framework for Finding Audio-Visual Needles by Reasoning over Multi-Video Haystacks Retrieval-Augmented Generation for Natural Language Processing: A Survey

Reference 73

Resolution
unresolved
no resolver link, observed 2026-08-07T05:49:53.447621Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:49:53.447621Z digest=sha256:6d223cace802bc6ecf3b9b9f5f7d9258ae89b9f52c1f20045919e037b0d47bd4

Observation 3de5b9a1-3313-48b8-a56f-8d2929889cf1 · outbound

This paper cites A Survey on Retrieval-Augmented Text Generation for Large Language Models.

MAGNET: A Multi-agent Framework for Finding Audio-Visual Needles by Reasoning over Multi-Video Haystacks A Survey on Retrieval-Augmented Text Generation for Large Language Models

Reference 74

Resolution
unresolved
no resolver link, observed 2026-08-07T05:49:53.450836Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:49:53.450836Z digest=sha256:c055b21b8f7cb3bceffa52018395142c12fa9633ef718fd21c338c4f58e09c83

Observation cea28d7a-102d-4cb3-856a-122697f58eb1 · outbound

This paper cites Retrieval-Augmented Generation for AI-Generated Content: A Survey.

MAGNET: A Multi-agent Framework for Finding Audio-Visual Needles by Reasoning over Multi-Video Haystacks Retrieval-Augmented Generation for AI-Generated Content: A Survey

Reference 75

Resolution
unresolved
no resolver link, observed 2026-08-07T05:49:53.453946Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:49:53.453946Z digest=sha256:018793ff7bc125f5fe373cc6bfd4e50d9ee8669b038ab168d02ae30195940280

Observation bf8931c0-9ddf-442f-9271-4b1c7f1279ff · outbound

This paper cites Self-RAG: Learning to Retrieve, Generate, and Critique through Self-Reflection.

MAGNET: A Multi-agent Framework for Finding Audio-Visual Needles by Reasoning over Multi-Video Haystacks Self-RAG: Learning to Retrieve, Generate, and Critique through Self-Reflection

Reference 76

Resolution
unresolved
no resolver link, observed 2026-08-07T05:49:53.457193Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:49:53.457193Z digest=sha256:b4c9a295e32f7c3e4509b192ef12c4af3461e53f8668805c10dbf08f20686049

Observation 82e24110-d08c-4dbe-9695-4b2b63acb26f · outbound

This paper cites Realm: Retrieval- augmented language model pre-training.

MAGNET: A Multi-agent Framework for Finding Audio-Visual Needles by Reasoning over Multi-Video Haystacks Realm: Retrieval- augmented language model pre-training

Reference 77

Resolution
unresolved
no resolver link, observed 2026-08-07T05:49:53.460122Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:49:53.460122Z digest=sha256:902e7c3f8a4cc30666b5133f566fd241626ab013e0a7ca2cecb3f955d1e2a9d0

Observation a1e3442f-a6d3-46a9-8103-6820d27ca7e6 · outbound

This paper cites SAIL: Search-Augmented Instruction Learning.

MAGNET: A Multi-agent Framework for Finding Audio-Visual Needles by Reasoning over Multi-Video Haystacks SAIL: Search-Augmented Instruction Learning

Reference 78

Resolution
unresolved
no resolver link, observed 2026-08-07T05:49:53.463160Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:49:53.463160Z digest=sha256:51592a8ea4e691f566089ffe48a012d7acb95612154e996735f0da10616e9e7f

Observation f6995908-3bb9-49a2-8ee2-fe0fc8b4e038 · outbound

This paper cites Dense passage retrieval for open-domain question answering.

MAGNET: A Multi-agent Framework for Finding Audio-Visual Needles by Reasoning over Multi-Video Haystacks Dense passage retrieval for open-domain question answering

Reference 79

Resolution
unresolved
no resolver link, observed 2026-08-07T05:49:53.465906Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:49:53.465906Z digest=sha256:e245c8c8546d1ea162ba2473ffc48ede781be5147f85be8ac6cb5cb4e65cf914

Observation 9d1c0ee8-f18d-4add-88fc-4f40f10f5aea · outbound

This paper cites Document haystacks: Vision-language reasoning over piles of 1000+ documents.

MAGNET: A Multi-agent Framework for Finding Audio-Visual Needles by Reasoning over Multi-Video Haystacks Document haystacks: Vision-language reasoning over piles of 1000+ documents

Reference 80

Resolution
unresolved
no resolver link, observed 2026-08-07T05:49:53.469041Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:49:53.469041Z digest=sha256:452bf9eb32e305f86d0e02494d990a743ad2b85fa6274f74a82c8cbd8a98710a

Observation ecef7520-2cff-45d5-9b6d-d3d4fb5e731a · outbound

This paper cites an unresolved cited work.

MAGNET: A Multi-agent Framework for Finding Audio-Visual Needles by Reasoning over Multi-Video Haystacks Unresolved cited work

Reference 81

Resolution
unresolved
no resolver link, observed 2026-08-07T05:49:53.472251Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:49:53.472251Z digest=sha256:1428bfaad1017a198ae0d82a9c6c08c04321bb94a0448b963a5dd0db4c57dc25

Observation 03f03589-0a7e-44de-a1b4-023e511b63e0 · outbound

This paper cites Visual Haystacks: A Vision-Centric Needle-In-A-Haystack Benchmark.

MAGNET: A Multi-agent Framework for Finding Audio-Visual Needles by Reasoning over Multi-Video Haystacks Visual Haystacks: A Vision-Centric Needle-In-A-Haystack Benchmark

Reference 82

Resolution
unresolved
no resolver link, observed 2026-08-07T05:49:53.475228Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:49:53.475228Z digest=sha256:e350344423a6128029dcd2cc5a615fa982680f2d678114894384dcbdea6cd8ab

Observation b601c34f-c801-408f-b65c-b77ebb3da369 · outbound

This paper cites VisRAG: Vision-based Retrieval-augmented Generation on Multi-modality Documents.

MAGNET: A Multi-agent Framework for Finding Audio-Visual Needles by Reasoning over Multi-Video Haystacks VisRAG: Vision-based Retrieval-augmented Generation on Multi-modality Documents

Reference 83

Resolution
unresolved
no resolver link, observed 2026-08-07T05:49:53.478155Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:49:53.478155Z digest=sha256:309d0b55e1611fe3d6149b6eea847413697506c853387b85217f74c69ae2254f

Observation 326e72fd-02cb-4c9e-924a-f48db5577fca · outbound

This paper cites Visual-RAG: Benchmarking Text-to-Image Retrieval Augmented Generation for Visual Knowledge Intensive Queries.

MAGNET: A Multi-agent Framework for Finding Audio-Visual Needles by Reasoning over Multi-Video Haystacks Visual-RAG: Benchmarking Text-to-Image Retrieval Augmented Generation for Visual Knowledge Intensive Queries

Reference 84

Resolution
unresolved
no resolver link, observed 2026-08-07T05:49:53.481450Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:49:53.481450Z digest=sha256:663286b0cf11fc2be6aa38e4f66c9831c9d1f8d14833d6b30adf629cf51e495d

Observation ffd5fd2a-22ec-48b0-bdf5-303b0fafa39b · outbound

This paper cites Vdocrag: Retrieval-augmented generation over visually-rich documents.

MAGNET: A Multi-agent Framework for Finding Audio-Visual Needles by Reasoning over Multi-Video Haystacks Vdocrag: Retrieval-augmented generation over visually-rich documents

Reference 85

Resolution
unresolved
no resolver link, observed 2026-08-07T05:49:53.484538Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:49:53.484538Z digest=sha256:f7645fba740e884d5e4b00ca625a2eb4d074d01498f9819dc71099c0ec8b8e8f

Observation cc435739-a5a2-4dbf-bfbf-127ac913f76d · outbound

This paper cites Retrieval Augmented Visual Question Answering with Outside Knowledge.

MAGNET: A Multi-agent Framework for Finding Audio-Visual Needles by Reasoning over Multi-Video Haystacks Retrieval Augmented Visual Question Answering with Outside Knowledge

Reference 86

Resolution
unresolved
no resolver link, observed 2026-08-07T05:49:53.487275Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:49:53.487275Z digest=sha256:39256400420e0e1738d1d87d011bc6cafe3d10c816a1fb91239d1a1458a37172

Observation 13bf5ae5-ae74-47dc-b20a-7252476bba83 · outbound

This paper cites Fact-Aware Multimodal Retrieval Augmentation for Accurate Medical Radiology Report Generation.

MAGNET: A Multi-agent Framework for Finding Audio-Visual Needles by Reasoning over Multi-Video Haystacks Fact-Aware Multimodal Retrieval Augmentation for Accurate Medical Radiology Report Generation

Reference 87

Resolution
unresolved
no resolver link, observed 2026-08-07T05:49:53.490369Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:49:53.490369Z digest=sha256:da1880940d9e1bc42d63a75f8079224906bb3167b143eb2180d5606480087116

Observation 3a3c1727-6ef2-48e4-88fb-1affe423f913 · outbound

This paper cites Rule: Reliable multimodal rag for factuality in medical vision language models.

MAGNET: A Multi-agent Framework for Finding Audio-Visual Needles by Reasoning over Multi-Video Haystacks Rule: Reliable multimodal rag for factuality in medical vision language models

Reference 88

Resolution
unresolved
no resolver link, observed 2026-08-07T05:49:53.493573Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:49:53.493573Z digest=sha256:bc79187e2c5d6bdaed49e02b452b8d9fcb9deab9942fb18e44a2df9f02bf08ec

Observation fe7e62fc-4e71-4d67-8fa1-ece2db426d5e · outbound

This paper cites Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution.

MAGNET: A Multi-agent Framework for Finding Audio-Visual Needles by Reasoning over Multi-Video Haystacks Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution

Reference 89

Resolution
unresolved
no resolver link, observed 2026-08-07T05:49:53.496853Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:49:53.496853Z digest=sha256:734486683b337d0fcf71b064d416d8382c8afd9d9bfe6d1dee4dc813797ef91f

Observation 27713dc0-bfae-43f4-8ab4-52514fc9136e · outbound

This paper cites Gpt-4o: Enhanced multimodal language model.

MAGNET: A Multi-agent Framework for Finding Audio-Visual Needles by Reasoning over Multi-Video Haystacks Gpt-4o: Enhanced multimodal language model

Reference 90

Resolution
unresolved
no resolver link, observed 2026-08-07T05:49:53.499783Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:49:53.499783Z digest=sha256:5ca8b2bcd43d97206cb4828858344515b551957ce17948aba7115b6325bc3b18

Observation 740f4cb2-fa07-4a92-9683-d5d6d5972ec0 · outbound

This paper cites Minigpt-4: Enhancing vision-language understanding with advanced large language models.

MAGNET: A Multi-agent Framework for Finding Audio-Visual Needles by Reasoning over Multi-Video Haystacks Minigpt-4: Enhancing vision-language understanding with advanced large language models

Reference 91

Resolution
unresolved
no resolver link, observed 2026-08-07T05:49:53.502745Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:49:53.502745Z digest=sha256:1686792469ed091a3facc8bb37c1b040f49aa13e3f061bdc48e822f140cbdf43

Observation 8129eeaa-5041-402e-90f5-2c55f99f7835 · outbound

This paper cites Unified-IO: A Unified Model for Vision, Language, and Multi-Modal Tasks.

MAGNET: A Multi-agent Framework for Finding Audio-Visual Needles by Reasoning over Multi-Video Haystacks Unified-IO: A Unified Model for Vision, Language, and Multi-Modal Tasks

Reference 92

Resolution
unresolved
no resolver link, observed 2026-08-07T05:49:53.505625Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:49:53.505625Z digest=sha256:5479bfc9539054e900cea6210d7e97308032f2a1e953fd0ae36e1d08e9589742

Observation e9e009aa-e610-432d-a75e-7b9b36171587 · outbound

This paper cites MMMU: A Massive Multi-discipline Multimodal Understanding and Reasoning Benchmark for Expert AGI.

MAGNET: A Multi-agent Framework for Finding Audio-Visual Needles by Reasoning over Multi-Video Haystacks MMMU: A Massive Multi-discipline Multimodal Understanding and Reasoning Benchmark for Expert AGI

Reference 93

Resolution
unresolved
no resolver link, observed 2026-08-07T05:49:53.509123Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:49:53.509123Z digest=sha256:3e2875d7a4ba453747ed5c4c589c1cae9ea1f49287df76e88dcfcf87f6c4706a

Observation f397db9d-cebe-4215-ac62-ec113453365f · outbound

This paper cites Making the V in VQA matter: Elevating the role of image understanding in Visual Question Answering.

MAGNET: A Multi-agent Framework for Finding Audio-Visual Needles by Reasoning over Multi-Video Haystacks Making the V in VQA matter: Elevating the role of image understanding in Visual Question Answering

Reference 94

Resolution
unresolved
no resolver link, observed 2026-08-07T05:49:53.511845Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:49:53.511845Z digest=sha256:ca335586cdd72fe4d9ef664293d82a2b28b891771123778d10503d278eedcd91

Observation fd642cb6-bd08-4575-a76d-a41451365bf0 · outbound

This paper cites MathVerse: Does Your Multi-modal LLM Truly See the Diagrams in Visual Math Problems?.

MAGNET: A Multi-agent Framework for Finding Audio-Visual Needles by Reasoning over Multi-Video Haystacks MathVerse: Does Your Multi-modal LLM Truly See the Diagrams in Visual Math Problems?

Reference 95

Resolution
unresolved
no resolver link, observed 2026-08-07T05:49:53.514613Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:49:53.514613Z digest=sha256:74cc48aa2ef299e2735803766c3fcf62297eb774cc48244ea710ac64d4fbdd02

Observation 10e2c647-8096-410e-bede-17e3d463131c · outbound

This paper cites MathVista: Evaluating Mathematical Reasoning of Foundation Models in Visual Contexts.

MAGNET: A Multi-agent Framework for Finding Audio-Visual Needles by Reasoning over Multi-Video Haystacks MathVista: Evaluating Mathematical Reasoning of Foundation Models in Visual Contexts

Reference 96

Resolution
unresolved
no resolver link, observed 2026-08-07T05:49:53.517861Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:49:53.517861Z digest=sha256:11a1ed186aa2b292d704703e5d96cb1f8845115ca87e82adcb33b08e176db32c

Observation 8712754c-c291-4b51-9ca2-36203477e74a · outbound

This paper cites VideoRAG: Retrieval-Augmented Generation with Extreme Long-Context Videos.

MAGNET: A Multi-agent Framework for Finding Audio-Visual Needles by Reasoning over Multi-Video Haystacks VideoRAG: Retrieval-Augmented Generation with Extreme Long-Context Videos

Reference 97

Resolution
unresolved
no resolver link, observed 2026-08-07T05:49:53.521074Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:49:53.521074Z digest=sha256:0bcefdc6fe6ad7a4bd1076ae39a202c6f48fbe053add3bc31d722ef098622dfb

Observation 94a4611f-a6a2-4342-a19f-f5cbfc809bc6 · outbound

This paper cites V-desirr: Very fast deep embedded single image reflection removal.

MAGNET: A Multi-agent Framework for Finding Audio-Visual Needles by Reasoning over Multi-Video Haystacks V-desirr: Very fast deep embedded single image reflection removal

Reference 98

Resolution
unresolved
no resolver link, observed 2026-08-07T05:49:53.523879Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:49:53.523879Z digest=sha256:a9985a31e9efe3f5fdff9a7ec80e18ed1699a1101002463cced2c26c3e8987aa

Observation cb64ee82-bb51-4f1b-a730-c2ccfc1ca690 · outbound

This paper cites Measured albedo in the wild: Filling the gap in intrinsics evaluation.

MAGNET: A Multi-agent Framework for Finding Audio-Visual Needles by Reasoning over Multi-Video Haystacks Measured albedo in the wild: Filling the gap in intrinsics evaluation

Reference 99

Resolution
unresolved
no resolver link, observed 2026-08-07T05:49:53.526839Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:49:53.526839Z digest=sha256:7e585620fd329ad47dd1494f80e8f0d3546dc6957c79b4a2ac7c3ab7f053801a

Observation ccf0e7a5-89b1-422e-b981-79f3699f5364 · outbound

This paper cites Adverb: Visually guided audio dereverberation.

MAGNET: A Multi-agent Framework for Finding Audio-Visual Needles by Reasoning over Multi-Video Haystacks Adverb: Visually guided audio dereverberation

Reference 100

Resolution
unresolved
no resolver link, observed 2026-08-07T05:49:53.530209Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:49:53.530209Z digest=sha256:c9028e6f1ff8009eb2707a1c3073b2057968dd214ef328697f4d37d0c36ad9fb

Observation 24a613e3-e1f0-43fe-89fb-6bf65e1a63a2 · outbound

This paper cites Melfusion: Synthesizing music from image and language cues using diffusion models.

MAGNET: A Multi-agent Framework for Finding Audio-Visual Needles by Reasoning over Multi-Video Haystacks Melfusion: Synthesizing music from image and language cues using diffusion models

Reference 101

Resolution
unresolved
no resolver link, observed 2026-08-07T05:49:53.533720Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:49:53.533720Z digest=sha256:d41b2a8cc8d1af77f3917611ced5bbe9c12e2664a56b1929bbb5de14e14bb98a

Observation ece97e52-8192-434d-8c9a-7279ccfaf2c2 · outbound

This paper cites Foleygen: Visually-guided audio generation.

MAGNET: A Multi-agent Framework for Finding Audio-Visual Needles by Reasoning over Multi-Video Haystacks Foleygen: Visually-guided audio generation

Reference 102

Resolution
unresolved
no resolver link, observed 2026-08-07T05:49:53.536648Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:49:53.536648Z digest=sha256:e7b575bce7823076beb90cc736ad58a273b3273286fad5a6690468d24203539e

Observation e61e0399-ea2b-4839-8675-591d47f90ef7 · outbound

This paper cites Codi-2: In-context interleaved and interactive any-to-any generation.

MAGNET: A Multi-agent Framework for Finding Audio-Visual Needles by Reasoning over Multi-Video Haystacks Codi-2: In-context interleaved and interactive any-to-any generation

Reference 103

Resolution
unresolved
no resolver link, observed 2026-08-07T05:49:53.539318Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:49:53.539318Z digest=sha256:281613229f5674fa30147f69c44204674d4b51d5290fb686c31ebb49a6579228

Observation 95e292b2-e55e-489f-8221-178eeb0a9785 · outbound

This paper cites Listen to the pixels.

MAGNET: A Multi-agent Framework for Finding Audio-Visual Needles by Reasoning over Multi-Video Haystacks Listen to the pixels

Reference 104

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:49:54.393061Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:49:53.542548Z digest=sha256:80561a00740992d7e640cba0bead300c1290bcc57fb25e865d043c468a21a748

Observation 8df9726a-60c8-4102-83a5-372640adcbd6 · outbound

This paper cites Audvisum: Self- supervised deep reinforcement learning for diverse audio-visual summary generation.

MAGNET: A Multi-agent Framework for Finding Audio-Visual Needles by Reasoning over Multi-Video Haystacks Audvisum: Self- supervised deep reinforcement learning for diverse audio-visual summary generation

Reference 105

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:49:54.385050Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:49:53.545730Z digest=sha256:e5aa20533e2010669ef984c1b80f5fae2e8df3a50fcde8a3c80a12f8f09effba

Pith citing papers

Observation b7e657f6-9660-4ad7-a8ed-513e02acdef8 · inbound

EgoAdapt: Adaptive Multisensory Distillation and Policy Learning for Efficient Egocentric Perception cites this paper.

EgoAdapt: Adaptive Multisensory Distillation and Policy Learning for Efficient Egocentric Perception MAGNET: A Multi-agent Framework for Finding Audio-Visual Needles by Reasoning over Multi-Video Haystacks

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-06T22:41:05.001058Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:41:05.001058Z digest=sha256:2cf29c69601ed9bb73868ead6e932b790a9d3cbbbf5c0e2e4f12cc29a1ba1e82

Observation dbe2db90-2cb4-4672-b286-ff4a14a70b56 · inbound

EgoSound: Benchmarking Sound Understanding in Egocentric Videos cites this paper.

EgoSound: Benchmarking Sound Understanding in Egocentric Videos MAGNET: A Multi-agent Framework for Finding Audio-Visual Needles by Reasoning over Multi-Video Haystacks

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-05-15T21:46:42.485110Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-15T21:44:47.636912Z digest=sha256:be4e28ca796633662ad7ad3cb34178b29c0e6bd5d1cb1e17e41d871f0736ec71

Observation ea59bb54-f6e6-48c1-b475-63155e155898 · inbound

Through the PRISM: Principle-Aware, Interpretable, and Multi-Scale Evaluation of Visual Designs cites this paper.

Through the PRISM: Principle-Aware, Interpretable, and Multi-Scale Evaluation of Visual Designs MAGNET: A Multi-agent Framework for Finding Audio-Visual Needles by Reasoning over Multi-Video Haystacks

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-06-28T19:22:34.455260Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-28T19:19:09.199731Z digest=sha256:06daa5e02c8c2f925c8ee0b9d84428f312606a2cd024f8c9e67914fb4c1a0428