Pith. sign in

Paper Citation Record · LEDGER

MR. Video: "MapReduce" is the Principle for Long Video Understanding

As of 20 August 2026, this Paper Citation Record lists 100 of 148 outbound references and 4 inbound Pith citation observations for arXiv:2504.16082.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2504.16082 v1

Coverage vector

measured 100 of 148 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-16T11:14:20.027276Z

measured 104 of 104 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-20T06:33:59.587034+00:00

measured 4 of 4 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-06-28T06:33:32.090913Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-02T07:56:47.335540Z

Reference resolution

100 of 148 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved100
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation c1092f00-5e5c-4fb7-a37a-d7f1c007282f · outbound

This paper cites GPT-4 Technical Report.

MR. Video: "MapReduce" is the Principle for Long Video Understanding GPT-4 Technical Report

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-16T11:14:19.582563Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:14:19.582563Z digest=sha256:6fc46989b00013f98950f6fcd26f8d02f124fbf479d7a2329baeac8011c4f1ff

Observation 33642b88-0fab-4c1c-919c-3fa238d837a8 · outbound

This paper cites Qwen2.5-VL Technical Report.

MR. Video: "MapReduce" is the Principle for Long Video Understanding Qwen2.5-VL Technical Report

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-16T11:14:19.588215Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:14:19.588215Z digest=sha256:6bac8ae20b3469966b8177da59f12f9c119c138a0528d68411a85050bc356bda

Observation 647c0f4c-cd02-47b7-9f82-f91f92bb0ec4 · outbound

This paper cites AuroraCap: Efficient, Performant Video Detailed Captioning and a New Benchmark.

MR. Video: "MapReduce" is the Principle for Long Video Understanding AuroraCap: Efficient, Performant Video Detailed Captioning and a New Benchmark

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-16T11:14:19.593130Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:14:19.593130Z digest=sha256:9ed6a7a6e4e92c2dc63c0b2e91e2c1fe9b8c3e46f80b5bf0cebe998ddb7132b3

Observation 646c31ea-ce45-47d4-a619-77cc32fd244c · outbound

This paper cites An image is worth 1/2 tokens after layer 2: Plug-and-play inference acceleration for large vision-language models.

MR. Video: "MapReduce" is the Principle for Long Video Understanding An image is worth 1/2 tokens after layer 2: Plug-and-play inference acceleration for large vision-language models

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-16T11:14:19.598473Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:14:19.598473Z digest=sha256:bb21a1b11000cd8344a567d9874bff1a00fe23db88b38ab0a97200782c0c3a68

Observation 53603afc-4afc-44c7-8c83-35fec21371ac · outbound

This paper cites TimeMarker: A Versatile Video-LLM for Long and Short Video Understanding with Superior Temporal Localization Ability.

MR. Video: "MapReduce" is the Principle for Long Video Understanding TimeMarker: A Versatile Video-LLM for Long and Short Video Understanding with Superior Temporal Localization Ability

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-16T11:14:19.602914Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:14:19.602914Z digest=sha256:e4b6666508ab15e95f3ebffcea2b3efad7bb70b6f697d2c2d0bca2ea12489475

Observation 40466cbc-85d2-454c-8611-bf7245309421 · outbound

This paper cites InternVL: Scaling up vision foundation models and aligning for generic visual-linguistic tasks.

MR. Video: "MapReduce" is the Principle for Long Video Understanding InternVL: Scaling up vision foundation models and aligning for generic visual-linguistic tasks

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-16T11:14:19.607385Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:14:19.607385Z digest=sha256:e5c5fa72c93abd139ec85cda4b23a3e37a4d3442bdff602d4f9c94890d9c1064

Observation 044f8aa0-8947-47cd-a3af-430818a1af6e · outbound

This paper cites MapReduce: simplified data processing on large clusters.

MR. Video: "MapReduce" is the Principle for Long Video Understanding MapReduce: simplified data processing on large clusters

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-16T11:14:19.611681Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:14:19.611681Z digest=sha256:65f8f96264e510f601640d32399780c21b4fec784337fb95935e9e4b915b6a5d

Observation 25d57e1b-059c-4663-b110-7506926c0ef9 · outbound

This paper cites VideoAgent: A memory-augmented mul- timodal agent for video understanding.

MR. Video: "MapReduce" is the Principle for Long Video Understanding VideoAgent: A memory-augmented mul- timodal agent for video understanding

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-16T11:14:19.615823Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:14:19.615823Z digest=sha256:d769eb43babb7dbf408acb21f9e111e8e21ad97ba9cab3a2ccdf2285489a0b43

Observation 920db2e5-7d9e-4302-8dee-ba661337c856 · outbound

This paper cites Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis.

MR. Video: "MapReduce" is the Principle for Long Video Understanding Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-16T11:14:19.620433Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:14:19.620433Z digest=sha256:cbd9c7a9091dcf66a2d4a309f2ad4fa27ff002f9b8b3cd2eafb70ed363c5a8eb

Observation 1b507845-0c48-40e7-b30c-10e946c75054 · outbound

This paper cites ChatGLM: A Family of Large Language Models from GLM-130B to GLM-4 All Tools.

MR. Video: "MapReduce" is the Principle for Long Video Understanding ChatGLM: A Family of Large Language Models from GLM-130B to GLM-4 All Tools

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-16T11:14:19.625228Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:14:19.625228Z digest=sha256:9994e16d11926cb81dc16e4e4d9ac75da7b9ab7961314f739a053affa409a7e6

Observation f58cfd34-1c92-42f1-a391-341a333d5e34 · outbound

This paper cites Visual program- ming: Compositional visual reasoning without training.

MR. Video: "MapReduce" is the Principle for Long Video Understanding Visual program- ming: Compositional visual reasoning without training

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-16T11:14:19.629719Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:14:19.629719Z digest=sha256:23b540c3de29b2f5e1434171d6651f0ca23dc1caadb6e0076df38108b75fdba1

Observation 0d1fc84e-f9af-4a15-9e36-c6e51ece875a · outbound

This paper cites Perceiver: General perception with iterative attention.

MR. Video: "MapReduce" is the Principle for Long Video Understanding Perceiver: General perception with iterative attention

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-16T11:14:19.633906Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:14:19.633906Z digest=sha256:d5e1684508597b275159b8b648206d588aec3a46192fe9ac9dee21fc2a7a32cc

Observation 6f686851-14c1-4309-b03d-55fa62707c2c · outbound

This paper cites Perceiver IO: A general architecture for structured inputs & outputs.

MR. Video: "MapReduce" is the Principle for Long Video Understanding Perceiver IO: A general architecture for structured inputs & outputs

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-16T11:14:19.637981Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:14:19.637981Z digest=sha256:084b408bf669e23d61a9ba164d2e718b54849e7fbdeb4c0c3b9d13c2fcf5234e

Observation 130fe568-ca69-44f7-823f-3b2c2e6c4b16 · outbound

This paper cites SWE- Bench: Can language models resolve real-world github is- sues? In ICLR, 2024.

MR. Video: "MapReduce" is the Principle for Long Video Understanding SWE- Bench: Can language models resolve real-world github is- sues? In ICLR, 2024

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-16T11:14:19.642214Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:14:19.642214Z digest=sha256:6122968be0c84f1a7d3781701c4752d5974915baa34f7c0a970092e3b7d0d80b

Observation 3db1e49b-7a5c-4c55-93f7-4ff2c2d4f706 · outbound

This paper cites LLaVA-OneVision: Easy Visual Task Transfer.

MR. Video: "MapReduce" is the Principle for Long Video Understanding LLaVA-OneVision: Easy Visual Task Transfer

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-16T11:14:19.646991Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:14:19.646991Z digest=sha256:d05f7af80b8e8123f690b278703798259fd9361398d2f827a14f5df50bb74dbb

Observation d4b7b2d6-89d2-4c76-809a-f8873361e3d4 · outbound

This paper cites LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models.

MR. Video: "MapReduce" is the Principle for Long Video Understanding LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-16T11:14:19.651623Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:14:19.651623Z digest=sha256:cc8f2038409d3746d20145e482dc2bb167cca67ab02c37bd2ba9ff43f4432cd7

Observation e07fd8e3-a97e-40e5-9e84-3ad949312a3f · outbound

This paper cites BLIP: Bootstrapping language-image pre-training for uni- fied vision-language understanding and generation.

MR. Video: "MapReduce" is the Principle for Long Video Understanding BLIP: Bootstrapping language-image pre-training for uni- fied vision-language understanding and generation

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-16T11:14:19.656024Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:14:19.656024Z digest=sha256:e662ce60a1f7697dc0144ff71c2c14495a68926cfd8bdce46f63c112487811cf

Observation 2f325a6f-7e85-4d76-bd63-4401f0c36c38 · outbound

This paper cites BLIP-2: Bootstrapping language-image pre-training with frozen image encoders and large language models.

MR. Video: "MapReduce" is the Principle for Long Video Understanding BLIP-2: Bootstrapping language-image pre-training with frozen image encoders and large language models

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-16T11:14:19.660349Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:14:19.660349Z digest=sha256:66a783214b718e6b3d6e2533b0255c06c4d7494e45e529c5d035f766a302f5be

Observation e14c2bd4-2ab7-4d13-8ad6-b9816ba29287 · outbound

This paper cites VideoChat: Chat-Centric Video Understanding.

MR. Video: "MapReduce" is the Principle for Long Video Understanding VideoChat: Chat-Centric Video Understanding

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-16T11:14:19.664419Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:14:19.664419Z digest=sha256:6a11616d850a48e5eb7eab469fa8a5f06ef35a848fc47b37148ad4fa8f2d3857

Observation a09e42ab-705b-4d78-91d8-6d93e72fbbaa · outbound

This paper cites Temporal Preference Optimization for Long-Form Video Understanding.

MR. Video: "MapReduce" is the Principle for Long Video Understanding Temporal Preference Optimization for Long-Form Video Understanding

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-16T11:14:19.668886Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:14:19.668886Z digest=sha256:d9a1388009e60eea48dd42006b6e4f99698312397195fd06b3bcd55478bd1947

Observation 565c56fc-a2a4-4cb9-9369-3c2a3140fe2e · outbound

This paper cites VideoChat-Flash: Hierarchical Compression for Long-Context Video Modeling.

MR. Video: "MapReduce" is the Principle for Long Video Understanding VideoChat-Flash: Hierarchical Compression for Long-Context Video Modeling

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-16T11:14:19.673130Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:14:19.673130Z digest=sha256:2ebc6e75a6dde94b0bd11c32fb62a2a167a85eeab0891d5a3b1476df1ba20d92

Observation 3ff42e3f-7bae-4540-8ace-e38ba110efef · outbound

This paper cites Llama-VID: An image is worth 2 tokens in large language models.

MR. Video: "MapReduce" is the Principle for Long Video Understanding Llama-VID: An image is worth 2 tokens in large language models

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-16T11:14:19.677426Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:14:19.677426Z digest=sha256:dbbd388fa1b2da24953158a810def2877237cb2e642d421392c11d45b650cc64

Observation f5bb5563-1585-412d-823c-6fa687f70f84 · outbound

This paper cites Video-LLaV A: Learning united visual rep- resentation by alignment before projection.

MR. Video: "MapReduce" is the Principle for Long Video Understanding Video-LLaV A: Learning united visual rep- resentation by alignment before projection

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-16T11:14:19.681805Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:14:19.681805Z digest=sha256:0b275c680391b7fda12402eaacc73b5efc87f9df5a77480e763cbdc3a80fc783

Observation f71adb37-4c22-44ea-8147-9622b7c3fa46 · outbound

This paper cites MM-Embed: Universal multimodal retrieval with multimodal LLMs.

MR. Video: "MapReduce" is the Principle for Long Video Understanding MM-Embed: Universal multimodal retrieval with multimodal LLMs

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-16T11:14:19.686041Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:14:19.686041Z digest=sha256:4ed1dbaf67e9e3f768f2e0f9e22318b07b572fd5aa28124c6c3a4750b7f1c492

Observation fd158bb0-d796-461a-bc07-2bf36b37103a · outbound

This paper cites Improved Baselines with Visual Instruction Tuning.

MR. Video: "MapReduce" is the Principle for Long Video Understanding Improved Baselines with Visual Instruction Tuning

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-16T11:14:19.690520Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:14:19.690520Z digest=sha256:b05a81d47176b2b00973da72c58efbd2df08a8b7f4bca0ced4991ede0cae11cc

Observation 4895218d-029d-4923-9f96-4481941b6cf2 · outbound

This paper cites Visual instruction tuning.

MR. Video: "MapReduce" is the Principle for Long Video Understanding Visual instruction tuning

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-16T11:14:19.695119Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:14:19.695119Z digest=sha256:d02c0a591a5f43d51d5147df26f858d70fc555ab04af6f4fd04635e1d27a4f54

Observation 36346025-23e5-4f5d-be8c-d2f61fe60eed · outbound

This paper cites LLaV A-NeXT: Im- proved reasoning, ocr, and world knowledge, 2024.

MR. Video: "MapReduce" is the Principle for Long Video Understanding LLaV A-NeXT: Im- proved reasoning, ocr, and world knowledge, 2024

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-16T11:14:19.699628Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:14:19.699628Z digest=sha256:dc49629851405edfb1ea758a2f82e6229a71658ca5101e03c367fb083e647491

Observation b5c13e66-7bb2-4e3e-b177-f9af96ab17eb · outbound

This paper cites NVILA: Efficient Frontier Visual Language Models.

MR. Video: "MapReduce" is the Principle for Long Video Understanding NVILA: Efficient Frontier Visual Language Models

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-16T11:14:19.703851Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:14:19.703851Z digest=sha256:ddb0362f675809f9dfb9be6d329a207617ec5c5e7cf052ffa34e088be4310d95

Observation 06dcb4a4-afc2-4527-9f21-f89c9179bd94 · outbound

This paper cites Oryx MLLM: On-demand spatial- temporal understanding at arbitrary resolution.

MR. Video: "MapReduce" is the Principle for Long Video Understanding Oryx MLLM: On-demand spatial- temporal understanding at arbitrary resolution

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-16T11:14:19.708339Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:14:19.708339Z digest=sha256:02ca177cb1d21b0b859e6e1fcf79353b2e98a347c5afd0b3ba484fd4f482b327

Observation b4a5e991-59a3-493e-9e92-685b3f934419 · outbound

This paper cites EgoSchema: A diagnostic benchmark for very long- form video language understanding.

MR. Video: "MapReduce" is the Principle for Long Video Understanding EgoSchema: A diagnostic benchmark for very long- form video language understanding

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-16T11:14:19.713041Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:14:19.713041Z digest=sha256:99aa0041886c5cbeca1e1e9b8f76387f19a992dc80bba756ab74df030a3120c5

Observation 330ec540-4a94-4708-b257-b0d74d2a36fb · outbound

This paper cites TimeChat: A time-sensitive multimodal large language model for long video understanding.

MR. Video: "MapReduce" is the Principle for Long Video Understanding TimeChat: A time-sensitive multimodal large language model for long video understanding

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-16T11:14:19.717062Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:14:19.717062Z digest=sha256:8c7dd4b2244ded25cdd8736737b02e0d93618bcc92d05633168c98639a78a942

Observation 872ad005-aa6b-4ccb-90a8-45ff89ab1737 · outbound

This paper cites LLaV A-PruMerge: Adaptive token reduc- tion for efficient large multimodal models.

MR. Video: "MapReduce" is the Principle for Long Video Understanding LLaV A-PruMerge: Adaptive token reduc- tion for efficient large multimodal models

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-16T11:14:19.721193Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:14:19.721193Z digest=sha256:4690c7c3918cded6e8485dfe7c429bb3b4690a6bbc782b6b23979f2bf08e6d52

Observation cee9a261-f156-4da4-b1e5-65296ba20d83 · outbound

This paper cites LongVU: Spatiotemporal Adaptive Compression for Long Video-Language Understanding.

MR. Video: "MapReduce" is the Principle for Long Video Understanding LongVU: Spatiotemporal Adaptive Compression for Long Video-Language Understanding

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-16T11:14:19.725198Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:14:19.725198Z digest=sha256:e0906a73de117edcf96707b6707478fa26b9e51dfa7edabc017979088fa357e0

Observation 92cd8083-6679-4c6c-89a6-bc79d22f6b21 · outbound

This paper cites MovieChat: From dense token to sparse memory for long video understanding.

MR. Video: "MapReduce" is the Principle for Long Video Understanding MovieChat: From dense token to sparse memory for long video understanding

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-16T11:14:19.729858Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:14:19.729858Z digest=sha256:caa5bb327e716319f48e4a681ddb04979509cc9e097cf481689dd2dd9f8a73db

Observation 3d39cd2f-a0be-4c6f-a916-3c842c5be408 · outbound

This paper cites ViperGPT: Visual inference via python execution for reasoning.

MR. Video: "MapReduce" is the Principle for Long Video Understanding ViperGPT: Visual inference via python execution for reasoning

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-16T11:14:19.734099Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:14:19.734099Z digest=sha256:0081b129fa00c372129f14222b411deb2b51beef0952dda1ec7d5e63c5d3f995

Observation fe553232-1f03-471e-b811-0b0615b95bf9 · outbound

This paper cites Gemini: A Family of Highly Capable Multimodal Models.

MR. Video: "MapReduce" is the Principle for Long Video Understanding Gemini: A Family of Highly Capable Multimodal Models

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-16T11:14:19.738048Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:14:19.738048Z digest=sha256:f9b4c5457559c54553945ee2536d9f2bb47a4748caa057f9ed967686fca76940

Observation e46b2d20-0deb-4627-8f67-74315e2db408 · outbound

This paper cites Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution.

MR. Video: "MapReduce" is the Principle for Long Video Understanding Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-16T11:14:19.742286Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:14:19.742286Z digest=sha256:7bc6ee4ce9c09d9d9a999726371152ef404631e9923448688ba47a5689681b0f

Observation a7d04396-d7b3-44ea-892d-a20693979b5b · outbound

This paper cites LVBench: An Extreme Long Video Understanding Benchmark.

MR. Video: "MapReduce" is the Principle for Long Video Understanding LVBench: An Extreme Long Video Understanding Benchmark

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-16T11:14:19.746574Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:14:19.746574Z digest=sha256:ae4a1345e51cd98babeb83ccba154dc4a6e9313affe9ee44d02b9581ced61d36

Observation 8298e034-14e3-4253-b1f1-3ba82ca623d7 · outbound

This paper cites ReTaKe: Reducing Temporal and Knowledge Redundancy for Long Video Understanding.

MR. Video: "MapReduce" is the Principle for Long Video Understanding ReTaKe: Reducing Temporal and Knowledge Redundancy for Long Video Understanding

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-16T11:14:19.751131Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:14:19.751131Z digest=sha256:668d1a8b2692804a8364783c72401afd01b8c0f6d53dff6390e620da401d8f5f

Observation 0cfc33df-fa4d-4996-83c1-34b9513aa4f0 · outbound

This paper cites LongLLaV A: Scaling multi-modal LLMs to 1000 images efficiently via a hybrid architecture.

MR. Video: "MapReduce" is the Principle for Long Video Understanding LongLLaV A: Scaling multi-modal LLMs to 1000 images efficiently via a hybrid architecture

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-16T11:14:19.755983Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:14:19.755983Z digest=sha256:0fa98215c9e1947fd72eda89f7b7e766ca00ea660ce927bf49eeff14624bbb54

Observation d300838f-fb52-4546-a345-221d53434497 · outbound

This paper cites VideoAgent: Long-form video understanding with large language model as agent.

MR. Video: "MapReduce" is the Principle for Long Video Understanding VideoAgent: Long-form video understanding with large language model as agent

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-16T11:14:19.760324Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:14:19.760324Z digest=sha256:05120ac58ab3b94949309ccc91a73abe848bde72cf7d488636917a059b66207a

Observation bd4a64fa-0b05-4b3d-b302-095a289f5579 · outbound

This paper cites InternVideo2.5: Empowering Video MLLMs with Long and Rich Context Modeling.

MR. Video: "MapReduce" is the Principle for Long Video Understanding InternVideo2.5: Empowering Video MLLMs with Long and Rich Context Modeling

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-16T11:14:19.765679Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:14:19.765679Z digest=sha256:92910252f83ec606d0dded011ce122a8ba843c013969054015ecdafdf217e4dd

Observation b40d3b04-17b8-4128-930d-09555e3d088c · outbound

This paper cites VideoTree: Adaptive Tree-based Video Representation for LLM Reasoning on Long Videos.

MR. Video: "MapReduce" is the Principle for Long Video Understanding VideoTree: Adaptive Tree-based Video Representation for LLM Reasoning on Long Videos

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-16T11:14:19.770287Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:14:19.770287Z digest=sha256:5215ac95eca9ebe4c11a15fcb2d5f794997a654d58801b66cf494c5f5e404282

Observation 32a72efd-32a9-4b5e-a15c-c514fbb23d26 · outbound

This paper cites Chain-of-thought prompting elicits reasoning in large lan- guage models.

MR. Video: "MapReduce" is the Principle for Long Video Understanding Chain-of-thought prompting elicits reasoning in large lan- guage models

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-16T11:14:19.774505Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:14:19.774505Z digest=sha256:eb3266e3a7d264b078cdb794abbc1cae4d7e1c426282ebfbf0d4afb1804bbc66

Observation 0881e717-e72e-4fc6-a042-58c45af04b66 · outbound

This paper cites LongVLM: Efficient long video understand- ing via large language models.

MR. Video: "MapReduce" is the Principle for Long Video Understanding LongVLM: Efficient long video understand- ing via large language models

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-16T11:14:19.779557Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:14:19.779557Z digest=sha256:de9cd6a47d1eb08bfe49f73d54995f1861e6da730e1259cddb198f56e9994b02

Observation bb3598c7-5a5b-45be-b697-da995de695ae · outbound

This paper cites Longvideobench: A benchmark for long-context interleaved video-language understanding.

MR. Video: "MapReduce" is the Principle for Long Video Understanding Longvideobench: A benchmark for long-context interleaved video-language understanding

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-16T11:14:19.783702Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:14:19.783702Z digest=sha256:9d0af55d41de003a4f33e033d4f54412a66716b36738ce05ab20b37a4d1d31e8

Observation 54199bcc-917c-4bfa-8915-248812d00afc · outbound

This paper cites Pyramiddrop: Accelerating your large vision-language models via pyramid visual redundancy re- duction.

MR. Video: "MapReduce" is the Principle for Long Video Understanding Pyramiddrop: Accelerating your large vision-language models via pyramid visual redundancy re- duction

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-16T11:14:19.787961Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:14:19.787961Z digest=sha256:ae274819cc7c785770f5754b37244c96577913460533388ae1ab89a7cd19a85c

Observation 33f83d22-0245-48ea-86b1-1551a5849916 · outbound

This paper cites SlowFast-LLaVA: A Strong Training-Free Baseline for Video Large Language Models.

MR. Video: "MapReduce" is the Principle for Long Video Understanding SlowFast-LLaVA: A Strong Training-Free Baseline for Video Large Language Models

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-16T11:14:19.792206Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:14:19.792206Z digest=sha256:2ce4298a188e32eba78761663d195e1af73e81ffeb26cbacaba6aae42cdc3262

Observation 4de1b057-12c3-4f1e-b9a5-9a0ebe6bb377 · outbound

This paper cites LongVILA: Scaling Long-Context Visual Language Models for Long Videos.

MR. Video: "MapReduce" is the Principle for Long Video Understanding LongVILA: Scaling Long-Context Visual Language Models for Long Videos

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-16T11:14:19.796447Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:14:19.796447Z digest=sha256:88d735d7f09b87fe166980d0ebbbc13bd300e8bc0f473dd09dedaa787c6ddfbe

Observation 6d766a48-3598-4139-954b-14a64ca56cf3 · outbound

This paper cites Qwen2.5 Technical Report.

MR. Video: "MapReduce" is the Principle for Long Video Understanding Qwen2.5 Technical Report

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-16T11:14:19.800900Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:14:19.800900Z digest=sha256:b65a9a06e2b83cd4176427fcf68f2d4fe4a6d6bb3c4ea7e51ec97d5a1955b53c

Observation f7212029-79a4-4d1f-89ae-bb0c8a72fb50 · outbound

This paper cites SWE- Agent: Agent-computer interfaces enable automated soft- ware engineering.

MR. Video: "MapReduce" is the Principle for Long Video Understanding SWE- Agent: Agent-computer interfaces enable automated soft- ware engineering

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-16T11:14:19.805192Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:14:19.805192Z digest=sha256:1f906d14e48b57ac721a9e3f924abe6a620bc1b47a35debd920acb9d047f655d

Observation 91b4fb99-caa2-4321-91bc-fdaa74ab55fd · outbound

This paper cites HotpotQA: A Dataset for Diverse, Explainable Multi-hop Question Answering.

MR. Video: "MapReduce" is the Principle for Long Video Understanding HotpotQA: A Dataset for Diverse, Explainable Multi-hop Question Answering

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-16T11:14:19.809225Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:14:19.809225Z digest=sha256:9f46a7953a95bde106cf326bba0d919cd27b9aa1a6514b6fc5e08cd1c220e032

Observation c15a3738-63ea-440d-8e9f-9a3f1d1c60da · outbound

This paper cites VCA: Video Curious Agent for Long Video Understanding.

MR. Video: "MapReduce" is the Principle for Long Video Understanding VCA: Video Curious Agent for Long Video Understanding

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-16T11:14:19.813990Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:14:19.813990Z digest=sha256:5ba98535b7ecb155aa4ed6e588dcb514915e2eec19e16950018bddf5f08798fb

Observation 4c7e8a2a-baa2-477f-b0a8-5ed4ca5d3938 · outbound

This paper cites ReAct: Synergizing reasoning and acting in language models.

MR. Video: "MapReduce" is the Principle for Long Video Understanding ReAct: Synergizing reasoning and acting in language models

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-16T11:14:19.818233Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:14:19.818233Z digest=sha256:486962dd6845fd03ca8d93a276d19812235f4149feceb6517f17d6626d2462e6

Observation d275e97a-2fce-4806-9c91-5f9d6900bd55 · outbound

This paper cites mPLUG- OWL3: Towards long image-sequence understanding in multi-modal large language models.

MR. Video: "MapReduce" is the Principle for Long Video Understanding mPLUG- OWL3: Towards long image-sequence understanding in multi-modal large language models

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-16T11:14:19.822257Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:14:19.822257Z digest=sha256:7ed78fd45a1e7a304d98debe2fc3fff1abc917e45ab583f04d2742dd756debd3

Observation 59742282-5e13-4c3b-ab04-f50be511f0ff · outbound

This paper cites VideoLLaMA 3: Frontier Multimodal Foundation Models for Image and Video Understanding.

MR. Video: "MapReduce" is the Principle for Long Video Understanding VideoLLaMA 3: Frontier Multimodal Foundation Models for Image and Video Understanding

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-16T11:14:19.830458Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:14:19.830458Z digest=sha256:7638c219332d60827dbde3fe82c4d6736adb28b899767b0456c23e195d7a4914

Observation 37451518-0140-4205-bde1-471fbf9a7401 · outbound

This paper cites A simple LLM framework for long-range video question-answering.

MR. Video: "MapReduce" is the Principle for Long Video Understanding A simple LLM framework for long-range video question-answering

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-16T11:14:19.834901Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:14:19.834901Z digest=sha256:297f14816f310c00de9690702616a6798f898256a71731888dd9a741c9d5e645

Observation 88f4134a-6376-4a23-84e8-649c40e784c7 · outbound

This paper cites LLaV A-NeXT: A strong zero-shot video understanding model, 2024.

MR. Video: "MapReduce" is the Principle for Long Video Understanding LLaV A-NeXT: A strong zero-shot video understanding model, 2024

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-16T11:14:19.839194Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:14:19.839194Z digest=sha256:d1dd16147d457242a8707acb7b9f6f61ebe20e8c02773521c71fe90d3ba73528

Observation 92d9400d-f343-4c4c-b24e-a6fb2824f3c0 · outbound

This paper cites Language agent tree search unifies reasoning acting and planning in language models.

MR. Video: "MapReduce" is the Principle for Long Video Understanding Language agent tree search unifies reasoning acting and planning in language models

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-16T11:14:19.843436Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:14:19.843436Z digest=sha256:1d28c847ff129e3bca769d2321f7894bb10087fbed8778ece34621afb6e9bd65

Observation 21db97ff-f3fb-41ae-8e51-a1146d116a98 · outbound

This paper cites MLVU: Benchmarking Multi-task Long Video Understanding.

MR. Video: "MapReduce" is the Principle for Long Video Understanding MLVU: Benchmarking Multi-task Long Video Understanding

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-16T11:14:19.848274Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:14:19.848274Z digest=sha256:b9708ed090f41b9d063f30c12f9a5bb65daa5e8b563c647e6a96c5a7f5e9711e

Observation 05b07131-905d-4b5f-9ab1-88579892eaed · outbound

This paper cites Apollo: An Exploration of Video Understanding in Large Multimodal Models.

MR. Video: "MapReduce" is the Principle for Long Video Understanding Apollo: An Exploration of Video Understanding in Large Multimodal Models

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-16T11:14:19.852868Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:14:19.852868Z digest=sha256:ea2b99421f29e8ea14f9c7b8ceb91e0890788a3ce40c28f14411747d3e2c2e79

Observation cd6a3520-631c-423d-b5a4-ca589a53efa3 · outbound

This paper cites Scene Merging.

MR. Video: "MapReduce" is the Principle for Long Video Understanding Scene Merging

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-16T11:14:19.858786Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:14:19.858786Z digest=sha256:b20dc760188834ebd76a2b7b7e850a4b7a845be5901aa6cdc7171c921a8c7d10

Observation acc6f42e-bf58-43e1-9831-85da297a9585 · outbound

This paper cites The prompts are in Table C.

MR. Video: "MapReduce" is the Principle for Long Video Understanding The prompts are in Table C

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-16T11:14:19.863370Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:14:19.863370Z digest=sha256:6ec7e40b663debf670db7ee0724f2f54050e74b17e69854cb94cfb0f8da3eed5

Observation ec5cd9dd-512c-40ed-8ddb-7160f7bee4dd · outbound

This paper cites Reduce: Consistent Characters and Objects.

MR. Video: "MapReduce" is the Principle for Long Video Understanding Reduce: Consistent Characters and Objects

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-16T11:14:19.867509Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:14:19.867509Z digest=sha256:0abac24b17e1f8528bbadfbd6f9e9ee4434c7b5e2bb7f6c0ab7db711c9c47696

Observation 85fbb2ac-147b-4767-bc6b-551b92346321 · outbound

This paper cites an unresolved cited work.

MR. Video: "MapReduce" is the Principle for Long Video Understanding Unresolved cited work

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-16T11:14:19.871945Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:14:19.871945Z digest=sha256:2ed09c0b0c8c122a9c7d303be027c2bef3014f254c67903e59010abc56c1d65a

Observation b12c936a-2ff6-44c2-ba6b-bd61fb2a6bae · outbound

This paper cites an unresolved cited work.

MR. Video: "MapReduce" is the Principle for Long Video Understanding Unresolved cited work

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-16T11:14:19.876430Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:14:19.876430Z digest=sha256:13f9c07098cb8f2f06857ae8681be2b33210b8bbadfe0407874f92e0542d3834

Observation d3751bd4-06f4-4bb4-899f-8af4c29b9153 · outbound

This paper cites [1. Description]:.

MR. Video: "MapReduce" is the Principle for Long Video Understanding [1. Description]:

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-16T11:14:19.881006Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:14:19.881006Z digest=sha256:e3aa027ea1504ab8a5b9462300f8e86544648eb22ea7e6a89c4a7f89b7fb90cd

Observation 69c2a5c7-0be9-4c59-b1e2-f057e589a776 · outbound

This paper cites Is this video segment a single scene or a combination of multiple scenes?.

MR. Video: "MapReduce" is the Principle for Long Video Understanding Is this video segment a single scene or a combination of multiple scenes?

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-16T11:14:19.886414Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:14:19.886414Z digest=sha256:13d3ff5e6d074f06a50211018ab35217c6e84d91af85152811fb6269040e0130

Observation 01e668f1-f9ef-4662-ab1a-1455fe5c613c · outbound

This paper cites no", please provide the index of frame(s) separating the scenes from the given frame. Your answer should come with a header:.

MR. Video: "MapReduce" is the Principle for Long Video Understanding no", please provide the index of frame(s) separating the scenes from the given frame. Your answer should come with a header:

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-16T11:14:19.890835Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:14:19.890835Z digest=sha256:8b4f79b8b86b620f8c4f70b47ec26a9c9599eafa243f02f3c463bd89e3133760

Observation 437d2a4f-b3e6-4824-acdf-5148dc9a9b1b · outbound

This paper cites Try to be rigorous and faithful to the video without making assumptions.

MR. Video: "MapReduce" is the Principle for Long Video Understanding Try to be rigorous and faithful to the video without making assumptions

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-16T11:14:19.896280Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:14:19.896280Z digest=sha256:3ae5fbd499c070cded5d0d54513022048306194edb436bc75fafbd9b264fb3f2

Observation e9728f37-3494-44b6-8c4c-9ceb3022b2b4 · outbound

This paper cites an unresolved cited work.

MR. Video: "MapReduce" is the Principle for Long Video Understanding Unresolved cited work

Reference 72

Resolution
unresolved
no resolver link, observed 2026-08-16T11:14:19.900527Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:14:19.900527Z digest=sha256:81407ccec8d522dc2e9387c382945d5c6f46d54472778c4d3f531da71caface5

Observation 0819d016-f84c-4d7f-a745-1c11ca072cac · outbound

This paper cites person a.

MR. Video: "MapReduce" is the Principle for Long Video Understanding person a

Reference 73

Resolution
unresolved
no resolver link, observed 2026-08-16T11:14:19.904614Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:14:19.904614Z digest=sha256:74ed6b03f6226f2fbdd56843a868e9a66ec1369d415af16c7e5321ebcd7be36f

Observation ab82b672-d59c-4cd2-ae9c-49c053dec7d4 · outbound

This paper cites an unresolved cited work.

MR. Video: "MapReduce" is the Principle for Long Video Understanding Unresolved cited work

Reference 74

Resolution
unresolved
no resolver link, observed 2026-08-16T11:14:19.909673Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:14:19.909673Z digest=sha256:931a05e88fbece29c2e196567f6e0def5e652180dbccd96ec2595759b3b096cc

Observation a17f75b7-cc26-4323-8a35-4a5238b16621 · outbound

This paper cites It could be a person in the movie, an animal in the documentary or cartoon, etc.

MR. Video: "MapReduce" is the Principle for Long Video Understanding It could be a person in the movie, an animal in the documentary or cartoon, etc

Reference 75

Resolution
unresolved
no resolver link, observed 2026-08-16T11:14:19.913885Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:14:19.913885Z digest=sha256:c38449c3181ebfb9289ab68fb3a787b19a46d6a0b600ff6c2cc3b29c44415c5d

Observation 4ae66a10-2ecf-40f7-af46-3009069ff4e3 · outbound

This paper cites Please keep the strings in identical formattings to ensure smooth post-processing.

MR. Video: "MapReduce" is the Principle for Long Video Understanding Please keep the strings in identical formattings to ensure smooth post-processing

Reference 76

Resolution
unresolved
no resolver link, observed 2026-08-16T11:14:19.918285Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:14:19.918285Z digest=sha256:35075e1d83e23ee49a4368d961ad9ee8d40b3d31a5e84f1219351273dfc4cb6e

Observation e826a54e-340a-4a7d-b9e5-1550d71b75ec · outbound

This paper cites Character Selection.

MR. Video: "MapReduce" is the Principle for Long Video Understanding Character Selection

Reference 77

Resolution
unresolved
no resolver link, observed 2026-08-16T11:14:19.922457Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:14:19.922457Z digest=sha256:c191878f70d4d64a3c204b3343ffb772bc938752669eb2ae30df586ce36aee56

Observation ccebd5ff-889b-4e5a-816b-12db6a89b9fe · outbound

This paper cites an unresolved cited work.

MR. Video: "MapReduce" is the Principle for Long Video Understanding Unresolved cited work

Reference 78

Resolution
unresolved
no resolver link, observed 2026-08-16T11:14:19.926656Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:14:19.926656Z digest=sha256:cd7d12e6cf679f3d3ba3b1f6256f8472030cd0fc90fc7689d02ac9000581a248

Observation bb6e644d-be11-4cd5-b166-f08e1fac7d17 · outbound

This paper cites an unresolved cited work.

MR. Video: "MapReduce" is the Principle for Long Video Understanding Unresolved cited work

Reference 79

Resolution
unresolved
no resolver link, observed 2026-08-16T11:14:19.930670Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:14:19.930670Z digest=sha256:ac1bc9263c691871463a106c3862dfc99f9fe5b778d3d13a3eff47348362b652

Observation 15af2956-cbdb-48e6-a04e-239ce4a1ba3f · outbound

This paper cites an unresolved cited work.

MR. Video: "MapReduce" is the Principle for Long Video Understanding Unresolved cited work

Reference 80

Resolution
unresolved
no resolver link, observed 2026-08-16T11:14:19.934742Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:14:19.934742Z digest=sha256:9c5455ba99e9e30484fc3096ff22a151cc1391bb60b4b6dd864bb233bbbbe909

Observation a77dbbc9-6a2b-4899-8c84-582a5aec5d53 · outbound

This paper cites an unresolved cited work.

MR. Video: "MapReduce" is the Principle for Long Video Understanding Unresolved cited work

Reference 81

Resolution
unresolved
no resolver link, observed 2026-08-16T11:14:19.938858Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:14:19.938858Z digest=sha256:882bae788f668aa5da98cd0facf99ec8efbb0caad741503c67baf4888f27f318

Observation 723026b6-4e2b-4be6-91f4-a96f3cd73642 · outbound

This paper cites If so, please list their name out.

MR. Video: "MapReduce" is the Principle for Long Video Understanding If so, please list their name out

Reference 82

Resolution
unresolved
no resolver link, observed 2026-08-16T11:14:19.943381Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:14:19.943381Z digest=sha256:055b5f8452bb3bdcbf04565c1812c020e0fd3b6fe30c09d484158c3a8f465d21

Observation d186c8d0-9601-459f-a9be-f017ba50f10d · outbound

This paper cites Some more detailed tips:.

MR. Video: "MapReduce" is the Principle for Long Video Understanding Some more detailed tips:

Reference 83

Resolution
unresolved
no resolver link, observed 2026-08-16T11:14:19.947634Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:14:19.947634Z digest=sha256:7277f057d4e651a9b43c5d574e01bd5afc2463372424c1aeecbae25bf0b6961c

Observation 8f4da4f3-e0c3-4159-a9ce-f1c7afe602e9 · outbound

This paper cites The goal is that a human should read your captions and feel like watching a continuous video.

MR. Video: "MapReduce" is the Principle for Long Video Understanding The goal is that a human should read your captions and feel like watching a continuous video

Reference 84

Resolution
unresolved
no resolver link, observed 2026-08-16T11:14:19.952037Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:14:19.952037Z digest=sha256:49b66427a81de32aa866c0c9c029224f6dd7d27178e67620a97eb5fb2c02e67c

Observation 2942e76c-f48d-4c7a-9328-80d8ac7d220a · outbound

This paper cites an unresolved cited work.

MR. Video: "MapReduce" is the Principle for Long Video Understanding Unresolved cited work

Reference 85

Resolution
unresolved
no resolver link, observed 2026-08-16T11:14:19.956100Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:14:19.956100Z digest=sha256:ca1d756ed9f5e095169d0552d726ac27f5be90c5f84cfefe152976b57b72034b

Observation dd50afd3-36df-460d-8dc5-482d90cfeb37 · outbound

This paper cites person a.

MR. Video: "MapReduce" is the Principle for Long Video Understanding person a

Reference 86

Resolution
unresolved
no resolver link, observed 2026-08-16T11:14:19.960533Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:14:19.960533Z digest=sha256:16a75de5a3fa3e50a08d63aa8e07be6f859f5aba2119a07c9abbfcbc03173b76

Observation e3506fac-bf4a-4e1a-aa43-6455aa342b87 · outbound

This paper cites an unresolved cited work.

MR. Video: "MapReduce" is the Principle for Long Video Understanding Unresolved cited work

Reference 87

Resolution
unresolved
no resolver link, observed 2026-08-16T11:14:19.965308Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:14:19.965308Z digest=sha256:0744154d14a7b1ae90234c516c077eda64d2e6910b3bf17987d28b69d02f2833

Observation 65767748-ecc0-41c8-860e-0b67ea170fee · outbound

This paper cites Dense Captioning.

MR. Video: "MapReduce" is the Principle for Long Video Understanding Dense Captioning

Reference 88

Resolution
unresolved
no resolver link, observed 2026-08-16T11:14:19.970068Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:14:19.970068Z digest=sha256:6aac6dfe5cf7957e862644d8d9e8701bf95f4ca1346d1587706875a172b821c1

Observation 8b58c023-235f-4897-aa09-9aec933aca4f · outbound

This paper cites an unresolved cited work.

MR. Video: "MapReduce" is the Principle for Long Video Understanding Unresolved cited work

Reference 89

Resolution
unresolved
no resolver link, observed 2026-08-16T11:14:19.974562Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:14:19.974562Z digest=sha256:f5fc4aac61e176fd8ad8e0d0dd03fd55d324aedc20bd558782a130cb5210dd57

Observation 547a60e2-4328-4ee2-a6e0-18accb8c50e1 · outbound

This paper cites an unresolved cited work.

MR. Video: "MapReduce" is the Principle for Long Video Understanding Unresolved cited work

Reference 90

Resolution
unresolved
no resolver link, observed 2026-08-16T11:14:19.979218Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:14:19.979218Z digest=sha256:9ca1668c92382085b2aa43e5239ff9aeb50d63136940bf720a1ec8e799d07773

Observation 2fb64f7b-888f-4a57-9211-2df026864d44 · outbound

This paper cites an unresolved cited work.

MR. Video: "MapReduce" is the Principle for Long Video Understanding Unresolved cited work

Reference 91

Resolution
unresolved
no resolver link, observed 2026-08-16T11:14:19.983545Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:14:19.983545Z digest=sha256:97b2674062809d4f39d920147e91d7590feff2d466a98f72e4df2a4d1089696b

Observation e0f22290-9772-4c17-b022-002571c640c4 · outbound

This paper cites an unresolved cited work.

MR. Video: "MapReduce" is the Principle for Long Video Understanding Unresolved cited work

Reference 92

Resolution
unresolved
no resolver link, observed 2026-08-16T11:14:19.988423Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:14:19.988423Z digest=sha256:79822e4b97d5f89d85b33d11b38f2044646d34769820e4d7e96d6e2a5f9929b3

Observation b2c06448-54ba-4e7e-bd37-caa872a49c2a · outbound

This paper cites an unresolved cited work.

MR. Video: "MapReduce" is the Principle for Long Video Understanding Unresolved cited work

Reference 93

Resolution
unresolved
no resolver link, observed 2026-08-16T11:14:19.992841Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:14:19.992841Z digest=sha256:3c94f7c7b9c35683ce330e04f27c15a634baa94b612f01e5698fe9d7c8e3eea2

Observation 542db1c6-6e2c-408a-917f-609eda5eef0a · outbound

This paper cites an unresolved cited work.

MR. Video: "MapReduce" is the Principle for Long Video Understanding Unresolved cited work

Reference 94

Resolution
unresolved
no resolver link, observed 2026-08-16T11:14:19.996813Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:14:19.996813Z digest=sha256:e1cc4ddd4d2dd276639a5f5f40d4181c3e6251010f46dcb1bf14048258827be8

Observation fda50274-fd07-45cb-abc7-4f13e3f645e9 · outbound

This paper cites an unresolved cited work.

MR. Video: "MapReduce" is the Principle for Long Video Understanding Unresolved cited work

Reference 95

Resolution
unresolved
no resolver link, observed 2026-08-16T11:14:20.000653Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:14:20.000653Z digest=sha256:9cf4be76b1a8fc3275c8f2c2929f7f4671425a0cb85a469a96fca729ad2404e3

Observation 3204551e-a098-4002-8e83-83705bb670bf · outbound

This paper cites an unresolved cited work.

MR. Video: "MapReduce" is the Principle for Long Video Understanding Unresolved cited work

Reference 96

Resolution
unresolved
no resolver link, observed 2026-08-16T11:14:20.004859Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:14:20.004859Z digest=sha256:9b58ed167ce172d08424864f6d1e99f03aabce25fe751f0dbf5a3ab1bdc6b3b4

Observation b707b233-79b9-4ebc-ac22-0a85f4b11191 · outbound

This paper cites Character Merging.

MR. Video: "MapReduce" is the Principle for Long Video Understanding Character Merging

Reference 97

Resolution
unresolved
no resolver link, observed 2026-08-16T11:14:20.009168Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:14:20.009168Z digest=sha256:d815a9669a80664c1b950e952f83707782fb26a460a98dcab806ed1f4ef8d827

Observation 0bb0c0d4-f311-4a52-8945-a45db7c2ba4f · outbound

This paper cites an unresolved cited work.

MR. Video: "MapReduce" is the Principle for Long Video Understanding Unresolved cited work

Reference 98

Resolution
unresolved
no resolver link, observed 2026-08-16T11:14:20.013506Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:14:20.013506Z digest=sha256:3aa38cba86d36aae55c81ca91dc02753dc58a91c33d50092a7e22a3fbb7360bb

Observation 8cf94b18-4fd2-4e7d-a21e-95f7a5c1db5b · outbound

This paper cites Output: Your output should be the modified description of the video clip strictly following the original format and contents, only with names changed.

MR. Video: "MapReduce" is the Principle for Long Video Understanding Output: Your output should be the modified description of the video clip strictly following the original format and contents, only with names changed

Reference 99

Resolution
unresolved
no resolver link, observed 2026-08-16T11:14:20.018995Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:14:20.018995Z digest=sha256:95a786c67cecbddb4c9c59c8aa944374de5223c9a2dcc6bdcb6fbb37701a5c5f

Observation c69f0c2b-e5f5-4303-9048-0f631db62514 · outbound

This paper cites an unresolved cited work.

MR. Video: "MapReduce" is the Principle for Long Video Understanding Unresolved cited work

Reference 100

Resolution
unresolved
raw_fallback, observed 2026-08-16T11:14:21.526142Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-16T11:14:20.023097Z digest=sha256:bb445eb3b2095a06a63d1ffdc7bb0a4bb3bac4cc89c9f5840192210aa652715c

Observation 7dd82448-1b05-4a3f-b431-64901403ac32 · outbound

This paper cites an unresolved cited work.

MR. Video: "MapReduce" is the Principle for Long Video Understanding Unresolved cited work

Reference 101

Resolution
unresolved
raw_fallback, observed 2026-08-16T11:14:21.512320Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-16T11:14:20.027276Z digest=sha256:5defc9cf51110ed675a335205367fe5ec9bd413dc57380bc5d7abc200f78acec

Pith citing papers

Observation 9eaf1eeb-ce67-45d9-bd88-2d53b82834e9 · inbound

LVBench: An Extreme Long Video Understanding Benchmark cites this paper.

LVBench: An Extreme Long Video Understanding Benchmark MR. Video: "MapReduce" is the Principle for Long Video Understanding

Reference 28

Resolution
verified exact
arxiv_id, observed 2026-05-19T11:55:30.086813Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-19T11:55:30.048525Z digest=sha256:60e9b7c0e1cf21ab709e80b40b32e2626e54bd39f7094ac7aec19b7cb1947953

Observation e2a32a97-bee5-4846-89e4-383a7787ce4e · inbound

Graph-to-Frame RAG: Visual-Space Knowledge Fusion for Training-Free and Auditable Video Reasoning cites this paper.

Graph-to-Frame RAG: Visual-Space Knowledge Fusion for Training-Free and Auditable Video Reasoning MR. Video: "MapReduce" is the Principle for Long Video Understanding

Reference 40

Resolution
verified exact
arxiv_id, observed 2026-05-10T22:30:52.648369Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-10T19:46:16.975267Z digest=sha256:dae59234631a9bdd2f0e02b8dc7e6a5ead69d55f16e05d0f1724b192f601d9e9

Observation e7919ba8-883e-46b6-9af3-a94637f69d8b · inbound

Scaling Video Understanding via Compact Latent Multi-Agent Collaboration cites this paper.

Scaling Video Understanding via Compact Latent Multi-Agent Collaboration MR. Video: "MapReduce" is the Principle for Long Video Understanding

Reference 21

Resolution
verified exact
arxiv_id, observed 2026-05-11T15:21:09.395092Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-09T20:11:11.410051Z digest=sha256:efc1a4d03bd42b802a835a8d5948060e343945949383c693c6ad7092021256d6

Observation b020e007-054f-4fde-a454-cbb756619e1f · inbound

GOPAgen: Motion-Aware and Efficient Agentic Long-Video Understanding with Structural Memory and Hierarchical Reasoning cites this paper.

GOPAgen: Motion-Aware and Efficient Agentic Long-Video Understanding with Structural Memory and Hierarchical Reasoning MR. Video: "MapReduce" is the Principle for Long Video Understanding

Reference 63

Resolution
verified exact
arxiv_id, observed 2026-07-02T07:56:47.336859Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-06-28T06:33:32.090913Z digest=sha256:c50421e59592643cf086aee1376f87d57614d2472ae9637943fb3fb4d77a7b48