Pith. sign in

Paper Citation Record · LEDGER

Moment-Video: Diagnosing Temporal Fidelity of Video MLLMs on Momentary Visual Events

As of 5 August 2026, this Paper Citation Record lists 57 of 57 outbound references and 1 inbound Pith citation observation for arXiv:2606.02522.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2606.02522 v1

Coverage vector

measured 57 of 57 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-06-28T14:50:02.159411Z

measured 58 of 58 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-05T06:32:48.257954+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-07-31T04:55:07.020202Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

57 of 57 outbound references displayed

  • verified exact17
  • verified fuzzy0
  • unresolved36
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch3

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 7ce63d8e-d541-44ba-913d-fd91d8e87c7b · outbound

This paper cites Qwen3-VL Technical Report.

Moment-Video: Diagnosing Temporal Fidelity of Video MLLMs on Momentary Visual Events Qwen3-VL Technical Report

Reference 1

Resolution
verified exact
local_arxiv, observed 2026-07-01T22:56:20.847215Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-28T14:50:02.159411Z digest=sha256:20c25ee80686c3ac32e09a01bb1debbc020022798d5e5eb87cfc98a9a5ac8cdb

Observation 08a64477-c3f5-4ecf-9a72-182d6117dcbf · outbound

This paper cites Qwen3.5-Omni Technical Report.

Moment-Video: Diagnosing Temporal Fidelity of Video MLLMs on Momentary Visual Events Qwen3.5-Omni Technical Report

Reference 2

Resolution
verified exact
local_arxiv, observed 2026-07-01T22:56:20.866115Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-28T14:50:02.159411Z digest=sha256:d19575aa49249732d4ab7ae290e5bab3ed5299bdcd46c2ceb5835f85f38dd7c9

Observation 3fd0a824-3534-4ce6-89f4-1b5d0ab61079 · outbound

This paper cites Animal kingdom: A large and diverse dataset for animal behavior understanding,.

Moment-Video: Diagnosing Temporal Fidelity of Video MLLMs on Momentary Visual Events Animal kingdom: A large and diverse dataset for animal behavior understanding,

Reference 3

Resolution
unresolved
no resolver link, observed 2026-06-28T14:50:02.159411Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T14:50:02.159411Z digest=sha256:6706147593e91c904c44c440a87b0efe8878645330e59fa6d71d91f5011c3f0b

Observation 7e9837e0-fefb-4b2f-88ff-a5b6aadca96f · outbound

This paper cites VideoLLaMA 3: Frontier Multimodal Foundation Models for Image and Video Understanding.

Moment-Video: Diagnosing Temporal Fidelity of Video MLLMs on Momentary Visual Events VideoLLaMA 3: Frontier Multimodal Foundation Models for Image and Video Understanding

Reference 4

Resolution
verified exact
local_arxiv, observed 2026-07-01T22:56:20.887187Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-28T14:50:02.159411Z digest=sha256:6cb5cb2439ed8c62551b7421d9504302d264c27d8bc68fcc5b9b678e7cc523dc

Observation 9433cb44-0d39-483f-aac0-417b46fac94f · outbound

This paper cites GLM-4.5V and GLM-4.1V-Thinking: Towards Versatile Multimodal Reasoning with Scalable Reinforcement Learning.

Moment-Video: Diagnosing Temporal Fidelity of Video MLLMs on Momentary Visual Events GLM-4.5V and GLM-4.1V-Thinking: Towards Versatile Multimodal Reasoning with Scalable Reinforcement Learning

Reference 5

Resolution
verified exact
local_arxiv, observed 2026-07-01T22:56:20.855054Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-28T14:50:02.159411Z digest=sha256:5424693ad5c930a0ac0185ff9c708c57bc386c10119c6e18b9d43e8b335214cc

Observation 7821c330-90e8-4408-b07d-c5b520b05166 · outbound

This paper cites Kimi-VL Technical Report.

Moment-Video: Diagnosing Temporal Fidelity of Video MLLMs on Momentary Visual Events Kimi-VL Technical Report

Reference 6

Resolution
verified exact
local_arxiv, observed 2026-07-01T22:56:20.894624Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-28T14:50:02.159411Z digest=sha256:33c38a2663911910ad67bbea60e3e5c9538467f39711540cb167ccf8f72e1486

Observation 6f17675a-a9cf-4893-930c-d022a3088706 · outbound

This paper cites Tracknetv2: Efficient shuttlecock tracking network,.

Moment-Video: Diagnosing Temporal Fidelity of Video MLLMs on Momentary Visual Events Tracknetv2: Efficient shuttlecock tracking network,

Reference 7

Resolution
unresolved
no resolver link, observed 2026-06-28T14:50:02.159411Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T14:50:02.159411Z digest=sha256:5357e1be8374d2745a6ad70b8bc4d2bc19c6a506000abdddda0978d2528c20c3

Observation b2dc98d0-796a-4fb1-b97b-2f910ae92e20 · outbound

This paper cites UCF101: A Dataset of 101 Human Actions Classes From Videos in The Wild.

Moment-Video: Diagnosing Temporal Fidelity of Video MLLMs on Momentary Visual Events UCF101: A Dataset of 101 Human Actions Classes From Videos in The Wild

Reference 8

Resolution
verified exact
local_arxiv, observed 2026-07-01T22:56:20.903412Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-28T14:50:02.159411Z digest=sha256:9df9f85045a14b20209281a6bd2c880fc6ed59998966905dd8a937c5b471b1e4

Observation 5c1fa725-b0ec-4bb9-b9ee-95baa3e53b83 · outbound

This paper cites Sports videos in the wild (svw): A video dataset for sports analysis,.

Moment-Video: Diagnosing Temporal Fidelity of Video MLLMs on Momentary Visual Events Sports videos in the wild (svw): A video dataset for sports analysis,

Reference 9

Resolution
unresolved
no resolver link, observed 2026-06-28T14:50:02.159411Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T14:50:02.159411Z digest=sha256:e850043adad599bb4f90250f91323caf7742b3b45f4199d7245a6090e4ea50b1

Observation 8dab6a8e-c68c-4971-9bb6-76e5b54d2d1a · outbound

This paper cites Multisports: A multi-person video dataset of spatio- temporally localized sports actions,.

Moment-Video: Diagnosing Temporal Fidelity of Video MLLMs on Momentary Visual Events Multisports: A multi-person video dataset of spatio- temporally localized sports actions,

Reference 10

Resolution
unresolved
no resolver link, observed 2026-06-28T14:50:02.159411Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T14:50:02.159411Z digest=sha256:936acba27c2a4eb2c3963687888e4e29d04e17243a8b1568bbe9121911e0759f

Observation 01a5f906-7144-4358-8e1c-892fdf347538 · outbound

This paper cites LLaVA-Octopus: Unlocking Instruction-Driven Adaptive Projector Fusion for Video Understanding.

Moment-Video: Diagnosing Temporal Fidelity of Video MLLMs on Momentary Visual Events LLaVA-Octopus: Unlocking Instruction-Driven Adaptive Projector Fusion for Video Understanding

Reference 11

Resolution
verified exact
local_arxiv, observed 2026-07-01T22:56:20.851381Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-28T14:50:02.159411Z digest=sha256:af6ae4d5e8e3cb04249d54c3c973be016488722b06505c9995c40e96b39da7e1

Observation c311c578-fa1e-45a7-bad3-e242998ad98d · outbound

This paper cites Video-R1: Reinforcing Video Reasoning in MLLMs.

Moment-Video: Diagnosing Temporal Fidelity of Video MLLMs on Momentary Visual Events Video-R1: Reinforcing Video Reasoning in MLLMs

Reference 12

Resolution
verified exact
local_arxiv, observed 2026-07-01T22:56:20.862878Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-28T14:50:02.159411Z digest=sha256:1b84f2ede05b2de7f50b4436a23333b8df0f23d9c334a41ee819fddd62ba064e

Observation f36497b5-429f-4978-bb95-23677da03585 · outbound

This paper cites VideoChat-R1: Enhancing Spatio-Temporal Perception via Reinforcement Fine-Tuning.

Moment-Video: Diagnosing Temporal Fidelity of Video MLLMs on Momentary Visual Events VideoChat-R1: Enhancing Spatio-Temporal Perception via Reinforcement Fine-Tuning

Reference 13

Resolution
verified exact
local_arxiv, observed 2026-07-01T22:56:20.898675Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-28T14:50:02.159411Z digest=sha256:d824367bf302dc96df88c0b7857c2f0ad88265d1b5a82876c7663d6a09ea761a

Observation 521b4e63-8d93-4afe-baeb-58a2b537e33d · outbound

This paper cites Videochat-r1.

Moment-Video: Diagnosing Temporal Fidelity of Video MLLMs on Momentary Visual Events Videochat-r1

Reference 14

Resolution
metadata mismatch
arxiv_id, observed 2026-07-01T22:56:20.840548Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-28T14:50:02.159411Z digest=sha256:76b9a940685b42309a3df2a46c47d3968892e7d07bf828ba9d5fd0196551fd8e

Observation ad3d368e-2a78-4611-80e2-b6d572f7d88c · outbound

This paper cites Mvbench: A comprehensive multi-modal video understanding benchmark,.

Moment-Video: Diagnosing Temporal Fidelity of Video MLLMs on Momentary Visual Events Mvbench: A comprehensive multi-modal video understanding benchmark,

Reference 15

Resolution
unresolved
no resolver link, observed 2026-06-28T14:50:02.159411Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T14:50:02.159411Z digest=sha256:5010cfe282f0f04e2a645efd37431bb12995cccab940fcb04735b11d774da4fd

Observation 88746ea4-69e3-4143-b76c-dae8301f75e1 · outbound

This paper cites Video-mme: The first-ever comprehensive evaluation benchmark of multi-modal llms in video analysis,.

Moment-Video: Diagnosing Temporal Fidelity of Video MLLMs on Momentary Visual Events Video-mme: The first-ever comprehensive evaluation benchmark of multi-modal llms in video analysis,

Reference 16

Resolution
unresolved
no resolver link, observed 2026-06-28T14:50:02.159411Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T14:50:02.159411Z digest=sha256:e764e3aad3054d0621516b7f4d935779fa8eef8664bc16d28b2b546779d25843

Observation 57b4fda6-02e1-424f-8cf3-a01012bba15a · outbound

This paper cites Video-MME-v2: Towards the Next Stage in Benchmarks for Comprehensive Video Understanding.

Moment-Video: Diagnosing Temporal Fidelity of Video MLLMs on Momentary Visual Events Video-MME-v2: Towards the Next Stage in Benchmarks for Comprehensive Video Understanding

Reference 17

Resolution
verified exact
local_arxiv, observed 2026-07-01T22:56:20.874736Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-28T14:50:02.159411Z digest=sha256:44de57ca3253a6108f079d5581812abcf2aac04b1c8dbd567bab800c0daf4a91

Observation 53216f08-b4cb-46cf-8f5f-9fe1345b2923 · outbound

This paper cites Motionbench: Benchmarking and improving fine-grained video motion understanding for vision language models,.

Moment-Video: Diagnosing Temporal Fidelity of Video MLLMs on Momentary Visual Events Motionbench: Benchmarking and improving fine-grained video motion understanding for vision language models,

Reference 18

Resolution
unresolved
no resolver link, observed 2026-06-28T14:50:02.159411Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T14:50:02.159411Z digest=sha256:ba6bccfb275fa253f91ac0b1ac14e8453d5bc07f17709282dc8db301a9d72a3d

Observation e72667ee-7ddd-435b-b475-43da0418d8b8 · outbound

This paper cites FAVOR-Bench: A Comprehensive Benchmark for Fine-Grained Video Motion Understanding.

Moment-Video: Diagnosing Temporal Fidelity of Video MLLMs on Momentary Visual Events FAVOR-Bench: A Comprehensive Benchmark for Fine-Grained Video Motion Understanding

Reference 19

Resolution
metadata mismatch
arxiv_id, observed 2026-07-01T22:56:20.891145Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-28T14:50:02.159411Z digest=sha256:82f11fe79d421b38970eef1d26abb2d33c78279b7243fef05d765b6d88e144b9

Observation b9cc5764-de71-4616-9c15-8bed25e4c86b · outbound

This paper cites Lvbench: An extreme long video understanding benchmark,.

Moment-Video: Diagnosing Temporal Fidelity of Video MLLMs on Momentary Visual Events Lvbench: An extreme long video understanding benchmark,

Reference 20

Resolution
unresolved
no resolver link, observed 2026-06-28T14:50:02.159411Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T14:50:02.159411Z digest=sha256:3fa2bce27c834ab235b46ae56ce3ef161bd5023a7105be7eeaacec2913c87f34

Observation a713c0db-e5f7-42ee-83e3-2a777922a005 · outbound

This paper cites Longvideobench: A benchmark for long-context interleaved video-language understanding,.

Moment-Video: Diagnosing Temporal Fidelity of Video MLLMs on Momentary Visual Events Longvideobench: A benchmark for long-context interleaved video-language understanding,

Reference 21

Resolution
unresolved
no resolver link, observed 2026-06-28T14:50:02.159411Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T14:50:02.159411Z digest=sha256:7f840c090b120ee38b807d75a23958d0876b4d4246bd76afbb04ef7eb014e6a4

Observation 4725545a-3bd9-42eb-9f07-79f12a045dd1 · outbound

This paper cites Mlvu: Benchmarking multi-task long video understanding,.

Moment-Video: Diagnosing Temporal Fidelity of Video MLLMs on Momentary Visual Events Mlvu: Benchmarking multi-task long video understanding,

Reference 22

Resolution
unresolved
no resolver link, observed 2026-06-28T14:50:02.159411Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T14:50:02.159411Z digest=sha256:0c1ec0b3d9f2f1995dd2db22c730074675a3a6ff994bca0db6373feeac0951c3

Observation 4b7fcc73-94ef-4ea3-b662-a697b800f01b · outbound

This paper cites Hourvideo: 1-hour video-language understanding,.

Moment-Video: Diagnosing Temporal Fidelity of Video MLLMs on Momentary Visual Events Hourvideo: 1-hour video-language understanding,

Reference 23

Resolution
unresolved
no resolver link, observed 2026-06-28T14:50:02.159411Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T14:50:02.159411Z digest=sha256:f1d7d7839f3ca9b65ac445c641e1ab85f92e2ce862d28ec89bb4904ac333fdb0

Observation 8aace875-383e-4e60-a3d6-bbadf57aab9e · outbound

This paper cites Seeing from another perspective: Evaluating multi-view understanding in mllms,.

Moment-Video: Diagnosing Temporal Fidelity of Video MLLMs on Momentary Visual Events Seeing from another perspective: Evaluating multi-view understanding in mllms,

Reference 24

Resolution
unresolved
no resolver link, observed 2026-06-28T14:50:02.159411Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T14:50:02.159411Z digest=sha256:a4f01f7face7a363ba4d14ecfc183cffa583a20aa943f5dcd7d89a5e6e7e844f

Observation 18a72315-5a40-452d-87d1-0c4142502165 · outbound

This paper cites Crossvid: A comprehensive benchmark for evaluating cross-video reasoning in multimodal large language models,.

Moment-Video: Diagnosing Temporal Fidelity of Video MLLMs on Momentary Visual Events Crossvid: A comprehensive benchmark for evaluating cross-video reasoning in multimodal large language models,

Reference 25

Resolution
unresolved
no resolver link, observed 2026-06-28T14:50:02.159411Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T14:50:02.159411Z digest=sha256:0b3f08b3c2c7fe36319274777ca45a1f53ddae16dfd6b4061db3c6fadd9d5df6

Observation 634962d4-5e78-45ee-ab15-8f79c5991c21 · outbound

This paper cites Videoreasonbench: Can mllms perform vision-centric complex video reasoning?.

Moment-Video: Diagnosing Temporal Fidelity of Video MLLMs on Momentary Visual Events Videoreasonbench: Can mllms perform vision-centric complex video reasoning?

Reference 26

Resolution
verified exact
arxiv_id, observed 2026-07-01T22:56:20.870527Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-28T14:50:02.159411Z digest=sha256:241f5aa560c1ff760b744b3e193cabecde361543d72bc7fd257e2483100a1030

Observation 242c2a57-807e-4572-a409-cb13252dbdfd · outbound

This paper cites Video-MMMU: Evaluating Knowledge Acquisition from Multi-Discipline Professional Videos.

Moment-Video: Diagnosing Temporal Fidelity of Video MLLMs on Momentary Visual Events Video-MMMU: Evaluating Knowledge Acquisition from Multi-Discipline Professional Videos

Reference 27

Resolution
verified exact
local_arxiv, observed 2026-07-01T22:56:20.859293Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-28T14:50:02.159411Z digest=sha256:c7a8eda4034951947e5c0a59894057b59433149073f6c24328eb40b4f8c6619e

Observation 7e2958a7-24d8-4644-a4fa-746c9615b66a · outbound

This paper cites Mmvu: Measuring expert-level multi-discipline video understanding,.

Moment-Video: Diagnosing Temporal Fidelity of Video MLLMs on Momentary Visual Events Mmvu: Measuring expert-level multi-discipline video understanding,

Reference 28

Resolution
unresolved
no resolver link, observed 2026-06-28T14:50:02.159411Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T14:50:02.159411Z digest=sha256:b2d014eec599b91bda3c64c9cdf8489be6e38f089c608ff531e08fd27c885096

Observation 6e370cfa-66a0-486f-b2ce-1fcd4605dc95 · outbound

This paper cites Videoads for fast-paced video understanding,.

Moment-Video: Diagnosing Temporal Fidelity of Video MLLMs on Momentary Visual Events Videoads for fast-paced video understanding,

Reference 29

Resolution
unresolved
no resolver link, observed 2026-06-28T14:50:02.159411Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T14:50:02.159411Z digest=sha256:c658551994adfb24f37a92dcccbfa3ddd486d11b2e5e218f1fdf31018447bd0b

Observation ca256f31-5a2d-482d-819e-bf6f8e3044d2 · outbound

This paper cites Egoschema: A diagnostic benchmark for very long-form video language understanding,.

Moment-Video: Diagnosing Temporal Fidelity of Video MLLMs on Momentary Visual Events Egoschema: A diagnostic benchmark for very long-form video language understanding,

Reference 30

Resolution
unresolved
no resolver link, observed 2026-06-28T14:50:02.159411Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T14:50:02.159411Z digest=sha256:72aab866783ea62ea82ecb04161ba0966e770058b9f262a5ba132901964f49d9

Observation 59bf114e-8823-4572-a5a5-825c58f22219 · outbound

This paper cites Llama-omni 2: Llm-based real-time spoken chatbot with autoregressive streaming speech synthesis,.

Moment-Video: Diagnosing Temporal Fidelity of Video MLLMs on Momentary Visual Events Llama-omni 2: Llm-based real-time spoken chatbot with autoregressive streaming speech synthesis,

Reference 31

Resolution
unresolved
no resolver link, observed 2026-06-28T14:50:02.159411Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T14:50:02.159411Z digest=sha256:91068c0d3eefbdb5ebb26e3300d47bdb2ac36a6ebd35e61067aacc0fdcce8aa0

Observation c0c69416-f179-4fe3-a134-47d7e2ec5a3f · outbound

This paper cites GLM-5: from Vibe Coding to Agentic Engineering.

Moment-Video: Diagnosing Temporal Fidelity of Video MLLMs on Momentary Visual Events GLM-5: from Vibe Coding to Agentic Engineering

Reference 32

Resolution
verified exact
local_arxiv, observed 2026-07-01T22:56:20.882937Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-28T14:50:02.159411Z digest=sha256:81fb974256873a240d439eb9e168304c44462a08a2034394d21c8e83c28f34dc

Observation a69b2502-b027-46d2-bb11-ac1bed800e82 · outbound

This paper cites Toward deep representation learning for event-enhanced visual autonomous perception: The eap dataset,.

Moment-Video: Diagnosing Temporal Fidelity of Video MLLMs on Momentary Visual Events Toward deep representation learning for event-enhanced visual autonomous perception: The eap dataset,

Reference 33

Resolution
unresolved
no resolver link, observed 2026-06-28T14:50:02.159411Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T14:50:02.159411Z digest=sha256:5458a30a4bca660933b8560b0f5f2b4fce70f6502343ab843f6cc0dbd6ed690e

Observation b30ea1fe-fd6e-4529-b1d0-795ba320217a · outbound

This paper cites Aisafety: An ai-based smart system for enhancing operator safety in production processes,.

Moment-Video: Diagnosing Temporal Fidelity of Video MLLMs on Momentary Visual Events Aisafety: An ai-based smart system for enhancing operator safety in production processes,

Reference 34

Resolution
unresolved
no resolver link, observed 2026-06-28T14:50:02.159411Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T14:50:02.159411Z digest=sha256:1b4508b33431eb8d339941fbe7e7424a14ae5a89a2a22654c41f4b602e494fbc

Observation 304f4478-bb3f-4ea5-a928-0049afab94a8 · outbound

This paper cites Introducing our most intelligent model yet. with state-of-the-art reasoning to help you learn, build, and plan anything,.

Moment-Video: Diagnosing Temporal Fidelity of Video MLLMs on Momentary Visual Events Introducing our most intelligent model yet. with state-of-the-art reasoning to help you learn, build, and plan anything,

Reference 35

Resolution
unresolved
no resolver link, observed 2026-06-28T14:50:02.159411Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T14:50:02.159411Z digest=sha256:0a2fdea9c304fc0c5f793eb5b6541e8306a9d054464cfbfc0185d721a3f53c5c

Observation 3e00c39b-620c-4abd-bab9-a911b168b803 · outbound

This paper cites Seed2.0 model card: Towards intelligence frontier for real-world complexity,.

Moment-Video: Diagnosing Temporal Fidelity of Video MLLMs on Momentary Visual Events Seed2.0 model card: Towards intelligence frontier for real-world complexity,

Reference 36

Resolution
unresolved
no resolver link, observed 2026-06-28T14:50:02.159411Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T14:50:02.159411Z digest=sha256:036cb67c483ea1d6ed329fae5cfb7ee5bba580818c6cccedec2f0154d65e6fa5

Observation 2a0c2852-8407-44cc-a48a-4aee4111ed30 · outbound

This paper cites OpenAI GPT-5 System Card.

Moment-Video: Diagnosing Temporal Fidelity of Video MLLMs on Momentary Visual Events OpenAI GPT-5 System Card

Reference 37

Resolution
verified exact
local_arxiv, observed 2026-07-01T22:56:20.879172Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-28T14:50:02.159411Z digest=sha256:70e1db4259521c8570de3253076d6377afa60fe8ef98c1dc1d3efa393bb55ea6

Observation 50d3c2aa-4a93-4d23-b619-9f786db7b748 · outbound

This paper cites Mimo-v2.5,.

Moment-Video: Diagnosing Temporal Fidelity of Video MLLMs on Momentary Visual Events Mimo-v2.5,

Reference 38

Resolution
unresolved
no resolver link, observed 2026-06-28T14:50:02.159411Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T14:50:02.159411Z digest=sha256:e8e2ea4810f395ce27ad514e83bef9dc24f28b4a8469b5f227f9a82493603471

Observation 15cafdbb-a7b7-4184-9880-be4582a54f53 · outbound

This paper cites Qwen3.5: Towards native multimodal agents,.

Moment-Video: Diagnosing Temporal Fidelity of Video MLLMs on Momentary Visual Events Qwen3.5: Towards native multimodal agents,

Reference 39

Resolution
unresolved
no resolver link, observed 2026-06-28T14:50:02.159411Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T14:50:02.159411Z digest=sha256:137c311993c9724a1e78a1bef1cfa387f677ad5dfd29d696596c230dc9eae685

Observation a0576681-6af3-46b8-b53f-ebf2fa3d8425 · outbound

This paper cites Qwen3.6-35B-A3B: Agentic coding power, now open to all,.

Moment-Video: Diagnosing Temporal Fidelity of Video MLLMs on Momentary Visual Events Qwen3.6-35B-A3B: Agentic coding power, now open to all,

Reference 40

Resolution
unresolved
no resolver link, observed 2026-06-28T14:50:02.159411Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T14:50:02.159411Z digest=sha256:afb51b82e12df73fb1ef6671e9a5eeaf5be7fcff932f994b676e547906c53078

Observation 6032dfa5-0746-4bc8-aabe-db52e4c7c1d1 · outbound

This paper cites Qwen3.6-27B: Flagship-level coding in a 27B dense model,.

Moment-Video: Diagnosing Temporal Fidelity of Video MLLMs on Momentary Visual Events Qwen3.6-27B: Flagship-level coding in a 27B dense model,

Reference 41

Resolution
unresolved
no resolver link, observed 2026-06-28T14:50:02.159411Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T14:50:02.159411Z digest=sha256:bf593bbcd0ed8a17deb58db98344e9bd8f82db98b391de7ca2a46891afd264ad

Observation 774d293e-f1ff-4f69-acd2-d22d96d6ed8b · outbound

This paper cites InternVL3.5: Advancing Open-Source Multimodal Models in Versatility, Reasoning, and Efficiency.

Moment-Video: Diagnosing Temporal Fidelity of Video MLLMs on Momentary Visual Events InternVL3.5: Advancing Open-Source Multimodal Models in Versatility, Reasoning, and Efficiency

Reference 42

Resolution
verified exact
local_arxiv, observed 2026-07-01T22:56:20.907805Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-28T14:50:02.159411Z digest=sha256:4249a1ac5bf7ca438358f2995f13705fc01341a21eedec14bd3061bb44c33f09

Observation 73d1f7c9-175c-4906-9556-0ebde69178ca · outbound

This paper cites LLaVA-Video: Video Instruction Tuning With Synthetic Data.

Moment-Video: Diagnosing Temporal Fidelity of Video MLLMs on Momentary Visual Events LLaVA-Video: Video Instruction Tuning With Synthetic Data

Reference 43

Resolution
verified exact
local_arxiv, observed 2026-07-01T22:56:20.829957Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-28T14:50:02.159411Z digest=sha256:5a1da553228085a170a3d9829f471d41d5c1708b36bb2afab7b574a47e09f6a1

Observation 5c3ad478-6fe2-4b04-890e-56dfe3337382 · outbound

This paper cites Kimi k2.6 tech blog: Advancing open-source coding,.

Moment-Video: Diagnosing Temporal Fidelity of Video MLLMs on Momentary Visual Events Kimi k2.6 tech blog: Advancing open-source coding,

Reference 44

Resolution
unresolved
no resolver link, observed 2026-06-28T14:50:02.159411Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T14:50:02.159411Z digest=sha256:05896fd415daffc538a784871387fbb68f5e69e21a76ef71151d60d867ea1199

Observation 015c5d72-5ac2-4882-bb41-e2088a24f2b4 · outbound

This paper cites Envisioning beyond the pixels: Benchmarking reasoning-informed visual editing,.

Moment-Video: Diagnosing Temporal Fidelity of Video MLLMs on Momentary Visual Events Envisioning beyond the pixels: Benchmarking reasoning-informed visual editing,

Reference 45

Resolution
unresolved
no resolver link, observed 2026-06-28T14:50:02.159411Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T14:50:02.159411Z digest=sha256:2d451154fdfed9105e7772367e3215192ebba43cf7dbcdddbc4aaacc74b5ddd2

Observation 4eb8595b-c4a9-4c92-88e5-d8a66d97e923 · outbound

This paper cites GUI-World: A Video Benchmark and Dataset for Multimodal GUI-oriented Understanding.

Moment-Video: Diagnosing Temporal Fidelity of Video MLLMs on Momentary Visual Events GUI-World: A Video Benchmark and Dataset for Multimodal GUI-oriented Understanding

Reference 46

Resolution
metadata mismatch
arxiv_id, observed 2026-07-01T22:56:20.833472Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-28T14:50:02.159411Z digest=sha256:77e26dfd51f4e7f557d82415b8485b98a26f0b4d5345b5d8758acf5ceb5a6146

Observation 34d31f1c-e049-4930-bded-aee7ec7e30ec · outbound

This paper cites Dataset of photographed lightning events attaching to and around the brixton tower, johannesburg, south africa for the 2015-2016 thunderstorm season.

Moment-Video: Diagnosing Temporal Fidelity of Video MLLMs on Momentary Visual Events Dataset of photographed lightning events attaching to and around the brixton tower, johannesburg, south africa for the 2015-2016 thunderstorm season

Reference 47

Resolution
unresolved
no resolver link, observed 2026-06-28T14:50:02.159411Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T14:50:02.159411Z digest=sha256:db1dc12ddfaa39354072863de8b01115c18745ade69d53184845f26e79d5757e

Observation 5bfeb4a9-a345-4709-8775-6bfb2ed68445 · outbound

This paper cites Egocentric-10k,.

Moment-Video: Diagnosing Temporal Fidelity of Video MLLMs on Momentary Visual Events Egocentric-10k,

Reference 48

Resolution
unresolved
no resolver link, observed 2026-06-28T14:50:02.159411Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T14:50:02.159411Z digest=sha256:99edcfdbdeab9a98adbcc5ad6b9133ea3d6ca3aba17e2153781bfe1e5d764524

Observation 14a41bf5-cd51-45b2-93b8-592a8a2535d8 · outbound

This paper cites Gemma 4: Our most capable open models to date,.

Moment-Video: Diagnosing Temporal Fidelity of Video MLLMs on Momentary Visual Events Gemma 4: Our most capable open models to date,

Reference 49

Resolution
unresolved
no resolver link, observed 2026-06-28T14:50:02.159411Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T14:50:02.159411Z digest=sha256:61fdc8ee4358574b109fc44ca9a4b8e530f8df4becc06fd9c09739983867c795

Observation b2f52786-ada2-4b3a-8b12-92cb7ab9a5ab · outbound

This paper cites Kwai Keye-VL 1.5 Technical Report.

Moment-Video: Diagnosing Temporal Fidelity of Video MLLMs on Momentary Visual Events Kwai Keye-VL 1.5 Technical Report

Reference 50

Resolution
verified exact
arxiv_id, observed 2026-07-01T22:56:20.836897Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-28T14:50:02.159411Z digest=sha256:e6627fd473b72d028bbab664c9d7b20ffcf259a64d9dd7412c9ffff70912db51

Observation bfee7568-6d3d-4d97-9298-e78e1bcb245f · outbound

This paper cites VITA-1.5: Towards GPT-4o Level Real-Time Vision and Speech Interaction.

Moment-Video: Diagnosing Temporal Fidelity of Video MLLMs on Momentary Visual Events VITA-1.5: Towards GPT-4o Level Real-Time Vision and Speech Interaction

Reference 51

Resolution
malformed identifier
local_arxiv, observed 2026-07-01T22:56:20.844108Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-28T14:50:02.159411Z digest=sha256:12dfb223a3c99660e1442dfde129cc0bd20699c01cb075c18129c166dcabff57

Observation 5cb731e8-540e-43e8-8195-34620b776c69 · outbound

This paper cites an unresolved cited work.

Moment-Video: Diagnosing Temporal Fidelity of Video MLLMs on Momentary Visual Events Unresolved cited work

Reference 52

Resolution
unresolved
no resolver link, observed 2026-06-28T14:50:02.159411Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T14:50:02.159411Z digest=sha256:3b39a3e1411ec96ee95928dacd23d5a34e66a9d2c2022b286f87cfe8ab9cd5df

Observation fe48c176-45d7-45ea-a12e-b94a27ad6864 · outbound

This paper cites an unresolved cited work.

Moment-Video: Diagnosing Temporal Fidelity of Video MLLMs on Momentary Visual Events Unresolved cited work

Reference 53

Resolution
unresolved
no resolver link, observed 2026-06-28T14:50:02.159411Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T14:50:02.159411Z digest=sha256:fa518c51b4c72d3b028f1f06a2f49b7448274c30a6e0a213abefcdbf7168860e

Observation c73410b3-4236-4c9e-9247-7b90c7102632 · outbound

This paper cites an unresolved cited work.

Moment-Video: Diagnosing Temporal Fidelity of Video MLLMs on Momentary Visual Events Unresolved cited work

Reference 54

Resolution
unresolved
no resolver link, observed 2026-06-28T14:50:02.159411Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T14:50:02.159411Z digest=sha256:12460aeb2b0bceab8199a3a04147415376aa4be08ed3d3137fc13e7574c07e08

Observation 8cc4afba-62ab-4525-aa13-4fac916de58b · outbound

This paper cites an unresolved cited work.

Moment-Video: Diagnosing Temporal Fidelity of Video MLLMs on Momentary Visual Events Unresolved cited work

Reference 55

Resolution
unresolved
no resolver link, observed 2026-06-28T14:50:02.159411Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T14:50:02.159411Z digest=sha256:02c57e46ccedfc62d8ccb70f5cae703d8c4ba90d9ffe43522027ad7e07304d04

Observation 55b66b22-502c-428d-b141-2651c19ac61a · outbound

This paper cites an unresolved cited work.

Moment-Video: Diagnosing Temporal Fidelity of Video MLLMs on Momentary Visual Events Unresolved cited work

Reference 56

Resolution
unresolved
no resolver link, observed 2026-06-28T14:50:02.159411Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T14:50:02.159411Z digest=sha256:e9f13e06d7174968b9a30a8fe37eb38b7c6a574e17de52be1cf84924cab07c68

Observation 9e773264-7cad-4340-9a60-973ba2f5c5cb · outbound

This paper cites is consistent.

Moment-Video: Diagnosing Temporal Fidelity of Video MLLMs on Momentary Visual Events is consistent

Reference 57

Resolution
unresolved
no resolver link, observed 2026-06-28T14:50:02.159411Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T14:50:02.159411Z digest=sha256:d2421d568294e2692804cf97eac5458b0000ddf4a0af5564a1eb6842bf5deb3e

Pith citing papers

Observation ad2aa71d-2926-493f-9a58-83fc99f8bb29 · inbound

Beyond Frame Selection: Generative Latent Evidence Aggregation for Long-Video Understanding cites this paper.

Beyond Frame Selection: Generative Latent Evidence Aggregation for Long-Video Understanding Moment-Video: Diagnosing Temporal Fidelity of Video MLLMs on Momentary Visual Events

Reference 15

Resolution
unresolved
no resolver link, observed 2026-07-31T04:55:07.020202Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T04:55:07.020202Z digest=sha256:6d27287ae6b246508b806f062069c0b37b514e342ff33aa5e5f8917be1d10e57