Pith. sign in

Paper Citation Record · LEDGER

Infinite Video Understanding

As of 7 August 2026, this Paper Citation Record lists 64 of 64 outbound references and 0 inbound Pith citation observations for arXiv:2507.09068.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.09068 v2

Coverage vector

measured 64 of 64 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T18:10:14.919272Z

measured 64 of 64 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-06T06:34:29.942622+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

64 of 64 outbound references displayed

  • verified exact8
  • verified fuzzy1
  • unresolved54
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch1

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation ff0a1ede-3856-4a74-ad1e-e3544cd0a8b4 · outbound

This paper cites CrossVideoMAE: Self-Supervised Image-Video Representation Learning with Masked Autoencoders.

Infinite Video Understanding CrossVideoMAE: Self-Supervised Image-Video Representation Learning with Masked Autoencoders

Reference 1

Resolution
verified exact
local_arxiv, observed 2026-08-06T18:10:16.731929Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T18:09:18.061419Z digest=sha256:3d2015520139ab13a7ef77510fe3bb8f7024162917feed4f3bc5dfc5806a33ca

Observation c25e915d-4ba9-493f-847b-51365e8f7fda · outbound

This paper cites an unresolved cited work.

Infinite Video Understanding Unresolved cited work

Reference 2

Resolution
unresolved
raw_fallback, observed 2026-08-06T18:10:18.299149Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T18:09:18.114356Z digest=sha256:056b0009665cdf1693ebc1b05deaf4bad5ea8b0a13091c5c6e538281623e15cd

Observation 01bcc5ab-e0c7-4408-a60e-48774759a8fe · outbound

This paper cites HourVideo: 1-Hour Video-Language Understanding.

Infinite Video Understanding HourVideo: 1-Hour Video-Language Understanding

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-06T18:09:18.211447Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:09:18.211447Z digest=sha256:bc2b1edaa5cafc73317e30020cc133412abfee815a7f6b7173e731f2fbdf2e1f

Observation 339e7bb3-dc67-4a6f-8af4-aff75a9dfa3c · outbound

This paper cites LongVILA: Scaling Long-Context Visual Language Models for Long Videos.

Infinite Video Understanding LongVILA: Scaling Long-Context Visual Language Models for Long Videos

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-06T18:09:18.313230Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:09:18.313230Z digest=sha256:16811c0c1d8d713076d806b9280efd69aac2489475ebb876388910cb0c571779

Observation 809b2dcf-8bb2-479b-964a-3488ccdb6144 · outbound

This paper cites Scaling Video-Language Models to 10K Frames via Hierarchical Differential Distillation.

Infinite Video Understanding Scaling Video-Language Models to 10K Frames via Hierarchical Differential Distillation

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-06T18:09:18.568750Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:09:18.568750Z digest=sha256:09e0101448fbe1fbe1f531e558ff96715f0b2c90f51cc42d7e480183558d7e20

Observation 8d053b18-1eb2-430a-9e88-2a6c782636e8 · outbound

This paper cites Video-Holmes: Can MLLM Think Like Holmes for Complex Video Reasoning?.

Infinite Video Understanding Video-Holmes: Can MLLM Think Like Holmes for Complex Video Reasoning?

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-06T18:09:18.954541Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:09:18.954541Z digest=sha256:dc24f6745244992b47846723d707db0cc59368c1bdceea51f73561335d12bb1b

Observation dbef02ca-88ff-488e-a914-02db772cb241 · outbound

This paper cites Zero-Shot Video Question Answering with Procedural Programs.

Infinite Video Understanding Zero-Shot Video Question Answering with Procedural Programs

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-06T18:09:19.409539Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:09:19.409539Z digest=sha256:e6b26e16a00b2fac5ac342c3b78a4b3d27864ca082cb3a2a657a944424d77ed7

Observation 68494767-b952-4b9b-bba1-bbddf06ed232 · outbound

This paper cites Project Aria: A New Tool for Egocentric Multi-Modal AI Research.

Infinite Video Understanding Project Aria: A New Tool for Egocentric Multi-Modal AI Research

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-06T18:10:11.725676Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:10:11.725676Z digest=sha256:54bde296047760742b21c9a7989d106304dbd47b58817ab7db880b1f7c3291e7

Observation 545305b2-f2bb-4383-9be8-639125a12ebf · outbound

This paper cites Efficient Image Pre-Training with Siamese Cropped Masked Autoencoders.

Infinite Video Understanding Efficient Image Pre-Training with Siamese Cropped Masked Autoencoders

Reference 9

Resolution
verified exact
local_arxiv, observed 2026-08-06T18:10:17.212457Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T18:10:11.854309Z digest=sha256:e3f2704b49772a2ed1c2e762ad4e5e0c503c73a3427a97741e6a2cfc648296c5

Observation 316ad4f0-026d-4451-84aa-0b1e8874df9e · outbound

This paper cites an unresolved cited work.

Infinite Video Understanding Unresolved cited work

Reference 10

Resolution
unresolved
raw_fallback, observed 2026-08-06T18:10:18.130805Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T18:10:11.899351Z digest=sha256:ddfc3af9aad93456b87d338e39b9622e5a439d5ef77ee84b1cddddc5021b95d4

Observation 276abd28-1608-4fa1-8d37-0830654d553a · outbound

This paper cites Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis.

Infinite Video Understanding Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-06T18:10:12.056934Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:10:12.056934Z digest=sha256:dfe09c029722f8a46f32ff126561d2fd94550b8394bf7ec2459aa1ba532f799e

Observation 4cbccd92-ffbe-4212-98e0-67fa7f83e1a5 · outbound

This paper cites The Ethics of Advanced AI Assistants.

Infinite Video Understanding The Ethics of Advanced AI Assistants

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-06T18:10:12.125589Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:10:12.125589Z digest=sha256:fc86d760a7bce91b35ab79c01cf8f2b47e966967a262c6bb37e2a72b643b3743

Observation 963c17ec-82c8-4343-91e9-d85d06aa17ca · outbound

This paper cites Siamese Masked Autoencoders.

Infinite Video Understanding Siamese Masked Autoencoders

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-06T18:10:12.195376Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:10:12.195376Z digest=sha256:466c73e7625e5ccf0ec721f733836a9d85bea17ce59623c0b3d0639cc0281e7b

Observation 12718065-f4c4-4ac8-9980-1d22368d0953 · outbound

This paper cites RULER: What's the Real Context Size of Your Long-Context Language Models?.

Infinite Video Understanding RULER: What's the Real Context Size of Your Long-Context Language Models?

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-06T18:10:12.307299Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:10:12.307299Z digest=sha256:b24a842291b01d4a68c524fe1bedf7c433f0a2c08d1bec480d7e5e3c043faa76

Observation 9110064f-0fc6-4002-a20d-c5fe49e3a397 · outbound

This paper cites Visual Representation Learning with Stochastic Frame Prediction.

Infinite Video Understanding Visual Representation Learning with Stochastic Frame Prediction

Reference 15

Resolution
verified exact
local_arxiv, observed 2026-08-06T18:10:16.421305Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T18:10:12.401921Z digest=sha256:c3f31251a4520e124a6b4f9b9b89e3100675af49f6df4022651840b578f5fb29

Observation 2b5d77ea-50d5-4d97-8f44-c7bd15d9693d · outbound

This paper cites Search-R1: Training LLMs to Reason and Leverage Search Engines with Reinforcement Learning.

Infinite Video Understanding Search-R1: Training LLMs to Reason and Leverage Search Engines with Reinforcement Learning

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-06T18:10:12.472087Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:10:12.472087Z digest=sha256:e42f0708c3a051aa9fd1c7fe1b8356182652999334a3bed679cddd7dc2a82704

Observation 0b728617-5100-420c-887d-8c52582f04f6 · outbound

This paper cites SNeRV: Spectra-preserving Neural Representation for Video.

Infinite Video Understanding SNeRV: Spectra-preserving Neural Representation for Video

Reference 17

Resolution
verified exact
local_arxiv, observed 2026-08-06T18:10:16.964469Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T18:10:12.519667Z digest=sha256:1d06bba76e8af5cbf3a74be54cb20244b9dae7eb56168fd8ff96227e81fe4589

Observation fb2b3eb7-0cd3-4be0-bd16-9ca2d015fd81 · outbound

This paper cites an unresolved cited work.

Infinite Video Understanding Unresolved cited work

Reference 18

Resolution
unresolved
raw_fallback, observed 2026-08-06T18:10:17.987768Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T18:10:12.559183Z digest=sha256:8bc841b7052637b212126067560e1c0186103ff62182c864dc7583045904e116

Observation 7b403170-dbfe-452a-bdbd-19522ed26d68 · outbound

This paper cites an unresolved cited work.

Infinite Video Understanding Unresolved cited work

Reference 19

Resolution
verified exact
doi, observed 2026-08-06T18:10:16.232019Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T18:10:12.643858Z digest=sha256:4948f20b9b9351d666606d82818223bf502807f831e4897e424b8637b0d5fea5

Observation 43e3fa0f-361d-4cba-9e31-bc1fd9e80412 · outbound

This paper cites START: Self-taught Reasoner with Tools.

Infinite Video Understanding START: Self-taught Reasoner with Tools

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-06T18:10:12.598324Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:10:12.598324Z digest=sha256:bdbe31902d8d109e1ff46ada1543ef2c659fe523ce96e2fb1d83d4081d0d1702

Observation 97ac225c-a44b-48e7-b367-d1e266ecfdf5 · outbound

This paper cites VideoChat-Flash: Hierarchical Compression for Long-Context Video Modeling.

Infinite Video Understanding VideoChat-Flash: Hierarchical Compression for Long-Context Video Modeling

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-06T18:10:12.714751Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:10:12.714751Z digest=sha256:01c590318787d7f2c007eb7c97e7f92e93be9d78babc0b1edd00620c63d09813

Observation cc6d9d5e-c795-46dd-b38d-5aa40d4a9983 · outbound

This paper cites MVBench: A Comprehensive Multi-modal Video Understanding Benchmark.

Infinite Video Understanding MVBench: A Comprehensive Multi-modal Video Understanding Benchmark

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-06T18:10:12.662954Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:10:12.662954Z digest=sha256:d27fcfa668073435775ddc7ce5156301ef311103cd385b0cba69b89edeb1cf47

Observation 3633c010-534b-47ba-b99c-22c9cd1f6f03 · outbound

This paper cites Video-XL-Pro: Reconstructive Token Compression for Extremely Long Video Understanding.

Infinite Video Understanding Video-XL-Pro: Reconstructive Token Compression for Extremely Long Video Understanding

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-06T18:10:12.773039Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:10:12.773039Z digest=sha256:1d37aa4f3a1e70a89fefa09ef12ef1a4febe075e7363f2e39a97eadd9742a18f

Observation 448a4907-dd0e-4419-9ed5-fba4759143ce · outbound

This paper cites World Model on Million-Length Video And Language With Blockwise RingAttention.

Infinite Video Understanding World Model on Million-Length Video And Language With Blockwise RingAttention

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-06T18:10:12.748486Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:10:12.748486Z digest=sha256:44a81476573d326791f28783f5de83087ab6caa6e5bcbb698a16e1a5af4af89d

Observation e0001093-1c8d-428c-8368-db9820138888 · outbound

This paper cites EgoSchema: A Diagnostic Benchmark for Very Long-form Video Language Understanding.

Infinite Video Understanding EgoSchema: A Diagnostic Benchmark for Very Long-form Video Language Understanding

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-06T18:10:12.839920Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:10:12.839920Z digest=sha256:67271d04b8b7f98aedcce5715f86ff8cea3ee376ea1f77678caf65a448a02559

Observation 6dcac48e-b7e6-4df2-9834-febd2b9658ff · outbound

This paper cites an unresolved cited work.

Infinite Video Understanding Unresolved cited work

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-06T18:10:12.786822Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:10:12.786822Z digest=sha256:bf741c139f4abd483ae34cd6135570be927dd4c7fa984dbe4d10bacb1e1f7af0

Observation 01d0693d-1a66-4237-85d4-664f63fa52b0 · outbound

This paper cites Cosmos World Foundation Model Platform for Physical AI.

Infinite Video Understanding Cosmos World Foundation Model Platform for Physical AI

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-06T18:10:12.942821Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:10:12.942821Z digest=sha256:d4acd87e45941f9f17dac66e77fe3dec4e6bc2014f2afecae62f2430a8e15154

Observation 20a4e7fb-271f-4a54-b77d-5c48378e0b2e · outbound

This paper cites an unresolved cited work.

Infinite Video Understanding Unresolved cited work

Reference 28

Resolution
unresolved
raw_fallback, observed 2026-08-06T18:10:17.807895Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T18:10:12.877764Z digest=sha256:b55a752503019723d7531f22ee266e88cfa31df023c67ceae6b11bfbb2f491ca

Observation af469056-2597-4916-802c-3e9a4e6eb44a · outbound

This paper cites Video-XL-2: Towards Very Long-Video Understanding Through Task-Aware KV Sparsification.

Infinite Video Understanding Video-XL-2: Towards Very Long-Video Understanding Through Task-Aware KV Sparsification

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-06T18:10:13.072860Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:10:13.072860Z digest=sha256:869befba16bd09ac95553fe97eb6bfdc0bb4141e7fda7c49a286c16916922852

Observation e76d6843-4626-47ad-ba97-c6f7c48102fe · outbound

This paper cites Streaming Long Video Understanding with Large Language Models.

Infinite Video Understanding Streaming Long Video Understanding with Large Language Models

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-06T18:10:13.022336Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:10:13.022336Z digest=sha256:c4e6b8e1d65c0d435650973506431d3bcb867b12591cf344c7778c9ac1b70db0

Observation e32d3cd9-dbef-4bfe-8da5-bd1a7f618ab4 · outbound

This paper cites Learning Transferable Visual Models From Natural Language Supervision.

Infinite Video Understanding Learning Transferable Visual Models From Natural Language Supervision

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-06T18:10:13.140709Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:10:13.140709Z digest=sha256:2e09d8b6970c91caab32f7327441543c0291fa3cd11b9e3cf1d2046a24b961fe

Observation 2aaa3216-7106-47ce-b2e1-40f190617bbc · outbound

This paper cites Recent Advances of Continual Learning in Computer Vision: An Overview.

Infinite Video Understanding Recent Advances of Continual Learning in Computer Vision: An Overview

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-06T18:10:13.104601Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:10:13.104601Z digest=sha256:0d90489b260259049bb21693b5869d7eae4ff5cb7860c67d263e86c6fd0eaef8

Observation ea010615-2ce1-47d8-9be6-4ea8823abe73 · outbound

This paper cites LongVU: Spatiotemporal Adaptive Compression for Long Video-Language Understanding.

Infinite Video Understanding LongVU: Spatiotemporal Adaptive Compression for Long Video-Language Understanding

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-06T18:10:13.260398Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:10:13.260398Z digest=sha256:d83911e68393ac99d3937fab31f39bc94699a0aefea186eef4b6aa0825a1a3a1

Observation edc308c6-cf1f-431d-b0d1-b121878aafd8 · outbound

This paper cites $\infty$-Video: A Training-Free Approach to Long Video Understanding via Continuous-Time Memory Consolidation.

Infinite Video Understanding $\infty$-Video: A Training-Free Approach to Long Video Understanding via Continuous-Time Memory Consolidation

Reference 34

Resolution
metadata mismatch
local_arxiv, observed 2026-08-06T18:10:15.807015Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T18:10:13.200174Z digest=sha256:60660f2f689e12f324f92bccb10c713863439b34f7b394e1fe3b93a9af25a594

Observation e76d195d-b50d-4df2-969f-288c3370d330 · outbound

This paper cites an unresolved cited work.

Infinite Video Understanding Unresolved cited work

Reference 35

Resolution
unresolved
raw_fallback, observed 2026-08-06T18:10:17.698230Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T18:10:13.328713Z digest=sha256:bc875ff4f4bab32059f99afdae6e9777287c722ee46c14e3533f5a99b9636df0

Observation c2df738d-2eb6-4bef-b7fd-a788422984d2 · outbound

This paper cites Video-XL: Extra-Long Vision Language Model for Hour-Scale Video Understanding.

Infinite Video Understanding Video-XL: Extra-Long Vision Language Model for Hour-Scale Video Understanding

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-06T18:10:13.296306Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:10:13.296306Z digest=sha256:bcabacca95b332e5d37a78af8f1184848e0eb3b447fd45cc4006ff0d41c82bba

Observation 6daaacba-cd5e-4c86-b121-09b740e1fcb6 · outbound

This paper cites Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context.

Infinite Video Understanding Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-06T18:10:13.446720Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:10:13.446720Z digest=sha256:049a68fe0f0c98ef87607acfb537a301e51f90fd8c2bcb6466dec6a8d71fc539

Observation 102ccd94-8ead-4ada-9021-3f44632e9396 · outbound

This paper cites MovieChat: From Dense Token to Sparse Memory for Long Video Understanding.

Infinite Video Understanding MovieChat: From Dense Token to Sparse Memory for Long Video Understanding

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-06T18:10:13.357356Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:10:13.357356Z digest=sha256:0afbf4ec2b00b98858487fd826a36e37d3fa1ddc6c6a45b1616f2e3dffa327ac

Observation 1d59b233-d2bb-4843-81a8-9b4565f0816f · outbound

This paper cites SigLIP 2: Multilingual Vision-Language Encoders with Improved Semantic Understanding, Localization, and Dense Features.

Infinite Video Understanding SigLIP 2: Multilingual Vision-Language Encoders with Improved Semantic Understanding, Localization, and Dense Features

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-06T18:10:13.651494Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:10:13.651494Z digest=sha256:22b0fb4df9573098ab5c399256ac15d57ae33287ae54a4cc6af3731fb304767f

Observation 8627080a-8146-4fd5-893f-d0b40c9a5e1c · outbound

This paper cites VideoMAE: Masked Autoencoders are Data-Efficient Learners for Self-Supervised Video Pre-Training.

Infinite Video Understanding VideoMAE: Masked Autoencoders are Data-Efficient Learners for Self-Supervised Video Pre-Training

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-06T18:10:13.574181Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:10:13.574181Z digest=sha256:b5bdb69abdc1ed9ee5dc4c2e54bf3d8e3eada9dcd8bc8a715cd6d69126ecbf57

Observation 5c8a1017-73f4-4f2c-9e4d-a780ab36782d · outbound

This paper cites A Comprehensive Survey of Continual Learning: Theory, Method and Application.

Infinite Video Understanding A Comprehensive Survey of Continual Learning: Theory, Method and Application

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-06T18:10:13.775787Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:10:13.775787Z digest=sha256:962fba57a42d2bbd95654e65d1148990ba0d17aa4a0ab99184fd04c9a7a9d9a2

Observation 06bfa572-b497-4805-9795-6453d48c3512 · outbound

This paper cites VideoMAE V2: Scaling Video Masked Autoencoders with Dual Masking.

Infinite Video Understanding VideoMAE V2: Scaling Video Masked Autoencoders with Dual Masking

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-06T18:10:13.711528Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:10:13.711528Z digest=sha256:c028899bc5d910e6fc297853b310deb05313ec212ef689cf9e53c09fb1f140bd

Observation ef828168-849f-4259-8d45-f7f2490571e5 · outbound

This paper cites VidTwin: Video VAE with Decoupled Structure and Dynamics.

Infinite Video Understanding VidTwin: Video VAE with Decoupled Structure and Dynamics

Reference 43

Resolution
verified exact
local_arxiv, observed 2026-08-06T18:10:15.572551Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T18:10:14.101777Z digest=sha256:e2b2f7bbc23314ba40ba4f04f8f572e04c8281a7c1eb6076d12aa18ccee49937

Observation a69f610b-c659-49d6-9892-4f1bf7f702e1 · outbound

This paper cites an unresolved cited work.

Infinite Video Understanding Unresolved cited work

Reference 44

Resolution
unresolved
raw_fallback, observed 2026-08-06T18:10:17.529367Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T18:10:13.937532Z digest=sha256:6aa2a6434d02645a41ec123cfc7b3d8323d2c3f7f0ec37e8096fb8c324f364ec

Observation f519b7f8-72c3-4347-acda-48ca2a008599 · outbound

This paper cites VideoRoPE: What Makes for Good Video Rotary Position Embedding?.

Infinite Video Understanding VideoRoPE: What Makes for Good Video Rotary Position Embedding?

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-06T18:10:14.257348Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:10:14.257348Z digest=sha256:fc680d7d8c309ecb4b1984e63889daca311bec37b9c1b18642ebe7397f71f22f

Observation a30160aa-9053-40af-9c42-bef00968d3a7 · outbound

This paper cites LongVLM: Efficient Long Video Understanding via Large Language Models.

Infinite Video Understanding LongVLM: Efficient Long Video Understanding via Large Language Models

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-06T18:10:14.298291Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:10:14.298291Z digest=sha256:be0f8d69ba229becc2abc6612d2a5db2c83501a745670f049b8ee89f4de59c32

Observation 5d65b882-b812-4eb2-9817-d7922d9a8c0b · outbound

This paper cites VideoTree: Adaptive Tree-based Video Representation for LLM Reasoning on Long Videos.

Infinite Video Understanding VideoTree: Adaptive Tree-based Video Representation for LLM Reasoning on Long Videos

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-06T18:10:14.223529Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:10:14.223529Z digest=sha256:e972e5445862cbf970b12a3851dfa14f34b49aeb0e59acd5b4f9f57940294627

Observation 88a9e838-e338-44d7-b0b5-2cabd1433fbc · outbound

This paper cites Agentic Reasoning: A Streamlined Framework for Enhancing LLM Reasoning with Agentic Tools.

Infinite Video Understanding Agentic Reasoning: A Streamlined Framework for Enhancing LLM Reasoning with Agentic Tools

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-06T18:10:14.365798Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:10:14.365798Z digest=sha256:6c3b42e572b89dd6430eb0c003e2c50ed802c7205951bdcfeabab77ab6733596

Observation aea3764e-83dd-40b8-9999-1eb2fca7070d · outbound

This paper cites LongViTU: Instruction Tuning for Long-Form Video Understanding.

Infinite Video Understanding LongViTU: Instruction Tuning for Long-Form Video Understanding

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-06T18:10:14.416661Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:10:14.416661Z digest=sha256:0c7ceaecda3c11967e58f891fa3afd00f0cfe7babd1e3eea600124701b3d5291

Observation b6143b35-bfaa-484d-9084-4430ec886d4b · outbound

This paper cites LongVideoBench: A Benchmark for Long-context Interleaved Video-Language Understanding.

Infinite Video Understanding LongVideoBench: A Benchmark for Long-context Interleaved Video-Language Understanding

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-06T18:10:14.340618Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:10:14.340618Z digest=sha256:e43c0966a63bc868ec2294b51154e6466557fbd5405e5981f1704698dbd59067

Observation 9dda5626-d89f-474a-92d8-7620a5c62f59 · outbound

This paper cites MAGVIT: Masked Generative Video Transformer.

Infinite Video Understanding MAGVIT: Masked Generative Video Transformer

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-06T18:10:14.507734Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:10:14.507734Z digest=sha256:b01131c721c0a7fa4cceec59d797678d1d592b89f85d27a56c124ca9466e5458

Observation a4883221-886c-418b-9b68-445e2b4784e5 · outbound

This paper cites Language Model Beats Diffusion -- Tokenizer is Key to Visual Generation.

Infinite Video Understanding Language Model Beats Diffusion -- Tokenizer is Key to Visual Generation

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-06T18:10:14.568887Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:10:14.568887Z digest=sha256:43f9b5daefe2bd3ef7cf9e6e52b0a9565f15a653ad5fb24844c8f89867c8998a

Observation 5ab42a16-8fdb-4d71-84eb-02b6ab549f18 · outbound

This paper cites an unresolved cited work.

Infinite Video Understanding Unresolved cited work

Reference 53

Resolution
verified exact
doi, observed 2026-08-06T18:10:15.381307Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T18:10:14.459822Z digest=sha256:cdd406dbf4793477254978cc402dcc4eb4d7ee4811a2bbd558acd026c4574b06

Observation ef731b3e-8c43-42d9-aa5e-109c6274611b · outbound

This paper cites an unresolved cited work.

Infinite Video Understanding Unresolved cited work

Reference 54

Resolution
verified exact
doi, observed 2026-08-06T18:10:15.203453Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T18:10:14.678140Z digest=sha256:3942cbf9afadd3cb4020078f6565978929701e79712dab8b463eae995c9ff78d

Observation 0c444863-559e-43e1-8eb9-3f9de9561d65 · outbound

This paper cites Gundavarapu, Liangzhe Yuan, Hao Zhou, Shen Yan, Jennifer J.

Infinite Video Understanding Gundavarapu, Liangzhe Yuan, Hao Zhou, Shen Yan, Jennifer J

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:10:17.373787Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T18:10:14.711077Z digest=sha256:f70a30c8d1dcfd39fe4eadf4260e9b137c70b64cbe096486b49d9d32759b01b1

Observation 631dce2d-c1ed-46e3-a663-7fb3e379b154 · outbound

This paper cites Sigmoid Loss for Language Image Pre-Training.

Infinite Video Understanding Sigmoid Loss for Language Image Pre-Training

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-06T18:10:14.635404Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:10:14.635404Z digest=sha256:d772d4c248e6d5dfa774f9486e7b849aca3eaac05f82ceb2a026179b75ccdae0

Observation bc707462-8bba-42d4-b63e-912ab0c3dfc9 · outbound

This paper cites Towards Lifelong Learning of Large Language Models: A Survey.

Infinite Video Understanding Towards Lifelong Learning of Large Language Models: A Survey

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-06T18:10:14.841489Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:10:14.841489Z digest=sha256:9668207c3c6b4ce4d66ed3a13110a0849dc65ce136928e3d537703c558ff5cf3

Observation 12c78463-a446-463e-8f21-9ae51e99a337 · outbound

This paper cites MLVU: Benchmarking Multi-task Long Video Understanding.

Infinite Video Understanding MLVU: Benchmarking Multi-task Long Video Understanding

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-06T18:10:14.864857Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:10:14.864857Z digest=sha256:5e5de16f863d3ef4e090b00a96ba927aca443954deff85352a5c9288151e3e9f

Observation 398801a5-873c-4169-abf0-73e544e5bb62 · outbound

This paper cites VideoPrism: A Foundational Visual Encoder for Video Understanding.

Infinite Video Understanding VideoPrism: A Foundational Visual Encoder for Video Understanding

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-06T18:10:14.749813Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:10:14.749813Z digest=sha256:7f21d823a3b43f19a370554b950b530201e1022f66dd0d0efe428aa3fb1b9273

Observation 5da26043-a850-4fdd-8944-8fc06fdf7b6a · outbound

This paper cites Needle In A Video Haystack: A Scalable Synthetic Evaluator for Video MLLMs.

Infinite Video Understanding Needle In A Video Haystack: A Scalable Synthetic Evaluator for Video MLLMs

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-06T18:10:14.777464Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:10:14.777464Z digest=sha256:deec322e9aadc535dc4ba83041ad2af8f0f3acbaacdf732c9a9ba4a438d0d6e8

Observation 168f0a8e-2993-48df-b4f0-266fb42fca75 · outbound

This paper cites an unresolved cited work.

Infinite Video Understanding Unresolved cited work

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-06T18:10:14.887539Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:10:14.887539Z digest=sha256:4b8ff838c473b8ea70bfd804b29a6c4e5aee75089219170db02667a319940175

Observation ff6bf875-6dfb-403e-aa65-9bd3504869d2 · outbound

This paper cites From Seconds to Hours: Reviewing MultiModal Large Language Models on Comprehensive Long Video Understanding.

Infinite Video Understanding From Seconds to Hours: Reviewing MultiModal Large Language Models on Comprehensive Long Video Understanding

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-06T18:10:14.919272Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:10:14.919272Z digest=sha256:2ccdfaceada137c4700d69d05708ab9b22d8f3ef961fc07cc3fbf41d6d6ff4c5

Observation 004a05d9-6451-40e2-afb9-1e60ec5aa9fb · outbound

This paper cites LVBench: An Extreme Long Video Understanding Benchmark.

Infinite Video Understanding LVBench: An Extreme Long Video Understanding Benchmark

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-06T18:10:13.986726Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:10:13.986726Z digest=sha256:837c5f3919d6177f73b675bd771abdc8f172f681c9fa5ae3abcdeadeee169bdb

Observation 5dfc5ac4-c02b-4cfa-9645-4613c33f3a71 · outbound

This paper cites Video-R1: Reinforcing Video Reasoning in MLLMs.

Infinite Video Understanding Video-R1: Reinforcing Video Reasoning in MLLMs

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-06T18:10:11.970598Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:10:11.970598Z digest=sha256:cfc2a14c4775a221aa0797b6246a9931ed8041d916681e6f13ac2463bf0aedc5

Pith citing papers

No inbound Pith citation observations are available.