Pith. sign in

Paper Citation Record · LEDGER

Infinite Video Understanding

As of 7 August 2026, this Paper Citation Record lists 64 of 64 outbound references and 0 inbound Pith citation observations for arXiv:2507.09068.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.09068 v2

Coverage vector

measured 64 of 64 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T18:10:14.919272Z

measured 64 of 64 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

64 of 64 outbound references displayed

  • verified exact8
  • verified fuzzy1
  • unresolved54
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch1

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation ff0a1ede-3856-4a74-ad1e-e3544cd0a8b4 · outbound

This paper cites CrossVideoMAE: Self-Supervised Image-Video Representation Learning with Masked Autoencoders.

Infinite Video Understanding CrossVideoMAE: Self-Supervised Image-Video Representation Learning with Masked Autoencoders

Reference 1

Resolution
verified exact
local_arxiv, observed 2026-08-06T18:10:16.731929Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T18:09:18.061419Z digest=sha256:cb18d402b88d54699412f3bf8a5c41d82a8b53a1c8b8b1c138cbed005264d62e

Observation c25e915d-4ba9-493f-847b-51365e8f7fda · outbound

This paper cites an unresolved cited work.

Infinite Video Understanding Unresolved cited work

Reference 2

Resolution
unresolved
raw_fallback, observed 2026-08-06T18:10:18.299149Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T18:09:18.114356Z digest=sha256:d6b1ee48572b15d3ef7e1b831660ca897e3bba7d2001837aefb519a5c48f5f2a

Observation 01bcc5ab-e0c7-4408-a60e-48774759a8fe · outbound

This paper cites HourVideo: 1-Hour Video-Language Understanding.

Infinite Video Understanding HourVideo: 1-Hour Video-Language Understanding

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-06T18:09:18.211447Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:09:18.211447Z digest=sha256:8a8f5fb80c10423b57ec8cea84153e16e6fb2edc084df5861b59c9705a2b4551

Observation 339e7bb3-dc67-4a6f-8af4-aff75a9dfa3c · outbound

This paper cites LongVILA: Scaling Long-Context Visual Language Models for Long Videos.

Infinite Video Understanding LongVILA: Scaling Long-Context Visual Language Models for Long Videos

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-06T18:09:18.313230Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:09:18.313230Z digest=sha256:f119e54f21f87277f448ac64cb373764621a4a5a0ffb36e85a7042ed4ee63b16

Observation 809b2dcf-8bb2-479b-964a-3488ccdb6144 · outbound

This paper cites Scaling Video-Language Models to 10K Frames via Hierarchical Differential Distillation.

Infinite Video Understanding Scaling Video-Language Models to 10K Frames via Hierarchical Differential Distillation

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-06T18:09:18.568750Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:09:18.568750Z digest=sha256:32aa4880a4321978e793e16f3efa719d513345f4bf4d413ad0a9e39ec9d72bed

Observation 8d053b18-1eb2-430a-9e88-2a6c782636e8 · outbound

This paper cites Video-Holmes: Can MLLM Think Like Holmes for Complex Video Reasoning?.

Infinite Video Understanding Video-Holmes: Can MLLM Think Like Holmes for Complex Video Reasoning?

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-06T18:09:18.954541Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:09:18.954541Z digest=sha256:ba7eea2e797fdfd07b9e442e80b39d65b45f0bdb3feff021a21b2118c6c54165

Observation dbef02ca-88ff-488e-a914-02db772cb241 · outbound

This paper cites Zero-Shot Video Question Answering with Procedural Programs.

Infinite Video Understanding Zero-Shot Video Question Answering with Procedural Programs

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-06T18:09:19.409539Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:09:19.409539Z digest=sha256:690631fc3779d1b48211957782838b8e06a36addfa596d86cc136f57d656a188

Observation 68494767-b952-4b9b-bba1-bbddf06ed232 · outbound

This paper cites Project Aria: A New Tool for Egocentric Multi-Modal AI Research.

Infinite Video Understanding Project Aria: A New Tool for Egocentric Multi-Modal AI Research

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-06T18:10:11.725676Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:10:11.725676Z digest=sha256:8b52db6830c8daa880b38335b6a5658ef9ac91e77d0889eab1363b9034652acc

Observation 545305b2-f2bb-4383-9be8-639125a12ebf · outbound

This paper cites Efficient Image Pre-Training with Siamese Cropped Masked Autoencoders.

Infinite Video Understanding Efficient Image Pre-Training with Siamese Cropped Masked Autoencoders

Reference 9

Resolution
verified exact
local_arxiv, observed 2026-08-06T18:10:17.212457Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T18:10:11.854309Z digest=sha256:ee1be4c2d5d7c7b944d8a4fd4c1b0200c8e5ce748d9eaf7bf9e4d00e4ac7b9e2

Observation 316ad4f0-026d-4451-84aa-0b1e8874df9e · outbound

This paper cites an unresolved cited work.

Infinite Video Understanding Unresolved cited work

Reference 10

Resolution
unresolved
raw_fallback, observed 2026-08-06T18:10:18.130805Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T18:10:11.899351Z digest=sha256:8d5152c18cfe10c2d87ce214a8feade28e4c2a183a5485d26574586feb702dce

Observation 276abd28-1608-4fa1-8d37-0830654d553a · outbound

This paper cites Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis.

Infinite Video Understanding Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-06T18:10:12.056934Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:10:12.056934Z digest=sha256:b695fa5970901c184d128bd47e45ebae98219a907395d98ec7638103179d0f5e

Observation 4cbccd92-ffbe-4212-98e0-67fa7f83e1a5 · outbound

This paper cites The Ethics of Advanced AI Assistants.

Infinite Video Understanding The Ethics of Advanced AI Assistants

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-06T18:10:12.125589Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:10:12.125589Z digest=sha256:b611bbdc2e9d2577c019dcd415fc8e6b26d707d41f43959c10a118ff350039a6

Observation 963c17ec-82c8-4343-91e9-d85d06aa17ca · outbound

This paper cites Siamese Masked Autoencoders.

Infinite Video Understanding Siamese Masked Autoencoders

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-06T18:10:12.195376Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:10:12.195376Z digest=sha256:e855101ce39dcb4d1236d9be83537490dc3d16f2651d709fb89e2147ee6a6a58

Observation 12718065-f4c4-4ac8-9980-1d22368d0953 · outbound

This paper cites RULER: What's the Real Context Size of Your Long-Context Language Models?.

Infinite Video Understanding RULER: What's the Real Context Size of Your Long-Context Language Models?

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-06T18:10:12.307299Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:10:12.307299Z digest=sha256:ffa775454cbdaa647282fa29e4e0f13aa150d1c58b5701d314b180daa5db39fb

Observation 9110064f-0fc6-4002-a20d-c5fe49e3a397 · outbound

This paper cites Visual Representation Learning with Stochastic Frame Prediction.

Infinite Video Understanding Visual Representation Learning with Stochastic Frame Prediction

Reference 15

Resolution
verified exact
local_arxiv, observed 2026-08-06T18:10:16.421305Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T18:10:12.401921Z digest=sha256:5f95ac49d37f8039cef8345d0a24a3484c5296df33a5f9c4672cad2633476d4a

Observation 2b5d77ea-50d5-4d97-8f44-c7bd15d9693d · outbound

This paper cites Search-R1: Training LLMs to Reason and Leverage Search Engines with Reinforcement Learning.

Infinite Video Understanding Search-R1: Training LLMs to Reason and Leverage Search Engines with Reinforcement Learning

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-06T18:10:12.472087Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:10:12.472087Z digest=sha256:a97fb0f2e8c2f59943f1c64473d27c8b423ea0406832d444b8c18cc14c0952a4

Observation 0b728617-5100-420c-887d-8c52582f04f6 · outbound

This paper cites SNeRV: Spectra-preserving Neural Representation for Video.

Infinite Video Understanding SNeRV: Spectra-preserving Neural Representation for Video

Reference 17

Resolution
verified exact
local_arxiv, observed 2026-08-06T18:10:16.964469Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T18:10:12.519667Z digest=sha256:5d6a3b57fb2549ba45b2d444a117946bdf9471814c70449ee786c4977e99fdbd

Observation fb2b3eb7-0cd3-4be0-bd16-9ca2d015fd81 · outbound

This paper cites an unresolved cited work.

Infinite Video Understanding Unresolved cited work

Reference 18

Resolution
unresolved
raw_fallback, observed 2026-08-06T18:10:17.987768Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T18:10:12.559183Z digest=sha256:ad7f734ef199b1abd435a1cc051a812447d91cbd730ef93d1192732256439137

Observation 7b403170-dbfe-452a-bdbd-19522ed26d68 · outbound

This paper cites an unresolved cited work.

Infinite Video Understanding Unresolved cited work

Reference 19

Resolution
verified exact
doi, observed 2026-08-06T18:10:16.232019Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T18:10:12.643858Z digest=sha256:50ddf734ad29fb9849ba13b18ea0258469b6cdc98f5fbad6424b76e1e7a006fd

Observation 43e3fa0f-361d-4cba-9e31-bc1fd9e80412 · outbound

This paper cites START: Self-taught Reasoner with Tools.

Infinite Video Understanding START: Self-taught Reasoner with Tools

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-06T18:10:12.598324Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:10:12.598324Z digest=sha256:0c567903d5d2a5da741aae1f7511ffe8efcd3de4fcab14d6c309c417c3a76e00

Observation 97ac225c-a44b-48e7-b367-d1e266ecfdf5 · outbound

This paper cites VideoChat-Flash: Hierarchical Compression for Long-Context Video Modeling.

Infinite Video Understanding VideoChat-Flash: Hierarchical Compression for Long-Context Video Modeling

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-06T18:10:12.714751Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:10:12.714751Z digest=sha256:b0aa9fd9525521d5c188cb3a255d4b8ed97fc219dd218f188b90aad2dd1a1143

Observation cc6d9d5e-c795-46dd-b38d-5aa40d4a9983 · outbound

This paper cites MVBench: A Comprehensive Multi-modal Video Understanding Benchmark.

Infinite Video Understanding MVBench: A Comprehensive Multi-modal Video Understanding Benchmark

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-06T18:10:12.662954Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:10:12.662954Z digest=sha256:cdad7215b80e77af7c29e11e662d6f4d8225b1ee6f4abeb5e2eda88abe2ddde1

Observation 3633c010-534b-47ba-b99c-22c9cd1f6f03 · outbound

This paper cites Video-XL-Pro: Reconstructive Token Compression for Extremely Long Video Understanding.

Infinite Video Understanding Video-XL-Pro: Reconstructive Token Compression for Extremely Long Video Understanding

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-06T18:10:12.773039Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:10:12.773039Z digest=sha256:45597df008272bff69f30c2ee1020cf856542d92cb319ff16d73447aa3c02ad8

Observation 448a4907-dd0e-4419-9ed5-fba4759143ce · outbound

This paper cites World Model on Million-Length Video And Language With Blockwise RingAttention.

Infinite Video Understanding World Model on Million-Length Video And Language With Blockwise RingAttention

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-06T18:10:12.748486Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:10:12.748486Z digest=sha256:c268cdde58e5433a0a0d243072a078bb79b775c7d185dceef372db75bd1aa6d3

Observation e0001093-1c8d-428c-8368-db9820138888 · outbound

This paper cites EgoSchema: A Diagnostic Benchmark for Very Long-form Video Language Understanding.

Infinite Video Understanding EgoSchema: A Diagnostic Benchmark for Very Long-form Video Language Understanding

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-06T18:10:12.839920Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:10:12.839920Z digest=sha256:bb56fa58581712d9f906736f74527dc511e28c6027f6b5ff0e7fbfddf6e8df83

Observation 6dcac48e-b7e6-4df2-9834-febd2b9658ff · outbound

This paper cites an unresolved cited work.

Infinite Video Understanding Unresolved cited work

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-06T18:10:12.786822Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:10:12.786822Z digest=sha256:8e000a89f8db19e1c558a27b7e019097a47c934cb64bd4e0ade664357104eaef

Observation 01d0693d-1a66-4237-85d4-664f63fa52b0 · outbound

This paper cites Cosmos World Foundation Model Platform for Physical AI.

Infinite Video Understanding Cosmos World Foundation Model Platform for Physical AI

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-06T18:10:12.942821Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:10:12.942821Z digest=sha256:3baa04cf61c97413bd60c9a4ee8a2281829a38295ecb6ab43a6f93141295c9b3

Observation 20a4e7fb-271f-4a54-b77d-5c48378e0b2e · outbound

This paper cites an unresolved cited work.

Infinite Video Understanding Unresolved cited work

Reference 28

Resolution
unresolved
raw_fallback, observed 2026-08-06T18:10:17.807895Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T18:10:12.877764Z digest=sha256:ddf7cfa69ff9270aa087119308a2848198975dc7f520c7baa24b4c1f3159d8a6

Observation af469056-2597-4916-802c-3e9a4e6eb44a · outbound

This paper cites Video-XL-2: Towards Very Long-Video Understanding Through Task-Aware KV Sparsification.

Infinite Video Understanding Video-XL-2: Towards Very Long-Video Understanding Through Task-Aware KV Sparsification

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-06T18:10:13.072860Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:10:13.072860Z digest=sha256:f04ac4a4a4c74773c95a48c4490ed9d8d7d72cce3c163f3e45725e3228e1b7fa

Observation e76d6843-4626-47ad-ba97-c6f7c48102fe · outbound

This paper cites Streaming Long Video Understanding with Large Language Models.

Infinite Video Understanding Streaming Long Video Understanding with Large Language Models

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-06T18:10:13.022336Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:10:13.022336Z digest=sha256:e8dee5bf9b5f2838de4f5fbfb6298fe7f0c20a88078381cf2a4dedaabd22d0c5

Observation e32d3cd9-dbef-4bfe-8da5-bd1a7f618ab4 · outbound

This paper cites Learning Transferable Visual Models From Natural Language Supervision.

Infinite Video Understanding Learning Transferable Visual Models From Natural Language Supervision

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-06T18:10:13.140709Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:10:13.140709Z digest=sha256:5106f976cc4f69cd9cc8423fe813f5c7e326d89d8626414b148d3214631e90eb

Observation 2aaa3216-7106-47ce-b2e1-40f190617bbc · outbound

This paper cites Recent Advances of Continual Learning in Computer Vision: An Overview.

Infinite Video Understanding Recent Advances of Continual Learning in Computer Vision: An Overview

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-06T18:10:13.104601Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:10:13.104601Z digest=sha256:52ae2d6213e53b151549bb4225a8790775fe31065b8482f00acdcc1ab969096e

Observation ea010615-2ce1-47d8-9be6-4ea8823abe73 · outbound

This paper cites LongVU: Spatiotemporal Adaptive Compression for Long Video-Language Understanding.

Infinite Video Understanding LongVU: Spatiotemporal Adaptive Compression for Long Video-Language Understanding

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-06T18:10:13.260398Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:10:13.260398Z digest=sha256:dd10d7d63245b7f6173c2845b42616356c3ef756cfe54014c4c1d698a35ca135

Observation edc308c6-cf1f-431d-b0d1-b121878aafd8 · outbound

This paper cites $\infty$-Video: A Training-Free Approach to Long Video Understanding via Continuous-Time Memory Consolidation.

Infinite Video Understanding $\infty$-Video: A Training-Free Approach to Long Video Understanding via Continuous-Time Memory Consolidation

Reference 34

Resolution
metadata mismatch
local_arxiv, observed 2026-08-06T18:10:15.807015Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T18:10:13.200174Z digest=sha256:636470456a367be9792aaf57a1d59c72f1f7a6a5f286f6c3314b867d610de2b8

Observation e76d195d-b50d-4df2-969f-288c3370d330 · outbound

This paper cites an unresolved cited work.

Infinite Video Understanding Unresolved cited work

Reference 35

Resolution
unresolved
raw_fallback, observed 2026-08-06T18:10:17.698230Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T18:10:13.328713Z digest=sha256:1bebcf247b25a70f894b97f4ef5b54d8ac08788b80b519171ba6315c4848a1c5

Observation c2df738d-2eb6-4bef-b7fd-a788422984d2 · outbound

This paper cites Video-XL: Extra-Long Vision Language Model for Hour-Scale Video Understanding.

Infinite Video Understanding Video-XL: Extra-Long Vision Language Model for Hour-Scale Video Understanding

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-06T18:10:13.296306Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:10:13.296306Z digest=sha256:3bf44d118a916e86401405413b380bed12d382ba31993d6b0eee8b2a6f63fff6

Observation 6daaacba-cd5e-4c86-b121-09b740e1fcb6 · outbound

This paper cites Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context.

Infinite Video Understanding Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-06T18:10:13.446720Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:10:13.446720Z digest=sha256:b2440b0e9e74971d7e7156ee8f17e678d5510d111b96328e25b6e8a95cea4b35

Observation 102ccd94-8ead-4ada-9021-3f44632e9396 · outbound

This paper cites MovieChat: From Dense Token to Sparse Memory for Long Video Understanding.

Infinite Video Understanding MovieChat: From Dense Token to Sparse Memory for Long Video Understanding

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-06T18:10:13.357356Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:10:13.357356Z digest=sha256:674a6e20d21cdee7dc2068d25cd638c7107c6bb83c576a6ce76cea31142c3c1c

Observation 1d59b233-d2bb-4843-81a8-9b4565f0816f · outbound

This paper cites SigLIP 2: Multilingual Vision-Language Encoders with Improved Semantic Understanding, Localization, and Dense Features.

Infinite Video Understanding SigLIP 2: Multilingual Vision-Language Encoders with Improved Semantic Understanding, Localization, and Dense Features

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-06T18:10:13.651494Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:10:13.651494Z digest=sha256:341871916ca4930151469218b156cee0a8dbc52daccd0eacd2a08b81ca6eac97

Observation 8627080a-8146-4fd5-893f-d0b40c9a5e1c · outbound

This paper cites VideoMAE: Masked Autoencoders are Data-Efficient Learners for Self-Supervised Video Pre-Training.

Infinite Video Understanding VideoMAE: Masked Autoencoders are Data-Efficient Learners for Self-Supervised Video Pre-Training

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-06T18:10:13.574181Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:10:13.574181Z digest=sha256:a5c7ea3ab6d004ee82c33d775dea402e57035a20747042da53ef1cb9b2b72b4c

Observation 5c8a1017-73f4-4f2c-9e4d-a780ab36782d · outbound

This paper cites A Comprehensive Survey of Continual Learning: Theory, Method and Application.

Infinite Video Understanding A Comprehensive Survey of Continual Learning: Theory, Method and Application

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-06T18:10:13.775787Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:10:13.775787Z digest=sha256:4d947d47e4bdc265ebbd98659fbf35fe92af187a38843c30bcacbb22720bacaf

Observation 06bfa572-b497-4805-9795-6453d48c3512 · outbound

This paper cites VideoMAE V2: Scaling Video Masked Autoencoders with Dual Masking.

Infinite Video Understanding VideoMAE V2: Scaling Video Masked Autoencoders with Dual Masking

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-06T18:10:13.711528Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:10:13.711528Z digest=sha256:e3f9355b8058fe54ae4975aa2554c1409ea1433d5b0dd023e421c888e32d4e84

Observation ef828168-849f-4259-8d45-f7f2490571e5 · outbound

This paper cites VidTwin: Video VAE with Decoupled Structure and Dynamics.

Infinite Video Understanding VidTwin: Video VAE with Decoupled Structure and Dynamics

Reference 43

Resolution
verified exact
local_arxiv, observed 2026-08-06T18:10:15.572551Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T18:10:14.101777Z digest=sha256:d0d5a1760807c5c059d7efc49df55eb9fa2adec7741a686855c803e3348fcb6a

Observation a69f610b-c659-49d6-9892-4f1bf7f702e1 · outbound

This paper cites an unresolved cited work.

Infinite Video Understanding Unresolved cited work

Reference 44

Resolution
unresolved
raw_fallback, observed 2026-08-06T18:10:17.529367Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T18:10:13.937532Z digest=sha256:8658bde5fb91dfc83ced2f1b48294bb923bd34cd00186e1d096a176564f03b20

Observation f519b7f8-72c3-4347-acda-48ca2a008599 · outbound

This paper cites VideoRoPE: What Makes for Good Video Rotary Position Embedding?.

Infinite Video Understanding VideoRoPE: What Makes for Good Video Rotary Position Embedding?

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-06T18:10:14.257348Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:10:14.257348Z digest=sha256:03cdbac653d4efeecb475e8ab066563daab05d65178f56e955e42db9a5c0a344

Observation a30160aa-9053-40af-9c42-bef00968d3a7 · outbound

This paper cites LongVLM: Efficient Long Video Understanding via Large Language Models.

Infinite Video Understanding LongVLM: Efficient Long Video Understanding via Large Language Models

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-06T18:10:14.298291Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:10:14.298291Z digest=sha256:e0a611ebc50a35611633d9b9f57587dee8510c5b84075957e7bfde1dd976eb33

Observation 5d65b882-b812-4eb2-9817-d7922d9a8c0b · outbound

This paper cites VideoTree: Adaptive Tree-based Video Representation for LLM Reasoning on Long Videos.

Infinite Video Understanding VideoTree: Adaptive Tree-based Video Representation for LLM Reasoning on Long Videos

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-06T18:10:14.223529Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:10:14.223529Z digest=sha256:f6313190c7e67be09364397b8a5b56dd557afbac39b9e9c888a177b8aee2584a

Observation 88a9e838-e338-44d7-b0b5-2cabd1433fbc · outbound

This paper cites Agentic Reasoning: A Streamlined Framework for Enhancing LLM Reasoning with Agentic Tools.

Infinite Video Understanding Agentic Reasoning: A Streamlined Framework for Enhancing LLM Reasoning with Agentic Tools

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-06T18:10:14.365798Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:10:14.365798Z digest=sha256:615302c5f5c0cc71dd667370d6237b95de7b3bc22f7c33c086f8a5128e97f8f0

Observation aea3764e-83dd-40b8-9999-1eb2fca7070d · outbound

This paper cites LongViTU: Instruction Tuning for Long-Form Video Understanding.

Infinite Video Understanding LongViTU: Instruction Tuning for Long-Form Video Understanding

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-06T18:10:14.416661Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:10:14.416661Z digest=sha256:cad9538f3809abb717b76e96c99a9870375803706f6aa9eeb3d64763ca9b97ca

Observation b6143b35-bfaa-484d-9084-4430ec886d4b · outbound

This paper cites LongVideoBench: A Benchmark for Long-context Interleaved Video-Language Understanding.

Infinite Video Understanding LongVideoBench: A Benchmark for Long-context Interleaved Video-Language Understanding

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-06T18:10:14.340618Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:10:14.340618Z digest=sha256:51ea3c187e5af9bae95d3055f41bb74df75e1a0b93301e75250320de9047b626

Observation 9dda5626-d89f-474a-92d8-7620a5c62f59 · outbound

This paper cites MAGVIT: Masked Generative Video Transformer.

Infinite Video Understanding MAGVIT: Masked Generative Video Transformer

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-06T18:10:14.507734Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:10:14.507734Z digest=sha256:74fd391c9cd659031b669f83bac439d80b9c51a08822bb2fefcf8c51e53cb42c

Observation a4883221-886c-418b-9b68-445e2b4784e5 · outbound

This paper cites Language Model Beats Diffusion -- Tokenizer is Key to Visual Generation.

Infinite Video Understanding Language Model Beats Diffusion -- Tokenizer is Key to Visual Generation

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-06T18:10:14.568887Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:10:14.568887Z digest=sha256:91d102493132e7889562566186a7b345283a243a1be48dbbf7daa44c23dc430f

Observation 5ab42a16-8fdb-4d71-84eb-02b6ab549f18 · outbound

This paper cites an unresolved cited work.

Infinite Video Understanding Unresolved cited work

Reference 53

Resolution
verified exact
doi, observed 2026-08-06T18:10:15.381307Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T18:10:14.459822Z digest=sha256:6bf77abfd65eb3401704590e092a3aee5766424f7e71b5eda45309c15760e6a9

Observation ef731b3e-8c43-42d9-aa5e-109c6274611b · outbound

This paper cites an unresolved cited work.

Infinite Video Understanding Unresolved cited work

Reference 54

Resolution
verified exact
doi, observed 2026-08-06T18:10:15.203453Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T18:10:14.678140Z digest=sha256:46402dc1711727dce400f9232c4f1eb5e062c3432ca8e290a93d5c9acbaf7815

Observation 0c444863-559e-43e1-8eb9-3f9de9561d65 · outbound

This paper cites Gundavarapu, Liangzhe Yuan, Hao Zhou, Shen Yan, Jennifer J.

Infinite Video Understanding Gundavarapu, Liangzhe Yuan, Hao Zhou, Shen Yan, Jennifer J

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:10:17.373787Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T18:10:14.711077Z digest=sha256:24843e40163a24f318c0789fdc7d822a3f64f959a0a32bdd6920bb1f5be80274

Observation 631dce2d-c1ed-46e3-a663-7fb3e379b154 · outbound

This paper cites Sigmoid Loss for Language Image Pre-Training.

Infinite Video Understanding Sigmoid Loss for Language Image Pre-Training

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-06T18:10:14.635404Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:10:14.635404Z digest=sha256:70749e3136eb74c90fa005ea0a36366bd4d40fd8a3b575a6b11270413778a7eb

Observation bc707462-8bba-42d4-b63e-912ab0c3dfc9 · outbound

This paper cites Towards Lifelong Learning of Large Language Models: A Survey.

Infinite Video Understanding Towards Lifelong Learning of Large Language Models: A Survey

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-06T18:10:14.841489Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:10:14.841489Z digest=sha256:e2077446b68ac0ca49ce45694ef6acb8aa748815f9235533b37c0982be3ffeaa

Observation 12c78463-a446-463e-8f21-9ae51e99a337 · outbound

This paper cites MLVU: Benchmarking Multi-task Long Video Understanding.

Infinite Video Understanding MLVU: Benchmarking Multi-task Long Video Understanding

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-06T18:10:14.864857Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:10:14.864857Z digest=sha256:5c1f5623d384cb15203a0c744bb698e9f4b571d7d46ff458ffb9cd5d20a801b9

Observation 398801a5-873c-4169-abf0-73e544e5bb62 · outbound

This paper cites VideoPrism: A Foundational Visual Encoder for Video Understanding.

Infinite Video Understanding VideoPrism: A Foundational Visual Encoder for Video Understanding

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-06T18:10:14.749813Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:10:14.749813Z digest=sha256:4794b8fa152cfddc715e4f3081d849e135726a21ad5c69873972476ccb9cfc7f

Observation 5da26043-a850-4fdd-8944-8fc06fdf7b6a · outbound

This paper cites Needle In A Video Haystack: A Scalable Synthetic Evaluator for Video MLLMs.

Infinite Video Understanding Needle In A Video Haystack: A Scalable Synthetic Evaluator for Video MLLMs

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-06T18:10:14.777464Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:10:14.777464Z digest=sha256:9bc423b59c70b4a53e669ff5823f02f7a164e55a6e8fcd8414e83bd8a68a39eb

Observation 168f0a8e-2993-48df-b4f0-266fb42fca75 · outbound

This paper cites an unresolved cited work.

Infinite Video Understanding Unresolved cited work

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-06T18:10:14.887539Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:10:14.887539Z digest=sha256:a35a8a73ea7cf08680c2bb4b04b62d1bb23c0a045cce9f9a31861ec1fa668119

Observation ff6bf875-6dfb-403e-aa65-9bd3504869d2 · outbound

This paper cites From Seconds to Hours: Reviewing MultiModal Large Language Models on Comprehensive Long Video Understanding.

Infinite Video Understanding From Seconds to Hours: Reviewing MultiModal Large Language Models on Comprehensive Long Video Understanding

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-06T18:10:14.919272Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:10:14.919272Z digest=sha256:e6d08d90be5c5970059d10c447281edeb1dbe2d305e8a70870892bb8f564ff52

Observation 004a05d9-6451-40e2-afb9-1e60ec5aa9fb · outbound

This paper cites LVBench: An Extreme Long Video Understanding Benchmark.

Infinite Video Understanding LVBench: An Extreme Long Video Understanding Benchmark

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-06T18:10:13.986726Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:10:13.986726Z digest=sha256:0ba0e42c091cb3ea3e9836fbc0ef3c009db27af11f5a6b4086150fb29f032d27

Observation 5dfc5ac4-c02b-4cfa-9645-4613c33f3a71 · outbound

This paper cites Video-R1: Reinforcing Video Reasoning in MLLMs.

Infinite Video Understanding Video-R1: Reinforcing Video Reasoning in MLLMs

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-06T18:10:11.970598Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:10:11.970598Z digest=sha256:152b03e71bd267b814f31e22b95d28254452980d95534a846767ba8bd7dede57

Pith citing papers

No inbound Pith citation observations are available.