Pith. sign in

Paper Citation Record · LEDGER

Infinite Video Understanding

As of 23 August 2026, this Paper Citation Record lists 64 of 64 outbound references and 0 inbound Pith citation observations for arXiv:2507.09068.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.09068 v2

Coverage vector

measured 64 of 64 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T18:10:14.919272Z

measured 64 of 64 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-23T06:30:58.430688+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

64 of 64 outbound references displayed

  • verified exact8
  • verified fuzzy1
  • unresolved54
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch1

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation ff0a1ede-3856-4a74-ad1e-e3544cd0a8b4 · outbound

This paper cites CrossVideoMAE: Self-Supervised Image-Video Representation Learning with Masked Autoencoders.

Infinite Video Understanding CrossVideoMAE: Self-Supervised Image-Video Representation Learning with Masked Autoencoders

Reference 1

Resolution
verified exact
local_arxiv, observed 2026-08-06T18:10:16.731929Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-06T18:09:18.061419Z digest=sha256:629e76022bfca2ae7359ce9c5dcd6dd4f75f516cbd2c3247a13fa18ce0ae0b7a

Observation c25e915d-4ba9-493f-847b-51365e8f7fda · outbound

This paper cites an unresolved cited work.

Infinite Video Understanding Unresolved cited work

Reference 2

Resolution
unresolved
raw_fallback, observed 2026-08-06T18:10:18.299149Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-06T18:09:18.114356Z digest=sha256:66340b00413f0da1c22939d6267d230eb90d7fc582fed43affaaac985bbba542

Observation 01bcc5ab-e0c7-4408-a60e-48774759a8fe · outbound

This paper cites HourVideo: 1-Hour Video-Language Understanding.

Infinite Video Understanding HourVideo: 1-Hour Video-Language Understanding

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-06T18:09:18.211447Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:09:18.211447Z digest=sha256:2d24202136ca44c0f79432fbe0aa29fe50295862ea109b897d436b9a5c25db57

Observation 339e7bb3-dc67-4a6f-8af4-aff75a9dfa3c · outbound

This paper cites LongVILA: Scaling Long-Context Visual Language Models for Long Videos.

Infinite Video Understanding LongVILA: Scaling Long-Context Visual Language Models for Long Videos

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-06T18:09:18.313230Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:09:18.313230Z digest=sha256:5f29ea66e1797de7d741c9606c5f813eac77978f6af66640e22dbfe5c7254be7

Observation 809b2dcf-8bb2-479b-964a-3488ccdb6144 · outbound

This paper cites Scaling Video-Language Models to 10K Frames via Hierarchical Differential Distillation.

Infinite Video Understanding Scaling Video-Language Models to 10K Frames via Hierarchical Differential Distillation

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-06T18:09:18.568750Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:09:18.568750Z digest=sha256:9ec0d55512a798ac1b5e2ba1ee872b69a1da01b6a64da4d3dd29edc84019b2b4

Observation 8d053b18-1eb2-430a-9e88-2a6c782636e8 · outbound

This paper cites Video-Holmes: Can MLLM Think Like Holmes for Complex Video Reasoning?.

Infinite Video Understanding Video-Holmes: Can MLLM Think Like Holmes for Complex Video Reasoning?

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-06T18:09:18.954541Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:09:18.954541Z digest=sha256:0e57651339d9e839cc3630ad9aea2ddd22926ed79a085ea9e56ab49f3bc987fd

Observation dbef02ca-88ff-488e-a914-02db772cb241 · outbound

This paper cites Zero-Shot Video Question Answering with Procedural Programs.

Infinite Video Understanding Zero-Shot Video Question Answering with Procedural Programs

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-06T18:09:19.409539Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:09:19.409539Z digest=sha256:2aad319af2c3e7c20b6215753bd298e0c2b3cf6caf9aa76a2702f6e67f956a88

Observation 68494767-b952-4b9b-bba1-bbddf06ed232 · outbound

This paper cites Project Aria: A New Tool for Egocentric Multi-Modal AI Research.

Infinite Video Understanding Project Aria: A New Tool for Egocentric Multi-Modal AI Research

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-06T18:10:11.725676Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:10:11.725676Z digest=sha256:ef6fe6a9c4e41871e9dc8acbf5a289dea5ed01d394e059a130e80ad4c6e989d6

Observation 545305b2-f2bb-4383-9be8-639125a12ebf · outbound

This paper cites Efficient Image Pre-Training with Siamese Cropped Masked Autoencoders.

Infinite Video Understanding Efficient Image Pre-Training with Siamese Cropped Masked Autoencoders

Reference 9

Resolution
verified exact
local_arxiv, observed 2026-08-06T18:10:17.212457Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-06T18:10:11.854309Z digest=sha256:4cf186434837080b6f0eaff1e70a9c6e3d12187aef1205282b2ff6dabb8ce8ad

Observation 316ad4f0-026d-4451-84aa-0b1e8874df9e · outbound

This paper cites an unresolved cited work.

Infinite Video Understanding Unresolved cited work

Reference 10

Resolution
unresolved
raw_fallback, observed 2026-08-06T18:10:18.130805Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-06T18:10:11.899351Z digest=sha256:8c5a5c02d377e88604e5b4245795abea660c59f678f654c9ce3b9375d0af032e

Observation 276abd28-1608-4fa1-8d37-0830654d553a · outbound

This paper cites Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis.

Infinite Video Understanding Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-06T18:10:12.056934Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:10:12.056934Z digest=sha256:c24ff0d8272e2d7edc3205a52adf49ff6b93325e693018f8e7def0601eeb960f

Observation 4cbccd92-ffbe-4212-98e0-67fa7f83e1a5 · outbound

This paper cites The Ethics of Advanced AI Assistants.

Infinite Video Understanding The Ethics of Advanced AI Assistants

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-06T18:10:12.125589Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:10:12.125589Z digest=sha256:3f67b47678eeee161deda5f7adfa6b84c0463986216f8c41a4c909fccdaeca57

Observation 963c17ec-82c8-4343-91e9-d85d06aa17ca · outbound

This paper cites Siamese Masked Autoencoders.

Infinite Video Understanding Siamese Masked Autoencoders

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-06T18:10:12.195376Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:10:12.195376Z digest=sha256:2a6835d2f02b15964a0898ba2ab1c111671f585a2eb5c91730037cb3c4c953a7

Observation 12718065-f4c4-4ac8-9980-1d22368d0953 · outbound

This paper cites RULER: What's the Real Context Size of Your Long-Context Language Models?.

Infinite Video Understanding RULER: What's the Real Context Size of Your Long-Context Language Models?

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-06T18:10:12.307299Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:10:12.307299Z digest=sha256:ecd994af1cccd836d3ee9656aa4664b0919212d73bf6ea8c25ed14d1d99999cc

Observation 9110064f-0fc6-4002-a20d-c5fe49e3a397 · outbound

This paper cites Visual Representation Learning with Stochastic Frame Prediction.

Infinite Video Understanding Visual Representation Learning with Stochastic Frame Prediction

Reference 15

Resolution
verified exact
local_arxiv, observed 2026-08-06T18:10:16.421305Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-06T18:10:12.401921Z digest=sha256:e3401a9ef1b72d9d26a63fef0bf3c3e57fda2e45f9f74728a8e5591d531dbf6f

Observation 2b5d77ea-50d5-4d97-8f44-c7bd15d9693d · outbound

This paper cites Search-R1: Training LLMs to Reason and Leverage Search Engines with Reinforcement Learning.

Infinite Video Understanding Search-R1: Training LLMs to Reason and Leverage Search Engines with Reinforcement Learning

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-06T18:10:12.472087Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:10:12.472087Z digest=sha256:e2c401ee34fedb010f11205e231f95b68930d5cd2213840ba1c5b1f6b9e3a18b

Observation 0b728617-5100-420c-887d-8c52582f04f6 · outbound

This paper cites SNeRV: Spectra-preserving Neural Representation for Video.

Infinite Video Understanding SNeRV: Spectra-preserving Neural Representation for Video

Reference 17

Resolution
verified exact
local_arxiv, observed 2026-08-06T18:10:16.964469Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-06T18:10:12.519667Z digest=sha256:7912c652d96b8ed9d719871dc2c4df1a792b117e1c1d76d9a263b83d8169ec3f

Observation fb2b3eb7-0cd3-4be0-bd16-9ca2d015fd81 · outbound

This paper cites an unresolved cited work.

Infinite Video Understanding Unresolved cited work

Reference 18

Resolution
unresolved
raw_fallback, observed 2026-08-06T18:10:17.987768Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-06T18:10:12.559183Z digest=sha256:f5bbbc83fa8bf90c9525843e8e41e13455f9b9807f7e2ee04c7423f02c9703f9

Observation 7b403170-dbfe-452a-bdbd-19522ed26d68 · outbound

This paper cites an unresolved cited work.

Infinite Video Understanding Unresolved cited work

Reference 19

Resolution
verified exact
doi, observed 2026-08-06T18:10:16.232019Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-06T18:10:12.643858Z digest=sha256:19574b78252659adff78e0ec418d608194e49ba24cdab47a316fb0865f7df1ab

Observation 43e3fa0f-361d-4cba-9e31-bc1fd9e80412 · outbound

This paper cites START: Self-taught Reasoner with Tools.

Infinite Video Understanding START: Self-taught Reasoner with Tools

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-06T18:10:12.598324Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:10:12.598324Z digest=sha256:5727f84120c5e48b640519b76b534bd2776677e8a3f1a1ad72dd4a36b1caa522

Observation 97ac225c-a44b-48e7-b367-d1e266ecfdf5 · outbound

This paper cites VideoChat-Flash: Hierarchical Compression for Long-Context Video Modeling.

Infinite Video Understanding VideoChat-Flash: Hierarchical Compression for Long-Context Video Modeling

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-06T18:10:12.714751Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:10:12.714751Z digest=sha256:ba11b8a26ecec3415a8e4bdf28362b028d6f6b350d3d1eb8d58c0cb2d9974099

Observation cc6d9d5e-c795-46dd-b38d-5aa40d4a9983 · outbound

This paper cites MVBench: A Comprehensive Multi-modal Video Understanding Benchmark.

Infinite Video Understanding MVBench: A Comprehensive Multi-modal Video Understanding Benchmark

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-06T18:10:12.662954Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:10:12.662954Z digest=sha256:9bf813f8daa38223eb6e7d8c5f218be225875dc43429b5c30005097746863b56

Observation 3633c010-534b-47ba-b99c-22c9cd1f6f03 · outbound

This paper cites Video-XL-Pro: Reconstructive Token Compression for Extremely Long Video Understanding.

Infinite Video Understanding Video-XL-Pro: Reconstructive Token Compression for Extremely Long Video Understanding

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-06T18:10:12.773039Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:10:12.773039Z digest=sha256:1801785729600be80740cc8dbb44887550d20507fe0dff62a6dd403a0377948a

Observation 448a4907-dd0e-4419-9ed5-fba4759143ce · outbound

This paper cites World Model on Million-Length Video And Language With Blockwise RingAttention.

Infinite Video Understanding World Model on Million-Length Video And Language With Blockwise RingAttention

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-06T18:10:12.748486Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:10:12.748486Z digest=sha256:ef23d6f05531dc5d468a774e1c43b0ddef3364564193830163d2a7cbc5ca3aac

Observation e0001093-1c8d-428c-8368-db9820138888 · outbound

This paper cites EgoSchema: A Diagnostic Benchmark for Very Long-form Video Language Understanding.

Infinite Video Understanding EgoSchema: A Diagnostic Benchmark for Very Long-form Video Language Understanding

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-06T18:10:12.839920Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:10:12.839920Z digest=sha256:a764cfb1a2a91b45323cef4ccfbd998e6ff67afab7d14048c23d31d22155e0b2

Observation 6dcac48e-b7e6-4df2-9834-febd2b9658ff · outbound

This paper cites an unresolved cited work.

Infinite Video Understanding Unresolved cited work

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-06T18:10:12.786822Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:10:12.786822Z digest=sha256:fda05562c460160809b08a7edf3633fa204351364b8a76855aca3752b0698200

Observation 01d0693d-1a66-4237-85d4-664f63fa52b0 · outbound

This paper cites Cosmos World Foundation Model Platform for Physical AI.

Infinite Video Understanding Cosmos World Foundation Model Platform for Physical AI

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-06T18:10:12.942821Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:10:12.942821Z digest=sha256:150a26855342b55478c12f68a154f9f7ba378ff45bf3209f37f084ae854dbe27

Observation 20a4e7fb-271f-4a54-b77d-5c48378e0b2e · outbound

This paper cites an unresolved cited work.

Infinite Video Understanding Unresolved cited work

Reference 28

Resolution
unresolved
raw_fallback, observed 2026-08-06T18:10:17.807895Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-06T18:10:12.877764Z digest=sha256:c0bf691dd8ebded60b488a5e0880759ad4e5bc4b18a618f809bc2418272a12f3

Observation af469056-2597-4916-802c-3e9a4e6eb44a · outbound

This paper cites Video-XL-2: Towards Very Long-Video Understanding Through Task-Aware KV Sparsification.

Infinite Video Understanding Video-XL-2: Towards Very Long-Video Understanding Through Task-Aware KV Sparsification

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-06T18:10:13.072860Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:10:13.072860Z digest=sha256:954c8b4aa40943c9abd3cf9e772c9ed1e58bf92fc5b1979e1f2a04b5e385643c

Observation e76d6843-4626-47ad-ba97-c6f7c48102fe · outbound

This paper cites Streaming Long Video Understanding with Large Language Models.

Infinite Video Understanding Streaming Long Video Understanding with Large Language Models

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-06T18:10:13.022336Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:10:13.022336Z digest=sha256:b98b17322d1a62d9b0ec937c6ffbc7d32d14ef6821fb86969454f96d85631f9f

Observation e32d3cd9-dbef-4bfe-8da5-bd1a7f618ab4 · outbound

This paper cites Learning Transferable Visual Models From Natural Language Supervision.

Infinite Video Understanding Learning Transferable Visual Models From Natural Language Supervision

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-06T18:10:13.140709Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:10:13.140709Z digest=sha256:f439c06be1490b5bc39b02dae3dcdd1b934dd89ea4cb74d0d380b7f1aca515fc

Observation 2aaa3216-7106-47ce-b2e1-40f190617bbc · outbound

This paper cites Recent Advances of Continual Learning in Computer Vision: An Overview.

Infinite Video Understanding Recent Advances of Continual Learning in Computer Vision: An Overview

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-06T18:10:13.104601Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:10:13.104601Z digest=sha256:36707050604c0ec168056a3d608ffcc296783c0261b2f07d009b19a315aafb82

Observation ea010615-2ce1-47d8-9be6-4ea8823abe73 · outbound

This paper cites LongVU: Spatiotemporal Adaptive Compression for Long Video-Language Understanding.

Infinite Video Understanding LongVU: Spatiotemporal Adaptive Compression for Long Video-Language Understanding

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-06T18:10:13.260398Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:10:13.260398Z digest=sha256:fa7123b181545fadba46c54acf63d307ebde57436bc9a0a9ae9d004a9aa67b34

Observation edc308c6-cf1f-431d-b0d1-b121878aafd8 · outbound

This paper cites $\infty$-Video: A Training-Free Approach to Long Video Understanding via Continuous-Time Memory Consolidation.

Infinite Video Understanding $\infty$-Video: A Training-Free Approach to Long Video Understanding via Continuous-Time Memory Consolidation

Reference 34

Resolution
metadata mismatch
local_arxiv, observed 2026-08-06T18:10:15.807015Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-06T18:10:13.200174Z digest=sha256:a1154e1f817fce9adb24c040b9a38dfe0738b64d24a9d4aad9b209ae1f3ae8c2

Observation e76d195d-b50d-4df2-969f-288c3370d330 · outbound

This paper cites an unresolved cited work.

Infinite Video Understanding Unresolved cited work

Reference 35

Resolution
unresolved
raw_fallback, observed 2026-08-06T18:10:17.698230Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-06T18:10:13.328713Z digest=sha256:abf2c59e8e20ca594541fa2fb0d541fb345086feaa1a2e811eaeb06866ebd451

Observation c2df738d-2eb6-4bef-b7fd-a788422984d2 · outbound

This paper cites Video-XL: Extra-Long Vision Language Model for Hour-Scale Video Understanding.

Infinite Video Understanding Video-XL: Extra-Long Vision Language Model for Hour-Scale Video Understanding

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-06T18:10:13.296306Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:10:13.296306Z digest=sha256:a177cc5227a3c3a498bd7a8f5f3ea1218b2a209f1544091b3dc42ed29f1411a3

Observation 6daaacba-cd5e-4c86-b121-09b740e1fcb6 · outbound

This paper cites Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context.

Infinite Video Understanding Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-06T18:10:13.446720Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:10:13.446720Z digest=sha256:c4e16bd9e84f459c6edb00c0338bf1750fed64339d73237f6ea29d08d202c7d0

Observation 102ccd94-8ead-4ada-9021-3f44632e9396 · outbound

This paper cites MovieChat: From Dense Token to Sparse Memory for Long Video Understanding.

Infinite Video Understanding MovieChat: From Dense Token to Sparse Memory for Long Video Understanding

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-06T18:10:13.357356Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:10:13.357356Z digest=sha256:29a3dc32909f8deacfd6299bf935b8c301784a24ed8f70833cee74f8026cd3c0

Observation 1d59b233-d2bb-4843-81a8-9b4565f0816f · outbound

This paper cites SigLIP 2: Multilingual Vision-Language Encoders with Improved Semantic Understanding, Localization, and Dense Features.

Infinite Video Understanding SigLIP 2: Multilingual Vision-Language Encoders with Improved Semantic Understanding, Localization, and Dense Features

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-06T18:10:13.651494Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:10:13.651494Z digest=sha256:2c0f27a69ceb39897c6878c46d3da844b24d3384ff993154029cfa7b059bb0a1

Observation 8627080a-8146-4fd5-893f-d0b40c9a5e1c · outbound

This paper cites VideoMAE: Masked Autoencoders are Data-Efficient Learners for Self-Supervised Video Pre-Training.

Infinite Video Understanding VideoMAE: Masked Autoencoders are Data-Efficient Learners for Self-Supervised Video Pre-Training

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-06T18:10:13.574181Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:10:13.574181Z digest=sha256:8748f93d52e8c63272d3e3201e05909f5e888b7637133582e55653ec8bac773f

Observation 5c8a1017-73f4-4f2c-9e4d-a780ab36782d · outbound

This paper cites A Comprehensive Survey of Continual Learning: Theory, Method and Application.

Infinite Video Understanding A Comprehensive Survey of Continual Learning: Theory, Method and Application

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-06T18:10:13.775787Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:10:13.775787Z digest=sha256:7f521e36098c1239f6111d03e664303055bec6d7e59304941567d92ed71315ac

Observation 06bfa572-b497-4805-9795-6453d48c3512 · outbound

This paper cites VideoMAE V2: Scaling Video Masked Autoencoders with Dual Masking.

Infinite Video Understanding VideoMAE V2: Scaling Video Masked Autoencoders with Dual Masking

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-06T18:10:13.711528Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:10:13.711528Z digest=sha256:562b78ee92f7d2fbf0484d3b21d6ee79b60a7937c5f4ee0b9e11afb36bf8b0a1

Observation ef828168-849f-4259-8d45-f7f2490571e5 · outbound

This paper cites VidTwin: Video VAE with Decoupled Structure and Dynamics.

Infinite Video Understanding VidTwin: Video VAE with Decoupled Structure and Dynamics

Reference 43

Resolution
verified exact
local_arxiv, observed 2026-08-06T18:10:15.572551Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-06T18:10:14.101777Z digest=sha256:88290c086c523fd3b212400d1bcf1904b0e606b8a1b753ca50cec010f49c125e

Observation a69f610b-c659-49d6-9892-4f1bf7f702e1 · outbound

This paper cites an unresolved cited work.

Infinite Video Understanding Unresolved cited work

Reference 44

Resolution
unresolved
raw_fallback, observed 2026-08-06T18:10:17.529367Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-06T18:10:13.937532Z digest=sha256:7e454a68c97bdf5169ec354563c11c66fc6215d36c366874557cb39f5cea9eb0

Observation f519b7f8-72c3-4347-acda-48ca2a008599 · outbound

This paper cites VideoRoPE: What Makes for Good Video Rotary Position Embedding?.

Infinite Video Understanding VideoRoPE: What Makes for Good Video Rotary Position Embedding?

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-06T18:10:14.257348Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:10:14.257348Z digest=sha256:fdda3b15658abb2d01e1043623313bb940ee007dd405b44ebaf11d1b30b1fde1

Observation a30160aa-9053-40af-9c42-bef00968d3a7 · outbound

This paper cites LongVLM: Efficient Long Video Understanding via Large Language Models.

Infinite Video Understanding LongVLM: Efficient Long Video Understanding via Large Language Models

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-06T18:10:14.298291Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:10:14.298291Z digest=sha256:d39a8c0a7a1b0b29cc245f4be0c87a205792da5ab93155c21b21ef84f6dbce87

Observation 5d65b882-b812-4eb2-9817-d7922d9a8c0b · outbound

This paper cites VideoTree: Adaptive Tree-based Video Representation for LLM Reasoning on Long Videos.

Infinite Video Understanding VideoTree: Adaptive Tree-based Video Representation for LLM Reasoning on Long Videos

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-06T18:10:14.223529Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:10:14.223529Z digest=sha256:c984de2b08f6042d18e942e6eb814967e909fcc0bf69e7ef81cac70eec7e6371

Observation 88a9e838-e338-44d7-b0b5-2cabd1433fbc · outbound

This paper cites Agentic Reasoning: A Streamlined Framework for Enhancing LLM Reasoning with Agentic Tools.

Infinite Video Understanding Agentic Reasoning: A Streamlined Framework for Enhancing LLM Reasoning with Agentic Tools

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-06T18:10:14.365798Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:10:14.365798Z digest=sha256:d8311ac4e832e561d1e676d9cf4d26eca392aa5e4f2d04637972e37b9916912e

Observation aea3764e-83dd-40b8-9999-1eb2fca7070d · outbound

This paper cites LongViTU: Instruction Tuning for Long-Form Video Understanding.

Infinite Video Understanding LongViTU: Instruction Tuning for Long-Form Video Understanding

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-06T18:10:14.416661Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:10:14.416661Z digest=sha256:ef66710ceedbb3f9c2514483bc8f9f0b0044a6abf877df9dd13c0c37432a9b38

Observation b6143b35-bfaa-484d-9084-4430ec886d4b · outbound

This paper cites LongVideoBench: A Benchmark for Long-context Interleaved Video-Language Understanding.

Infinite Video Understanding LongVideoBench: A Benchmark for Long-context Interleaved Video-Language Understanding

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-06T18:10:14.340618Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:10:14.340618Z digest=sha256:07798b82d03ae3b0ed465efc36f572c62a76b9bac6e28fc12f506ae2481ffbc4

Observation 9dda5626-d89f-474a-92d8-7620a5c62f59 · outbound

This paper cites MAGVIT: Masked Generative Video Transformer.

Infinite Video Understanding MAGVIT: Masked Generative Video Transformer

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-06T18:10:14.507734Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:10:14.507734Z digest=sha256:1bc8b754b77899ee7d3717aaa967e95996253bf7347a800c4145c8a776dbaad3

Observation a4883221-886c-418b-9b68-445e2b4784e5 · outbound

This paper cites Language Model Beats Diffusion -- Tokenizer is Key to Visual Generation.

Infinite Video Understanding Language Model Beats Diffusion -- Tokenizer is Key to Visual Generation

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-06T18:10:14.568887Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:10:14.568887Z digest=sha256:2fa46d314e3d205d15ecd99fa3376c58b5b6d2f302f727d326cf57c774c7eaa5

Observation 5ab42a16-8fdb-4d71-84eb-02b6ab549f18 · outbound

This paper cites an unresolved cited work.

Infinite Video Understanding Unresolved cited work

Reference 53

Resolution
verified exact
doi, observed 2026-08-06T18:10:15.381307Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-06T18:10:14.459822Z digest=sha256:457c457139f8d46c98a3ac28f4cc8f823e4ad4617aaf79603250d1e33b9a4c91

Observation ef731b3e-8c43-42d9-aa5e-109c6274611b · outbound

This paper cites an unresolved cited work.

Infinite Video Understanding Unresolved cited work

Reference 54

Resolution
verified exact
doi, observed 2026-08-06T18:10:15.203453Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-06T18:10:14.678140Z digest=sha256:c3cd619158da4bfc5cad3da6afdcf15abf3b60a5eaa8a0f0d0ff15ffe53ba83e

Observation 0c444863-559e-43e1-8eb9-3f9de9561d65 · outbound

This paper cites Gundavarapu, Liangzhe Yuan, Hao Zhou, Shen Yan, Jennifer J.

Infinite Video Understanding Gundavarapu, Liangzhe Yuan, Hao Zhou, Shen Yan, Jennifer J

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:10:17.373787Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-06T18:10:14.711077Z digest=sha256:7acb2474a65c8b5dc980f5df954015c677f4d6a5127ab1706bc27691ae6be4f1

Observation 631dce2d-c1ed-46e3-a663-7fb3e379b154 · outbound

This paper cites Sigmoid Loss for Language Image Pre-Training.

Infinite Video Understanding Sigmoid Loss for Language Image Pre-Training

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-06T18:10:14.635404Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:10:14.635404Z digest=sha256:a4e15577c21d29c0327b3277a6e501f87fbf00d4f26fc66034f522c2a9772adc

Observation bc707462-8bba-42d4-b63e-912ab0c3dfc9 · outbound

This paper cites Towards Lifelong Learning of Large Language Models: A Survey.

Infinite Video Understanding Towards Lifelong Learning of Large Language Models: A Survey

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-06T18:10:14.841489Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:10:14.841489Z digest=sha256:d393800d03e03b789cf9f8835e09479d20dfadc163ad5f092a027c79a76143e7

Observation 12c78463-a446-463e-8f21-9ae51e99a337 · outbound

This paper cites MLVU: Benchmarking Multi-task Long Video Understanding.

Infinite Video Understanding MLVU: Benchmarking Multi-task Long Video Understanding

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-06T18:10:14.864857Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:10:14.864857Z digest=sha256:65bae9834f5f95909249ff8373ba4b21f0a2179c78b0cf4fdf33e5fa2b998f63

Observation 398801a5-873c-4169-abf0-73e544e5bb62 · outbound

This paper cites VideoPrism: A Foundational Visual Encoder for Video Understanding.

Infinite Video Understanding VideoPrism: A Foundational Visual Encoder for Video Understanding

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-06T18:10:14.749813Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:10:14.749813Z digest=sha256:a5dd2b136f53c69a47f45b2d73a0402b70c2081808f1c10a376f020b357dd928

Observation 5da26043-a850-4fdd-8944-8fc06fdf7b6a · outbound

This paper cites Needle In A Video Haystack: A Scalable Synthetic Evaluator for Video MLLMs.

Infinite Video Understanding Needle In A Video Haystack: A Scalable Synthetic Evaluator for Video MLLMs

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-06T18:10:14.777464Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:10:14.777464Z digest=sha256:66c52ef469e32ac4333b0521903f0bf3232918a8d30d93c1707096951a41b86a

Observation 168f0a8e-2993-48df-b4f0-266fb42fca75 · outbound

This paper cites an unresolved cited work.

Infinite Video Understanding Unresolved cited work

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-06T18:10:14.887539Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:10:14.887539Z digest=sha256:3bec4fbf6e2894bf1c82d3e990e34e6d869c2044e6901a46c6804bf29af72742

Observation ff6bf875-6dfb-403e-aa65-9bd3504869d2 · outbound

This paper cites From Seconds to Hours: Reviewing MultiModal Large Language Models on Comprehensive Long Video Understanding.

Infinite Video Understanding From Seconds to Hours: Reviewing MultiModal Large Language Models on Comprehensive Long Video Understanding

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-06T18:10:14.919272Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:10:14.919272Z digest=sha256:d4ad278ea052db6a7b8f09e1ce5a855e55d22edf5f76bcb5a60b03489fb7a732

Observation 004a05d9-6451-40e2-afb9-1e60ec5aa9fb · outbound

This paper cites LVBench: An Extreme Long Video Understanding Benchmark.

Infinite Video Understanding LVBench: An Extreme Long Video Understanding Benchmark

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-06T18:10:13.986726Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:10:13.986726Z digest=sha256:3575419417e17c2bcc84fd4cce4898eafba10a14bc0b47df75acab65dc73bb70

Observation 5dfc5ac4-c02b-4cfa-9645-4613c33f3a71 · outbound

This paper cites Video-R1: Reinforcing Video Reasoning in MLLMs.

Infinite Video Understanding Video-R1: Reinforcing Video Reasoning in MLLMs

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-06T18:10:11.970598Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:10:11.970598Z digest=sha256:c5129c78aa043ae8f77181f0d075f9a4aba25735a26bd8b1364de205f2be80cd

Pith citing papers

No inbound Pith citation observations are available.