Pith. sign in

Paper Citation Record · LEDGER

EgoSchema: A Diagnostic Benchmark for Very Long-form Video Language Understanding

As of 22 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 27 inbound Pith citation observations for arXiv:2308.09126.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2308.09126 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 27 of 27 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-21T06:32:19.484+00:00

measured 27 of 27 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-16T10:50:21.460707Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T16:39:57.306434Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 64afd474-cd84-470f-a8c7-086f8848c8d4 · inbound

MVBench: A Comprehensive Multi-modal Video Understanding Benchmark cites this paper.

MVBench: A Comprehensive Multi-modal Video Understanding Benchmark EgoSchema: A Diagnostic Benchmark for Very Long-form Video Language Understanding

Reference 52

Resolution
verified exact
arxiv_id, observed 2026-05-17T20:22:35.036418Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-17T20:22:34.954228Z digest=sha256:1bb56f316adf5e40d04d429a1532b78687a4b7bf50f63da097dfc61e64b606cd

Observation 46e6c3fb-224b-45ad-9dca-67673151045b · inbound

VCA: Video Curious Agent for Long Video Understanding cites this paper.

VCA: Video Curious Agent for Long Video Understanding EgoSchema: A Diagnostic Benchmark for Very Long-form Video Language Understanding

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-11T16:51:15.224560Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T16:51:15.224560Z digest=sha256:3902faab2b20925fb13637ac557b2e7a450f341e9add284f05fe9d17ade8e971

Observation 8e5744cd-a924-45b2-8ab4-c4eb5b1f5cca · inbound

Online Video Understanding: OVBench and VideoChat-Online cites this paper.

Online Video Understanding: OVBench and VideoChat-Online EgoSchema: A Diagnostic Benchmark for Very Long-form Video Language Understanding

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-10T22:57:40.076720Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:57:40.076720Z digest=sha256:78b3d3c4eccaa86ed6bc31476c7dc8db6d183e21287a9d80f0ef070ea575ef44

Observation 19a7a086-642e-4a1c-8375-70f6a74947f8 · inbound

$\infty$-Video: A Training-Free Approach to Long Video Understanding via Continuous-Time Memory Consolidation cites this paper.

$\infty$-Video: A Training-Free Approach to Long Video Understanding via Continuous-Time Memory Consolidation EgoSchema: A Diagnostic Benchmark for Very Long-form Video Language Understanding

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-09T21:21:45.710426Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T21:21:45.710426Z digest=sha256:4946d0c46d6ae42963d313216c1acaff01b0fea166533fd0669059478e688f6a

Observation 68a5db79-232f-41c0-be4a-2b388ff281c4 · inbound

MASR: Self-Reflective Reasoning through Multimodal Hierarchical Attention Focusing for Agent-based Video Understanding cites this paper.

MASR: Self-Reflective Reasoning through Multimodal Hierarchical Attention Focusing for Agent-based Video Understanding EgoSchema: A Diagnostic Benchmark for Very Long-form Video Language Understanding

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-16T10:50:21.460707Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T10:50:21.460707Z digest=sha256:79a284578e5b3dcafd8991a0998cf00086d3c68c820b0fd814e6157f76e8584b

Observation 2942772e-a2f3-4211-b5d8-6c2a222f5199 · inbound

VideoLLM Benchmarks and Evaluation: A Survey cites this paper.

VideoLLM Benchmarks and Evaluation: A Survey EgoSchema: A Diagnostic Benchmark for Very Long-form Video Language Understanding

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-16T04:10:57.295597Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T04:10:57.295597Z digest=sha256:4f48706642218bf7c3d18485543de6b8f3ad4d81b3b7af6c3f8c6de76b806649

Observation d3ebb7c6-b8f5-4554-81cc-2595dd0a4c2b · inbound

Movie Facts and Fibs (MF$^2$): A Benchmark for Long Movie Understanding cites this paper.

Movie Facts and Fibs (MF$^2$): A Benchmark for Long Movie Understanding EgoSchema: A Diagnostic Benchmark for Very Long-form Video Language Understanding

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T06:00:56.930356Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:00:56.930356Z digest=sha256:c192450fe953e27535dfc4d56501e7c0e1045480e5c3ef1c82301f0e0d4df14c

Observation 9e01ba51-5333-48b2-be59-6a13003a0f9b · inbound

CausalVQA: A Physically Grounded Causal Reasoning Benchmark for Video Models cites this paper.

CausalVQA: A Physically Grounded Causal Reasoning Benchmark for Video Models EgoSchema: A Diagnostic Benchmark for Very Long-form Video Language Understanding

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T04:45:00.265342Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:45:00.265342Z digest=sha256:a8e5f7d25b64b00e422804edd39f352d6db72b7c85db041535eeaec6efb7b15c

Observation 468d22ad-8ab8-4b9f-b59b-9d312cb0db29 · inbound

LaCo: Efficient Layer-wise Compression of Visual Tokens for Multimodal Large Language Models cites this paper.

LaCo: Efficient Layer-wise Compression of Visual Tokens for Multimodal Large Language Models EgoSchema: A Diagnostic Benchmark for Very Long-form Video Language Understanding

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-06T20:38:30.134630Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:38:30.134630Z digest=sha256:4652deacab25f05c891e3745a430cd3a22f7d64ba561f519c0d8690614de4c41

Observation e0001093-1c8d-428c-8368-db9820138888 · inbound

Infinite Video Understanding cites this paper.

Infinite Video Understanding EgoSchema: A Diagnostic Benchmark for Very Long-form Video Language Understanding

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-06T18:10:12.839920Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:10:12.839920Z digest=sha256:a764cfb1a2a91b45323cef4ccfbd998e6ff67afab7d14048c23d31d22155e0b2

Observation 86434552-e4a1-40e7-b298-5d29ddd11a74 · inbound

HERMES: KV Cache as Hierarchical Memory for Efficient Streaming Video Understanding cites this paper.

HERMES: KV Cache as Hierarchical Memory for Efficient Streaming Video Understanding EgoSchema: A Diagnostic Benchmark for Very Long-form Video Language Understanding

Reference 31

Resolution
verified exact
arxiv_id, observed 2026-05-16T12:57:53.948203Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-16T12:55:04.564442Z digest=sha256:f475240122c94987e159a5003dc624b43f2f93b94310d02f0bab9c9ae561009c

Observation 05cf38b2-0bf8-4475-b41a-9d86e78056f9 · inbound

Adaptive Greedy Frame Selection for Long Video Understanding cites this paper.

Adaptive Greedy Frame Selection for Long Video Understanding EgoSchema: A Diagnostic Benchmark for Very Long-form Video Language Understanding

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-05-15T08:25:18.980287Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-15T08:22:24.872665Z digest=sha256:3cca5fa6ede9a0724adbd2817a68810c59c7ce827729bd80f9d65ec1caa2449b

Observation e0e18d26-0bf9-486f-b1f2-9c81172d79f5 · inbound

MACS: Modality-Aware Capacity Scaling for Efficient Multimodal MoE Inference cites this paper.

MACS: Modality-Aware Capacity Scaling for Efficient Multimodal MoE Inference EgoSchema: A Diagnostic Benchmark for Very Long-form Video Language Understanding

Reference 35

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T01:45:52.094055Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-05-11T01:29:26.298131Z digest=sha256:5725c5565a8ba730c53a65c15280de73235bff124d618b3e2dfae57d4214d4ee

Observation c04ffac2-679e-424d-a546-0c9ed249b1ff · inbound

HumanNet: Scaling Human-centric Video Learning to One Million Hours cites this paper.

HumanNet: Scaling Human-centric Video Learning to One Million Hours EgoSchema: A Diagnostic Benchmark for Very Long-form Video Language Understanding

Reference 27

Resolution
verified exact
arxiv_id, observed 2026-05-11T05:10:54.191355Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-11T00:51:08.414394Z digest=sha256:59a1c4ba9b06679b36327447162804456aaa050b26146eeb7dae49de65454c09

Observation f0cc8be1-9168-4401-b590-c2240d6e1ad1 · inbound

EgoCoT-Bench: Benchmarking Grounded and Verifiable Operation-Centric Chain of Thought Reasoning for MLLMs cites this paper.

EgoCoT-Bench: Benchmarking Grounded and Verifiable Operation-Centric Chain of Thought Reasoning for MLLMs EgoSchema: A Diagnostic Benchmark for Very Long-form Video Language Understanding

Reference 23

Resolution
verified exact
arxiv_id, observed 2026-05-20T05:53:22.271439Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-20T05:53:05.450946Z digest=sha256:aed903641430dff4fce484c8e19e12295db7449fd015f2f6628428ca658cee0e

Observation ccad6fc8-7167-48d0-a6bc-5f44830d7c13 · inbound

EgoProx: Evaluating MLLMs on Egocentric 3D Proximity Reasoning Across a Cognitive Hierarchy cites this paper.

EgoProx: Evaluating MLLMs on Egocentric 3D Proximity Reasoning Across a Cognitive Hierarchy EgoSchema: A Diagnostic Benchmark for Very Long-form Video Language Understanding

Reference 43

Resolution
verified exact
arxiv_id, observed 2026-06-30T13:34:40.524149Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-06-30T13:28:43.538541Z digest=sha256:c3ca000ccc7ec5eb85f8211c64a20399adfe21338d590267f650e7d5edf6e6a9

Observation dc68e372-e42d-4d31-9ffa-227accf43bb7 · inbound

TeachObs: A Human-Validated Benchmark for Multimodal Teaching Observation and Model Evaluation cites this paper.

TeachObs: A Human-Validated Benchmark for Multimodal Teaching Observation and Model Evaluation EgoSchema: A Diagnostic Benchmark for Very Long-form Video Language Understanding

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-06-28T23:12:46.641053Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-06-28T23:11:43.712391Z digest=sha256:64490e1c9c5fd6da7bd3742d9563db92824e24973e92ad1bae4d4e2e200132a5

Observation 09494012-2fb3-4efb-8910-3abc33393d82 · inbound

Pause and Think: A Dataset and Benchmark for Video-Grounded Assistive Action Suggestion cites this paper.

Pause and Think: A Dataset and Benchmark for Video-Grounded Assistive Action Suggestion EgoSchema: A Diagnostic Benchmark for Very Long-form Video Language Understanding

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-06-28T19:12:34.434642Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-06-28T19:11:28.937135Z digest=sha256:08ac4a118fadf3506f1e9f4a4520a8e37f4407d2ec0702dd4b2c437b22c0f4be

Observation b38608ab-01fb-4bc1-b61a-febc772f53a0 · inbound

Pause and Think: A Dataset and Benchmark for Video-Grounded Assistive Action Suggestion cites this paper.

Pause and Think: A Dataset and Benchmark for Video-Grounded Assistive Action Suggestion EgoSchema: A Diagnostic Benchmark for Very Long-form Video Language Understanding

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-06-30T11:44:38.506790Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-06-30T11:37:50.291543Z digest=sha256:0a968d34324b77e351c7a70405c2901ca5ac32289384ebaec78398a23d66117a

Observation a35d243a-9411-42c6-a6eb-2415516408d5 · inbound

SuperMemory-VQA: An Egocentric Visual Question-Answering Benchmark for Long-Horizon Memory cites this paper.

SuperMemory-VQA: An Egocentric Visual Question-Answering Benchmark for Long-Horizon Memory EgoSchema: A Diagnostic Benchmark for Very Long-form Video Language Understanding

Reference 38

Resolution
verified exact
arxiv_id, observed 2026-06-28T19:02:34.145726Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-06-28T19:00:20.891691Z digest=sha256:d81aa1b460869538f05f171c61776b8fe92091e79bf201a62f93fac354569a09

Observation b33479a2-f95b-47b1-bb50-245f20f7b627 · inbound

NextMotionQA: Benchmarking and Judging Human Motion Understanding with Vision-Language Models cites this paper.

NextMotionQA: Benchmarking and Judging Human Motion Understanding with Vision-Language Models EgoSchema: A Diagnostic Benchmark for Very Long-form Video Language Understanding

Reference 44

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T07:26:46.156168Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-06-28T06:55:22.332372Z digest=sha256:142cc52a20051bc2de16b0727cb372179de132439e76b44e68a0fa3318a63e29

Observation 2ceec0df-85e7-4ce4-8fab-9895482f6318 · inbound

MOSS-Video-Preview: Toward Real-Time Video Understanding via Cross-Attention cites this paper.

MOSS-Video-Preview: Toward Real-Time Video Understanding via Cross-Attention EgoSchema: A Diagnostic Benchmark for Very Long-form Video Language Understanding

Reference 40

Resolution
verified exact
arxiv_id, observed 2026-07-01T22:26:17.890023Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-06-28T15:22:31.310003Z digest=sha256:057d0caa07ae3cf71c531ff2ab36bede72dc4061c9b121c8bfb387688903fa0d

Observation 4f2851b0-9617-47dc-92f0-6fb59e21afad · inbound

EgoSAT: A Comprehensive Benchmark of Egocentric Streaming Interaction Understanding cites this paper.

EgoSAT: A Comprehensive Benchmark of Egocentric Streaming Interaction Understanding EgoSchema: A Diagnostic Benchmark for Very Long-form Video Language Understanding

Reference 24

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T16:39:57.307946Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-06-26T00:26:33.306128Z digest=sha256:c3c61c4982dbe06e34de0399c61414ea37201a6eabd40d319d1e92b2ce563c99

Observation 365b3d41-34c0-4c51-bf09-e04b99c0b6c3 · inbound

From Accuracy to Visual Dependence: Auditing and Filtering Modality Collapse in Traffic VideoQA cites this paper.

From Accuracy to Visual Dependence: Auditing and Filtering Modality Collapse in Traffic VideoQA EgoSchema: A Diagnostic Benchmark for Very Long-form Video Language Understanding

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-06-30T06:14:18.581710Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-06-30T06:14:11.109840Z digest=sha256:cad0aa3f708a37de4774220f4f81caf65990f9674ced570fd3814a7f8c6dc87d

Observation dbf7d4f3-fae6-4381-8f1b-f91477ad19b1 · inbound

MEMORA: Embodied Action Memory from Egocentric Videos for Reasoning and Planning cites this paper.

MEMORA: Embodied Action Memory from Egocentric Videos for Reasoning and Planning EgoSchema: A Diagnostic Benchmark for Very Long-form Video Language Understanding

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-02T02:42:47.589671Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T02:42:47.589671Z digest=sha256:9c10ed62173450b4fe7ebf196496365d23d62d8b9086da351f5bd26de91ec42f

Observation b7e99be9-fe47-41e4-a095-7abe1408cc75 · inbound

ObjectStream: Latent Objects as Memory Anchors for Streaming Video Understanding cites this paper.

ObjectStream: Latent Objects as Memory Anchors for Streaming Video Understanding EgoSchema: A Diagnostic Benchmark for Very Long-form Video Language Understanding

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-04T03:21:46.535881Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T03:21:46.535881Z digest=sha256:4cef0d6620a8600359e235fc12cf6951dd7a4426cd8af56a791112c0eda3cac2

Observation 6e0885d0-099b-498c-9184-bfdcca6e8f75 · inbound

CRAFT: Compression via Recursive Adaptive Fusion of Video Tokens for Vision-Language Models cites this paper.

CRAFT: Compression via Recursive Adaptive Fusion of Video Tokens for Vision-Language Models EgoSchema: A Diagnostic Benchmark for Very Long-form Video Language Understanding

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-04T23:46:51.588231Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T23:46:51.588231Z digest=sha256:ef3de4825d8530bc4c9f5df83a48597fc33e9160fe13498c8a7c36b085235963