Pith. sign in

Paper Citation Record · LEDGER

E.T. Bench: Towards Open-Ended Event-Level Video-Language Understanding

As of 23 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 9 inbound Pith citation observations for arXiv:2409.18111.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2409.18111 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 9 of 9 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-23T06:30:58.430688+00:00

measured 9 of 9 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-16T04:40:19.121484Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-06-28T23:12:46.642138Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 86425931-5824-470a-b991-7e2b0337dc5a · inbound

LinVT: Empower Your Image-level Large Language Model to Understand Videos cites this paper.

LinVT: Empower Your Image-level Large Language Model to Understand Videos E.T. Bench: Towards Open-Ended Event-Level Video-Language Understanding

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-11T20:54:16.286499Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:54:16.286499Z digest=sha256:987c4c0faebec80f8260da77d0f824b9e7ab052416d89db8b33e93a4506f2b3f

Observation da1ebe5e-3684-46bb-a256-17da29bf03c6 · inbound

MINERVA: Evaluating Complex Video Reasoning cites this paper.

MINERVA: Evaluating Complex Video Reasoning E.T. Bench: Towards Open-Ended Event-Level Video-Language Understanding

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-16T04:40:19.121484Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T04:40:19.121484Z digest=sha256:7578668bbd1c1c41ad6db409cb389d63b5563ead88e02d66e4501d19523fb695

Observation 039de858-bae8-427b-a797-8f0160f23146 · inbound

DisTime: Distribution-based Time Representation for Video Large Language Models cites this paper.

DisTime: Distribution-based Time Representation for Video Large Language Models E.T. Bench: Towards Open-Ended Event-Level Video-Language Understanding

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T12:32:52.655077Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:32:52.655077Z digest=sha256:ff5f7f725435dec8ecb1e20d10408c0648d00356f1c50e9158534d747b2f1f32

Observation ca08579a-f769-47a9-87cc-a7bb7a4e650c · inbound

VRBench: A Benchmark for Multi-Step Reasoning in Long Narrative Videos cites this paper.

VRBench: A Benchmark for Multi-Step Reasoning in Long Narrative Videos E.T. Bench: Towards Open-Ended Event-Level Video-Language Understanding

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-07T04:22:55.922688Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:22:55.922688Z digest=sha256:846cd27b8d2d8978abbd8bb57a61fc2e28aa8a41f18a4c4654002eedeb216f38

Observation ec5d62f4-428d-4497-bbd9-08a8a5b08701 · inbound

AuroraLong: Bringing RNNs Back to Efficient Open-Ended Video Understanding cites this paper.

AuroraLong: Bringing RNNs Back to Efficient Open-Ended Video Understanding E.T. Bench: Towards Open-Ended Event-Level Video-Language Understanding

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-06T20:29:52.796096Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:29:52.796096Z digest=sha256:55693852e53b6491db3708741b3beaa586e0b69a0c599d6e2aff543e5ba96849

Observation 420a5d6b-82a0-4e9c-97f4-e00fa50e180c · inbound

ProactiveVideoQA: A Comprehensive Benchmark Evaluating Proactive Interactions in Video Large Language Models cites this paper.

ProactiveVideoQA: A Comprehensive Benchmark Evaluating Proactive Interactions in Video Large Language Models E.T. Bench: Towards Open-Ended Event-Level Video-Language Understanding

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-06T18:03:14.017689Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:03:14.017689Z digest=sha256:7ae49b4b4569b74564556b29cf65c54262d6ff5921a6af7ef44bac355dddff3e

Observation dbb30744-8dba-4a3a-814f-8920adb9b8f4 · inbound

Minerva-Ego: Spatiotemporal Hints for Egocentric Video Understanding cites this paper.

Minerva-Ego: Spatiotemporal Hints for Egocentric Video Understanding E.T. Bench: Towards Open-Ended Event-Level Video-Language Understanding

Reference 26

Resolution
verified exact
arxiv_id, observed 2026-05-19T16:03:08.056726Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-19T16:02:53.887605Z digest=sha256:c610fddb215233a9af08fc28d1b9cb2f7fdc9cee62b58c8bf1e758d01ac00594

Observation 08226c30-4fac-4025-8537-5582c9283b83 · inbound

An Efficient Streaming Video Understanding Framework with Agentic Control cites this paper.

An Efficient Streaming Video Understanding Framework with Agentic Control E.T. Bench: Towards Open-Ended Event-Level Video-Language Understanding

Reference 50

Resolution
verified exact
arxiv_id, observed 2026-05-20T11:33:14.447270Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-20T11:30:22.151045Z digest=sha256:cad364478533c8838556a4bb375a6fec7c9560bcbd41c4f00e23bebb0cf48d39

Observation ca3797d4-6e56-45df-9a2e-bf35d2f14a0e · inbound

TeachObs: A Human-Validated Benchmark for Multimodal Teaching Observation and Model Evaluation cites this paper.

TeachObs: A Human-Validated Benchmark for Multimodal Teaching Observation and Model Evaluation E.T. Bench: Towards Open-Ended Event-Level Video-Language Understanding

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-06-28T23:12:46.643590Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-06-28T23:11:43.712391Z digest=sha256:f18d3acc65f2e18492662e2481a8b3a4f4a45ab0d0bfcbd1a240febe0c29975e