Pith. sign in

Paper Citation Record · LEDGER

EASG-Bench: Video Q&A Benchmark with Egocentric Action Scene Graphs

As of 8 August 2026, this Paper Citation Record lists 17 of 17 outbound references and 1 inbound Pith citation observation for arXiv:2506.05787.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.05787 v2

Coverage vector

measured 17 of 17 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T10:17:17.863917Z

measured 18 of 18 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-05-15T21:44:47.636912Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-15T21:46:42.559417Z

Reference resolution

17 of 17 outbound references displayed

  • verified exact0
  • verified fuzzy7
  • unresolved10
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation ebd6a408-77b5-4c92-b811-b690bf1bca56 · outbound

This paper cites Qwen2.5-VL Technical Report.

EASG-Bench: Video Q&A Benchmark with Egocentric Action Scene Graphs Qwen2.5-VL Technical Report

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T10:17:16.931877Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:17:16.931877Z digest=sha256:6cb4893f774f55895cb3cfb76eb1e090da8850234a9fa39c249bcb8639895061

Observation 77936fc4-c253-46c7-a151-bae9430c7e99 · outbound

This paper cites Grounded multi- hop videoqa in long-form egocentric videos.

EASG-Bench: Video Q&A Benchmark with Egocentric Action Scene Graphs Grounded multi- hop videoqa in long-form egocentric videos

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:17:18.895134Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T10:17:16.991104Z digest=sha256:fb31e06e91e2f29d16c4fa0dd5d12f713f0a8abe6d0d6da5efae4e00dc32488e

Observation bb6c4bb6-bd45-47f3-aa1d-4a2e269afdae · outbound

This paper cites Egoplan- bench: Benchmarking egocentric embodied planning with multimodal large language models.

EASG-Bench: Video Q&A Benchmark with Egocentric Action Scene Graphs Egoplan- bench: Benchmarking egocentric embodied planning with multimodal large language models

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:17:18.837625Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T10:17:17.065212Z digest=sha256:93b77819e6b4c91252c050255a710951be81ca22e25b5d3ecb0f826a487bc01e

Observation 2b4fd427-484a-40b0-b351-97fc5b2e991e · outbound

This paper cites Egothink: Evalu- ating first-person perspective thinking capability of vision- language models.

EASG-Bench: Video Q&A Benchmark with Egocentric Action Scene Graphs Egothink: Evalu- ating first-person perspective thinking capability of vision- language models

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:17:18.669276Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T10:17:17.109460Z digest=sha256:e87cb615d4c2ae6bccd45b287e1a7c6496484ab5374aaf3a9993d9b1177bd30a

Observation 12e356a4-6fde-41f9-96d2-f2dda9ad7c58 · outbound

This paper cites Amego: Active memory from long egocentric videos.

EASG-Bench: Video Q&A Benchmark with Egocentric Action Scene Graphs Amego: Active memory from long egocentric videos

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:17:18.458621Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T10:17:17.161749Z digest=sha256:fcc4c164ea87b5c5ca8da2d54cbf8d2d699b012287550ff8d24f70471f27caf7

Observation b5a8d896-9e7b-4f8b-98ea-797f0f9a0ee8 · outbound

This paper cites The Llama 3 Herd of Models.

EASG-Bench: Video Q&A Benchmark with Egocentric Action Scene Graphs The Llama 3 Herd of Models

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T10:17:17.214584Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:17:17.214584Z digest=sha256:1959b5b3df5bb468b43ed0eace33d6d40d8df288d2a2af0aca835d7971faf950

Observation 48d545e1-e90e-4626-85f8-b7b0ea73850a · outbound

This paper cites Lita: Language instructed temporal-localization assistant.

EASG-Bench: Video Q&A Benchmark with Egocentric Action Scene Graphs Lita: Language instructed temporal-localization assistant

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:17:18.297147Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T10:17:17.276458Z digest=sha256:e2216fccf17afda63d35b1646b8005beb7bea67df2a9fa49034d1f89ec6a2a20

Observation 5f1e08e6-2724-4f8e-9511-050f66dfeda8 · outbound

This paper cites LLaVA-OneVision: Easy Visual Task Transfer.

EASG-Bench: Video Q&A Benchmark with Egocentric Action Scene Graphs LLaVA-OneVision: Easy Visual Task Transfer

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T10:17:17.306251Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:17:17.306251Z digest=sha256:ab06bc53f68c96e210514c03f996770a1e0ba4e7318d724c65d1842dbedaeb97

Observation 91176ecb-c72a-4fcb-99e3-a88da9192cc0 · outbound

This paper cites Egoschema: A diagnostic benchmark for very long- form video language understanding.

EASG-Bench: Video Q&A Benchmark with Egocentric Action Scene Graphs Egoschema: A diagnostic benchmark for very long- form video language understanding

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:17:18.176446Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T10:17:17.362153Z digest=sha256:cc765e1053ca40e0b6e0bc7f370028bc263d6b84912c29e386a45db9493a69f0

Observation 774a6b2c-af99-406f-9b2f-ea360eb6f438 · outbound

This paper cites Advancing Egocentric Video Question Answering with Multimodal Large Language Models.

EASG-Bench: Video Q&A Benchmark with Egocentric Action Scene Graphs Advancing Egocentric Video Question Answering with Multimodal Large Language Models

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T10:17:17.405246Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:17:17.405246Z digest=sha256:15d62eedf09e20c3a58c90aaec5c270a404932e889aa9ca644dc20e947ee712e

Observation 2b7c6d39-8b81-48b1-bfc0-ba90502dfb84 · outbound

This paper cites Omnia de EgoTempo: Benchmarking Temporal Understanding of Multi-Modal LLMs in Egocentric Videos.

EASG-Bench: Video Q&A Benchmark with Egocentric Action Scene Graphs Omnia de EgoTempo: Benchmarking Temporal Understanding of Multi-Modal LLMs in Egocentric Videos

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T10:17:17.470112Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:17:17.470112Z digest=sha256:bd6e27474d707934f006748029c9730b46655da414840e406a000c6dc85d389e

Observation 9be2a220-c8fb-407e-b465-0f72d7430153 · outbound

This paper cites SAM 2: Segment Anything in Images and Videos.

EASG-Bench: Video Q&A Benchmark with Egocentric Action Scene Graphs SAM 2: Segment Anything in Images and Videos

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T10:17:17.540048Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:17:17.540048Z digest=sha256:78c5147db0e1bf9b875eff8a029b093ffc077916933605a2d785664e5953eb6a

Observation db37397a-99a4-4b89-a1ef-8f05e1286840 · outbound

This paper cites TimeChat: A Time-sensitive Multimodal Large Language Model for Long Video Understanding.

EASG-Bench: Video Q&A Benchmark with Egocentric Action Scene Graphs TimeChat: A Time-sensitive Multimodal Large Language Model for Long Video Understanding

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T10:17:17.607966Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:17:17.607966Z digest=sha256:8215695c7ff573ca61fe33e4de98ef16edda74eb22407db1cdf73db788cc2481

Observation 5bbc6e66-1f33-4eec-a119-b43a4071736b · outbound

This paper cites Action scene graphs for long- form understanding of egocentric videos.

EASG-Bench: Video Q&A Benchmark with Egocentric Action Scene Graphs Action scene graphs for long- form understanding of egocentric videos

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:17:18.074491Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T10:17:17.670365Z digest=sha256:ab0650e4fb832542604abac7ce4432ee8c4d9ffe3cd921a6cb7a1d40973ffe27

Observation ec86de80-259e-4e91-b678-e1b0a80acb35 · outbound

This paper cites Gemma 2: Improving Open Language Models at a Practical Size.

EASG-Bench: Video Q&A Benchmark with Egocentric Action Scene Graphs Gemma 2: Improving Open Language Models at a Practical Size

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T10:17:17.765331Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:17:17.765331Z digest=sha256:bf9a323ff49a7289fcc0d774fef8d6bb466bd105691f8883f0cea07de68b6374

Observation 124428a9-3528-4ac4-a7de-a1c62531e500 · outbound

This paper cites Qwen3 Technical Report.

EASG-Bench: Video Q&A Benchmark with Egocentric Action Scene Graphs Qwen3 Technical Report

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T10:17:17.819501Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:17:17.819501Z digest=sha256:0ec85a33429f78e13adec199ad3132b4d2d3dc123120f49dc4465a56671fe403

Observation c1befbaa-7f4e-4627-a078-a06e67a9416b · outbound

This paper cites Sa2VA: Marrying SAM2 with LLaVA for Dense Grounded Understanding of Images and Videos.

EASG-Bench: Video Q&A Benchmark with Egocentric Action Scene Graphs Sa2VA: Marrying SAM2 with LLaVA for Dense Grounded Understanding of Images and Videos

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T10:17:17.863917Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:17:17.863917Z digest=sha256:523155406e70dda6bdfbe3ce62c627f87972bb43d91b6b7b942d564b1bc58cff

Pith citing papers

Observation 031158b4-8164-4661-a239-df46c1f64ed8 · inbound

EgoSound: Benchmarking Sound Understanding in Egocentric Videos cites this paper.

EgoSound: Benchmarking Sound Understanding in Egocentric Videos EASG-Bench: Video Q&A Benchmark with Egocentric Action Scene Graphs

Reference 30

Resolution
verified exact
arxiv_id, observed 2026-05-15T21:46:42.563537Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-15T21:44:47.636912Z digest=sha256:66adb7615c48ad80917754f86e8ff66920106079312251d44411cb13508d6fb0