Pith. sign in

Paper Citation Record · LEDGER

EASG-Bench: Video Q&A Benchmark with Egocentric Action Scene Graphs

As of 8 August 2026, this Paper Citation Record lists 17 of 17 outbound references and 1 inbound Pith citation observation for arXiv:2506.05787.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.05787 v2

Coverage vector

measured 17 of 17 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T10:17:17.863917Z

measured 18 of 18 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-05-15T21:44:47.636912Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-15T21:46:42.559417Z

Reference resolution

17 of 17 outbound references displayed

  • verified exact0
  • verified fuzzy7
  • unresolved10
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation ebd6a408-77b5-4c92-b811-b690bf1bca56 · outbound

This paper cites Qwen2.5-VL Technical Report.

EASG-Bench: Video Q&A Benchmark with Egocentric Action Scene Graphs Qwen2.5-VL Technical Report

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T10:17:16.931877Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:17:16.931877Z digest=sha256:6cb4893f774f55895cb3cfb76eb1e090da8850234a9fa39c249bcb8639895061

Observation 77936fc4-c253-46c7-a151-bae9430c7e99 · outbound

This paper cites Grounded multi- hop videoqa in long-form egocentric videos.

EASG-Bench: Video Q&A Benchmark with Egocentric Action Scene Graphs Grounded multi- hop videoqa in long-form egocentric videos

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:17:18.895134Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T10:17:16.991104Z digest=sha256:8c6fec71527b9f986a97cb12a69b4d7c460ef1652d6e99b7fae79e76651e96ab

Observation bb6c4bb6-bd45-47f3-aa1d-4a2e269afdae · outbound

This paper cites Egoplan- bench: Benchmarking egocentric embodied planning with multimodal large language models.

EASG-Bench: Video Q&A Benchmark with Egocentric Action Scene Graphs Egoplan- bench: Benchmarking egocentric embodied planning with multimodal large language models

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:17:18.837625Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T10:17:17.065212Z digest=sha256:1f64273f40dc230e6daa4f3412e8cafa192cb6c530d1bc929b903960bb9222f1

Observation 2b4fd427-484a-40b0-b351-97fc5b2e991e · outbound

This paper cites Egothink: Evalu- ating first-person perspective thinking capability of vision- language models.

EASG-Bench: Video Q&A Benchmark with Egocentric Action Scene Graphs Egothink: Evalu- ating first-person perspective thinking capability of vision- language models

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:17:18.669276Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T10:17:17.109460Z digest=sha256:cba3d860507f515d51c4ff50ded76f5c498b13d868d8599b0c9f0ba4046a7555

Observation 12e356a4-6fde-41f9-96d2-f2dda9ad7c58 · outbound

This paper cites Amego: Active memory from long egocentric videos.

EASG-Bench: Video Q&A Benchmark with Egocentric Action Scene Graphs Amego: Active memory from long egocentric videos

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:17:18.458621Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T10:17:17.161749Z digest=sha256:986e371f4aff4972cf7945f6953eefb1be4244c32fe06b0534578b6157ce8534

Observation b5a8d896-9e7b-4f8b-98ea-797f0f9a0ee8 · outbound

This paper cites The Llama 3 Herd of Models.

EASG-Bench: Video Q&A Benchmark with Egocentric Action Scene Graphs The Llama 3 Herd of Models

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T10:17:17.214584Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:17:17.214584Z digest=sha256:1959b5b3df5bb468b43ed0eace33d6d40d8df288d2a2af0aca835d7971faf950

Observation 48d545e1-e90e-4626-85f8-b7b0ea73850a · outbound

This paper cites Lita: Language instructed temporal-localization assistant.

EASG-Bench: Video Q&A Benchmark with Egocentric Action Scene Graphs Lita: Language instructed temporal-localization assistant

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:17:18.297147Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T10:17:17.276458Z digest=sha256:2f8146ea087cd8e22806864212dc83592b869b0c480b00067b8bcac380f71770

Observation 5f1e08e6-2724-4f8e-9511-050f66dfeda8 · outbound

This paper cites LLaVA-OneVision: Easy Visual Task Transfer.

EASG-Bench: Video Q&A Benchmark with Egocentric Action Scene Graphs LLaVA-OneVision: Easy Visual Task Transfer

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T10:17:17.306251Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:17:17.306251Z digest=sha256:ab06bc53f68c96e210514c03f996770a1e0ba4e7318d724c65d1842dbedaeb97

Observation 91176ecb-c72a-4fcb-99e3-a88da9192cc0 · outbound

This paper cites Egoschema: A diagnostic benchmark for very long- form video language understanding.

EASG-Bench: Video Q&A Benchmark with Egocentric Action Scene Graphs Egoschema: A diagnostic benchmark for very long- form video language understanding

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:17:18.176446Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T10:17:17.362153Z digest=sha256:dd213b99405450342ddf527d1eb5d403a57b8074b5c48c57d243bc6a71684f56

Observation 774a6b2c-af99-406f-9b2f-ea360eb6f438 · outbound

This paper cites Advancing Egocentric Video Question Answering with Multimodal Large Language Models.

EASG-Bench: Video Q&A Benchmark with Egocentric Action Scene Graphs Advancing Egocentric Video Question Answering with Multimodal Large Language Models

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T10:17:17.405246Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:17:17.405246Z digest=sha256:15d62eedf09e20c3a58c90aaec5c270a404932e889aa9ca644dc20e947ee712e

Observation 2b7c6d39-8b81-48b1-bfc0-ba90502dfb84 · outbound

This paper cites Omnia de EgoTempo: Benchmarking Temporal Understanding of Multi-Modal LLMs in Egocentric Videos.

EASG-Bench: Video Q&A Benchmark with Egocentric Action Scene Graphs Omnia de EgoTempo: Benchmarking Temporal Understanding of Multi-Modal LLMs in Egocentric Videos

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T10:17:17.470112Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:17:17.470112Z digest=sha256:bd6e27474d707934f006748029c9730b46655da414840e406a000c6dc85d389e

Observation 9be2a220-c8fb-407e-b465-0f72d7430153 · outbound

This paper cites SAM 2: Segment Anything in Images and Videos.

EASG-Bench: Video Q&A Benchmark with Egocentric Action Scene Graphs SAM 2: Segment Anything in Images and Videos

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T10:17:17.540048Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:17:17.540048Z digest=sha256:78c5147db0e1bf9b875eff8a029b093ffc077916933605a2d785664e5953eb6a

Observation db37397a-99a4-4b89-a1ef-8f05e1286840 · outbound

This paper cites TimeChat: A Time-sensitive Multimodal Large Language Model for Long Video Understanding.

EASG-Bench: Video Q&A Benchmark with Egocentric Action Scene Graphs TimeChat: A Time-sensitive Multimodal Large Language Model for Long Video Understanding

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T10:17:17.607966Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:17:17.607966Z digest=sha256:8215695c7ff573ca61fe33e4de98ef16edda74eb22407db1cdf73db788cc2481

Observation 5bbc6e66-1f33-4eec-a119-b43a4071736b · outbound

This paper cites Action scene graphs for long- form understanding of egocentric videos.

EASG-Bench: Video Q&A Benchmark with Egocentric Action Scene Graphs Action scene graphs for long- form understanding of egocentric videos

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:17:18.074491Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T10:17:17.670365Z digest=sha256:ab5e16c1dbd9ba7ff1023822404f6c798656e2da197fc9d4176539879d72c84e

Observation ec86de80-259e-4e91-b678-e1b0a80acb35 · outbound

This paper cites Gemma 2: Improving Open Language Models at a Practical Size.

EASG-Bench: Video Q&A Benchmark with Egocentric Action Scene Graphs Gemma 2: Improving Open Language Models at a Practical Size

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T10:17:17.765331Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:17:17.765331Z digest=sha256:bf9a323ff49a7289fcc0d774fef8d6bb466bd105691f8883f0cea07de68b6374

Observation 124428a9-3528-4ac4-a7de-a1c62531e500 · outbound

This paper cites Qwen3 Technical Report.

EASG-Bench: Video Q&A Benchmark with Egocentric Action Scene Graphs Qwen3 Technical Report

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T10:17:17.819501Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:17:17.819501Z digest=sha256:0ec85a33429f78e13adec199ad3132b4d2d3dc123120f49dc4465a56671fe403

Observation c1befbaa-7f4e-4627-a078-a06e67a9416b · outbound

This paper cites Sa2VA: Marrying SAM2 with LLaVA for Dense Grounded Understanding of Images and Videos.

EASG-Bench: Video Q&A Benchmark with Egocentric Action Scene Graphs Sa2VA: Marrying SAM2 with LLaVA for Dense Grounded Understanding of Images and Videos

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T10:17:17.863917Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:17:17.863917Z digest=sha256:523155406e70dda6bdfbe3ce62c627f87972bb43d91b6b7b942d564b1bc58cff

Pith citing papers

Observation 031158b4-8164-4661-a239-df46c1f64ed8 · inbound

EgoSound: Benchmarking Sound Understanding in Egocentric Videos cites this paper.

EgoSound: Benchmarking Sound Understanding in Egocentric Videos EASG-Bench: Video Q&A Benchmark with Egocentric Action Scene Graphs

Reference 30

Resolution
verified exact
arxiv_id, observed 2026-05-15T21:46:42.563537Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-15T21:44:47.636912Z digest=sha256:dde965355b07bb052ffd74488d42495160f29a4bb5ca53644238f7ccd9aa05e3