Pith. sign in

Paper Citation Record · LEDGER

From Pixels to Graphs: using Scene and Knowledge Graphs for HD-EPIC VQA Challenge

As of 7 August 2026, this Paper Citation Record lists 27 of 27 outbound references and 1 inbound Pith citation observation for arXiv:2506.08553.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.08553 v1

Coverage vector

measured 27 of 27 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T05:11:19.983039Z

measured 28 of 28 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-07-01T06:21:09.996386Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

27 of 27 outbound references displayed

  • verified exact0
  • verified fuzzy23
  • unresolved4
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

0
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

Observation c2f6542a-0d1b-4646-bb97-d9376f273ce1 · outbound

This paper cites Flamingo: a visual language model for few-shot learning.

From Pixels to Graphs: using Scene and Knowledge Graphs for HD-EPIC VQA Challenge Flamingo: a visual language model for few-shot learning

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:11:20.316550Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:11:19.878734Z digest=sha256:1f74281bf8fba56ae47bcb14b2c32979d79a126c26e4a686c3fffa8e2fd877dc

Observation be1adb73-2497-4037-9299-eb3f6afb6417 · outbound

This paper cites 3d scene graph: A structure for unified semantics, 3d space, and cam- era.

From Pixels to Graphs: using Scene and Knowledge Graphs for HD-EPIC VQA Challenge 3d scene graph: A structure for unified semantics, 3d space, and cam- era

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:11:20.304373Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:11:19.883010Z digest=sha256:a0b221e8569b60e5231c942eb00d32a99c024cc1b47b231c72bdda56f30f2c51

Observation 4c63f293-f3c6-48e6-bcf6-0f1373285c4c · outbound

This paper cites Lamb Artur d’Avila Garcez.

From Pixels to Graphs: using Scene and Knowledge Graphs for HD-EPIC VQA Challenge Lamb Artur d’Avila Garcez

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:11:20.291760Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:11:19.887059Z digest=sha256:44a3e21d6e04bdb24a21fef15c463b81a3fbae0a4c32747292307efed5a5585b

Observation 235cf7a3-ff0b-42d6-9078-43480f14a488 · outbound

This paper cites Comet: Com- monsense transformers for automatic knowledge graph con- struction.

From Pixels to Graphs: using Scene and Knowledge Graphs for HD-EPIC VQA Challenge Comet: Com- monsense transformers for automatic knowledge graph con- struction

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:11:20.280349Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:11:19.891088Z digest=sha256:82e89cfad85e00feca6781175efe1e76c965de6d01fe19b955c23b14c7f3563f

Observation 21240ddd-d8de-4b24-b25c-01d998c275d9 · outbound

This paper cites Towards Neuro- Symbolic Video Understanding, page 220–236.

From Pixels to Graphs: using Scene and Knowledge Graphs for HD-EPIC VQA Challenge Towards Neuro- Symbolic Video Understanding, page 220–236

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:11:20.268509Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:11:19.895044Z digest=sha256:61fddffcf40262efa4de99acec6dbd2baa2972702816f1f669c8c44d570cd860

Observation 73ff61ec-5855-43d9-b6b6-04b1491bb20a · outbound

This paper cites The epic-kitchens dataset: Collection, challenges and base- lines.

From Pixels to Graphs: using Scene and Knowledge Graphs for HD-EPIC VQA Challenge The epic-kitchens dataset: Collection, challenges and base- lines

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:11:20.256974Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:11:19.899376Z digest=sha256:695bb18568872db0b593cdf85bc7dd457d18a06967a12b9df0d35a4da40a4e2f

Observation 28a4f84c-5e7d-4872-9065-b822f31ad9fe · outbound

This paper cites HuggingGPT: Solving AI Tasks with ChatGPT and its Friends in Hugging Face.

From Pixels to Graphs: using Scene and Knowledge Graphs for HD-EPIC VQA Challenge HuggingGPT: Solving AI Tasks with ChatGPT and its Friends in Hugging Face

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T05:11:19.903965Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:11:19.903965Z digest=sha256:7bc42612860c551afbe2ed83381b683ae2982538bff98365deedc282f686f23f

Observation fa42f414-e5fd-44bc-8d4f-b057a07f5006 · outbound

This paper cites Gemini 2.0 flash.

From Pixels to Graphs: using Scene and Knowledge Graphs for HD-EPIC VQA Challenge Gemini 2.0 flash

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:11:20.245182Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:11:19.908376Z digest=sha256:3dcc1035ad2c529f7f8ba7516022081476c1b5f53958a659763e4166a9f14d58

Observation d8803587-dafd-4457-9d08-c24f064298b6 · outbound

This paper cites Ego4d: Around the world in 3,000 hours of egocentric video.

From Pixels to Graphs: using Scene and Knowledge Graphs for HD-EPIC VQA Challenge Ego4d: Around the world in 3,000 hours of egocentric video

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T05:11:19.912469Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:11:19.912469Z digest=sha256:e6485c69a9d23a031e5af6e450ee5a0426e6102b90ea5e0e013d1f0b59de419e

Observation bdd20501-19e8-4cda-9108-a5bf667b1c25 · outbound

This paper cites Learning by asking questions.

From Pixels to Graphs: using Scene and Knowledge Graphs for HD-EPIC VQA Challenge Learning by asking questions

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:11:20.226951Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:11:19.916164Z digest=sha256:55784c07ded6d94c487b80ebb4ec6a4dbaedcc1494e1b3e0126f91b123588b0f

Observation 81b095c9-5195-41f7-a11d-31a8aefc8eab · outbound

This paper cites Patching open- vocabulary vision models with commonsense.

From Pixels to Graphs: using Scene and Knowledge Graphs for HD-EPIC VQA Challenge Patching open- vocabulary vision models with commonsense

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:11:20.216102Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:11:19.920387Z digest=sha256:9fde1ab7236d5a2069f31fe14367ab066eb7c0b7d01b734e46a1262f8633606e

Observation 9465e5cb-9df3-4a3a-9914-43f908f42393 · outbound

This paper cites Shamma, Michael Bernstein, and Li Fei-Fei.

From Pixels to Graphs: using Scene and Knowledge Graphs for HD-EPIC VQA Challenge Shamma, Michael Bernstein, and Li Fei-Fei

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:11:20.205152Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:11:19.924190Z digest=sha256:d981a09ebadb5176914f10672a6105b8199c6593a0268865a976251878d6781d

Observation 53b74dbc-71cb-4134-8ecc-e7a59a79144a · outbound

This paper cites an unresolved cited work.

From Pixels to Graphs: using Scene and Knowledge Graphs for HD-EPIC VQA Challenge Unresolved cited work

Reference 13

Resolution
unresolved
raw_fallback, observed 2026-08-07T05:11:20.194397Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:11:19.927911Z digest=sha256:8dd6858441add42ec19498d54d60bd34fe811e00c4862fd9c4c389bf7fbfcd64

Observation e3bc46e9-c2d9-4ca7-a005-b8171900903c · outbound

This paper cites BLIP: Bootstrapping Language-Image Pre-training for Unified Vision-Language Understanding and Generation.

From Pixels to Graphs: using Scene and Knowledge Graphs for HD-EPIC VQA Challenge BLIP: Bootstrapping Language-Image Pre-training for Unified Vision-Language Understanding and Generation

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T05:11:19.932135Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:11:19.932135Z digest=sha256:2ce589393d6698945b5394ee95e6c797b4eaac040f74bcd64b4ff8d55b09ee0a

Observation f794b1ba-7f42-464f-a881-e32cc9fb03c8 · outbound

This paper cites Neuro-symbolic concept learner: Interpreting scenes, words, and sentences from nat- ural supervision.

From Pixels to Graphs: using Scene and Knowledge Graphs for HD-EPIC VQA Challenge Neuro-symbolic concept learner: Interpreting scenes, words, and sentences from nat- ural supervision

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:11:20.183469Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:11:19.936145Z digest=sha256:15202a521d855121be036bd66ca5078b8fa3526bafeac925b28b236f79e4cda6

Observation e30a909d-c479-445e-8313-a936e9425e0d · outbound

This paper cites Augmented common- sense knowledge for remote object grounding.

From Pixels to Graphs: using Scene and Knowledge Graphs for HD-EPIC VQA Challenge Augmented common- sense knowledge for remote object grounding

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:11:20.172186Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:11:19.939700Z digest=sha256:73d4b2e50f12cb12087aa966f632bea3620cc85566a173cf146e78fc3f8e1184

Observation 3a6262cb-d8cc-4fd7-8038-7e5750166e15 · outbound

This paper cites Towards unbiased and ro- bust spatio-temporal scene graph generation and anticipa- tion, 2025.

From Pixels to Graphs: using Scene and Knowledge Graphs for HD-EPIC VQA Challenge Towards unbiased and ro- bust spatio-temporal scene graph generation and anticipa- tion, 2025

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:11:20.160515Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:11:19.943784Z digest=sha256:dd992f9f68ab42028a289884f1c443c9c5eb209f619ea6fd9934c83e2d10a952

Observation 2576c7b1-7f67-475a-8946-afb9727e12cd · outbound

This paper cites Pedregosa, G.

From Pixels to Graphs: using Scene and Knowledge Graphs for HD-EPIC VQA Challenge Pedregosa, G

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:11:20.149184Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:11:19.947463Z digest=sha256:e64718e07632d14f2159d6df0d0aff8b004530c6c9f0f7f18f59e40b85679d40

Observation eaf39c6b-bb7a-49b8-a893-0c96c71070af · outbound

This paper cites Sentence-bert: Sentence embeddings using siamese bert-networks.

From Pixels to Graphs: using Scene and Knowledge Graphs for HD-EPIC VQA Challenge Sentence-bert: Sentence embeddings using siamese bert-networks

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:11:20.137882Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:11:19.951109Z digest=sha256:308ea0c339ac54b166970424556d33b389795b085d48767ad4b9caa3ebf2d39a

Observation b7d9afb5-6413-4072-9713-88931f8cef68 · outbound

This paper cites Action scene graphs for long- form understanding of egocentric videos, 2023.

From Pixels to Graphs: using Scene and Knowledge Graphs for HD-EPIC VQA Challenge Action scene graphs for long- form understanding of egocentric videos, 2023

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:11:20.125175Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:11:19.954515Z digest=sha256:91946ab2f185ac76e72b7c0f0f78f4c545bb8d55aba2be7927363473b9888f82

Observation 7063de9e-9a9b-43dd-be6c-d82e2d42bdb8 · outbound

This paper cites Atomic: An atlas of ma- chine commonsense for if-then reasoning.

From Pixels to Graphs: using Scene and Knowledge Graphs for HD-EPIC VQA Challenge Atomic: An atlas of ma- chine commonsense for if-then reasoning

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:11:20.113042Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:11:19.959085Z digest=sha256:17647ebcb8b957b134f172c01a2854f62234d3ca56176c82d78daee81c9dfbdc

Observation eb06ebe3-fb83-4fbe-a318-3a0aee9005a0 · outbound

This paper cites Concept- net 5.5: An open multilingual graph of general knowledge.

From Pixels to Graphs: using Scene and Knowledge Graphs for HD-EPIC VQA Challenge Concept- net 5.5: An open multilingual graph of general knowledge

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:11:20.100797Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:11:19.963013Z digest=sha256:04a8dfc3d7f55cab2b1394707618c4fb47fabcca3a94c56e9b223bc9103e276b

Observation 61b922e8-d390-45f0-918f-9cc63241f79a · outbound

This paper cites Concept- net 5.5: An open multilingual graph of general knowledge,.

From Pixels to Graphs: using Scene and Knowledge Graphs for HD-EPIC VQA Challenge Concept- net 5.5: An open multilingual graph of general knowledge,

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:11:20.087645Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:11:19.966591Z digest=sha256:5e193fd799c600bd97a846137414fa6857113864ca050d2e75c5322bd53374af

Observation 26dc9ebc-4dac-454d-b468-42cee761571f · outbound

This paper cites Learning situation hyper-graphs for video question answer- ing.

From Pixels to Graphs: using Scene and Knowledge Graphs for HD-EPIC VQA Challenge Learning situation hyper-graphs for video question answer- ing

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:11:20.074898Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:11:19.970671Z digest=sha256:4a76538d45edf8d4290c34e1299ed9fec3dc566d367da57e29a50cb1a15d8432

Observation 4f776e03-62e6-4fe9-b272-7deba1911577 · outbound

This paper cites Scene graph generation by iterative message passing.

From Pixels to Graphs: using Scene and Knowledge Graphs for HD-EPIC VQA Challenge Scene graph generation by iterative message passing

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:11:20.062960Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:11:19.975080Z digest=sha256:a143b794f0d8e4681c080e36306a1d489181dc01044c02330cc63f2acdc265a6

Observation 957fbb91-98dc-4674-b2db-2fe9e6a755cd · outbound

This paper cites Neuro-symbolic visual reasoning: Disentangling ”vi- sual” from ”reasoning”.

From Pixels to Graphs: using Scene and Knowledge Graphs for HD-EPIC VQA Challenge Neuro-symbolic visual reasoning: Disentangling ”vi- sual” from ”reasoning”

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:11:20.050199Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:11:19.978859Z digest=sha256:e2688884b7b76630526db17c0a9bd98b453fef451dd7268f2479749c7fc69e7f

Observation 692377ea-c82e-4a12-bfc1-c2e2479fe93c · outbound

This paper cites Jasper and stella: distillation of sota embedding models,.

From Pixels to Graphs: using Scene and Knowledge Graphs for HD-EPIC VQA Challenge Jasper and stella: distillation of sota embedding models,

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:11:20.037717Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:11:19.983039Z digest=sha256:925b4c45a77a0e919f0b055197165241bc6ee58f82e3bde441728e18d0aff664

Pith citing papers

Observation 2f522f67-b64e-4061-b04d-5487d7ea1536 · inbound

Learning to Deny: Action Denial in Multimodal Large Language Models cites this paper.

Learning to Deny: Action Denial in Multimodal Large Language Models From Pixels to Graphs: using Scene and Knowledge Graphs for HD-EPIC VQA Challenge

Reference 67

Resolution
verified exact
arxiv_id, observed 2026-07-01T06:25:26.952478Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-07-01T06:21:09.996386Z digest=sha256:24a6ba57dfa22d102e58c12e54f275de804ce0f7d99e96b807f3ee2fa4f5707d