Pith. sign in

Paper Citation Record · LEDGER

From Pixels to Graphs: using Scene and Knowledge Graphs for HD-EPIC VQA Challenge

As of 18 August 2026, this Paper Citation Record lists 27 of 27 outbound references and 1 inbound Pith citation observation for arXiv:2506.08553.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.08553 v1

Coverage vector

measured 27 of 27 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T05:11:19.983039Z

measured 28 of 28 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-07-01T06:21:09.996386Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

27 of 27 outbound references displayed

  • verified exact0
  • verified fuzzy23
  • unresolved4
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

0
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

Observation c2f6542a-0d1b-4646-bb97-d9376f273ce1 · outbound

This paper cites Flamingo: a visual language model for few-shot learning.

From Pixels to Graphs: using Scene and Knowledge Graphs for HD-EPIC VQA Challenge Flamingo: a visual language model for few-shot learning

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:11:20.316550Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T05:11:19.878734Z digest=sha256:94b43e3479bf5a9fc2aa72507f65703a5dd477f65cf954b5423f81946586d9ee

Observation be1adb73-2497-4037-9299-eb3f6afb6417 · outbound

This paper cites 3d scene graph: A structure for unified semantics, 3d space, and cam- era.

From Pixels to Graphs: using Scene and Knowledge Graphs for HD-EPIC VQA Challenge 3d scene graph: A structure for unified semantics, 3d space, and cam- era

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:11:20.304373Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T05:11:19.883010Z digest=sha256:785f6275ebd4300bb74bb35814beea47b3e48804ee92c81bb54c1ad2dad60a83

Observation 4c63f293-f3c6-48e6-bcf6-0f1373285c4c · outbound

This paper cites Lamb Artur d’Avila Garcez.

From Pixels to Graphs: using Scene and Knowledge Graphs for HD-EPIC VQA Challenge Lamb Artur d’Avila Garcez

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:11:20.291760Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T05:11:19.887059Z digest=sha256:b7bbaaa4708490907647d246798240ffcdc26811787607561ec08466cd172396

Observation 235cf7a3-ff0b-42d6-9078-43480f14a488 · outbound

This paper cites Comet: Com- monsense transformers for automatic knowledge graph con- struction.

From Pixels to Graphs: using Scene and Knowledge Graphs for HD-EPIC VQA Challenge Comet: Com- monsense transformers for automatic knowledge graph con- struction

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:11:20.280349Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T05:11:19.891088Z digest=sha256:9377c39be3ae6f4ec4ba140e7db9d6c9e8f0d565c5bc3414a7781bb097e4429d

Observation 21240ddd-d8de-4b24-b25c-01d998c275d9 · outbound

This paper cites Towards Neuro- Symbolic Video Understanding, page 220–236.

From Pixels to Graphs: using Scene and Knowledge Graphs for HD-EPIC VQA Challenge Towards Neuro- Symbolic Video Understanding, page 220–236

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:11:20.268509Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T05:11:19.895044Z digest=sha256:e0ff250daae53528cd11daf7825c68df9ab62404a3b5fb2dc2d2b01e29077f62

Observation 73ff61ec-5855-43d9-b6b6-04b1491bb20a · outbound

This paper cites The epic-kitchens dataset: Collection, challenges and base- lines.

From Pixels to Graphs: using Scene and Knowledge Graphs for HD-EPIC VQA Challenge The epic-kitchens dataset: Collection, challenges and base- lines

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:11:20.256974Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T05:11:19.899376Z digest=sha256:f5301db567ea4bc95ddc0a70dddfae0262492811fd60a312b0e557cb62f27b00

Observation 28a4f84c-5e7d-4872-9065-b822f31ad9fe · outbound

This paper cites HuggingGPT: Solving AI Tasks with ChatGPT and its Friends in Hugging Face.

From Pixels to Graphs: using Scene and Knowledge Graphs for HD-EPIC VQA Challenge HuggingGPT: Solving AI Tasks with ChatGPT and its Friends in Hugging Face

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T05:11:19.903965Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:11:19.903965Z digest=sha256:bf78802f5998f44551ef71a5a10b010c211163e3facb5f64f5afd32696b8643b

Observation fa42f414-e5fd-44bc-8d4f-b057a07f5006 · outbound

This paper cites Gemini 2.0 flash.

From Pixels to Graphs: using Scene and Knowledge Graphs for HD-EPIC VQA Challenge Gemini 2.0 flash

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:11:20.245182Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T05:11:19.908376Z digest=sha256:ba5f222e0c8c079e82c0414472cea3839a08846124ac3e31f221f3eada8c8ed0

Observation d8803587-dafd-4457-9d08-c24f064298b6 · outbound

This paper cites Ego4d: Around the world in 3,000 hours of egocentric video.

From Pixels to Graphs: using Scene and Knowledge Graphs for HD-EPIC VQA Challenge Ego4d: Around the world in 3,000 hours of egocentric video

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T05:11:19.912469Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:11:19.912469Z digest=sha256:9c16e5537729d50720a2e55b4953fadd78ca9ea4edf2628aaa078e1aca1707b7

Observation bdd20501-19e8-4cda-9108-a5bf667b1c25 · outbound

This paper cites Learning by asking questions.

From Pixels to Graphs: using Scene and Knowledge Graphs for HD-EPIC VQA Challenge Learning by asking questions

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:11:20.226951Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T05:11:19.916164Z digest=sha256:e396bde7d545a4d6dfb6f941a9ef3ae60274d12645d0a32a8d556525c763c0e7

Observation 81b095c9-5195-41f7-a11d-31a8aefc8eab · outbound

This paper cites Patching open- vocabulary vision models with commonsense.

From Pixels to Graphs: using Scene and Knowledge Graphs for HD-EPIC VQA Challenge Patching open- vocabulary vision models with commonsense

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:11:20.216102Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T05:11:19.920387Z digest=sha256:0ccc7ff9d4afe2028a9f2c82574260df2750cf8f5e45ae3eca93dbb28df9fe81

Observation 9465e5cb-9df3-4a3a-9914-43f908f42393 · outbound

This paper cites Shamma, Michael Bernstein, and Li Fei-Fei.

From Pixels to Graphs: using Scene and Knowledge Graphs for HD-EPIC VQA Challenge Shamma, Michael Bernstein, and Li Fei-Fei

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:11:20.205152Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T05:11:19.924190Z digest=sha256:eb585aa1b3ad6a0f4aeaa2d53b913a8fe5c5163375254bed35235c94d20313e6

Observation 53b74dbc-71cb-4134-8ecc-e7a59a79144a · outbound

This paper cites an unresolved cited work.

From Pixels to Graphs: using Scene and Knowledge Graphs for HD-EPIC VQA Challenge Unresolved cited work

Reference 13

Resolution
unresolved
raw_fallback, observed 2026-08-07T05:11:20.194397Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T05:11:19.927911Z digest=sha256:824aa7c46dbf2f5d3b5792fee74dad41115618c4b2080862557cc56e1e4320e8

Observation e3bc46e9-c2d9-4ca7-a005-b8171900903c · outbound

This paper cites BLIP: Bootstrapping Language-Image Pre-training for Unified Vision-Language Understanding and Generation.

From Pixels to Graphs: using Scene and Knowledge Graphs for HD-EPIC VQA Challenge BLIP: Bootstrapping Language-Image Pre-training for Unified Vision-Language Understanding and Generation

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T05:11:19.932135Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:11:19.932135Z digest=sha256:7b579d7124a4bdb305aaaf098bcd8c8bbb0008307d21c89b9487c818a14d37af

Observation f794b1ba-7f42-464f-a881-e32cc9fb03c8 · outbound

This paper cites Neuro-symbolic concept learner: Interpreting scenes, words, and sentences from nat- ural supervision.

From Pixels to Graphs: using Scene and Knowledge Graphs for HD-EPIC VQA Challenge Neuro-symbolic concept learner: Interpreting scenes, words, and sentences from nat- ural supervision

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:11:20.183469Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T05:11:19.936145Z digest=sha256:2890b9295e6dcf01216e3c271030938b0d675850041af733c998ee8643fb9d6d

Observation e30a909d-c479-445e-8313-a936e9425e0d · outbound

This paper cites Augmented common- sense knowledge for remote object grounding.

From Pixels to Graphs: using Scene and Knowledge Graphs for HD-EPIC VQA Challenge Augmented common- sense knowledge for remote object grounding

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:11:20.172186Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T05:11:19.939700Z digest=sha256:e1abcac8531a638985afd7d6aca05e12e9cf88caa6f4bb9fb322233aacffd51c

Observation 3a6262cb-d8cc-4fd7-8038-7e5750166e15 · outbound

This paper cites Towards unbiased and ro- bust spatio-temporal scene graph generation and anticipa- tion, 2025.

From Pixels to Graphs: using Scene and Knowledge Graphs for HD-EPIC VQA Challenge Towards unbiased and ro- bust spatio-temporal scene graph generation and anticipa- tion, 2025

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:11:20.160515Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T05:11:19.943784Z digest=sha256:ec7795d45015516a3d5509443f483de49def318bbc9625ae784353431d3be443

Observation 2576c7b1-7f67-475a-8946-afb9727e12cd · outbound

This paper cites Pedregosa, G.

From Pixels to Graphs: using Scene and Knowledge Graphs for HD-EPIC VQA Challenge Pedregosa, G

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:11:20.149184Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T05:11:19.947463Z digest=sha256:726b99a3523abb812b38fae1096f9129b4677128e32f2bc48019942c0381033c

Observation eaf39c6b-bb7a-49b8-a893-0c96c71070af · outbound

This paper cites Sentence-bert: Sentence embeddings using siamese bert-networks.

From Pixels to Graphs: using Scene and Knowledge Graphs for HD-EPIC VQA Challenge Sentence-bert: Sentence embeddings using siamese bert-networks

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:11:20.137882Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T05:11:19.951109Z digest=sha256:f1ad4e82ea2d0ff2f0d49455639c63368f23c4cd375f5e7d0b2d341fa899e33d

Observation b7d9afb5-6413-4072-9713-88931f8cef68 · outbound

This paper cites Action scene graphs for long- form understanding of egocentric videos, 2023.

From Pixels to Graphs: using Scene and Knowledge Graphs for HD-EPIC VQA Challenge Action scene graphs for long- form understanding of egocentric videos, 2023

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:11:20.125175Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T05:11:19.954515Z digest=sha256:7100e13e0eff161a16c2805dee98e0c832bab3b675f243ad40bb47280cf05af2

Observation 7063de9e-9a9b-43dd-be6c-d82e2d42bdb8 · outbound

This paper cites Atomic: An atlas of ma- chine commonsense for if-then reasoning.

From Pixels to Graphs: using Scene and Knowledge Graphs for HD-EPIC VQA Challenge Atomic: An atlas of ma- chine commonsense for if-then reasoning

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:11:20.113042Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T05:11:19.959085Z digest=sha256:cf9b36d6fde6be827b7b5fa111896a10c4bb75e39a8f6d2b5691d46c7b02679a

Observation eb06ebe3-fb83-4fbe-a318-3a0aee9005a0 · outbound

This paper cites Concept- net 5.5: An open multilingual graph of general knowledge.

From Pixels to Graphs: using Scene and Knowledge Graphs for HD-EPIC VQA Challenge Concept- net 5.5: An open multilingual graph of general knowledge

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:11:20.100797Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T05:11:19.963013Z digest=sha256:c90b31d09d2d659d3b12afe857e227f944bd786b487a2060c68368c28e424a38

Observation 61b922e8-d390-45f0-918f-9cc63241f79a · outbound

This paper cites Concept- net 5.5: An open multilingual graph of general knowledge,.

From Pixels to Graphs: using Scene and Knowledge Graphs for HD-EPIC VQA Challenge Concept- net 5.5: An open multilingual graph of general knowledge,

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:11:20.087645Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T05:11:19.966591Z digest=sha256:be6de8ade30a8f0096e812b0dc53d87659949f002b3a3f552584d31a0b26df3f

Observation 26dc9ebc-4dac-454d-b468-42cee761571f · outbound

This paper cites Learning situation hyper-graphs for video question answer- ing.

From Pixels to Graphs: using Scene and Knowledge Graphs for HD-EPIC VQA Challenge Learning situation hyper-graphs for video question answer- ing

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:11:20.074898Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T05:11:19.970671Z digest=sha256:15ce9323f67eb3a11994f38d70557fdcb84e2fe962411edd65affb792bce73f3

Observation 4f776e03-62e6-4fe9-b272-7deba1911577 · outbound

This paper cites Scene graph generation by iterative message passing.

From Pixels to Graphs: using Scene and Knowledge Graphs for HD-EPIC VQA Challenge Scene graph generation by iterative message passing

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:11:20.062960Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T05:11:19.975080Z digest=sha256:816b0a7782c2b837a0d3c4c96e8f32e7dae2c347d5aac8b62605150e23e93ec7

Observation 957fbb91-98dc-4674-b2db-2fe9e6a755cd · outbound

This paper cites Neuro-symbolic visual reasoning: Disentangling ”vi- sual” from ”reasoning”.

From Pixels to Graphs: using Scene and Knowledge Graphs for HD-EPIC VQA Challenge Neuro-symbolic visual reasoning: Disentangling ”vi- sual” from ”reasoning”

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:11:20.050199Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T05:11:19.978859Z digest=sha256:6235b998f848da710bf8cab6dac12db1990a4a213e1e7dd46f7f4b92dd0dd1ad

Observation 692377ea-c82e-4a12-bfc1-c2e2479fe93c · outbound

This paper cites Jasper and stella: distillation of sota embedding models,.

From Pixels to Graphs: using Scene and Knowledge Graphs for HD-EPIC VQA Challenge Jasper and stella: distillation of sota embedding models,

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:11:20.037717Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T05:11:19.983039Z digest=sha256:89cdb68b232fbd39757493814657fe4dede879d7e6e88e64d8a55cb1ac4e9984

Pith citing papers

Observation 2f522f67-b64e-4061-b04d-5487d7ea1536 · inbound

Learning to Deny: Action Denial in Multimodal Large Language Models cites this paper.

Learning to Deny: Action Denial in Multimodal Large Language Models From Pixels to Graphs: using Scene and Knowledge Graphs for HD-EPIC VQA Challenge

Reference 67

Resolution
verified exact
arxiv_id, observed 2026-07-01T06:25:26.952478Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-07-01T06:21:09.996386Z digest=sha256:b9e9af35448a345a2c0f403223a09d6d067ca35cbdc5f4fca7ea04b2a24ca364