Pith. sign in

Paper Citation Record · LEDGER

ENTER: Event Based Interpretable Reasoning for VideoQA

As of 5 August 2026, this Paper Citation Record lists 54 of 54 outbound references and 1 inbound Pith citation observation for arXiv:2501.14194.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2501.14194 v2

Coverage vector

measured 54 of 54 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-05-23T05:01:06.299758Z

measured 55 of 55 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-05T06:32:48.257954+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-05-08T08:41:42.061219Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-05-11T20:31:13.191851Z

Reference resolution

54 of 54 outbound references displayed

  • verified exact1
  • verified fuzzy50
  • unresolved3
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation d07d8b03-e851-4f23-96c4-c44d70652789 · outbound

This paper cites Neural module networks.

ENTER: Event Based Interpretable Reasoning for VideoQA Neural module networks

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T05:05:24.990893Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-23T05:01:06.299758Z digest=sha256:ea4dea9d18f25e8f2f5ded0be2746c384c24da5c43866b50127b15bdd0dfa9a2

Observation 8699864a-d3a2-4c52-a355-91571959b530 · outbound

This paper cites Hiervl: Learning hierarchical video-language embeddings.

ENTER: Event Based Interpretable Reasoning for VideoQA Hiervl: Learning hierarchical video-language embeddings

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T05:05:24.994488Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-23T05:01:06.299758Z digest=sha256:31e8c80e58abfec58a47b9fc2b8b4bd3475bf1175a045409011eb21f60ea0e57

Observation de589788-3fb3-4e9c-bb5b-9447450f977e · outbound

This paper cites an unresolved cited work.

ENTER: Event Based Interpretable Reasoning for VideoQA Unresolved cited work

Reference 3

Resolution
unresolved
raw_fallback, observed 2026-05-23T05:05:24.985455Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-23T05:01:06.299758Z digest=sha256:661ad70b609fd28dde404ea8d5d22df4f66b92d6cb4a6e1c8413c5a6a275a4e5

Observation 6a054b05-a283-419b-bd53-51d98f322e95 · outbound

This paper cites Grounded multi- hop videoqa in long-form egocentric videos.

ENTER: Event Based Interpretable Reasoning for VideoQA Grounded multi- hop videoqa in long-form egocentric videos

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T05:02:37.270670Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-23T05:01:06.299758Z digest=sha256:f46343b8ccd5502c4c51478b53b6c1d47061404564c66d30bb5e6ad2c35812d3

Observation d32f3b1f-6ffa-46ce-91fd-e4dae113264e · outbound

This paper cites Kitani, and László A.

ENTER: Event Based Interpretable Reasoning for VideoQA Kitani, and László A

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T05:05:24.988331Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-23T05:01:06.299758Z digest=sha256:1001e601efb17332bd5236d56c882c39270d964d83500b947962995c0f747a36

Observation 6e6bd0ec-03a1-4329-b41a-038edf6a51d9 · outbound

This paper cites The llama 3 herd of models.

ENTER: Event Based Interpretable Reasoning for VideoQA The llama 3 herd of models

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T05:05:24.979794Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-23T05:01:06.299758Z digest=sha256:0e37e32bbe72c18d316dd223ec11045cc65ed484b52e2b7af7c0ebd91761084c

Observation 15544bb3-8e15-4064-aece-eaee78c0082f · outbound

This paper cites Videoagent: A memory-augmented multi- modal agent for video understanding.

ENTER: Event Based Interpretable Reasoning for VideoQA Videoagent: A memory-augmented multi- modal agent for video understanding

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T05:05:24.997794Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-23T05:01:06.299758Z digest=sha256:24339506d3671334ebc18122df39264306f1decdb5877c86b324405e6501e1b5

Observation ddef7a3a-d5c7-4f5c-b674-2fa461a00acb · outbound

This paper cites Video-of-thought: Step-by-step video reasoning from perception to cognition.

ENTER: Event Based Interpretable Reasoning for VideoQA Video-of-thought: Step-by-step video reasoning from perception to cognition

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T05:02:37.274890Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-23T05:01:06.299758Z digest=sha256:ef56922a1808770c36f3bd6ba121a2c18f7cf4a3022d9f54ae5805b3c55e2cf9

Observation dc1f27ce-957e-4b37-bedc-182de65930f0 · outbound

This paper cites Visual program- ming: Compositional visual reasoning without training.

ENTER: Event Based Interpretable Reasoning for VideoQA Visual program- ming: Compositional visual reasoning without training

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T05:05:24.982429Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-23T05:01:06.299758Z digest=sha256:5b425e9a52ff2fa6e5e623a3a72731114052961e833732aa194f7502c059342c

Observation e990e8d3-d1ea-418f-b7ea-f5a5329c3250 · outbound

This paper cites Free video-llm: Prompt-guided visual perception for efficient training-free video llms.

ENTER: Event Based Interpretable Reasoning for VideoQA Free video-llm: Prompt-guided visual perception for efficient training-free video llms

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T05:05:24.949958Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-23T05:01:06.299758Z digest=sha256:120c5b9834fb478e1941a063e70b8b13c4016770ccbcf3ad4edd040a77d04a08

Observation 896c9d79-5826-414d-8021-02c1fa92a8e5 · outbound

This paper cites Video recap: Recursive captioning of hour-long videos.

ENTER: Event Based Interpretable Reasoning for VideoQA Video recap: Recursive captioning of hour-long videos

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T05:05:24.982252Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-23T05:01:06.299758Z digest=sha256:f53f88435d0c10e4eecb97858cc7accf835d9a1da1490746eeb616315239e1ee

Observation cd80d71d-2e03-4afb-b129-522ffd695730 · outbound

This paper cites an unresolved cited work.

ENTER: Event Based Interpretable Reasoning for VideoQA Unresolved cited work

Reference 12

Resolution
unresolved
raw_fallback, observed 2026-05-23T05:05:24.911485Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-23T05:01:06.299758Z digest=sha256:cf07453f6c1aec9b097df0a550232264f89138869bd75ed2f5629c657790a694

Observation 90b8005a-a565-4389-80dc-f0838b7a2a14 · outbound

This paper cites Intentqa: Context-aware video intent reasoning.

ENTER: Event Based Interpretable Reasoning for VideoQA Intentqa: Context-aware video intent reasoning

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T05:05:24.946657Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-23T05:01:06.299758Z digest=sha256:9abf97061f537c7b6fac94084ca576af819624c4772c326bd74c24cdf51c5eb4

Observation 3171786e-4afe-40f0-815d-4b0b656c054d · outbound

This paper cites Intentqa: Context-aware video intent reasoning.

ENTER: Event Based Interpretable Reasoning for VideoQA Intentqa: Context-aware video intent reasoning

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T05:05:24.952755Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-23T05:01:06.299758Z digest=sha256:ec17272f4a08e314c14674ff2b48a72d2781556f23bbe2aa51391ac0d4c08757

Observation b15977f6-c395-45a1-8392-193ebd2a4982 · outbound

This paper cites Videochat: Chat-centric video understanding.

ENTER: Event Based Interpretable Reasoning for VideoQA Videochat: Chat-centric video understanding

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T05:05:24.936590Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-23T05:01:06.299758Z digest=sha256:f0a32896cfb96631370e2376dc0cad0aee8c5a3cfd5dc351ea0b29528be47704

Observation 9440fa02-5773-45f2-b853-e3c76e922f59 · outbound

This paper cites Mvbench: A comprehensive multi-modal video understanding benchmark.

ENTER: Event Based Interpretable Reasoning for VideoQA Mvbench: A comprehensive multi-modal video understanding benchmark

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T05:05:24.956512Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-23T05:01:06.299758Z digest=sha256:77e0e6703f82fa063872d58a7b7dadd91b5fd8469c04f67223b52cb2d9fa0ab8

Observation 58dc5d70-4ba1-4ab3-ae60-dec492c3812e · outbound

This paper cites End-to-end video question answering with frame scoring mechanisms and adaptive sam- pling.

ENTER: Event Based Interpretable Reasoning for VideoQA End-to-end video question answering with frame scoring mechanisms and adaptive sam- pling

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T05:05:24.959343Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-23T05:01:06.299758Z digest=sha256:47ba0d87944543ce960348c25d5b7581b1ab218e43a0771e0a7764d66765b197

Observation c6ea5862-fb7a-4aef-a489-18ff8209eaef · outbound

This paper cites VideoIN- STA: Zero-shot long video understanding via informative spatial-temporal reasoning with LLMs.

ENTER: Event Based Interpretable Reasoning for VideoQA VideoIN- STA: Zero-shot long video understanding via informative spatial-temporal reasoning with LLMs

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T05:05:24.937177Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-23T05:01:06.299758Z digest=sha256:43dd1588991e8d50c558bf0fc532c04d19b5b0d3cab5ab45e5193d5db6d3539b

Observation c5a1ddf9-00b0-4bae-a3aa-e0dfa638d972 · outbound

This paper cites Vx2text: End-to-end learning of video-based text generation from multimodal in- puts.

ENTER: Event Based Interpretable Reasoning for VideoQA Vx2text: End-to-end learning of video-based text generation from multimodal in- puts

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T05:05:24.939673Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-23T05:01:06.299758Z digest=sha256:05a6a51b3edd2596e18e1a2a4e7639c4d6b3da81a8d5b7ebd8db51b78cb7719e

Observation a5be5e2c-e8a5-4970-88e5-a648e50a850d · outbound

This paper cites Towards fast adaptation of pretrained contrastive models for multi-channel video-language retrieval.

ENTER: Event Based Interpretable Reasoning for VideoQA Towards fast adaptation of pretrained contrastive models for multi-channel video-language retrieval

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T05:05:24.873754Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-23T05:01:06.299758Z digest=sha256:4229a561a682f3ad4e179e32700cbef0a1e4e065809a4be4e15c56e2fed02c3d

Observation f7d45b12-2098-4069-a197-646927d4d0bd · outbound

This paper cites Training-free deep concept injection enables lan- guage models for video question answering.

ENTER: Event Based Interpretable Reasoning for VideoQA Training-free deep concept injection enables lan- guage models for video question answering

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T05:05:25.030597Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-23T05:01:06.299758Z digest=sha256:db737ac54a21621f3a881512c060aea26191451a29e5e2e7f202469887aa5693

Observation be98f818-536d-411e-b3a8-5d232fa30cdb · outbound

This paper cites Hair: Hierarchical visual-semantic relational reasoning for video question answering.

ENTER: Event Based Interpretable Reasoning for VideoQA Hair: Hierarchical visual-semantic relational reasoning for video question answering

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T05:05:24.922613Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-23T05:01:06.299758Z digest=sha256:538a30f7afb0c580b1420b8b544f48933ce7c4ddb38a4f4d32f36865c4ead443

Observation 44b67fe6-a40f-441f-9a58-bcf85655c5dd · outbound

This paper cites Video-chatgpt: Towards detailed video understanding via large vision and language models.

ENTER: Event Based Interpretable Reasoning for VideoQA Video-chatgpt: Towards detailed video understanding via large vision and language models

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T05:05:24.926521Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-23T05:01:06.299758Z digest=sha256:1eb684c64b454d1a238e0e165efffc3bc65c67af6c1fa7a5dce9af1df79b5c7a

Observation 65fa6e4c-3285-4513-b444-d4ea0482cf19 · outbound

This paper cites Egoschema: A diagnostic benchmark for very long- form video language understanding.

ENTER: Event Based Interpretable Reasoning for VideoQA Egoschema: A diagnostic benchmark for very long- form video language understanding

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T05:05:24.919876Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-23T05:01:06.299758Z digest=sha256:d2d53590f06687d0f1283e3488e5d6faf49f0e7210f88820f2a4f03d283fb272

Observation 00c61d5e-95d3-4662-9a73-c563c5fe94b4 · outbound

This paper cites Morevqa: Exploring modular reasoning models for video question answering.

ENTER: Event Based Interpretable Reasoning for VideoQA Morevqa: Exploring modular reasoning models for video question answering

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T05:05:25.033628Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-23T05:01:06.299758Z digest=sha256:120f4d3289291b0717000072f5e170d30f16791b0abc70b68b07ffd8fbbac3a7

Observation 7f61adff-ff87-4656-99b4-fc396cff212e · outbound

This paper cites Question-instructed visual descriptions for zero-shot video answering.

ENTER: Event Based Interpretable Reasoning for VideoQA Question-instructed visual descriptions for zero-shot video answering

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T05:05:24.988970Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-23T05:01:06.299758Z digest=sha256:5080a07c45255df4eea9aedad558c186b733fa241e2ee90e4362de184cc4e1d3

Observation d4485737-3411-4d71-bbf1-95e07bf34cb8 · outbound

This paper cites Correlation-Guided Query-Dependency Calibration for Video Temporal Grounding.

ENTER: Event Based Interpretable Reasoning for VideoQA Correlation-Guided Query-Dependency Calibration for Video Temporal Grounding

Reference 27

Resolution
verified exact
arxiv_id, observed 2026-05-23T05:02:35.938663Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-23T05:01:06.299758Z digest=sha256:27180608a2e7ae24b38049b2067df6d2642c5f6f1d83d6e7852ee9ad8419d6eb

Observation af6a8a11-0812-44d9-a385-0462a68822e6 · outbound

This paper cites Gpt-4 technical report.

ENTER: Event Based Interpretable Reasoning for VideoQA Gpt-4 technical report

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T05:05:24.962422Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-23T05:01:06.299758Z digest=sha256:bfbfd69c4bd8842e54a7f934a0c684470e467e4ae676bb0ac67b84fbc31b7bc4

Observation 594ff0a8-6fc3-4906-ad49-cd492f120dad · outbound

This paper cites Retrieving-to-answer: Zero-shot video question answering with frozen large lan- guage models.

ENTER: Event Based Interpretable Reasoning for VideoQA Retrieving-to-answer: Zero-shot video question answering with frozen large lan- guage models

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T05:05:25.012879Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-23T05:01:06.299758Z digest=sha256:5fdbb462e8f07da9a82af5f7812248ab369e0a4c0f6c409685cdebda16c84ca3

Observation be66a808-9b32-4fbf-9d6e-ec488c361027 · outbound

This paper cites an unresolved cited work.

ENTER: Event Based Interpretable Reasoning for VideoQA Unresolved cited work

Reference 30

Resolution
unresolved
raw_fallback, observed 2026-05-23T05:05:24.949307Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-23T05:01:06.299758Z digest=sha256:56660e4ba47f87229224f5c742ea996dd804a75ac61dea1caf2cc0e540d91678

Observation 07ac8de7-76db-47f6-a70a-6b9171a03db6 · outbound

This paper cites Micap: A unified model for identity- aware movie descriptions.

ENTER: Event Based Interpretable Reasoning for VideoQA Micap: A unified model for identity- aware movie descriptions

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T05:05:24.979491Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-23T05:01:06.299758Z digest=sha256:2c7e706344e0233554a9d8897818cf719374c1a9dcc609a4d9e0f9001b4898e9

Observation b9e2e6c4-83a0-4678-9b8e-e55ca3d93f79 · outbound

This paper cites Traveler: A modular multi-lmm agent framework for video question-answering.

ENTER: Event Based Interpretable Reasoning for VideoQA Traveler: A modular multi-lmm agent framework for video question-answering

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T05:05:25.027244Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-23T05:01:06.299758Z digest=sha256:47bae242d824fa9c39f9a37d5579bd5c9fe9db2fc6ef1b2c770cbacefed02c8c

Observation dbd6406f-eede-4283-87aa-2c0473cbc8fe · outbound

This paper cites Kim, Bilge Soran, Raghuraman Krishnamoorthi, Mohamed Elhoseiny, and Vikas Chandra.

ENTER: Event Based Interpretable Reasoning for VideoQA Kim, Bilge Soran, Raghuraman Krishnamoorthi, Mohamed Elhoseiny, and Vikas Chandra

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T05:05:24.929695Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-23T05:01:06.299758Z digest=sha256:796e73d69d901bf6490b9b4977a0558c25b4bca656ec6a15a2faa95749153700

Observation a905d258-c2e8-4213-91e6-089eb8141f0e · outbound

This paper cites Progprompt: Generating situated robot task plans using large language models.

ENTER: Event Based Interpretable Reasoning for VideoQA Progprompt: Generating situated robot task plans using large language models

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T05:05:24.992579Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-23T05:01:06.299758Z digest=sha256:6bdc58c968891d939358b78c7d7d7efb5eaf59e5b204e36bac4950c8e54c2aeb

Observation 635f295a-feae-430a-b6c7-081446c85ff8 · outbound

This paper cites Moviechat: From dense token to sparse memory for long video understanding.

ENTER: Event Based Interpretable Reasoning for VideoQA Moviechat: From dense token to sparse memory for long video understanding

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T05:05:25.009653Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-23T05:01:06.299758Z digest=sha256:d43900be2a5c095db54e326f9ee337d4f41be96abf621fb45e43eb9fde83dfad

Observation dab82b8d-8f7f-4d33-b091-7e7b9e46447d · outbound

This paper cites Modular visual question answering via code generation.

ENTER: Event Based Interpretable Reasoning for VideoQA Modular visual question answering via code generation

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T05:05:24.870947Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-23T05:01:06.299758Z digest=sha256:608ace6cb672eff5a4aebe928aef3b2b907144168d14f55e0001682bccb9811a

Observation 1facce90-297f-455a-a708-7ea961c17ae1 · outbound

This paper cites Vipergpt: Visual inference via python execution for reasoning.

ENTER: Event Based Interpretable Reasoning for VideoQA Vipergpt: Visual inference via python execution for reasoning

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T05:05:24.965623Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-23T05:01:06.299758Z digest=sha256:98e74a0580f2e8a6437419dc01751f3aed304be9582271045bde8666df2a363e

Observation 9099eb23-eb5e-4ff5-9e9e-b47ebb119794 · outbound

This paper cites Videoagent: Long-form video understanding with large language model as agent.

ENTER: Event Based Interpretable Reasoning for VideoQA Videoagent: Long-form video understanding with large language model as agent

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T05:05:25.006658Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-23T05:01:06.299758Z digest=sha256:6cf4a1ddee340ab053eb36f5e180d5ebfc0bb43aa37c3740f5267cd4226a1965

Observation 9a661515-dfdb-41d3-9f13-36c6a28d6340 · outbound

This paper cites Internvideo: General video foundation models via generative and discriminative learning.

ENTER: Event Based Interpretable Reasoning for VideoQA Internvideo: General video foundation models via generative and discriminative learning

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T05:05:25.016989Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-23T05:01:06.299758Z digest=sha256:e007414e3e6be972315261b345f41d736e79fb6cad5579866ee4e42f7bd0229f

Observation 49b4b98c-63ed-4160-b8e6-65a74784b721 · outbound

This paper cites Stair: Spatial-temporal reasoning with auditable intermediate results for video question answering.

ENTER: Event Based Interpretable Reasoning for VideoQA Stair: Spatial-temporal reasoning with auditable intermediate results for video question answering

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T05:05:24.881287Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-23T05:01:06.299758Z digest=sha256:b87811edf149434e6060521a1ce82309952853b2946631694567c9a0327b717f

Observation 8771a3d3-480f-4c70-b78c-9842f3196bb3 · outbound

This paper cites Videotree: Adaptive tree-based video representation for llm reasoning on long videos.

ENTER: Event Based Interpretable Reasoning for VideoQA Videotree: Adaptive tree-based video representation for llm reasoning on long videos

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T05:05:25.023611Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-23T05:01:06.299758Z digest=sha256:ca32cae0ca31cf4c13c674cc89f396d2fb364f8a390c3586369080fd42f645da

Observation be423fa0-f946-4500-907f-1795e19430c9 · outbound

This paper cites Freeva: Offline mllm as training-free video assistant.

ENTER: Event Based Interpretable Reasoning for VideoQA Freeva: Offline mllm as training-free video assistant

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T05:05:24.995907Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-23T05:01:06.299758Z digest=sha256:fe9c14062a0e77a351c5045eb5f12de25ee80f724e9c5c8af332a25917d537ac

Observation b2282c50-4b72-46c9-9da6-e55757e63cff · outbound

This paper cites Next-qa:next phase of question-answering to explaining tem- poral actions.

ENTER: Event Based Interpretable Reasoning for VideoQA Next-qa:next phase of question-answering to explaining tem- poral actions

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T05:05:24.976858Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-23T05:01:06.299758Z digest=sha256:fd9451bf2f28d3ff501f8a534c65f5b1edd6599b4c38a2637f69422970bc7475

Observation 775fe7ea-d826-41cc-933c-ec4acb7e6195 · outbound

This paper cites Next-qa:next phase of question-answering to explaining tem- poral actions.

ENTER: Event Based Interpretable Reasoning for VideoQA Next-qa:next phase of question-answering to explaining tem- poral actions

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T05:05:24.955730Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-23T05:01:06.299758Z digest=sha256:5b8eea9d844e8c06b053ed0e51ddbdd0179a8098d1f93a535205ef85432cd38d

Observation 660226c7-617c-45c3-a678-7215f34516a9 · outbound

This paper cites Slowfast-llava: A strong training-free baseline for video large language models.

ENTER: Event Based Interpretable Reasoning for VideoQA Slowfast-llava: A strong training-free baseline for video large language models

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T05:05:24.969341Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-23T05:01:06.299758Z digest=sha256:2d0bed9a3815aeb8537bb8324213a3b588e87c020781deb33846d711110b9717

Observation 5e76f76e-502a-4372-948e-95ed9fa1725a · outbound

This paper cites Panoptic video scene graph generation.

ENTER: Event Based Interpretable Reasoning for VideoQA Panoptic video scene graph generation

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T05:05:24.999054Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-23T05:01:06.299758Z digest=sha256:0145b4f3cedffb7cb7d2f49a57cf2013aa3db9560536dc8ea42b5c40fa064b52

Observation 31b81c97-78c9-4879-9f44-5ce2586c6d0d · outbound

This paper cites Mm-react: Prompting chatgpt for multimodal reasoning and action.

ENTER: Event Based Interpretable Reasoning for VideoQA Mm-react: Prompting chatgpt for multimodal reasoning and action

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T05:05:24.959141Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-23T05:01:06.299758Z digest=sha256:ec3feffcc0f6337e1904df79a89b3e4db67179d89db7e21e694a954c73a75149

Observation 54c34af3-7868-4e7b-9ba1-0e3e9e8c4dbf · outbound

This paper cites Ayyubi, Kai-Wei Chang, and Shih-Fu Chang.

ENTER: Event Based Interpretable Reasoning for VideoQA Ayyubi, Kai-Wei Chang, and Shih-Fu Chang

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T05:05:25.003171Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-23T05:01:06.299758Z digest=sha256:f2eda74e7a8c99d3ba43f8d8ab0d880e5f93344323198be361a3bfeeb36dae72

Observation 9bf4f48a-f555-4c4a-bbe0-fea2236f7e89 · outbound

This paper cites Self-chained image-language model for video localization and question answering.

ENTER: Event Based Interpretable Reasoning for VideoQA Self-chained image-language model for video localization and question answering

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T05:05:25.020206Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-23T05:01:06.299758Z digest=sha256:a61bbac187016532219da3ca84cdb64d2fe8572ce00ee0dc1759f99e0f3e46a6

Observation c1a98543-ef4a-4ec6-be37-0ba32ee7884c · outbound

This paper cites Activitynet-qa: A dataset for understanding complex web videos via question answering.

ENTER: Event Based Interpretable Reasoning for VideoQA Activitynet-qa: A dataset for understanding complex web videos via question answering

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T05:05:24.975881Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-23T05:01:06.299758Z digest=sha256:c037c16b72675b086a11be20f360f23008774cd9ad40ae5345d27caccd1c2a51

Observation 0566ea3d-888b-49b9-9fe6-484a693fc0b0 · outbound

This paper cites A simple llm framework for long-range video question-answering.

ENTER: Event Based Interpretable Reasoning for VideoQA A simple llm framework for long-range video question-answering

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T05:05:24.962199Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-23T05:01:06.299758Z digest=sha256:c0705f10167287b4c8cff72ca7a3214e7e53005ae10f682a6b0cec39f40e02b1

Observation 50ebef7d-9e3f-4eae-bf1b-6ce8e5466c5d · outbound

This paper cites A simple llm framework for long-range video question-answering.

ENTER: Event Based Interpretable Reasoning for VideoQA A simple llm framework for long-range video question-answering

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T05:05:24.968679Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-23T05:01:06.299758Z digest=sha256:f5e727eb980d178db2a9872dbc91027eac5aaa5060bfffc07ba1427c99ff0a3c

Observation 1dc9fa09-9994-4065-b641-d050ffdbabe1 · outbound

This paper cites Video-llama: An instruction-tuned audio-visual language model for video un- derstanding.

ENTER: Event Based Interpretable Reasoning for VideoQA Video-llama: An instruction-tuned audio-visual language model for video un- derstanding

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T05:05:24.985768Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-23T05:01:06.299758Z digest=sha256:26aaa79df5c5bc1ccced7a992b03415137f0eb7fee832e27eb058666e6adb008

Observation a4353d16-04fb-4e7c-81b0-313e5c1399cf · outbound

This paper cites Base" component refers to cases where the generated code operates without triggering addi- tional modules like the.

ENTER: Event Based Interpretable Reasoning for VideoQA Base" component refers to cases where the generated code operates without triggering addi- tional modules like the

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T05:05:24.972387Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-23T05:01:06.299758Z digest=sha256:9b08de375f51dd177f99e1e38eb9c738a6b86e2be105aaacca13b903f3572c1c

Pith citing papers

Observation e091819e-610d-4270-a066-990a1831f0e7 · inbound

UpstreamQA: A Modular Framework for Explicit Reasoning on Video Question Answering Tasks cites this paper.

UpstreamQA: A Modular Framework for Explicit Reasoning on Video Question Answering Tasks ENTER: Event Based Interpretable Reasoning for VideoQA

Reference 14

Resolution
metadata mismatch
local_arxiv, observed 2026-05-11T20:31:13.198639Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-08T08:41:42.061219Z digest=sha256:9adeb3452996cff4cf45230dd154a6c0c35df14210f4bdbf714d22c8b73120a2