Pith. sign in

Paper Citation Record · LEDGER

Perceive, Query & Reason: Enhancing Video QA with Question-Guided Temporal Queries

As of 21 August 2026, this Paper Citation Record lists 58 of 58 outbound references and 0 inbound Pith citation observations for arXiv:2412.19304.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2412.19304 v1

Coverage vector

measured 58 of 58 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-11T00:48:43.194779Z

measured 58 of 58 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-20T06:33:59.587034+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

58 of 58 outbound references displayed

  • verified exact0
  • verified fuzzy47
  • unresolved11
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 154d752b-cdaf-4547-95df-d3292b0c188c · outbound

This paper cites Flamingo: a visual language model for few-shot learning.NeurIPS, 2022.

Perceive, Query & Reason: Enhancing Video QA with Question-Guided Temporal Queries Flamingo: a visual language model for few-shot learning.NeurIPS, 2022

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T00:48:44.230547Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-11T00:48:42.266101Z digest=sha256:6257c016393d5ed5e90191b4bf3f855e3982e7eef5d9fec87955c9b99dba6ac4

Observation 3834a3f1-95d1-48f8-8a68-6252e6677220 · outbound

This paper cites Frozen in time: A joint video and image encoder for end-to-end retrieval.

Perceive, Query & Reason: Enhancing Video QA with Question-Guided Temporal Queries Frozen in time: A joint video and image encoder for end-to-end retrieval

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T00:48:44.221159Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-11T00:48:42.269674Z digest=sha256:165bfae825166e02bcdc737d9b2fc340e0c32c00dd2435c7d522bf24490c4478

Observation a2c0e827-b6c8-4902-ab8d-c77a0ccd632d · outbound

This paper cites Revisiting the” video” in video-language understanding.

Perceive, Query & Reason: Enhancing Video QA with Question-Guided Temporal Queries Revisiting the” video” in video-language understanding

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T00:48:44.211765Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-11T00:48:42.273627Z digest=sha256:29761a868a9b289c2de2b2599c269d323b9316053c600015468413bf5afe9133

Observation f46a1479-f897-460a-bd69-4eb73b0c87b9 · outbound

This paper cites The (r) evolution of multimodal large language models: A survey.

Perceive, Query & Reason: Enhancing Video QA with Question-Guided Temporal Queries The (r) evolution of multimodal large language models: A survey

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T00:48:44.201425Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-11T00:48:42.277666Z digest=sha256:5b5ff7be166306e84b9007b27bd8121007d5a4cb1e34df9a2b13c5660b00ffee

Observation 0ead5ca2-15f8-4d63-97e3-9916d54a91ca · outbound

This paper cites Pali: A jointly- scaled multilingual language-image model.

Perceive, Query & Reason: Enhancing Video QA with Question-Guided Temporal Queries Pali: A jointly- scaled multilingual language-image model

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T00:48:44.191889Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-11T00:48:42.282350Z digest=sha256:d9834443334634baeb54f2483dca59b2e1c818de1aa655f071755cb7cc5e6bed

Observation 93d004f6-3dbb-49e9-9092-5641b9b1fbea · outbound

This paper cites Scaling instruction- finetuned language models.

Perceive, Query & Reason: Enhancing Video QA with Question-Guided Temporal Queries Scaling instruction- finetuned language models

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T00:48:44.181850Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-11T00:48:42.286120Z digest=sha256:7fd3cd6fd19f77950092379eb86744c67921626c7517076c3a63284c69b3161b

Observation cf0827a6-ce71-40b2-a9e8-9d2227f22cfe · outbound

This paper cites Instructblip: Towards general- purpose vision-language models with instruction tuning.

Perceive, Query & Reason: Enhancing Video QA with Question-Guided Temporal Queries Instructblip: Towards general- purpose vision-language models with instruction tuning

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T00:48:44.171305Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-11T00:48:42.289481Z digest=sha256:12ccf4d2305afb11d94ece188d5dd027f3def176e6857fd237f25a9f88344985

Observation 0c251260-2d6c-46b9-aa5d-c19a614ee018 · outbound

This paper cites Eva: Exploring the limits of masked visual representa- tion learning at scale.

Perceive, Query & Reason: Enhancing Video QA with Question-Guided Temporal Queries Eva: Exploring the limits of masked visual representa- tion learning at scale

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T00:48:44.162785Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-11T00:48:42.292700Z digest=sha256:09d087db6020f1607c8e9987fcf05ff4cbeb1ab0451b94271e114791d5c46e11

Observation e47e539d-ae72-444c-8722-4c0424e1522a · outbound

This paper cites Mist: Multi-modal iterative spatial- temporal transformer for long-form video question answer- ing.

Perceive, Query & Reason: Enhancing Video QA with Question-Guided Temporal Queries Mist: Multi-modal iterative spatial- temporal transformer for long-form video question answer- ing

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T00:48:44.153607Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-11T00:48:42.295862Z digest=sha256:24c770163d6c8198d41c845fe281bb6ddf5097bab07db95810b813d3b9af4356

Observation ef91b739-7a94-4ec2-9c9a-4411cd3a51ad · outbound

This paper cites Agqa: A benchmark for compositional spatio-temporal reasoning.

Perceive, Query & Reason: Enhancing Video QA with Question-Guided Temporal Queries Agqa: A benchmark for compositional spatio-temporal reasoning

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T00:48:44.143283Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-11T00:48:42.299225Z digest=sha256:a3c0c4ba08b22129684a8fbdab3d1fca2ecc3a0f47e2b6ef9858f424d8444c37

Observation fa560ae0-4620-4ee8-9589-cdaa47c47993 · outbound

This paper cites Tada! temporally-adaptive convolutions for video understanding.

Perceive, Query & Reason: Enhancing Video QA with Question-Guided Temporal Queries Tada! temporally-adaptive convolutions for video understanding

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T00:48:44.132668Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-11T00:48:42.302448Z digest=sha256:639b90090d44ed032b4c21a32e4349d03f557e706fc0f25992d7016ee7761df7

Observation 300391c7-4964-4b96-b71c-401de2653535 · outbound

This paper cites Tgif-qa: Toward spatio-temporal reasoning in visual question answering.

Perceive, Query & Reason: Enhancing Video QA with Question-Guided Temporal Queries Tgif-qa: Toward spatio-temporal reasoning in visual question answering

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T00:48:44.119773Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-11T00:48:42.305360Z digest=sha256:d26c9d7690037def08a004e12f83358f0010674f936487ef64cf3b88083c0630

Observation 2755f11c-296a-4c16-870a-7a8b25413931 · outbound

This paper cites Large language models are tempo- ral and causal reasoners for video question answering.

Perceive, Query & Reason: Enhancing Video QA with Question-Guided Temporal Queries Large language models are tempo- ral and causal reasoners for video question answering

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T00:48:44.108171Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-11T00:48:42.308480Z digest=sha256:78e5d7c17cdaf19c77e0fc982cd20cde1f49a1c641b93009cc1bb8186a5074f4

Observation bcb966c4-866a-4942-9ccf-91192b14637c · outbound

This paper cites Large language models are zero-shot reasoners.

Perceive, Query & Reason: Enhancing Video QA with Question-Guided Temporal Queries Large language models are zero-shot reasoners

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T00:48:44.097181Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-11T00:48:42.316529Z digest=sha256:bd5030658b0e640a2a81e807d3b3c161023529fa1683336b2d778290df16ac29

Observation cf77afc9-9433-43bd-ad72-e9c591a4f4d1 · outbound

This paper cites Neural reasoning, fast and slow, for video question answering.

Perceive, Query & Reason: Enhancing Video QA with Question-Guided Temporal Queries Neural reasoning, fast and slow, for video question answering

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T00:48:44.086032Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-11T00:48:42.367581Z digest=sha256:aa64cdb1494ca0b6f4e7fd92d657960933af29317073ebc577466b342295d1e3

Observation 6a20bb95-1427-49d8-bec0-7565e8dd7b68 · outbound

This paper cites Revealing single frame bias for video-and-language learning.

Perceive, Query & Reason: Enhancing Video QA with Question-Guided Temporal Queries Revealing single frame bias for video-and-language learning

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T00:48:44.075261Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-11T00:48:42.424289Z digest=sha256:8ef669a6a02b9074123c95f5586b6ad72943166b371fc1c188536768a09ab268

Observation de36ac1d-2c98-4b61-a5de-908ad88ef7eb · outbound

This paper cites Less is more: Clipbert for video-and-language learning via sparse sampling.

Perceive, Query & Reason: Enhancing Video QA with Question-Guided Temporal Queries Less is more: Clipbert for video-and-language learning via sparse sampling

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-11T00:48:42.468319Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T00:48:42.468319Z digest=sha256:6e905d47b6af9e17758166a029079d7233bf9f52510be6e8d9e6ccd8cb7c26ef

Observation e57abf6b-173f-4c08-9d03-828dfbb7e30f · outbound

This paper cites Tvqa: Localized, compositional video question answering.

Perceive, Query & Reason: Enhancing Video QA with Question-Guided Temporal Queries Tvqa: Localized, compositional video question answering

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T00:48:44.057547Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-11T00:48:42.558423Z digest=sha256:fba8e759bad67549248169b2b65efc5c1de64321e28fbf98d66e5ee0b11749dc

Observation 9c99da35-7668-4bfe-9be9-8436efe586ee · outbound

This paper cites What is more likely to happen next? video-and-language future event prediction.

Perceive, Query & Reason: Enhancing Video QA with Question-Guided Temporal Queries What is more likely to happen next? video-and-language future event prediction

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T00:48:44.045644Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-11T00:48:42.614817Z digest=sha256:9f4ed3f852c9e7976874ca1930ce3a3d128efc24328d7d30a474d43a882986ce

Observation d6ec1f53-c1f9-47f1-85dc-8408e0b4f90d · outbound

This paper cites Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models.

Perceive, Query & Reason: Enhancing Video QA with Question-Guided Temporal Queries Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T00:48:44.034186Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-11T00:48:42.633527Z digest=sha256:defa0a49185883f2f311468ab35116643d2833d91474e6aefc89cd3de0fac026

Observation ffafda43-f63a-46b9-b5fe-8a0a6c098fb9 · outbound

This paper cites Hero: Hierarchical encoder for video+ lan- guage omni-representation pre-training.

Perceive, Query & Reason: Enhancing Video QA with Question-Guided Temporal Queries Hero: Hierarchical encoder for video+ lan- guage omni-representation pre-training

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T00:48:44.024193Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-11T00:48:42.693952Z digest=sha256:1e4f55954a22a2c0017415e2562562480bbe62b2c3de00889d3db3af6b1c104f

Observation 4b334a97-1f5e-4d14-8834-9658b91d9f4a · outbound

This paper cites Beyond rnns: Positional self-attention with co-attention for video question answering.

Perceive, Query & Reason: Enhancing Video QA with Question-Guided Temporal Queries Beyond rnns: Positional self-attention with co-attention for video question answering

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T00:48:44.013279Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-11T00:48:42.724768Z digest=sha256:e87e0be2f9bf2e5562b095f4afc81b72e5463d2f06c20aabd19b2b4fcb65e7e7

Observation 64c3a33d-a1ff-4585-89e4-350b79eacc31 · outbound

This paper cites Discovering spatio-temporal rationales for video question answering.

Perceive, Query & Reason: Enhancing Video QA with Question-Guided Temporal Queries Discovering spatio-temporal rationales for video question answering

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T00:48:44.001343Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-11T00:48:42.754821Z digest=sha256:47a56d6bdea3f615ced718b9da044514ad41f60a031f832f23e250cfe549f481

Observation b0507ba2-b27e-49de-aadb-6bca7632a0d1 · outbound

This paper cites Llava-next: Im- proved reasoning, ocr, and world knowledge, 2024.

Perceive, Query & Reason: Enhancing Video QA with Question-Guided Temporal Queries Llava-next: Im- proved reasoning, ocr, and world knowledge, 2024

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-11T00:48:42.757817Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T00:48:42.757817Z digest=sha256:a782121a470c35c21f27a26f826eb061ddc1b43374f14dd7320033ff8dfefeac

Observation d7a595b0-4a49-43cc-8b8a-55088a58e846 · outbound

This paper cites Visual instruction tuning, 2023.

Perceive, Query & Reason: Enhancing Video QA with Question-Guided Temporal Queries Visual instruction tuning, 2023

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-11T00:48:42.760865Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T00:48:42.760865Z digest=sha256:df90a65452c230915bc16c78fa444c77d0843ce6620ca14e50448ae8210c195c

Observation 322836d5-7009-4da1-b81e-40702cad1c3f · outbound

This paper cites Evaluating the Logical Reasoning Ability of ChatGPT and GPT-4.

Perceive, Query & Reason: Enhancing Video QA with Question-Guided Temporal Queries Evaluating the Logical Reasoning Ability of ChatGPT and GPT-4

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-11T00:48:42.763702Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T00:48:42.763702Z digest=sha256:60a98306533a187eee8d49f4b1c178aa51769c86fba44599cf98ef29454ba884

Observation 64770b96-7032-4c9d-bb6a-428ab7a56a96 · outbound

This paper cites Clip4clip: An empirical study of clip for end to end video clip retrieval.

Perceive, Query & Reason: Enhancing Video QA with Question-Guided Temporal Queries Clip4clip: An empirical study of clip for end to end video clip retrieval

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T00:48:43.977619Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-11T00:48:42.767179Z digest=sha256:2097d920175bd352c83f6d53883496c6669bf98db77ed6791bca52c220253b63

Observation a342ef9f-2bd8-40c0-b299-a33487b9fec0 · outbound

This paper cites Video-chatgpt: Towards detailed video understanding via large vision and language models.

Perceive, Query & Reason: Enhancing Video QA with Question-Guided Temporal Queries Video-chatgpt: Towards detailed video understanding via large vision and language models

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-11T00:48:42.770466Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T00:48:42.770466Z digest=sha256:b1a65823f24af4f502e98547f4b96e5dec5198ce91199effab6d3c6386c7f9bc

Observation b9a2f028-e551-4633-98f3-609507572a32 · outbound

This paper cites Pro- gressive graph attention network for video question answer- ing.

Perceive, Query & Reason: Enhancing Video QA with Question-Guided Temporal Queries Pro- gressive graph attention network for video question answer- ing

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T00:48:43.960052Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-11T00:48:42.774800Z digest=sha256:f01ed967071eaf4ab8b77fb8950e503d2f8a672fe38dc697446b147e479015c2

Observation d6909758-f515-4d46-b3ca-e1afa0dcec83 · outbound

This paper cites Learn- ing transferable visual models from natural language super- vision.

Perceive, Query & Reason: Enhancing Video QA with Question-Guided Temporal Queries Learn- ing transferable visual models from natural language super- vision

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-11T00:48:42.779289Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T00:48:42.779289Z digest=sha256:b6a59d2ecb42605309ef45e3eb50844e5953f0e45af96a521cdcd23195e1b9e9

Observation c4398a63-72a4-46f9-b1f1-e6f151c7c68a · outbound

This paper cites Video understanding with large language models: A survey.

Perceive, Query & Reason: Enhancing Video QA with Question-Guided Temporal Queries Video understanding with large language models: A survey

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-11T00:48:42.782371Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T00:48:42.782371Z digest=sha256:8d05103e4d3215830195ce0cb4c1be6e010f07af6bc13b13109a04559f16c49b

Observation 22f155f6-b2cb-42f8-b4eb-714eb29cc530 · outbound

This paper cites Hashimoto.

Perceive, Query & Reason: Enhancing Video QA with Question-Guided Temporal Queries Hashimoto

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T00:48:43.942313Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-11T00:48:42.786147Z digest=sha256:a82e1d2ea294b17821cc55457eed05314d67d3baafba10cacebe2da9299fc1e4

Observation 5df412a4-4f3c-4aa7-9797-3974c72cea3e · outbound

This paper cites Movieqa: Understanding stories in movies through question- answering.

Perceive, Query & Reason: Enhancing Video QA with Question-Guided Temporal Queries Movieqa: Understanding stories in movies through question- answering

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T00:48:43.930810Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-11T00:48:42.789881Z digest=sha256:a18983568e3c66b21e440c997fecd26c105a5cd973e6ae38bed17ba6bd2aa1c1

Observation 44e5ecbd-b7fe-48ec-95ed-40be015b2ccd · outbound

This paper cites LLaMA: Open and Efficient Foundation Language Models.

Perceive, Query & Reason: Enhancing Video QA with Question-Guided Temporal Queries LLaMA: Open and Efficient Foundation Language Models

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-11T00:48:42.793308Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T00:48:42.793308Z digest=sha256:932845e360c3deae869a2d17dcd26e0a28b1b6b890350942ab05226bfac20c65

Observation 02b72742-e925-47c2-9b47-ef1d1f87bef6 · outbound

This paper cites Attention is all you need.

Perceive, Query & Reason: Enhancing Video QA with Question-Guided Temporal Queries Attention is all you need

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T00:48:43.920119Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-11T00:48:42.797212Z digest=sha256:d781e353a2fb5da3a1ed0bf114c3a24584389ca6fd636192c5db2d2a0577ca1c

Observation a82233d4-d974-4666-b5f5-c0dec8021185 · outbound

This paper cites All in one: Exploring unified video-language pre-training.

Perceive, Query & Reason: Enhancing Video QA with Question-Guided Temporal Queries All in one: Exploring unified video-language pre-training

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T00:48:43.910065Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-11T00:48:42.800047Z digest=sha256:b0cc36fafb68762276cbe46e2c3374860219f6619e363db84e54696e98062307

Observation 47ee8998-e4aa-4b7f-befe-9a345cd6ab89 · outbound

This paper cites InternVideo: General Video Foundation Models via Generative and Discriminative Learning.

Perceive, Query & Reason: Enhancing Video QA with Question-Guided Temporal Queries InternVideo: General Video Foundation Models via Generative and Discriminative Learning

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-11T00:48:42.803699Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T00:48:42.803699Z digest=sha256:38348e57509358ced2074c5361bd6d55d50ccdc1e0f6cf7113bd278655588baf

Observation 49a5dc49-71c5-4077-bb71-d5d77a86ff7d · outbound

This paper cites Star: A benchmark for situated reasoning in real-world videos.

Perceive, Query & Reason: Enhancing Video QA with Question-Guided Temporal Queries Star: A benchmark for situated reasoning in real-world videos

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T00:48:43.899172Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-11T00:48:42.807161Z digest=sha256:33ed89b0f20587fc0459b05e78cc8a92c144076450c0973195f8e84c187c6d99

Observation e899e737-d376-4d22-9410-91d8dd9c058b · outbound

This paper cites Next-qa: Next phase of question-answering to explaining temporal actions.

Perceive, Query & Reason: Enhancing Video QA with Question-Guided Temporal Queries Next-qa: Next phase of question-answering to explaining temporal actions

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T00:48:43.888669Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-11T00:48:42.811012Z digest=sha256:328bfa16cf1fa2505df2e276d15a97c73191824e14c226a982b20aab73dd8224

Observation 58e6e453-1aef-42b1-b279-771b4597ac14 · outbound

This paper cites Video graph transformer for video question answering.

Perceive, Query & Reason: Enhancing Video QA with Question-Guided Temporal Queries Video graph transformer for video question answering

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-11T00:48:42.814578Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T00:48:42.814578Z digest=sha256:4bd92586cbe5187cf7df5f47199deae2f3a8e375bd0efafb78835b1d9385d231

Observation 983f1a51-e1a7-42b6-b11f-e745a9b29dda · outbound

This paper cites Video question answer- ing via gradually refined attention over appearance and mo- tion.

Perceive, Query & Reason: Enhancing Video QA with Question-Guided Temporal Queries Video question answer- ing via gradually refined attention over appearance and mo- tion

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T00:48:43.871833Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-11T00:48:42.819185Z digest=sha256:f7a9ac0491829019d9f66e8fbe27904f29e3b215587da30fa625ee486a27e147

Observation 8d6733f9-fb67-47ff-b6f8-f69e313f1942 · outbound

This paper cites Msr-vtt: A large video description dataset for bridging video and language.

Perceive, Query & Reason: Enhancing Video QA with Question-Guided Temporal Queries Msr-vtt: A large video description dataset for bridging video and language

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T00:48:43.861375Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-11T00:48:42.822541Z digest=sha256:e3b0416b5e3e386932553679b6418d37cd09f4e05eef96a08d82c10e019be138

Observation d6347ca2-222c-4724-8da1-ce2079e3a66b · outbound

This paper cites Multimodal learning with transformers: A survey.

Perceive, Query & Reason: Enhancing Video QA with Question-Guided Temporal Queries Multimodal learning with transformers: A survey

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T00:48:43.850152Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-11T00:48:42.826289Z digest=sha256:fefdb37ed3354d74dfca2d1adc974b295e0c63b35bf7ec9f551c0d7d63e0e8cf

Observation 731b8052-d4d5-405b-8006-f58cb30f4e1f · outbound

This paper cites Just ask: Learning to answer questions from millions of narrated videos.

Perceive, Query & Reason: Enhancing Video QA with Question-Guided Temporal Queries Just ask: Learning to answer questions from millions of narrated videos

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T00:48:43.838425Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-11T00:48:42.829623Z digest=sha256:5ff5d0780d11d9570d035853593efce06f0e68d3ddddba8522d1b1e47c25b50a

Observation 384b0dfa-afbe-4f4f-ac91-7d31dcc189a4 · outbound

This paper cites Zero-shot video question answering via frozen bidirectional language models.

Perceive, Query & Reason: Enhancing Video QA with Question-Guided Temporal Queries Zero-shot video question answering via frozen bidirectional language models

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T00:48:43.825656Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-11T00:48:42.833478Z digest=sha256:3c882873637fce1742b8ff6d34985bf84d891932be9088b9f6f247cb7fbe8fe9

Observation 21cfaa11-bbc0-412b-a9a1-ed7b07dc42d4 · outbound

This paper cites React: Synergizing rea- soning and acting in language models.

Perceive, Query & Reason: Enhancing Video QA with Question-Guided Temporal Queries React: Synergizing rea- soning and acting in language models

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T00:48:43.813855Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-11T00:48:42.836507Z digest=sha256:4489653636752fbc112516d05fa80080eef5561fa553cdcd7527de7ceb15cc52

Observation b6953f6d-3107-44a7-b82c-ac84c3a0862f · outbound

This paper cites Hitea: Hierarchical temporal-aware video-language pre-training.

Perceive, Query & Reason: Enhancing Video QA with Question-Guided Temporal Queries Hitea: Hierarchical temporal-aware video-language pre-training

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T00:48:43.802812Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-11T00:48:42.887332Z digest=sha256:f7ccedc149775cbac50f75c8c2a5d3dd03f3f7dd5da789a0c289533801ccf953

Observation 8c2857dd-6d7d-4a68-80f3-b733404f8f25 · outbound

This paper cites Self-chained image-language model for video localization and question answering.

Perceive, Query & Reason: Enhancing Video QA with Question-Guided Temporal Queries Self-chained image-language model for video localization and question answering

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T00:48:43.790221Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-11T00:48:42.980304Z digest=sha256:c5d9debf40f7c7e178cb41d0be5575ab56caa86f9b4b5e46201c7dd8c04bf189

Observation e16a10a6-fe6a-4f3b-aeb3-355b04e21120 · outbound

This paper cites Con- necting speech encoder and large language model for asr.

Perceive, Query & Reason: Enhancing Video QA with Question-Guided Temporal Queries Con- necting speech encoder and large language model for asr

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T00:48:43.779106Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-11T00:48:43.063331Z digest=sha256:9016edc2c8ed65a818244c25f1a2570ed94e9231bad2083095fae5ae6ea8ebe1

Observation 526f7e9e-da41-4601-bcc1-2d2bbebd5d6f · outbound

This paper cites Activitynet-qa: A dataset for understanding complex web videos via question answering.

Perceive, Query & Reason: Enhancing Video QA with Question-Guided Temporal Queries Activitynet-qa: A dataset for understanding complex web videos via question answering

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-11T00:48:43.166576Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T00:48:43.166576Z digest=sha256:03ee64c8adac88ab864a573aace56cf782a1cfd2fb08d5f50de1e97642637096

Observation 678e6af7-8511-453d-8b9a-495f8846e56c · outbound

This paper cites Multi-event video-text retrieval.

Perceive, Query & Reason: Enhancing Video QA with Question-Guided Temporal Queries Multi-event video-text retrieval

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T00:48:43.701148Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-11T00:48:43.169582Z digest=sha256:f167abdec40cfc136bd5771f821f28f1ca3880500d19d071af890edc28c83261

Observation 05e6c6ee-bbef-4f72-942f-085e1c2f890f · outbound

This paper cites Video-llama: An instruction-tuned audio-visual language model for video un- derstanding.

Perceive, Query & Reason: Enhancing Video QA with Question-Guided Temporal Queries Video-llama: An instruction-tuned audio-visual language model for video un- derstanding

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T00:48:43.592092Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-11T00:48:43.173692Z digest=sha256:47b8325a6c4de669ddf34950e58d6e4e18eb52d88ee216441241f0d856e064c0

Observation 62d222aa-82e5-41e3-bcca-d374ac017ad9 · outbound

This paper cites Mmicl: Empowering vision-language model with multi-modal in-context learning.

Perceive, Query & Reason: Enhancing Video QA with Question-Guided Temporal Queries Mmicl: Empowering vision-language model with multi-modal in-context learning

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T00:48:43.426278Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-11T00:48:43.177412Z digest=sha256:93469e4f7dcb6073a86a6596422abf861c8a86f05ee03e5b3c327ff9ad9aa535

Observation b933c7ea-94f7-4b32-a7fc-dafdeb0e221d · outbound

This paper cites Judging llm-as-a-judge with mt-bench and chatbot arena.

Perceive, Query & Reason: Enhancing Video QA with Question-Guided Temporal Queries Judging llm-as-a-judge with mt-bench and chatbot arena

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T00:48:43.416013Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-11T00:48:43.180461Z digest=sha256:df365b24b5812e62f14900af021f732aee1043ba665e1f8932cd4ccb4c5cb7d5

Observation 58cdf1c8-74a0-432c-a84c-b08c34b6fc72 · outbound

This paper cites Video question answering: Datasets, algorithms and challenges.

Perceive, Query & Reason: Enhancing Video QA with Question-Guided Temporal Queries Video question answering: Datasets, algorithms and challenges

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T00:48:43.404911Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-11T00:48:43.184381Z digest=sha256:4fad380f42e6c95af9439ebc28c90998f2baa5553fe1e58b5bb9637453a68b55

Observation df27bd4a-3d11-4ade-8341-3811a1e38127 · outbound

This paper cites Adaptive pooling in multi-instance learning for web video annotation.

Perceive, Query & Reason: Enhancing Video QA with Question-Guided Temporal Queries Adaptive pooling in multi-instance learning for web video annotation

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T00:48:43.393996Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-11T00:48:43.188113Z digest=sha256:1201f6479388840ed10242c1c57b55d94cd8d262de4dd37a0f78b7a8a1abe6a7

Observation 3b214edd-3fe0-47cd-8002-aa70363495d0 · outbound

This paper cites Minigpt-4: Enhancing vision-language understanding with advanced large language models.

Perceive, Query & Reason: Enhancing Video QA with Question-Guided Temporal Queries Minigpt-4: Enhancing vision-language understanding with advanced large language models

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T00:48:43.384144Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-11T00:48:43.190999Z digest=sha256:c15f5efc709b2adc722e81d72be0f8cc2d2b7ee7c0a61af0458b8e27b241e2f0

Observation 838bb7e0-3451-4466-a06b-2308dfdf70dd · outbound

This paper cites blind guess.

Perceive, Query & Reason: Enhancing Video QA with Question-Guided Temporal Queries blind guess

Reference 2024

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T00:48:43.372823Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-11T00:48:43.194779Z digest=sha256:4608a8457c3404d278baf5dc9ef255cf76c56250c7693c97aa12f6de43b20d7a

Pith citing papers

No inbound Pith citation observations are available.