Pith. sign in

Paper Citation Record · LEDGER

Admitting Ignorance Helps the Video Question Answering Models to Answer

As of 12 August 2026, this Paper Citation Record lists 75 of 75 outbound references and 0 inbound Pith citation observations for arXiv:2501.08771.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2501.08771 v2

Coverage vector

measured 75 of 75 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-10T20:22:29.090022Z

measured 75 of 75 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-12T06:34:41.77262+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

75 of 75 outbound references displayed

  • verified exact1
  • verified fuzzy55
  • unresolved19
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 9532b078-bd9d-472e-8895-c0dcacd07458 · outbound

This paper cites Bilinear attention networks,.

Admitting Ignorance Helps the Video Question Answering Models to Answer Bilinear attention networks,

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:22:30.083794Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T20:22:28.815676Z digest=sha256:eaed13b5e18af0f2e27e0a2d139c2eecd320cce9bbf4b6f15bc697ebf1b8237c

Observation 78a8b23e-b5f6-4ecc-8f34-c9b7493bd253 · outbound

This paper cites Attend what you need: Motion-appearance synergistic networks for video question answering,.

Admitting Ignorance Helps the Video Question Answering Models to Answer Attend what you need: Motion-appearance synergistic networks for video question answering,

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:22:30.072790Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T20:22:28.820087Z digest=sha256:fcb4c6bf69587bafd792131f4b5f77f2dd20aead5b1b8feee7c4dc8d4db9995e

Observation 58cab2bd-070b-4bfe-807c-6733d6ab7df8 · outbound

This paper cites Video as conditional graph hierarchy for multi-granular question answering.

Admitting Ignorance Helps the Video Question Answering Models to Answer Video as conditional graph hierarchy for multi-granular question answering

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:22:30.061788Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T20:22:28.824049Z digest=sha256:4aa73ef869baddf711002eb575c56f2dbe15e5be29cd3ca03966a25eebc6f169

Observation 852a8b05-c6fc-4fb9-8212-9d3c23425d1b · outbound

This paper cites Merlot: Multimodal neural script knowledge models,.

Admitting Ignorance Helps the Video Question Answering Models to Answer Merlot: Multimodal neural script knowledge models,

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:22:30.052125Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T20:22:28.828251Z digest=sha256:b3fb87b2c9afb8ec4fedd0ec757092bb96cf5d20fbda53fb8b617f45fd03fa20

Observation 21caaec5-68af-4a53-bed9-af6c1a4b862a · outbound

This paper cites VIOLET : End-to-End Video-Language Transformers with Masked Visual-token Modeling.

Admitting Ignorance Helps the Video Question Answering Models to Answer VIOLET : End-to-End Video-Language Transformers with Masked Visual-token Modeling

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-10T20:22:28.832209Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:22:28.832209Z digest=sha256:14678ecf24ecd316f69310aa894adc03ebd5bd5229071c2c9cc0b004bf4d87da

Observation 803b90c5-5350-4593-ad3e-b099f4f80ad1 · outbound

This paper cites X$^2$-VLM: All-In-One Pre-trained Model For Vision-Language Tasks.

Admitting Ignorance Helps the Video Question Answering Models to Answer X$^2$-VLM: All-In-One Pre-trained Model For Vision-Language Tasks

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-10T20:22:28.836606Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:22:28.836606Z digest=sha256:0ca5467884f541a9d21e192a090a20379324f310713bb990281695374a213ac7

Observation a1bf4bc0-51da-4e06-92df-b148982abc98 · outbound

This paper cites InternVideo: General Video Foundation Models via Generative and Discriminative Learning.

Admitting Ignorance Helps the Video Question Answering Models to Answer InternVideo: General Video Foundation Models via Generative and Discriminative Learning

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-10T20:22:28.840944Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:22:28.840944Z digest=sha256:ecfa96bf2c8fc07aa8147e81fd5ab22786aec764821f2d32c09b83b5f913c88b

Observation 9b75616c-ba59-4a81-b431-8e48192bd12a · outbound

This paper cites All in one: Exploring unified video-language pre-training,.

Admitting Ignorance Helps the Video Question Answering Models to Answer All in one: Exploring unified video-language pre-training,

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:22:30.041510Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T20:22:28.845487Z digest=sha256:0d44df8d4e4ef2fc34d8ece587c3f7fd50daaa51bd25940a2746ad8b1b086292

Observation 1eb77c64-db60-464f-bb67-2227cdef559f · outbound

This paper cites Invariant grounding for video question answering,.

Admitting Ignorance Helps the Video Question Answering Models to Answer Invariant grounding for video question answering,

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:22:30.029231Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T20:22:28.849971Z digest=sha256:41038445df534ec750c9151758a5d5a598c337fca1ef3df25a537cbba2f26d1f

Observation bd64f643-ca09-41ba-a6c7-a19c30af739e · outbound

This paper cites Equivariant and invariant grounding for video question answering,.

Admitting Ignorance Helps the Video Question Answering Models to Answer Equivariant and invariant grounding for video question answering,

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:22:30.016176Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T20:22:28.853927Z digest=sha256:90a5acdd7e018a19287be76ced9c84424b3ded152fc2e8f1e41c965bf3fe21d8

Observation 31bda1b2-d1e2-41f1-8610-04c2a0ef86fb · outbound

This paper cites Transformer-empowered invariant grounding for video question answering,.

Admitting Ignorance Helps the Video Question Answering Models to Answer Transformer-empowered invariant grounding for video question answering,

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:22:30.000451Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T20:22:28.858221Z digest=sha256:5d66f014bb16643df399da79ecb4af54b7f793d67816a5ed6681f42493fccde9

Observation dcbdd53f-f305-4ab8-926c-31a3dd24f80a · outbound

This paper cites Discovering spatio- temporal rationales for video question answering,.

Admitting Ignorance Helps the Video Question Answering Models to Answer Discovering spatio- temporal rationales for video question answering,

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:22:29.988421Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T20:22:28.862463Z digest=sha256:a4af6ed40dd62b059f65f9deb99cb9f9c6e24e2311d77d88be718d8477d6e88f

Observation 4819a8d9-4faa-4d3d-9909-24a637d781d7 · outbound

This paper cites Adversarial vqa: A new benchmark for evaluating the robustness of vqa models,.

Admitting Ignorance Helps the Video Question Answering Models to Answer Adversarial vqa: A new benchmark for evaluating the robustness of vqa models,

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:22:29.976532Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T20:22:28.866406Z digest=sha256:784fdaf998f9e919bf039c52d9b854a8060808979fa3df9c77487e7c869051e6

Observation 51a192e9-c3b7-48e5-acff-21624f1775ce · outbound

This paper cites Discovering the real association: Multimodal causal reasoning in video question answering,.

Admitting Ignorance Helps the Video Question Answering Models to Answer Discovering the real association: Multimodal causal reasoning in video question answering,

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:22:29.964326Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T20:22:28.869197Z digest=sha256:6a710872d4a21b76dd2a3b3a305677cc3005781852cb0d0400c41783038ead14

Observation a49e0de7-7419-402d-a33b-e09219841169 · outbound

This paper cites Coun- terfactual vqa: A cause-effect look at language bias,.

Admitting Ignorance Helps the Video Question Answering Models to Answer Coun- terfactual vqa: A cause-effect look at language bias,

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:22:29.951666Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T20:22:28.872228Z digest=sha256:a70aaf1f22a08eb863287644f610ef8ea4d1bdf2854a265e5db782d9bca0c212

Observation e8864a2e-5808-40c6-b518-77ab0c08d48f · outbound

This paper cites Beyond question- based biases: Assessing multimodal shortcut learning in visual question answering,.

Admitting Ignorance Helps the Video Question Answering Models to Answer Beyond question- based biases: Assessing multimodal shortcut learning in visual question answering,

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:22:29.938674Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T20:22:28.875299Z digest=sha256:93886e231dfab06d19eea8484008f5685ed7cf3babb19ff8e2703868994610bb

Observation d575a6f5-4b9f-4d92-b8cc-39f2a7f3baa7 · outbound

This paper cites Roses are red, violets are blue... but should vqa expect them to?.

Admitting Ignorance Helps the Video Question Answering Models to Answer Roses are red, violets are blue... but should vqa expect them to?

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:22:29.926218Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T20:22:28.878700Z digest=sha256:914d7815e2afffa753d84b46f4557a2366355412e5bc31d77b0e60ca4482a7da

Observation b2019ef5-2c84-4f92-a84e-c247a7e3a4d3 · outbound

This paper cites Human-adversarial visual question answer- ing,.

Admitting Ignorance Helps the Video Question Answering Models to Answer Human-adversarial visual question answer- ing,

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:22:29.912538Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T20:22:28.882176Z digest=sha256:d719bc7b80a96fee73a51400bcd905f333f92fe8c1917cf3c131a56a913af1d4

Observation 4e026506-d2ad-4521-ae3e-6cad9a0e5ec7 · outbound

This paper cites A survey on curriculum learning,.

Admitting Ignorance Helps the Video Question Answering Models to Answer A survey on curriculum learning,

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:22:29.898588Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T20:22:28.886366Z digest=sha256:fd9ae13dd355b52fed87be47e9e884bd6af50db5aa1c0a255a3c4ac30a4b34c3

Observation 3e8ded24-e28a-4311-8830-2aeb40d69a6b · outbound

This paper cites Curriculum learning: A survey,.

Admitting Ignorance Helps the Video Question Answering Models to Answer Curriculum learning: A survey,

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-10T20:22:28.889818Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:22:28.889818Z digest=sha256:45db58c7ce54ffb263678993e0b286951c5f33ce7593fde645d94df80e366b43

Observation d0b6f7db-8ade-4e59-ac02-3bccd89b6f9b · outbound

This paper cites Tgif-qa: Toward spatio- temporal reasoning in visual question answering,.

Admitting Ignorance Helps the Video Question Answering Models to Answer Tgif-qa: Toward spatio- temporal reasoning in visual question answering,

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:22:29.878483Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T20:22:28.893376Z digest=sha256:fe8f0922fe80d2fda749b050c7ca79faa8f89cf5d3b1bea6f826d16d5078d540

Observation bb3893a4-465a-41e6-85a1-5026b5dd85fa · outbound

This paper cites Revisiting the.

Admitting Ignorance Helps the Video Question Answering Models to Answer Revisiting the

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:22:29.867380Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T20:22:28.897187Z digest=sha256:06303685d0b38ad1a32857a613bc470d2df2a9ef797d15c8eb98ab6b561e2176

Observation b2da382d-c2bd-44d2-9171-27d8ee531a5f · outbound

This paper cites Answering from sure to uncertain: Uncertainty-aware curriculum learning for video question answering,.

Admitting Ignorance Helps the Video Question Answering Models to Answer Answering from sure to uncertain: Uncertainty-aware curriculum learning for video question answering,

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-10T20:22:28.901073Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:22:28.901073Z digest=sha256:522f2ec62ef7843231765a35b63412caa60a29a29c96d38d6cae7d298c52bf56

Observation e2a152b2-a62d-42ff-86cd-e1328346f135 · outbound

This paper cites Question-guided erasing-based spatiotemporal attention learning for video question answering,.

Admitting Ignorance Helps the Video Question Answering Models to Answer Question-guided erasing-based spatiotemporal attention learning for video question answering,

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:22:29.857861Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T20:22:28.904741Z digest=sha256:e15c452f3b0447df38bb23329b72a921e5ecdef9a57adbb02608c6e0371fd72e

Observation 06093558-5918-49fd-9669-f538d0ec5d5f · outbound

This paper cites Memory augmented deep recurrent neural network for video question answering,.

Admitting Ignorance Helps the Video Question Answering Models to Answer Memory augmented deep recurrent neural network for video question answering,

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:22:29.846939Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T20:22:28.909220Z digest=sha256:5b4f3838dd67c2ed8aa5fb71b1d7cc1f4effed3488bd52ee366792cb0d1b7fd6

Observation 431db470-deaf-48f4-a4fe-75b4f3d4abc5 · outbound

This paper cites Knowledge-routed visual question reasoning: Challenges for deep representation embedding,.

Admitting Ignorance Helps the Video Question Answering Models to Answer Knowledge-routed visual question reasoning: Challenges for deep representation embedding,

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:22:29.835410Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T20:22:28.913520Z digest=sha256:0ea3dc53360f1948db293e4740fc5d91dc7007eb565b4c422912a21ead8362ea

Observation b3be0ff1-18e9-4203-8bcf-5a412b7c57d0 · outbound

This paper cites Multitask learning for visual question answering,.

Admitting Ignorance Helps the Video Question Answering Models to Answer Multitask learning for visual question answering,

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:22:29.823215Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T20:22:28.917336Z digest=sha256:eacc5f34a37b51bc0a437de812197790d1ee46057c987a94ef21f25c05c0caaa

Observation b830ec11-203f-4216-b2f6-8cabae86c334 · outbound

This paper cites Bilinear graph networks for visual ques- tion answering,.

Admitting Ignorance Helps the Video Question Answering Models to Answer Bilinear graph networks for visual ques- tion answering,

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:22:29.811946Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T20:22:28.922428Z digest=sha256:bc6feba06f843d357aed445113e3a4a87a01a6b78e43630562f0c02261e3af2f

Observation a36e2dd3-5034-466a-8c1f-e7fd57e0df58 · outbound

This paper cites Bilateral cross-modality graph matching attention for feature fusion in visual question answering,.

Admitting Ignorance Helps the Video Question Answering Models to Answer Bilateral cross-modality graph matching attention for feature fusion in visual question answering,

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:22:29.800548Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T20:22:28.925840Z digest=sha256:e4ca653e9ec4f5c89c2e768f3963a057cb631b46f931d6110b27c527d56d864c

Observation 45178bf5-0eb8-414d-9906-9416d950a396 · outbound

This paper cites Bridging the cross- modality semantic gap in visual question answering,.

Admitting Ignorance Helps the Video Question Answering Models to Answer Bridging the cross- modality semantic gap in visual question answering,

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:22:29.787587Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T20:22:28.929126Z digest=sha256:3fe2d3c1a1ccb6beedabed252c36d1208db87d14e018595257212e3cdb05538d

Observation cc25a2d1-e088-44e7-9b65-4fd7ae5029a7 · outbound

This paper cites Latent attention network with position perception for visual question answering,.

Admitting Ignorance Helps the Video Question Answering Models to Answer Latent attention network with position perception for visual question answering,

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:22:29.776048Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T20:22:28.933066Z digest=sha256:6294fae947f0c4dd58cbc87f9654bfcffd047e598d50faace9f8e1497753722c

Observation aac3d65f-0e70-4238-9cc7-daabc6288571 · outbound

This paper cites Webly supervised knowledge-embedded model for visual reasoning,.

Admitting Ignorance Helps the Video Question Answering Models to Answer Webly supervised knowledge-embedded model for visual reasoning,

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:22:29.764674Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T20:22:28.936324Z digest=sha256:2147a6796aee16c9489a5b9b881a4521b214ac025b9a55403ee09bc9dce324aa

Observation c7110ae4-0f09-434c-9d73-442c8d76815b · outbound

This paper cites Uncovering the temporal context for video question answering,.

Admitting Ignorance Helps the Video Question Answering Models to Answer Uncovering the temporal context for video question answering,

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:22:29.753036Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T20:22:28.939590Z digest=sha256:38be8fcb9fb37c2a819c17ca63a68d6710929de78b67461d90561361e6f97b5c

Observation 8d3ecdae-de28-48e3-8d52-27ad057b8c44 · outbound

This paper cites Video question answering via gradually refined attention over appearance and motion,.

Admitting Ignorance Helps the Video Question Answering Models to Answer Video question answering via gradually refined attention over appearance and motion,

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:22:29.740341Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T20:22:28.943213Z digest=sha256:33fbaf5b0c02bb5cc0910cb74f5e333dc79ccbdf56bb9586493ecc6b077435e9

Observation 863e8001-2d72-42b1-ad53-1526236563e3 · outbound

This paper cites Divide and conquer: Question-guided spatio-temporal contextual attention for video question answering,.

Admitting Ignorance Helps the Video Question Answering Models to Answer Divide and conquer: Question-guided spatio-temporal contextual attention for video question answering,

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:22:29.728219Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T20:22:28.946761Z digest=sha256:29bbb7a5608d8948325342818cc486c1a952e18b63fd8f5b68a7e40983a30ae0

Observation 8dc85746-341a-442d-9ad1-b34a4b8fd5e1 · outbound

This paper cites Video question answering with spatio-temporal reasoning,.

Admitting Ignorance Helps the Video Question Answering Models to Answer Video question answering with spatio-temporal reasoning,

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:22:29.716195Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T20:22:28.950509Z digest=sha256:fe59ffb4ab4fab9bc2d469ae3822db612af6a8b4afda1a530e3c97a23cdab2af

Observation 69e3270d-7c37-4141-9d1a-424baa1332d6 · outbound

This paper cites Dualvgr: A dual-visual graph reasoning unit for video question answering,.

Admitting Ignorance Helps the Video Question Answering Models to Answer Dualvgr: A dual-visual graph reasoning unit for video question answering,

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:22:29.704012Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T20:22:28.953955Z digest=sha256:0d2b4e271b41dd073f0c424924156848390bb3f4d8cbff261b6b62a13de15c8a

Observation 3cd45be9-6a74-4205-af87-22150a4fa54e · outbound

This paper cites Bridge to answer: Structure-aware graph interaction network for video question answering,.

Admitting Ignorance Helps the Video Question Answering Models to Answer Bridge to answer: Structure-aware graph interaction network for video question answering,

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:22:29.691371Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T20:22:28.957975Z digest=sha256:e8fb80eb170beb686fd3bbdbad3692d4c545142b0a02ce5425d8d737b93ed6b9

Observation bd0de087-f9af-47aa-a621-98dddf70774b · outbound

This paper cites Motion-appearance co-memory networks for video question answering,.

Admitting Ignorance Helps the Video Question Answering Models to Answer Motion-appearance co-memory networks for video question answering,

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:22:29.679267Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T20:22:28.961511Z digest=sha256:1839393102f00eb9288fcca946319173d480b23a593d89387feb27fc52b9583d

Observation c1149717-3af7-4511-80d4-ea9312ec6448 · outbound

This paper cites Heterogeneous memory enhanced multimodal attention model for video question answering,.

Admitting Ignorance Helps the Video Question Answering Models to Answer Heterogeneous memory enhanced multimodal attention model for video question answering,

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:22:29.668248Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T20:22:28.964910Z digest=sha256:63a5472541d0a82e5684e8c07f012ced8948dff01c4f0bb0adc4a54d2dd052ea

Observation 0fc7a196-c301-474e-a175-c42904a5edbf · outbound

This paper cites Hierarchical conditional relation networks for video question answering,.

Admitting Ignorance Helps the Video Question Answering Models to Answer Hierarchical conditional relation networks for video question answering,

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:22:29.656871Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T20:22:28.968357Z digest=sha256:1794812d2188d8b89041ece9263d39cb53e0cc00b297bfc79b98ea370e66450a

Observation e3cfb448-3188-4f39-a469-c80dc57ae182 · outbound

This paper cites Verbs in action: Improving verb understanding in video-language models,.

Admitting Ignorance Helps the Video Question Answering Models to Answer Verbs in action: Improving verb understanding in video-language models,

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:22:29.645191Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T20:22:28.971679Z digest=sha256:0e9d116ff9ff4a3c29b6c1e5e162eaf29bdcf4e6c026b7f0d5b7407d42f1b8d1

Observation 42900346-0920-434b-ad7a-25087c9f9ff9 · outbound

This paper cites When and why vision-language models behave like bags-of-words, and what to do about it?.

Admitting Ignorance Helps the Video Question Answering Models to Answer When and why vision-language models behave like bags-of-words, and what to do about it?

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:22:29.633656Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T20:22:28.974963Z digest=sha256:7b716d57b97f55fc792307274f24c8e4d5b5277f5bf520e8c1fb8d3d97d17ced

Observation aebf50f1-0815-48f5-981b-127bb796945f · outbound

This paper cites Teaching structured vision & language concepts to vision & language models,.

Admitting Ignorance Helps the Video Question Answering Models to Answer Teaching structured vision & language concepts to vision & language models,

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-10T20:22:28.978593Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:22:28.978593Z digest=sha256:a2d4da68e8345d2e996835fea194f6a8328bce622ae6ce2898d0b317f2b49a94

Observation 3a7f542e-94e2-4463-b39b-ce6ffd6f613d · outbound

This paper cites Rubi: Reducing unimodal biases for visual question answering,.

Admitting Ignorance Helps the Video Question Answering Models to Answer Rubi: Reducing unimodal biases for visual question answering,

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:22:29.614069Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T20:22:28.981805Z digest=sha256:7ec1cb7c447b2785db8d328b7d5f18b7d6578d764c99372bd66003c817c8d9d8

Observation 0390a37c-da9f-41b4-b8a0-51f726c37a2b · outbound

This paper cites Don’t just assume; look and answer: Overcoming priors for visual question answering,.

Admitting Ignorance Helps the Video Question Answering Models to Answer Don’t just assume; look and answer: Overcoming priors for visual question answering,

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:22:29.600848Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T20:22:28.985272Z digest=sha256:8ddac42fb9e9873c8b54dcdd1adefd0d7e5c4e5a366e0434d84e937fc4884d52

Observation 0ea48d93-d1da-411f-9aa2-eb29c500de50 · outbound

This paper cites Overcoming language priors in visual question answering with adversarial regularization,.

Admitting Ignorance Helps the Video Question Answering Models to Answer Overcoming language priors in visual question answering with adversarial regularization,

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:22:29.588740Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T20:22:28.988924Z digest=sha256:14dcd3bfce811b928d4ab0b455e90ac35ce39006e7984648058a94273a86ae65

Observation 2297cfd9-6abf-4134-9808-5ac885c7dc23 · outbound

This paper cites Reliable visual question answering: Abstain rather than answer incorrectly,.

Admitting Ignorance Helps the Video Question Answering Models to Answer Reliable visual question answering: Abstain rather than answer incorrectly,

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:22:29.577485Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T20:22:28.992385Z digest=sha256:41fd4a866903b6c55639c25780315b7c1022016b17fd8436d7081dde16196604

Observation 889e1f67-fff4-4e59-aada-db3fe4670164 · outbound

This paper cites Addressing failure prediction by learning model confidence,.

Admitting Ignorance Helps the Video Question Answering Models to Answer Addressing failure prediction by learning model confidence,

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-10T20:22:28.996874Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:22:28.996874Z digest=sha256:af2c922b0f0ae5ad0fcbc67f7d8fb899b862932e84edf72844f35bab29d67347

Observation 4b57a7cf-c70b-4472-909e-31cc9f10b7f6 · outbound

This paper cites Combating Label Noise in Deep Learning Using Abstention.

Admitting Ignorance Helps the Video Question Answering Models to Answer Combating Label Noise in Deep Learning Using Abstention

Reference 50

Resolution
verified exact
local_arxiv, observed 2026-08-10T20:22:29.182883Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T20:22:29.001019Z digest=sha256:400583e815be4b75e8a4f0e7530f7993d31d1f8667930f0aee86d4f8d621b2c9

Observation ddf71bdb-f7d6-42df-ae1f-a2794b0077d5 · outbound

This paper cites The art of abstention: Selective prediction and error regularization for natural language processing,.

Admitting Ignorance Helps the Video Question Answering Models to Answer The art of abstention: Selective prediction and error regularization for natural language processing,

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:22:29.559720Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T20:22:29.004953Z digest=sha256:1bd08c2709a43bcae3258cb5e866b466d4b4b8dbb44b1ccfae7344fc88b6e88b

Observation 7a613141-c7e7-461b-b55f-889d64a619b3 · outbound

This paper cites On the foundations of noise-free selective classifi- cation.

Admitting Ignorance Helps the Video Question Answering Models to Answer On the foundations of noise-free selective classifi- cation

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:22:29.547982Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T20:22:29.008397Z digest=sha256:37b06b73d31ffce8c040d68369fe7a738e075544b967da00c4bf7fb8463352d5

Observation 1379861a-8ac1-4286-ba1b-c9a08c13a3e8 · outbound

This paper cites Investigating Selective Prediction Approaches Across Several Tasks in IID, OOD, and Adversarial Settings.

Admitting Ignorance Helps the Video Question Answering Models to Answer Investigating Selective Prediction Approaches Across Several Tasks in IID, OOD, and Adversarial Settings

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-10T20:22:29.011509Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:22:29.011509Z digest=sha256:6ecfdf98edd3b4baace07dcf26041173ae4f78dcc0ec6629f2e4a46ac6f1b2a3

Observation 3341b454-363a-46cb-ab49-62b3041d7861 · outbound

This paper cites Selectivenet: A deep neural network with an integrated reject option,.

Admitting Ignorance Helps the Video Question Answering Models to Answer Selectivenet: A deep neural network with an integrated reject option,

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:22:29.535581Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T20:22:29.015510Z digest=sha256:a099f87ef7ac2402c76732f727688704b99743964631d9f9a4a5a7f24954ebdc

Observation e151b0e8-61a6-4d4a-87b5-20442f4a437d · outbound

This paper cites Selective question answering under domain shift,.

Admitting Ignorance Helps the Video Question Answering Models to Answer Selective question answering under domain shift,

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:22:29.523427Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T20:22:29.018882Z digest=sha256:c6eb6333c04eb3f76cf339c148204a32846fc45144a3dd9bf52cbc8917625937

Observation 1325e5ac-d15d-4ada-9316-a8365ef8b309 · outbound

This paper cites Attention is all you need,.

Admitting Ignorance Helps the Video Question Answering Models to Answer Attention is all you need,

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-10T20:22:29.022646Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:22:29.022646Z digest=sha256:06f37ce47bf311bb8dc51afa3dbb7a26910e06b42744f7fee33274b42e9f80da

Observation 0f449ef2-86c9-4a81-9fda-b7d70ecf11fb · outbound

This paper cites Learning transferable visual models from natural language supervision,.

Admitting Ignorance Helps the Video Question Answering Models to Answer Learning transferable visual models from natural language supervision,

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-10T20:22:29.026236Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:22:29.026236Z digest=sha256:19b38e1c9708144bfeb63566864814312c6824defd7fa95e4ddae731b196d79a

Observation 2f469043-f51d-4d8e-8e30-0a21a2fb13e7 · outbound

This paper cites Quo vadis, action recognition? a new model and the kinetics dataset,.

Admitting Ignorance Helps the Video Question Answering Models to Answer Quo vadis, action recognition? a new model and the kinetics dataset,

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:22:29.496535Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T20:22:29.029742Z digest=sha256:4fed32dedb26c6a98804c9e1de4023d864ef4e71570991395a91447112ec7485

Observation 99d0fe4a-e38b-4ac0-8e79-f35b8e9f8117 · outbound

This paper cites Howto100m: Learning a text-video embedding by watching hundred million narrated video clips,.

Admitting Ignorance Helps the Video Question Answering Models to Answer Howto100m: Learning a text-video embedding by watching hundred million narrated video clips,

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:22:29.484728Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T20:22:29.033198Z digest=sha256:14afc6a288109ed841a5ee6d13ca0fd9004ba981e55153c6589ae45fd75bb2b9

Observation 094aa94b-d13c-4553-bd4a-22f3d69d4f8f · outbound

This paper cites Ava: A video dataset of spatio-temporally localized atomic visual actions,.

Admitting Ignorance Helps the Video Question Answering Models to Answer Ava: A video dataset of spatio-temporally localized atomic visual actions,

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:22:29.473291Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T20:22:29.036935Z digest=sha256:f356a9a5ace9cad43db8b9ea9e3254a338604f3871cc529edd11d8087a755543

Observation 999faf4b-2dce-45c4-9888-1b0b1c780e61 · outbound

This paper cites Frozen in time: A joint video and image encoder for end-to-end retrieval,.

Admitting Ignorance Helps the Video Question Answering Models to Answer Frozen in time: A joint video and image encoder for end-to-end retrieval,

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:22:29.461656Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T20:22:29.040850Z digest=sha256:34177d6b751fb5111facda8c67527ea74466cd51bc9caeebf9ff6bf3c4408c2e

Observation edf1e6c2-5324-4e8d-a7e5-b820685d25e4 · outbound

This paper cites Next-qa: Next phase of question-answering to explaining temporal actions,.

Admitting Ignorance Helps the Video Question Answering Models to Answer Next-qa: Next phase of question-answering to explaining temporal actions,

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:22:29.450128Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T20:22:29.044658Z digest=sha256:3f9edc7993efb98a7e6927855d0331cc3bce9ca5b3cad61de3de7457b704d4fb

Observation 21650f3e-d479-4a7f-9edc-e1f891b68af4 · outbound

This paper cites Hierarchical Object-oriented Spatio-Temporal Reasoning for Video Question Answering.

Admitting Ignorance Helps the Video Question Answering Models to Answer Hierarchical Object-oriented Spatio-Temporal Reasoning for Video Question Answering

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-10T20:22:29.048239Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:22:29.048239Z digest=sha256:e9f27d297166efa9439f9d549576cc811398f1e7b4cecd95b345b9f1f576b54e

Observation ad5ab970-1c3d-4c30-b94a-cfb9e4e2cd4f · outbound

This paper cites Less is more: Clipbert for video-and-language learning via sparse sampling,.

Admitting Ignorance Helps the Video Question Answering Models to Answer Less is more: Clipbert for video-and-language learning via sparse sampling,

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:22:29.438477Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T20:22:29.052228Z digest=sha256:d521201d9772a5039ff64e8e3dc6aacf165d5f8ee5d8719757cc1ea185f1517e

Observation adbb3e57-2533-42ff-9de2-c5893e946466 · outbound

This paper cites Causal inference in natural language processing: Estimation, prediction, interpretation and beyond,.

Admitting Ignorance Helps the Video Question Answering Models to Answer Causal inference in natural language processing: Estimation, prediction, interpretation and beyond,

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-10T20:22:29.055862Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:22:29.055862Z digest=sha256:35a7981455cba33fc5104e70348ec1a5aedc32a3e1df2d96805115bbd291a89e

Observation 758ebd2b-997d-491a-9a43-aa119f008327 · outbound

This paper cites Docogen: Domain counterfactual generation for low resource domain adaptation,.

Admitting Ignorance Helps the Video Question Answering Models to Answer Docogen: Domain counterfactual generation for low resource domain adaptation,

Reference 66

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:22:29.419757Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T20:22:29.059587Z digest=sha256:e81d350b5d8ea04096472bdf4aad53433a7a20603593d5beb315d3b5f562f788

Observation 4967e378-4854-41ca-bcab-a0c2b723200c · outbound

This paper cites Polyjuice: Generating Counterfactuals for Explaining, Evaluating, and Improving Models.

Admitting Ignorance Helps the Video Question Answering Models to Answer Polyjuice: Generating Counterfactuals for Explaining, Evaluating, and Improving Models

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-10T20:22:29.063214Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:22:29.063214Z digest=sha256:584da84b9af92772b0af5bcbf661ac3887d63a02c67be262c72be5165e472a14

Observation 3c9abe9d-38ec-438e-a2fc-2ae6b0974985 · outbound

This paper cites Language models are unsupervised multitask learners,.

Admitting Ignorance Helps the Video Question Answering Models to Answer Language models are unsupervised multitask learners,

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-10T20:22:29.067137Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:22:29.067137Z digest=sha256:bd8599c2374df5acf004ad8472b2e809b1cd8fc25c870180206cfafcb5e1d0d9

Observation e7017573-60f0-4102-aeb3-ff0fc7b0940b · outbound

This paper cites Sentence-BERT: Sentence Embeddings using Siamese BERT-Networks.

Admitting Ignorance Helps the Video Question Answering Models to Answer Sentence-BERT: Sentence Embeddings using Siamese BERT-Networks

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-10T20:22:29.070599Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:22:29.070599Z digest=sha256:473dc2b5c1fa8b60d030f3ad177b0edd5ac979779603cfdf32a7720736938fdd

Observation 3c5a5043-d24b-441b-9a3e-7133004cc6ac · outbound

This paper cites Stacked attention networks for image question answering,.

Admitting Ignorance Helps the Video Question Answering Models to Answer Stacked attention networks for image question answering,

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-10T20:22:29.074601Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:22:29.074601Z digest=sha256:7964985a7c6f12c362b88ed17f5a5457f924a1455c290b90c4ead1cdebc88e15

Observation f693dfea-2e83-4e8b-9452-b39d8a32a65d · outbound

This paper cites Multimodal compact bilinear pooling for visual question answering and visual grounding,.

Admitting Ignorance Helps the Video Question Answering Models to Answer Multimodal compact bilinear pooling for visual question answering and visual grounding,

Reference 71

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:22:29.392912Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T20:22:29.078064Z digest=sha256:0e5211b23c066af18b74dcc21fd256c975fbc789a55ec7a84c81b13cbe163a96

Observation d2d1af64-df40-4420-9a50-aa37acd053a1 · outbound

This paper cites Vqa: Visual question answering,.

Admitting Ignorance Helps the Video Question Answering Models to Answer Vqa: Visual question answering,

Reference 72

Resolution
unresolved
no resolver link, observed 2026-08-10T20:22:29.081042Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:22:29.081042Z digest=sha256:b8a744514777bfe15c80b5f5848bc90802ab63220913d00e16581c3432657e0b

Observation eceb8198-17c8-4dd0-993a-e8d068f6e578 · outbound

This paper cites Making the v in vqa matter: Elevating the role of image understanding in visual question answering,.

Admitting Ignorance Helps the Video Question Answering Models to Answer Making the v in vqa matter: Elevating the role of image understanding in visual question answering,

Reference 73

Resolution
unresolved
no resolver link, observed 2026-08-10T20:22:29.083891Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:22:29.083891Z digest=sha256:d7ac126f79a80f82b23909018c74a41d093d0e636801b86b33b2e170ecc7a839

Observation 2abc7e2f-be50-408b-a452-3a954810dc27 · outbound

This paper cites Mvbench: A comprehensive multi-modal video understanding benchmark,.

Admitting Ignorance Helps the Video Question Answering Models to Answer Mvbench: A comprehensive multi-modal video understanding benchmark,

Reference 74

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:22:29.368636Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-10T20:22:29.087051Z digest=sha256:155ebb4c77760fdddbd7b4965fb46c0bbc26f160f38a5fff747beb290eecb717

Observation 2938132f-fb81-4bf3-b4bd-7aa0e143015e · outbound

This paper cites VideoLLaMA 2: Advancing Spatial-Temporal Modeling and Audio Understanding in Video-LLMs.

Admitting Ignorance Helps the Video Question Answering Models to Answer VideoLLaMA 2: Advancing Spatial-Temporal Modeling and Audio Understanding in Video-LLMs

Reference 75

Resolution
unresolved
no resolver link, observed 2026-08-10T20:22:29.090022Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:22:29.090022Z digest=sha256:ea7fbf052a13d6572f6c53b913633f9c500c6efa9da59b46d63f84e6e21a5751

Pith citing papers

No inbound Pith citation observations are available.