Pith. sign in

Paper Citation Record · LEDGER

EgoTextVQA: Towards Egocentric Scene-Text Aware Video Question Answering

As of 16 August 2026, this Paper Citation Record lists 87 of 87 outbound references and 0 inbound Pith citation observations for arXiv:2502.07411.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2502.07411 v2

Coverage vector

measured 87 of 87 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-08T12:54:56.474300Z

measured 87 of 87 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-16T06:30:59.297886+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

87 of 87 outbound references displayed

  • verified exact1
  • verified fuzzy42
  • unresolved43
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 20a166fa-0211-4747-8913-f525b65029fa · outbound

This paper cites Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond.

EgoTextVQA: Towards Egocentric Scene-Text Aware Video Question Answering Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-08T12:54:56.114885Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T12:54:56.114885Z digest=sha256:f220d4d94701bfa0d7fc9e4f14544196c4784c85e3449b746001690780c22a59

Observation 70b606bd-fcfd-4e71-ba3c-c5b514aa235a · outbound

This paper cites Where did i leave my keys? - episodic-memory-based question answering on egocentric videos.

EgoTextVQA: Towards Egocentric Scene-Text Aware Video Question Answering Where did i leave my keys? - episodic-memory-based question answering on egocentric videos

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-08T12:54:56.120306Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T12:54:56.120306Z digest=sha256:f2c4ee9de3d5132fe20a02b85e759cb94212a504912ee457a72aee9e651a2ae1

Observation 12c78ca8-94ce-4a70-bc4c-046db810186f · outbound

This paper cites Scene text visual question answering.

EgoTextVQA: Towards Egocentric Scene-Text Aware Video Question Answering Scene text visual question answering

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-08T12:54:56.124737Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T12:54:56.124737Z digest=sha256:2d86d91e33a40eb35491172f462b2cdce817738cba134640ebd66b6ea8028cb4

Observation 701d60ec-cf0d-4114-9199-9c1eea17b542 · outbound

This paper cites Nougat: Neural Optical Understanding for Academic Documents.

EgoTextVQA: Towards Egocentric Scene-Text Aware Video Question Answering Nougat: Neural Optical Understanding for Academic Documents

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-08T12:54:56.129066Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T12:54:56.129066Z digest=sha256:cb2eef72e64b9290ef710a63d942d03835bcbe86b14be57035cab23be0e637ad

Observation 6b7ebe36-a71a-44df-b74c-9fb4f39b937c · outbound

This paper cites InternLM2 Technical Report.

EgoTextVQA: Towards Egocentric Scene-Text Aware Video Question Answering InternLM2 Technical Report

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-08T12:54:56.133631Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T12:54:56.133631Z digest=sha256:b1083a6c64879f86914c771555d34ebf630a5c0e8aee17565fc7839844fa5f1e

Observation 6a3ad532-9edd-4f91-a7da-72196dccc641 · outbound

This paper cites ShareGPT4Video: Improving Video Understanding and Generation with Better Captions.

EgoTextVQA: Towards Egocentric Scene-Text Aware Video Question Answering ShareGPT4Video: Improving Video Understanding and Generation with Better Captions

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-08T12:54:56.138227Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T12:54:56.138227Z digest=sha256:c651daa5c6f4ab3b7bc609bd15003050695c3a14ab4f8b182146c549382d2d13

Observation adb75ff9-37fb-458b-a65c-3d2cbbcc3648 · outbound

This paper cites How Far Are We to GPT-4V? Closing the Gap to Commercial Multimodal Models with Open-Source Suites.

EgoTextVQA: Towards Egocentric Scene-Text Aware Video Question Answering How Far Are We to GPT-4V? Closing the Gap to Commercial Multimodal Models with Open-Source Suites

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-08T12:54:56.142716Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T12:54:56.142716Z digest=sha256:87ed18a4fd7b92cee471e46bec5e861e63399d454984f43b9df4942227d8c7c7

Observation d9661cac-6181-4b7a-bee4-f2a91c45d520 · outbound

This paper cites Internvl: Scaling up vision foundation mod- els and aligning for generic visual-linguistic tasks.

EgoTextVQA: Towards Egocentric Scene-Text Aware Video Question Answering Internvl: Scaling up vision foundation mod- els and aligning for generic visual-linguistic tasks

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-08T12:54:56.147036Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T12:54:56.147036Z digest=sha256:1b2d5515d4bbce3324723d1255f695a706ba3cf1cbc38f7a08a56f1d85dfe082

Observation 4b68f2c0-dcc7-4251-8541-cfa9be868bbd · outbound

This paper cites VidEgoThink: Assessing Egocentric Video Understanding Capabilities for Embodied AI.

EgoTextVQA: Towards Egocentric Scene-Text Aware Video Question Answering VidEgoThink: Assessing Egocentric Video Understanding Capabilities for Embodied AI

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-08T12:54:56.151522Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T12:54:56.151522Z digest=sha256:e03c5efbb87862ec53c9c97c4db3a7c6ad1728a46bc33df91c113c60e9e44ee5

Observation 5b98d8e6-29cf-4455-bb5c-51f67345d832 · outbound

This paper cites Egothink: Evalu- ating first-person perspective thinking capability of vision- language models.

EgoTextVQA: Towards Egocentric Scene-Text Aware Video Question Answering Egothink: Evalu- ating first-person perspective thinking capability of vision- language models

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-08T12:54:56.156323Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T12:54:56.156323Z digest=sha256:15f647f7219cb6839bd08456564cc044d23c79ba4f1b77c1e5266d9a14113cc8

Observation cdef0dc8-631b-43ab-86e0-740dde11ec80 · outbound

This paper cites Grounded question-answering in long egocentric videos.

EgoTextVQA: Towards Egocentric Scene-Text Aware Video Question Answering Grounded question-answering in long egocentric videos

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-08T12:54:56.160240Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T12:54:56.160240Z digest=sha256:5890964dcf4ad65dd690218cc23fdad0b030794907f3e7b7d69f031aca86c84c

Observation 060f0cb2-c727-4ea5-a003-85cc2ce94246 · outbound

This paper cites Egovqa-an egocentric video question answer- ing benchmark dataset.

EgoTextVQA: Towards Egocentric Scene-Text Aware Video Question Answering Egovqa-an egocentric video question answer- ing benchmark dataset

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-08T12:54:56.164622Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T12:54:56.164622Z digest=sha256:35a77d88da746edd3eecb5c6f5821998ae9ce22bcb303e9cbd5d88bc837742f3

Observation 320ac191-1522-4b26-8062-0dac98c0c302 · outbound

This paper cites Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis.

EgoTextVQA: Towards Egocentric Scene-Text Aware Video Question Answering Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-08T12:54:56.168704Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T12:54:56.168704Z digest=sha256:719133178733aab3a44bab1765cdd28ceb68929cc5f832f1224049a88a00d41c

Observation 4a1f8428-c2b7-4097-9d5c-e43294e33dfe · outbound

This paper cites Ego4d: Around the world in 3,000 hours of egocentric video.

EgoTextVQA: Towards Egocentric Scene-Text Aware Video Question Answering Ego4d: Around the world in 3,000 hours of egocentric video

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-08T12:54:56.173248Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T12:54:56.173248Z digest=sha256:f8428301a1596a2d0c8920738664c3b8063e4af423395b2c3cd74c5f5c965e6e

Observation cd64bf27-6376-4f56-853d-a0148fe9ab1b · outbound

This paper cites Context-aware graph inference with knowledge distillation for visual dialog.IEEE TPAMI, 44(10):6056–6073, 2021.

EgoTextVQA: Towards Egocentric Scene-Text Aware Video Question Answering Context-aware graph inference with knowledge distillation for visual dialog.IEEE TPAMI, 44(10):6056–6073, 2021

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T12:54:57.611618Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-08T12:54:56.177341Z digest=sha256:2fe7057ca5f574d9b7cb05caf071c035d4a6a66d3a5516610be83de70bceb312

Observation 9cf3228b-03da-46f6-970d-c11f31eb9879 · outbound

This paper cites Vizwiz grand challenge: Answering visual questions from blind people.

EgoTextVQA: Towards Egocentric Scene-Text Aware Video Question Answering Vizwiz grand challenge: Answering visual questions from blind people

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T12:54:57.598545Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-08T12:54:56.181352Z digest=sha256:c83e9e42192425def9e5feec3678b728d493c9ea76337574be1a3d3fe7083f71

Observation 9a8936b8-d246-4859-9ec9-d5c652b389c6 · outbound

This paper cites GoMatching: A Simple Baseline for Video Text Spotting via Long and Short Term Matching.

EgoTextVQA: Towards Egocentric Scene-Text Aware Video Question Answering GoMatching: A Simple Baseline for Video Text Spotting via Long and Short Term Matching

Reference 17

Resolution
verified exact
local_arxiv, observed 2026-08-08T12:54:56.771045Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-08T12:54:56.185230Z digest=sha256:35ebbcbbcb030505dfa89e665cc47ed26f2dee60c8c8e2722fc7d84a8bdbda7c

Observation b749407b-50fe-48b0-a978-ce8c3d8f9694 · outbound

This paper cites CogVLM2: Visual Language Models for Image and Video Understanding.

EgoTextVQA: Towards Egocentric Scene-Text Aware Video Question Answering CogVLM2: Visual Language Models for Image and Video Understanding

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-08T12:54:56.189601Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T12:54:56.189601Z digest=sha256:9ea0b0841cfdec4bfabfa712167af1f3bd66ef8217a627db9a3278fa99dbb2a5

Observation a9b38a20-4d6d-4bd2-b78b-d9538c0a23b8 · outbound

This paper cites Understanding video scenes through text: Insights from text-based video question answering.

EgoTextVQA: Towards Egocentric Scene-Text Aware Video Question Answering Understanding video scenes through text: Insights from text-based video question answering

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T12:54:57.585323Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-08T12:54:56.193845Z digest=sha256:bf07e0511e720eb848a5674e253aab19a2a68aa38889256fd24fcbddfb9f0541

Observation 1e950c31-aac4-469b-89d8-369b270623c6 · outbound

This paper cites Egotaskqa: Understanding human tasks in egocentric videos.

EgoTextVQA: Towards Egocentric Scene-Text Aware Video Question Answering Egotaskqa: Understanding human tasks in egocentric videos

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T12:54:57.571518Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-08T12:54:56.197727Z digest=sha256:c13825907006cb36d6954ea6552afccdba2f67558df16f4a5302ca4f34d7e26b

Observation de1cdeb5-a53e-4964-9007-ac5b3d5ba440 · outbound

This paper cites Llava-next: Stronger llms supercharge multimodal capa- bilities in the wild, 2024.

EgoTextVQA: Towards Egocentric Scene-Text Aware Video Question Answering Llava-next: Stronger llms supercharge multimodal capa- bilities in the wild, 2024

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T12:54:57.558677Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-08T12:54:56.201725Z digest=sha256:1f5780927b861032540560557d08853aec60a5d38af6be203b8028576149908c

Observation 7ddea211-36e8-41b0-a85b-d97653ced7d2 · outbound

This paper cites PP-OCRv3: More Attempts for the Improvement of Ultra Lightweight OCR System.

EgoTextVQA: Towards Egocentric Scene-Text Aware Video Question Answering PP-OCRv3: More Attempts for the Improvement of Ultra Lightweight OCR System

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-08T12:54:56.205791Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T12:54:56.205791Z digest=sha256:7870982cdc53c4cf12b5b0fade739f198c440ae6763c28a1e8f4ba3f8c903f8a

Observation bc3bc08a-889d-45db-ab84-ad27feba1338 · outbound

This paper cites Flex- attention for efficient high-resolution vision-language mod- els.

EgoTextVQA: Towards Egocentric Scene-Text Aware Video Question Answering Flex- attention for efficient high-resolution vision-language mod- els

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T12:54:57.546062Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-08T12:54:56.209825Z digest=sha256:28f9ec89c87bb7a0d84ddd69a322f3026d5a077481655e69c266568a59a166aa

Observation e573457e-0ffe-4b46-9ccb-2873f6e870fd · outbound

This paper cites Mvbench: A comprehensive multi-modal video understand- ing benchmark.

EgoTextVQA: Towards Egocentric Scene-Text Aware Video Question Answering Mvbench: A comprehensive multi-modal video understand- ing benchmark

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T12:54:57.532537Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-08T12:54:56.213744Z digest=sha256:27944aceeff9392eb8e94ede586b69d4f5dc2f66a9c723313b90a8295f850260

Observation ad03128f-81c0-439f-b859-7bbc660cd8ab · outbound

This paper cites Invariant grounding for video question answering.

EgoTextVQA: Towards Egocentric Scene-Text Aware Video Question Answering Invariant grounding for video question answering

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T12:54:57.519362Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-08T12:54:56.217734Z digest=sha256:8deab28dab719b0b1fa4b355c461b272f37dc07703ade9401e33ae5946958a69

Observation 11d2956c-034d-4817-8ebd-446f2ec62029 · outbound

This paper cites Transformer-empowered invariant grounding for video question answering.

EgoTextVQA: Towards Egocentric Scene-Text Aware Video Question Answering Transformer-empowered invariant grounding for video question answering

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T12:54:57.506144Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-08T12:54:56.221976Z digest=sha256:974571a506ac6e708cb3d02f3bc0906ae7ef700f13f7a8ac666b3e35ed81b450

Observation 832a7826-8ac0-4a40-9743-48d429db36a9 · outbound

This paper cites Vila: On pre-training for visual language models.

EgoTextVQA: Towards Egocentric Scene-Text Aware Video Question Answering Vila: On pre-training for visual language models

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T12:54:57.492342Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-08T12:54:56.225880Z digest=sha256:292ddd9d933ed1b215f60a9fa1301d201a638636849565909ed5a001656590e6

Observation 35ec54d9-b5ae-410c-a685-7a899fc47138 · outbound

This paper cites Egocentric video-language pretraining.

EgoTextVQA: Towards Egocentric Scene-Text Aware Video Question Answering Egocentric video-language pretraining

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T12:54:57.478968Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-08T12:54:56.229826Z digest=sha256:1f5a25e9cbd8c8b6b6c8f51ed33e25ebc94647be1f962ef412cefa16d791374f

Observation afa12dbe-bca5-4a39-8c19-f12ce12055ac · outbound

This paper cites Llava-next: Im- proved reasoning, ocr, and world knowledge, 2024.

EgoTextVQA: Towards Egocentric Scene-Text Aware Video Question Answering Llava-next: Im- proved reasoning, ocr, and world knowledge, 2024

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T12:54:57.465382Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-08T12:54:56.233949Z digest=sha256:fac71cfd279dfbd62c7346119c95ea073694b51eb6395974fa8e97033c794954

Observation 3f9cd5da-9089-41e9-9dc0-b25fbb69f521 · outbound

This paper cites Ocrbench: On the hidden mystery of ocr in large multimodal models, 2024.

EgoTextVQA: Towards Egocentric Scene-Text Aware Video Question Answering Ocrbench: On the hidden mystery of ocr in large multimodal models, 2024

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T12:54:57.451768Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-08T12:54:56.238033Z digest=sha256:35469227ac4c5a1f9597d86acc53da6d9cf124a062f81f657e6e40af2242415f

Observation 479d4a04-a8e8-44c7-89c6-c9552dd96ba4 · outbound

This paper cites Feast Your Eyes: Mixture-of-Resolution Adaptation for Multimodal Large Language Models.

EgoTextVQA: Towards Egocentric Scene-Text Aware Video Question Answering Feast Your Eyes: Mixture-of-Resolution Adaptation for Multimodal Large Language Models

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-08T12:54:56.241899Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T12:54:56.241899Z digest=sha256:c822759d7c920520accb8a6166eb71b861a2ffeef5e95806a422a6ddc696fd73

Observation 021e9583-38fd-4bca-9973-fa36d7b69e7d · outbound

This paper cites Video-ChatGPT: Towards Detailed Video Understanding via Large Vision and Language Models.

EgoTextVQA: Towards Egocentric Scene-Text Aware Video Question Answering Video-ChatGPT: Towards Detailed Video Understanding via Large Vision and Language Models

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-08T12:54:56.246341Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T12:54:56.246341Z digest=sha256:75be1620bf8792592ed37d91962bf55d2de8cb35117f064d8e28ad26f1cbee73

Observation 47d7ef9f-b74f-4650-85d7-479402dc0a49 · outbound

This paper cites Egoschema: A diagnostic benchmark for very long- form video language understanding.

EgoTextVQA: Towards Egocentric Scene-Text Aware Video Question Answering Egoschema: A diagnostic benchmark for very long- form video language understanding

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T12:54:57.437157Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-08T12:54:56.250465Z digest=sha256:8c863416bca9e10c15b02a63807b0f0c3d54619c0e884b2016b8a9fd72a7f72c

Observation 6e4a0a78-fd94-4af5-9336-0400dd24e2b4 · outbound

This paper cites Docvqa: A dataset for vqa on document images.

EgoTextVQA: Towards Egocentric Scene-Text Aware Video Question Answering Docvqa: A dataset for vqa on document images

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T12:54:57.423250Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-08T12:54:56.254547Z digest=sha256:ea2df69f402c6ec9fae06695fa0d8c6e0e5e1bfe7090348d0b47ac2719a731d3

Observation 078de082-67ba-4d79-abcc-6bc671999581 · outbound

This paper cites Infographicvqa.

EgoTextVQA: Towards Egocentric Scene-Text Aware Video Question Answering Infographicvqa

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T12:54:57.409491Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-08T12:54:56.258520Z digest=sha256:bf1a6670ff0e1168475c5e2a615eb0713739aab7f28dce863b1111ea3d8ee02d

Observation 78603dc3-316d-4193-9e7b-c32dfcbcb91e · outbound

This paper cites Ocr-vqa: Visual question answering by reading text in images.

EgoTextVQA: Towards Egocentric Scene-Text Aware Video Question Answering Ocr-vqa: Visual question answering by reading text in images

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T12:54:57.395858Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-08T12:54:56.262488Z digest=sha256:f9fdf26150819b66db6214187d1f11fa7d093a6bec58434c9837f8cd2dfc7247

Observation ee178bbd-26bf-4bf4-ab7a-c2d1842cf9e8 · outbound

This paper cites Gpt-4o system card.

EgoTextVQA: Towards Egocentric Scene-Text Aware Video Question Answering Gpt-4o system card

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T12:54:57.382022Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-08T12:54:56.266298Z digest=sha256:6eef8b2aee623b2ad03859b1b61b3f750ae13cb369204b8b53212007ef227cb1

Observation d88a5540-81ff-4582-ba0e-db44de5c68fd · outbound

This paper cites Gpt-4o mini: advancing cost-efficient intelligence.

EgoTextVQA: Towards Egocentric Scene-Text Aware Video Question Answering Gpt-4o mini: advancing cost-efficient intelligence

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-08T12:54:56.270226Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T12:54:56.270226Z digest=sha256:041867913060c82938c657e50f11239626156535eed308c89af9f777463d5e2c

Observation 2b8c4d99-595b-4aa8-b55d-507fefc2139a · outbound

This paper cites Egovlpv2: Egocentric video-language pre-training with fusion in the backbone.

EgoTextVQA: Towards Egocentric Scene-Text Aware Video Question Answering Egovlpv2: Egocentric video-language pre-training with fusion in the backbone

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T12:54:57.359826Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-08T12:54:56.274244Z digest=sha256:3d94a6c5ad6ba8d20d7fc1350cf11676ecc0b79194c5f28854a6ea768b88b2ae

Observation 1a549ecd-6650-4be7-97ab-26169dfee7ff · outbound

This paper cites Learn- ing transferable visual models from natural language super- vision.

EgoTextVQA: Towards Egocentric Scene-Text Aware Video Question Answering Learn- ing transferable visual models from natural language super- vision

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T12:54:57.346472Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-08T12:54:56.278216Z digest=sha256:6b8047049ad3ed7d56745882f466f48759d1e01846aa3a0f639403b7bc1c23fc

Observation b40979d0-1bb6-450f-af94-f324be730f5a · outbound

This paper cites Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context.

EgoTextVQA: Towards Egocentric Scene-Text Aware Video Question Answering Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-08T12:54:56.282084Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T12:54:56.282084Z digest=sha256:db9e1233e9b9cf326d88a31e5448fbea2290559545123a1f9efb28b6b450f323

Observation 79a9a262-70b9-4325-b797-1de2079f524e · outbound

This paper cites LAION-400M: Open Dataset of CLIP-Filtered 400 Million Image-Text Pairs.

EgoTextVQA: Towards Egocentric Scene-Text Aware Video Question Answering LAION-400M: Open Dataset of CLIP-Filtered 400 Million Image-Text Pairs

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-08T12:54:56.286985Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T12:54:56.286985Z digest=sha256:61a341ca16de285f514250791e9e1f1cace65b66543a4cc063912759e96cf0d6

Observation bd87f972-824e-4eb2-943b-4e0dad9ca5a9 · outbound

This paper cites Annotating objects and relations in user- generated videos.

EgoTextVQA: Towards Egocentric Scene-Text Aware Video Question Answering Annotating objects and relations in user- generated videos

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T12:54:57.333415Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-08T12:54:56.291155Z digest=sha256:880510db950064fdc603b28d78e3969d49e862029a0a6e7e7c40d0c4cebb7366

Observation ad865063-2f7f-46ff-8da0-c522eb334193 · outbound

This paper cites GLU Variants Improve Transformer.

EgoTextVQA: Towards Egocentric Scene-Text Aware Video Question Answering GLU Variants Improve Transformer

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-08T12:54:56.295240Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T12:54:56.295240Z digest=sha256:57a1c15e40c6d0a0716ef25a4e51bc9d04613b40cc6cef9907a87f29843ae746

Observation 1f35c8e7-ee3a-4a8a-b152-e8ba7621aba7 · outbound

This paper cites Towards vqa models that can read.

EgoTextVQA: Towards Egocentric Scene-Text Aware Video Question Answering Towards vqa models that can read

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T12:54:57.319919Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-08T12:54:56.299537Z digest=sha256:a83090b6c10355ecd15396dcac23cacf4da3036181e3dfa33df05f979ddf6e84

Observation 7f75d1ff-79d6-49ef-ad9d-e34462632914 · outbound

This paper cites EVA-CLIP: Improved Training Techniques for CLIP at Scale.

EgoTextVQA: Towards Egocentric Scene-Text Aware Video Question Answering EVA-CLIP: Improved Training Techniques for CLIP at Scale

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-08T12:54:56.303789Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T12:54:56.303789Z digest=sha256:a9d2bd0027bd8c20b8d2520839914feb345d75f5030aa2bb973da57cce35189f

Observation 322010f4-f1b3-4cb8-b534-b2a38c6f2186 · outbound

This paper cites Gemini: A Family of Highly Capable Multimodal Models.

EgoTextVQA: Towards Egocentric Scene-Text Aware Video Question Answering Gemini: A Family of Highly Capable Multimodal Models

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-08T12:54:56.308154Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T12:54:56.308154Z digest=sha256:cc82b6c43be9118978ec006324b69909bf043010d802a540526c0fe27b033ae5

Observation 1c5febe1-94b6-4db5-9200-00f1c50ca31b · outbound

This paper cites Reading between the lanes: Text videoqa on the road.

EgoTextVQA: Towards Egocentric Scene-Text Aware Video Question Answering Reading between the lanes: Text videoqa on the road

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T12:54:57.306916Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-08T12:54:56.312592Z digest=sha256:ce17efc7cb7dc14e01ae4dc925ad5897b50b24fc1df4499ca7a44b74e0c5d6f1

Observation ac6caca4-760a-4449-ae7f-b41e840b3571 · outbound

This paper cites Llama 2: Open Foundation and Fine-Tuned Chat Models.

EgoTextVQA: Towards Egocentric Scene-Text Aware Video Question Answering Llama 2: Open Foundation and Fine-Tuned Chat Models

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-08T12:54:56.317080Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T12:54:56.317080Z digest=sha256:9a8fb0f4da17495a0282534a5d95466c8236ec28c311c287529c2ef90609425f

Observation 505a1408-9d54-4ed5-a34d-ef627ddcb886 · outbound

This paper cites Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution.

EgoTextVQA: Towards Egocentric Scene-Text Aware Video Question Answering Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-08T12:54:56.321415Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T12:54:56.321415Z digest=sha256:20877697ec0858843ae6175d1eb59bf88ae22941e193feaf884385945e728546

Observation e0dce2cd-fd45-4e5c-8037-27e7602e7d54 · outbound

This paper cites CogVLM: Visual Expert for Pretrained Language Models.

EgoTextVQA: Towards Egocentric Scene-Text Aware Video Question Answering CogVLM: Visual Expert for Pretrained Language Models

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-08T12:54:56.325636Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T12:54:56.325636Z digest=sha256:8e4693e99960d2103aba2f14f67a26631dc402b6ddc7368a5636059ac9497b72

Observation ce44bf7e-1841-465e-b4d6-66bba7d0d2bb · outbound

This paper cites On the general value of ev- idence, and bilingual scene-text visual question answering.

EgoTextVQA: Towards Egocentric Scene-Text Aware Video Question Answering On the general value of ev- idence, and bilingual scene-text visual question answering

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T12:54:57.293858Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-08T12:54:56.329931Z digest=sha256:0e43686350dd31cbc9e4f76134ff1d1e6e875a7830a4595d45dc4130cef0188d

Observation 3c6960a3-522f-49e1-924d-5beab171956a · outbound

This paper cites Assistq: Affordance-centric question-driven task completion for ego- centric assistant.

EgoTextVQA: Towards Egocentric Scene-Text Aware Video Question Answering Assistq: Affordance-centric question-driven task completion for ego- centric assistant

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T12:54:57.280040Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-08T12:54:56.334020Z digest=sha256:716d52090c0d83994174c72b1eccb9ba9d7bc7b5b2f52f8768028849f06617fc

Observation 21504019-c8a1-4b2a-b024-8bcaaa3c39ae · outbound

This paper cites Next-qa: Next phase of question-answering to explaining temporal actions.

EgoTextVQA: Towards Egocentric Scene-Text Aware Video Question Answering Next-qa: Next phase of question-answering to explaining temporal actions

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T12:54:57.266470Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-08T12:54:56.337940Z digest=sha256:fd36b255cf46184d9a802172a66e495ea84bb8c638d6242fbf6583230ed5d6c7

Observation bc512d6f-164a-4ff4-91ac-0044c19ddbab · outbound

This paper cites VideoQA in the Era of LLMs: An Empirical Study.

EgoTextVQA: Towards Egocentric Scene-Text Aware Video Question Answering VideoQA in the Era of LLMs: An Empirical Study

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-08T12:54:56.342107Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T12:54:56.342107Z digest=sha256:4b65e645642408eacab25bb6d9777198e31961bb7409e26e22c67998db27f351

Observation de7ca57f-0f03-4635-ba80-185cf6a168b2 · outbound

This paper cites Deconfounded video moment retrieval with causal intervention.

EgoTextVQA: Towards Egocentric Scene-Text Aware Video Question Answering Deconfounded video moment retrieval with causal intervention

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T12:54:57.252848Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-08T12:54:56.346273Z digest=sha256:a064139be697c1d9c2f22860d48c2d9e36e8a7154f03d14e659576d4536d2d04

Observation 579d69eb-bbc1-4f7b-b01e-272581b454eb · outbound

This paper cites Video moment retrieval with cross-modal neural architecture search.

EgoTextVQA: Towards Egocentric Scene-Text Aware Video Question Answering Video moment retrieval with cross-modal neural architecture search

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T12:54:57.239274Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-08T12:54:56.350327Z digest=sha256:a192f17fc4a99a8e3dfb57e9e7e854c063670a40fd1ee0b1db975fd022fde0a5

Observation dcb06e1c-a964-4ef3-bab3-768a016de2e9 · outbound

This paper cites Robust video question answer- ing via contrastive cross-modality representation learning.

EgoTextVQA: Towards Egocentric Scene-Text Aware Video Question Answering Robust video question answer- ing via contrastive cross-modality representation learning

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T12:54:57.225574Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-08T12:54:56.355088Z digest=sha256:ca0d8cf177f1101aca776e7f5e05f908adcb98906da0d33422efd8aba8b263f8

Observation 028fa900-a965-4130-b355-e1a364996573 · outbound

This paper cites MiniCPM-V: A GPT-4V Level MLLM on Your Phone.

EgoTextVQA: Towards Egocentric Scene-Text Aware Video Question Answering MiniCPM-V: A GPT-4V Level MLLM on Your Phone

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-08T12:54:56.359154Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T12:54:56.359154Z digest=sha256:f95d3d2fe5e960708f585e6c46ba62a43e4a149039c4ceb93c1a3fc46daf350f

Observation eb601c86-bfa8-4d66-b3ba-f2c41546f759 · outbound

This paper cites MM-Ego: Towards Building Egocentric Multimodal LLMs for Video QA.

EgoTextVQA: Towards Egocentric Scene-Text Aware Video Question Answering MM-Ego: Towards Building Egocentric Multimodal LLMs for Video QA

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-08T12:54:56.363353Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T12:54:56.363353Z digest=sha256:4e2844247f0bdd98d88d69229bd660acf90cba6e81e8030fb78b3e9502fb88b8

Observation 9437a4c9-1fae-4824-a387-cdecb0997592 · outbound

This paper cites Sigmoid loss for language image pre-training.

EgoTextVQA: Towards Egocentric Scene-Text Aware Video Question Answering Sigmoid loss for language image pre-training

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T12:54:57.211122Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-08T12:54:56.368065Z digest=sha256:2f67e3d0253302b1db058e07080e1278e1efa7e34ab4bdf3b36678ff3f44bd90

Observation 743af68e-86a0-4e98-a5c4-6662b4f6faf3 · outbound

This paper cites Multi-factor adaptive vision selec- tion for egocentric video question answering.

EgoTextVQA: Towards Egocentric Scene-Text Aware Video Question Answering Multi-factor adaptive vision selec- tion for egocentric video question answering

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T12:54:57.197499Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-08T12:54:56.372009Z digest=sha256:360ebee3312572bd37982880c24fd4211dd5d60dcd18246cb3331c0746e4e160

Observation 03acbac3-de00-4c32-926f-a2c60e18e971 · outbound

This paper cites LLaVA-Read: Enhancing Reading Ability of Multimodal Language Models.

EgoTextVQA: Towards Egocentric Scene-Text Aware Video Question Answering LLaVA-Read: Enhancing Reading Ability of Multimodal Language Models

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-08T12:54:56.375610Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T12:54:56.375610Z digest=sha256:2c93c6523d536352bf6ecfd684baf5baf004b1c4d0a40084d2f87ac7fe412f1c

Observation c964b9ad-5945-4630-94c1-15f72c0224b5 · outbound

This paper cites Llava- next: A strong zero-shot video understanding model, 2024.

EgoTextVQA: Towards Egocentric Scene-Text Aware Video Question Answering Llava- next: A strong zero-shot video understanding model, 2024

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T12:54:57.183592Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-08T12:54:56.379684Z digest=sha256:9cf26054ac281a25483e980a78dd487b58eb19a0fe596ee416acf99429b7ceed

Observation 6490b45e-29a4-45bb-a171-635ebdb7bd6c · outbound

This paper cites Diffusion-based blind text image super-resolution.

EgoTextVQA: Towards Egocentric Scene-Text Aware Video Question Answering Diffusion-based blind text image super-resolution

Reference 65

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T12:54:57.169518Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-08T12:54:56.383600Z digest=sha256:e3247386591208f3425095fe66a9af70501b8ece80e5494718f71286b61a22db

Observation 8c8b5f18-f5f1-4841-8f7e-5927283508ab · outbound

This paper cites Beyond LLaVA-HD: Diving into High-Resolution Large Multimodal Models.

EgoTextVQA: Towards Egocentric Scene-Text Aware Video Question Answering Beyond LLaVA-HD: Diving into High-Resolution Large Multimodal Models

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-08T12:54:56.387608Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T12:54:56.387608Z digest=sha256:31c02e456a0b565c36b56b518efd7a8444eefc706384bf49a1c4baca368e0aaf

Observation 0dd917b1-93d0-4ce9-94f5-f25fc4a8ae42 · outbound

This paper cites Towards video text visual question answering: Benchmark and baseline.

EgoTextVQA: Towards Egocentric Scene-Text Aware Video Question Answering Towards video text visual question answering: Benchmark and baseline

Reference 67

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T12:54:57.156198Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-08T12:54:56.391728Z digest=sha256:197a0ade1d0efa25edeb737551464414740ec2242f4c19973c5640e1bb27edab

Observation 72714bec-23ec-4e34-a8fa-b3ab4e74f246 · outbound

This paper cites Exploring sparse spatial relation in graph inference for text- based vqa.

EgoTextVQA: Towards Egocentric Scene-Text Aware Video Question Answering Exploring sparse spatial relation in graph inference for text- based vqa

Reference 68

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T12:54:57.142122Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-08T12:54:56.395299Z digest=sha256:d37b04ca29f89dd272566dc004f9191c54703e8f54d39ab0c5e96f2f44d8c148

Observation 68071ac0-de17-4ee3-9072-9ca247484a51 · outbound

This paper cites Scene-Text Grounding for Text-Based Video Question Answering.

EgoTextVQA: Towards Egocentric Scene-Text Aware Video Question Answering Scene-Text Grounding for Text-Based Video Question Answering

Reference 69

Resolution
malformed identifier
local_arxiv, observed 2026-08-08T12:54:56.516262Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-08T12:54:56.399357Z digest=sha256:da15aeb73a50a82f42da3a817895b321c8eb2a7e320b07fba72ac2b0ea7d77a9

Observation 5cc3978e-f96d-4cba-a5aa-c51c18902e0e · outbound

This paper cites I” should be used appropriately. Requirement 5: The questions should be of moderate length. When announcing the question please label each question as “Question 1, 2, 3: {question}.

EgoTextVQA: Towards Egocentric Scene-Text Aware Video Question Answering I” should be used appropriately. Requirement 5: The questions should be of moderate length. When announcing the question please label each question as “Question 1, 2, 3: {question}

Reference 70

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T12:54:57.128528Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-08T12:54:56.404806Z digest=sha256:f2c4ed6f260e9392e7b2acaaf08f6319e137a57c6887c0e983259fbc6b0ecb9a

Observation e2e47122-0bb7-4a7c-b3bd-dfa58fb80fa8 · outbound

This paper cites For example:.

EgoTextVQA: Towards Egocentric Scene-Text Aware Video Question Answering For example:

Reference 71

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T12:54:57.114701Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-08T12:54:56.408988Z digest=sha256:d54d8326ac56bf76a40a8ecf4a2e6ed0fd72cbdd199a0e03cc0fb43b8bf0823d

Observation 2b9872bd-9eab-43fe-9ccb-b8c93b47c37c · outbound

This paper cites an unresolved cited work.

EgoTextVQA: Towards Egocentric Scene-Text Aware Video Question Answering Unresolved cited work

Reference 72

Resolution
unresolved
raw_fallback, observed 2026-08-08T12:54:57.100405Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-08T12:54:56.413149Z digest=sha256:02e072ef5d1db517e347f510b69e4bf767edf5050b56d53db1d3071393999360

Observation 458701d6-fef6-4b9e-9a3f-844702db5b7d · outbound

This paper cites an unresolved cited work.

EgoTextVQA: Towards Egocentric Scene-Text Aware Video Question Answering Unresolved cited work

Reference 73

Resolution
unresolved
raw_fallback, observed 2026-08-08T12:54:57.086757Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-08T12:54:56.417105Z digest=sha256:869eb03c64880b271deeee75e71773bae1cb70ccc406592228336d37e483c23c

Observation 54dec6c6-efc5-4252-97c2-4b18cbce608f · outbound

This paper cites For example:.

EgoTextVQA: Towards Egocentric Scene-Text Aware Video Question Answering For example:

Reference 74

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T12:54:57.073350Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-08T12:54:56.421837Z digest=sha256:1e7aefd8ea1d0e4f9d6aec29e7484e50263ac7f3eeec2f5639199f80e73e246d

Observation 82e9b652-e3b4-40f4-aa59-b4cecc0acba3 · outbound

This paper cites an unresolved cited work.

EgoTextVQA: Towards Egocentric Scene-Text Aware Video Question Answering Unresolved cited work

Reference 75

Resolution
unresolved
raw_fallback, observed 2026-08-08T12:54:57.059763Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-08T12:54:56.425714Z digest=sha256:e977a59822208e77a3d0bce9e2c7fe80da1fd0c7d8fcdd406b83f52cf0925291

Observation a1ed87b2-e23a-473e-a0a2-10f9877f776c · outbound

This paper cites an unresolved cited work.

EgoTextVQA: Towards Egocentric Scene-Text Aware Video Question Answering Unresolved cited work

Reference 76

Resolution
unresolved
raw_fallback, observed 2026-08-08T12:54:57.046256Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-08T12:54:56.429588Z digest=sha256:7d4755ce996208b53dd4578fc1fa3b931548ee847bd6310e5c26623a991403e2

Observation a97fc98a-d790-4ae9-96b9-b64303857f51 · outbound

This paper cites an unresolved cited work.

EgoTextVQA: Towards Egocentric Scene-Text Aware Video Question Answering Unresolved cited work

Reference 77

Resolution
unresolved
raw_fallback, observed 2026-08-08T12:54:57.032389Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-08T12:54:56.433279Z digest=sha256:a91c95a8e73bd7a8245c4d715be06390c0a4b70fce76720b273f8687dd28af8a

Observation 0a2249fe-e1cd-4213-91f8-40bdbd6c6474 · outbound

This paper cites For example:.

EgoTextVQA: Towards Egocentric Scene-Text Aware Video Question Answering For example:

Reference 78

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T12:54:57.018820Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-08T12:54:56.437721Z digest=sha256:32eb6708640af820233ab9f79adf93521cbc0a60ec4c9f8d18dc4d88917cc60d

Observation 5b4ca62c-4072-42f2-acf9-3f8579a0c459 · outbound

This paper cites an unresolved cited work.

EgoTextVQA: Towards Egocentric Scene-Text Aware Video Question Answering Unresolved cited work

Reference 79

Resolution
unresolved
raw_fallback, observed 2026-08-08T12:54:57.004118Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-08T12:54:56.442124Z digest=sha256:d22766731a73f697277cd3492b184236fb67c864a403433e7f9922e06df143fa

Observation 22df1ca5-ea0d-4615-beff-1780dbe5dafa · outbound

This paper cites an unresolved cited work.

EgoTextVQA: Towards Egocentric Scene-Text Aware Video Question Answering Unresolved cited work

Reference 80

Resolution
unresolved
raw_fallback, observed 2026-08-08T12:54:56.989360Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-08T12:54:56.445784Z digest=sha256:0aad6bb3b70fe399c0260c4e3faab71e49c02b77a11e5e88f0032a389d4da64d

Observation c703cccc-3d52-47c6-aa33-797bd13546f4 · outbound

This paper cites an unresolved cited work.

EgoTextVQA: Towards Egocentric Scene-Text Aware Video Question Answering Unresolved cited work

Reference 81

Resolution
unresolved
raw_fallback, observed 2026-08-08T12:54:56.975681Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-08T12:54:56.449378Z digest=sha256:a8953f36d0bae9cb2819c1de9a922d1c1d852dba37684fd02d35bd43092a7646

Observation 9c8d3008-3003-4077-b3ff-77a4a12865ad · outbound

This paper cites For example:.

EgoTextVQA: Towards Egocentric Scene-Text Aware Video Question Answering For example:

Reference 82

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T12:54:56.961082Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-08T12:54:56.453801Z digest=sha256:9f310b112a955ca5bd47f1de635d58c6391f1cba0126550d5a582bc58eef2837

Observation 86832389-b959-4be6-add2-26981f4cbd4b · outbound

This paper cites an unresolved cited work.

EgoTextVQA: Towards Egocentric Scene-Text Aware Video Question Answering Unresolved cited work

Reference 83

Resolution
unresolved
raw_fallback, observed 2026-08-08T12:54:56.945488Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-08T12:54:56.458157Z digest=sha256:87e73123dd842802fa34695e2de6d4558d9d744fb0adf39e61e49457f93beaf8

Observation 34ead312-430b-4a56-b89c-d83a5f1eb0a1 · outbound

This paper cites an unresolved cited work.

EgoTextVQA: Towards Egocentric Scene-Text Aware Video Question Answering Unresolved cited work

Reference 84

Resolution
unresolved
raw_fallback, observed 2026-08-08T12:54:56.930849Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-08T12:54:56.462137Z digest=sha256:81956b1184b1ae2654162cb65d1aa6c246c51a1d2e945df6a7b0a202f27b4d2b

Observation 6a77d3a1-a090-4eed-89fd-57360faa4222 · outbound

This paper cites For example:.

EgoTextVQA: Towards Egocentric Scene-Text Aware Video Question Answering For example:

Reference 85

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T12:54:56.916036Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-08T12:54:56.465973Z digest=sha256:218ddbd7763432d595ab21870c4238478e3cffbc027a2a13a2db8cdc1d39d48e

Observation 13fb9799-69ef-463f-8e69-265da5d13af4 · outbound

This paper cites an unresolved cited work.

EgoTextVQA: Towards Egocentric Scene-Text Aware Video Question Answering Unresolved cited work

Reference 86

Resolution
unresolved
raw_fallback, observed 2026-08-08T12:54:56.901345Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-08T12:54:56.470475Z digest=sha256:fdb82ad92411978560dabda99afe8e968baab8520f1d7a483ef67f85031f3679

Observation c69b76a0-2013-4a16-bcb6-6290905136b1 · outbound

This paper cites Unanswerable.

EgoTextVQA: Towards Egocentric Scene-Text Aware Video Question Answering Unanswerable

Reference 87

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T12:54:56.886834Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-08T12:54:56.474300Z digest=sha256:429b5e1803bf19884d004e86f158410ceb8ff0f22d384d9047e8b09b6e31eb5e

Pith citing papers

No inbound Pith citation observations are available.