Pith. sign in

Paper Citation Record · LEDGER

HAUR: Human Annotation Understanding and Recognition Through Text-Heavy Images

As of 20 August 2026, this Paper Citation Record lists 21 of 21 outbound references and 0 inbound Pith citation observations for arXiv:2412.18327.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2412.18327 v1

Coverage vector

measured 21 of 21 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-11T04:51:41.504443Z

measured 21 of 21 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-20T06:33:59.587034+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

21 of 21 outbound references displayed

  • verified exact0
  • verified fuzzy10
  • unresolved11
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation b13f1780-b852-4c66-81f7-29c62b1203d3 · outbound

This paper cites Textcaps: a dataset for image captioning with reading compre- hension,.

HAUR: Human Annotation Understanding and Recognition Through Text-Heavy Images Textcaps: a dataset for image captioning with reading compre- hension,

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T04:51:41.736295Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-11T04:51:41.431664Z digest=sha256:aeac73ef2db64f7c6183b6c9b3fd25fa4e4f1157eba9027d044d07c29aeb0221

Observation 4d15c7f0-4683-4d32-a25d-048bd19f9614 · outbound

This paper cites Towards vqa models that can read,.

HAUR: Human Annotation Understanding and Recognition Through Text-Heavy Images Towards vqa models that can read,

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-11T04:51:41.435970Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:51:41.435970Z digest=sha256:00a8a3d623a22943542f778b3ca65b3e2b8eca29245bdf40b1f837a2e42ecaf4

Observation 0fce03ac-160f-474e-be42-57b4f41b65a5 · outbound

This paper cites Scene text visual question answering,.

HAUR: Human Annotation Understanding and Recognition Through Text-Heavy Images Scene text visual question answering,

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T04:51:41.719788Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-11T04:51:41.439636Z digest=sha256:b2ca16b0c35955590c70d4b1d4f313ba5f88bea52d863f1ea3af9ab2c0893414

Observation 947b8f29-7617-4143-81be-1827ba2f137e · outbound

This paper cites RecipeQA: A Challenge Dataset for Multimodal Comprehension of Cooking Recipes.

HAUR: Human Annotation Understanding and Recognition Through Text-Heavy Images RecipeQA: A Challenge Dataset for Multimodal Comprehension of Cooking Recipes

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-11T04:51:41.443604Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:51:41.443604Z digest=sha256:0dfc973541de68385f8d3703687d052856c2a61b97610098153145e48a3b1d04

Observation 6845cd48-f542-4935-a45d-6b161b81db68 · outbound

This paper cites Visualmrc: Machine reading comprehension on document images,.

HAUR: Human Annotation Understanding and Recognition Through Text-Heavy Images Visualmrc: Machine reading comprehension on document images,

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T04:51:41.709627Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-11T04:51:41.447736Z digest=sha256:c92076cefeddd0f211d80d984753d282ef007b65251511d5f26f64ac133345c2

Observation b9f4e125-59c9-4e21-a69d-a6509f3abd84 · outbound

This paper cites Textbook question answering under instructor guidance with memory networks,.

HAUR: Human Annotation Understanding and Recognition Through Text-Heavy Images Textbook question answering under instructor guidance with memory networks,

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T04:51:41.699323Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-11T04:51:41.451438Z digest=sha256:0bc68b780e37effa6201146b6e432c39a51c36d2a1ea64ebccdba5e9daf6a305

Observation 9641ece4-cc8e-448f-8941-46004f050281 · outbound

This paper cites ViOCRVQA: Novel Benchmark Dataset and Vision Reader for Visual Question Answering by Understanding Vietnamese Text in Images.

HAUR: Human Annotation Understanding and Recognition Through Text-Heavy Images ViOCRVQA: Novel Benchmark Dataset and Vision Reader for Visual Question Answering by Understanding Vietnamese Text in Images

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-11T04:51:41.455584Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:51:41.455584Z digest=sha256:a74ee0536dc5fc8713091b00a96c860e86acd4145669dab1e315dcecf4dc3129

Observation 0ad4ae96-5f80-4f18-82d8-fd7c1b2137ee · outbound

This paper cites Docvqa: A dataset for vqa on document images,.

HAUR: Human Annotation Understanding and Recognition Through Text-Heavy Images Docvqa: A dataset for vqa on document images,

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T04:51:41.687924Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-11T04:51:41.459665Z digest=sha256:59769d457e29d129655ce56b19065d34abfa898c62adc28199099fd15eae8584

Observation 4685e783-8a7c-43ea-9679-20d2ad89f37c · outbound

This paper cites Infographicvqa,.

HAUR: Human Annotation Understanding and Recognition Through Text-Heavy Images Infographicvqa,

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T04:51:41.675490Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-11T04:51:41.463469Z digest=sha256:6c25d3410318452ca6065125d13e0ce2cf0125c45a332736c291ab718149a396

Observation cdc5accd-ca60-4c84-92e9-bf5a8ae19529 · outbound

This paper cites Iterative answer prediction with pointer-augmented multimodal trans- formers for textvqa,.

HAUR: Human Annotation Understanding and Recognition Through Text-Heavy Images Iterative answer prediction with pointer-augmented multimodal trans- formers for textvqa,

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T04:51:41.664809Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-11T04:51:41.466907Z digest=sha256:1ef9a048a1dded96157c670bc09ff73bb0104f1f8084bd0f3e5e972f3763174e

Observation 90cfb49b-6902-4e11-9324-50a9f6011988 · outbound

This paper cites Tap: Text- aware pre-training for text-vqa and text-caption,.

HAUR: Human Annotation Understanding and Recognition Through Text-Heavy Images Tap: Text- aware pre-training for text-vqa and text-caption,

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T04:51:41.653349Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-11T04:51:41.470323Z digest=sha256:bdd06798dfd2f8f8e74b7cf6f0e4af88ddb4b325d8fd396fb77aa0329ca00e38

Observation 764ffbff-1bd1-4b1c-a658-d51cfe931b0b · outbound

This paper cites Vilt: Vision-and-language transformer without convolution or region supervision,.

HAUR: Human Annotation Understanding and Recognition Through Text-Heavy Images Vilt: Vision-and-language transformer without convolution or region supervision,

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T04:51:41.640878Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-11T04:51:41.473362Z digest=sha256:4b61af3c9d808bf6c3e8be848a9daddcb7ac2567207ca5e546e7e470e744108f

Observation c417c1f9-1530-4def-a164-59ff8f68653c · outbound

This paper cites PaLI: A Jointly-Scaled Multilingual Language-Image Model.

HAUR: Human Annotation Understanding and Recognition Through Text-Heavy Images PaLI: A Jointly-Scaled Multilingual Language-Image Model

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-11T04:51:41.476261Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:51:41.476261Z digest=sha256:3d570841a9d082eb09bde141c84d781233aaaa9ce7ab9ba6f85cc7c175f8ef0f

Observation 41fc879e-a193-4ee4-a796-5aae5b8012bf · outbound

This paper cites PaLI-3 Vision Language Models: Smaller, Faster, Stronger.

HAUR: Human Annotation Understanding and Recognition Through Text-Heavy Images PaLI-3 Vision Language Models: Smaller, Faster, Stronger

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-11T04:51:41.479694Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:51:41.479694Z digest=sha256:129d5ba327c5b804cfd251a91bde83e918d139a570cb460b54460128f0176931

Observation 3ee8ae27-66d3-44e4-af24-911f9d44c533 · outbound

This paper cites ScreenAI: A Vision-Language Model for UI and Infographics Understanding.

HAUR: Human Annotation Understanding and Recognition Through Text-Heavy Images ScreenAI: A Vision-Language Model for UI and Infographics Understanding

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-11T04:51:41.483094Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:51:41.483094Z digest=sha256:1744d688cab4634d4de398b39a4cdae5fb56a3fb62bbbb90c2c74ac734e4a401

Observation b162b25a-c8ac-4f19-b3c6-a53157779eba · outbound

This paper cites Pix2struct: Screenshot parsing as pretraining for visual language understanding,.

HAUR: Human Annotation Understanding and Recognition Through Text-Heavy Images Pix2struct: Screenshot parsing as pretraining for visual language understanding,

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T04:51:41.629046Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-11T04:51:41.486401Z digest=sha256:295ceda445317450017adf0be5054c791294a180587a57c8c2ba011169bc8e7c

Observation e2eea61b-064d-4b38-ba10-864913a2286c · outbound

This paper cites Llava-next: Improved reasoning, ocr, and world knowledge,.

HAUR: Human Annotation Understanding and Recognition Through Text-Heavy Images Llava-next: Improved reasoning, ocr, and world knowledge,

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-11T04:51:41.489335Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:51:41.489335Z digest=sha256:5929d464e5439b9272c715852c043bcfea5708fbbc1f63e25bc71fa72f56aeee

Observation 6ce0abf7-27f0-4f9c-bbbc-b94bc4fb3574 · outbound

This paper cites Improved baselines with visual instruction tuning,.

HAUR: Human Annotation Understanding and Recognition Through Text-Heavy Images Improved baselines with visual instruction tuning,

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-11T04:51:41.492694Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:51:41.492694Z digest=sha256:088f5eb145bac8f1203f992fc44ac6aac6a0d20a48c690599f9caecf19f798f4

Observation caf4b0ec-05e8-4315-96cb-dd3096ae5d26 · outbound

This paper cites Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond.

HAUR: Human Annotation Understanding and Recognition Through Text-Heavy Images Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-11T04:51:41.496349Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:51:41.496349Z digest=sha256:a9c7bb242b889b294ad3d7ca732d627dac250b7bfa77ce099a6bd4ebe05e6f77

Observation c6f5a895-1ada-4aa3-95e1-5dca3a1ef7b4 · outbound

This paper cites Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution.

HAUR: Human Annotation Understanding and Recognition Through Text-Heavy Images Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-11T04:51:41.500696Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:51:41.500696Z digest=sha256:baa551cf8aa1ae4374ec92e465c7b3c4b5ccd157265f7590a7b6a99bdca6edf9

Observation 77b65599-22bc-4802-b0fd-8dc5b7e743b1 · outbound

This paper cites an unresolved cited work.

HAUR: Human Annotation Understanding and Recognition Through Text-Heavy Images Unresolved cited work

Reference 500

Resolution
unresolved
raw_fallback, observed 2026-08-11T04:51:41.603712Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-11T04:51:41.504443Z digest=sha256:20197741ee3ca6f9ab9c75f24a08c79dc56399f06e34cc039275500b403d0c86

Pith citing papers

No inbound Pith citation observations are available.