Pith. sign in

Paper Citation Record · LEDGER

TextHawk: Exploring Efficient Fine-Grained Perception of Multimodal Large Language Models

As of 9 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 11 inbound Pith citation observations for arXiv:2404.09204.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2404.09204 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 11 of 11 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 11 of 11 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T15:42:30.992620Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-06-29T08:13:15.344020Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 03c3bda6-d15d-4433-b521-9eefbf563a73 · inbound

PDF-WuKong: A Large Multimodal Model for Efficient Long PDF Reading with End-to-End Sparse Sampling cites this paper.

PDF-WuKong: A Large Multimodal Model for Efficient Long PDF Reading with End-to-End Sparse Sampling TextHawk: Exploring Efficient Fine-Grained Perception of Multimodal Large Language Models

Reference 62

Resolution
verified exact
arxiv_id, observed 2026-05-23T19:43:23.684066Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-23T19:39:35.147671Z digest=sha256:b06c47be63123e22615468d8657c4cba3bb6681dcd8b6dbd1710959af027ea67

Observation bab3b555-d38c-47f6-afe0-f98bd8f69ca1 · inbound

Dolphin: Document Image Parsing via Heterogeneous Anchor Prompting cites this paper.

Dolphin: Document Image Parsing via Heterogeneous Anchor Prompting TextHawk: Exploring Efficient Fine-Grained Perception of Multimodal Large Language Models

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-07T15:42:30.992620Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:42:30.992620Z digest=sha256:0d8719dc76c4292f7aceea627cf0b6988d324afe56bf136184a998fac987290d

Observation bd675666-0726-40a7-ada7-dcd14b2e0ccb · inbound

Native Visual Understanding: Resolving Resolution Dilemmas in Vision-Language Models cites this paper.

Native Visual Understanding: Resolving Resolution Dilemmas in Vision-Language Models TextHawk: Exploring Efficient Fine-Grained Perception of Multimodal Large Language Models

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-07T00:48:33.913826Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:48:33.913826Z digest=sha256:52969330fee99a402e695ec4dbf5e39d6f0bc188ce92b97eba6fa390e783ddc8

Observation da945551-70eb-4509-a555-6964945401cb · inbound

Single-to-mix Modality Alignment with Multimodal Large Language Model for Document Image Machine Translation cites this paper.

Single-to-mix Modality Alignment with Multimodal Large Language Model for Document Image Machine Translation TextHawk: Exploring Efficient Fine-Grained Perception of Multimodal Large Language Models

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-06T18:43:21.184881Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:43:21.184881Z digest=sha256:c04a0e0fb2e6a9ced8a255f04057e4e7dbb7bc38bdabb89d79ee094d3f6cc75c

Observation bf2c4e41-add8-435f-831c-2cee03d3ce58 · inbound

Improving MLLM's Document Image Machine Translation via Synchronously Self-reviewing Its OCR Proficiency cites this paper.

Improving MLLM's Document Image Machine Translation via Synchronously Self-reviewing Its OCR Proficiency TextHawk: Exploring Efficient Fine-Grained Perception of Multimodal Large Language Models

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-06T18:28:13.949209Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:28:13.949209Z digest=sha256:34a7a1c50f96fe8b836a520dc549e014994a5a83849c73feec44c2fb0ba453b6

Observation 45ded1b3-bba6-442c-9ca4-9370fca55ed5 · inbound

A Survey on MLLM-based Visually Rich Document Understanding: Methods, Challenges, and Emerging Trends cites this paper.

A Survey on MLLM-based Visually Rich Document Understanding: Methods, Challenges, and Emerging Trends TextHawk: Exploring Efficient Fine-Grained Perception of Multimodal Large Language Models

Reference 68

Resolution
verified exact
arxiv_id, observed 2026-05-19T04:42:04.336814Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-19T04:38:49.512293Z digest=sha256:ffbe382f378b0c3eeea8f509cf2056416112a73281786be400a3639131d53ec7

Observation d4dd574f-1f6e-4058-a219-a9e9662bea66 · inbound

Cognitive Mismatch in Multimodal Large Language Models for Discrete Symbol Understanding cites this paper.

Cognitive Mismatch in Multimodal Large Language Models for Discrete Symbol Understanding TextHawk: Exploring Efficient Fine-Grained Perception of Multimodal Large Language Models

Reference 27

Resolution
verified exact
arxiv_id, observed 2026-05-15T09:15:21.061230Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-15T09:11:31.870441Z digest=sha256:56fe03465c4ab02f4b0df4eea8b56106af25f48d3f059076aee04486505d329d

Observation b5d9ef0d-62c2-4bed-9758-5cc8ab1a50a9 · inbound

From Plausibility to Verifiability: Risk-Controlled Generative OCR with Vision-Language Models cites this paper.

From Plausibility to Verifiability: Risk-Controlled Generative OCR with Vision-Language Models TextHawk: Exploring Efficient Fine-Grained Perception of Multimodal Large Language Models

Reference 45

Resolution
metadata mismatch
arxiv_id, observed 2026-05-15T08:55:19.397669Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-15T08:53:18.268970Z digest=sha256:61ff9b54bbf3e800272dfc53511fb87534d9421bb1cb53f0a9e548ad6b7fad1a

Observation f2017440-6d04-4e6e-98ab-582e3ae656eb · inbound

ReAlign: Optimizing the Visual Document Retriever with Reasoning-Guided Fine-Grained Alignment cites this paper.

ReAlign: Optimizing the Visual Document Retriever with Reasoning-Guided Fine-Grained Alignment TextHawk: Exploring Efficient Fine-Grained Perception of Multimodal Large Language Models

Reference 95

Resolution
verified exact
arxiv_id, observed 2026-05-11T06:15:58.167006Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-10T17:43:15.630570Z digest=sha256:362a779311dd8c2fbced54dec9c9f915bbe6ec5d5cb533d20767993f6ad6a610

Observation c1e3a305-7ffb-4e9a-9c0a-8ffd881c53d2 · inbound

MetaRA: Metamorphic Robustness Assessment for Multimodal Large Language Model-based Visual Question Answering Systems cites this paper.

MetaRA: Metamorphic Robustness Assessment for Multimodal Large Language Model-based Visual Question Answering Systems TextHawk: Exploring Efficient Fine-Grained Perception of Multimodal Large Language Models

Reference 39

Resolution
verified exact
arxiv_id, observed 2026-05-20T07:13:06.655691Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-20T07:08:27.096029Z digest=sha256:c7c2fbf602547ab9184e6f3be26ba8cfee706dbdbec9c8a27cf5f20b04b62a52

Observation 619ea8c1-1590-46c9-a539-750845217293 · inbound

Grounded 3D-Aware Spatial Vision-Language Modeling cites this paper.

Grounded 3D-Aware Spatial Vision-Language Modeling TextHawk: Exploring Efficient Fine-Grained Perception of Multimodal Large Language Models

Reference 119

Resolution
verified exact
arxiv_id, observed 2026-06-29T08:13:15.353661Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-29T08:08:36.012761Z digest=sha256:ee20c107f26285a312e417cfe0af66e89cfdc76980f948ecacdf497c5e1b08f2