Pith. sign in

Paper Citation Record · LEDGER

TextHawk: Exploring Efficient Fine-Grained Perception of Multimodal Large Language Models

As of 19 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 16 inbound Pith citation observations for arXiv:2404.09204.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2404.09204 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 16 of 16 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00

measured 16 of 16 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-12T15:28:37.501180Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-06-29T08:13:15.344020Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 03c3bda6-d15d-4433-b521-9eefbf563a73 · inbound

PDF-WuKong: A Large Multimodal Model for Efficient Long PDF Reading with End-to-End Sparse Sampling cites this paper.

PDF-WuKong: A Large Multimodal Model for Efficient Long PDF Reading with End-to-End Sparse Sampling TextHawk: Exploring Efficient Fine-Grained Perception of Multimodal Large Language Models

Reference 62

Resolution
verified exact
arxiv_id, observed 2026-05-23T19:43:23.684066Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-23T19:39:35.147671Z digest=sha256:2b113794ea2bd4f60608a013632704a0da12570771e086605262e397df8125b3

Observation c4ce4de7-7a45-411c-a2a9-754906917c48 · inbound

FocusLLaVA: A Coarse-to-Fine Approach for Efficient and Effective Visual Token Compression cites this paper.

FocusLLaVA: A Coarse-to-Fine Approach for Efficient and Effective Visual Token Compression TextHawk: Exploring Efficient Fine-Grained Perception of Multimodal Large Language Models

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-12T15:28:37.501180Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:28:37.501180Z digest=sha256:f1c4e2fd0301fcf4db0c6828b698b38f1fd93d61e640b2b60b789c02be479b98

Observation 8c9bc405-f105-4bd6-abc5-9ae0a315a36c · inbound

Cross-Lingual Text-Rich Visual Comprehension: An Information Theory Perspective cites this paper.

Cross-Lingual Text-Rich Visual Comprehension: An Information Theory Perspective TextHawk: Exploring Efficient Fine-Grained Perception of Multimodal Large Language Models

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-11T05:14:27.980946Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T05:14:27.980946Z digest=sha256:c95b1190a949e2ada39314fcc781cc44590539bd695d904d5d9d167f4e5a460c

Observation 45f6205c-b1ea-48cc-969b-71dca94e7c7a · inbound

Survey on Question Answering over Visually Rich Documents: Methods, Challenges, and Trends cites this paper.

Survey on Question Answering over Visually Rich Documents: Methods, Challenges, and Trends TextHawk: Exploring Efficient Fine-Grained Perception of Multimodal Large Language Models

Reference 103

Resolution
unresolved
no resolver link, observed 2026-08-10T22:17:28.102708Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T22:17:28.102708Z digest=sha256:a05be7a97aecc4c899ca1b5666a0133e26a7e6f6010ad1393d74eb4aeacacc7d

Observation f857a2fb-c246-49ff-8285-92827b57a347 · inbound

Visual Large Language Models for Generalized and Specialized Applications cites this paper.

Visual Large Language Models for Generalized and Specialized Applications TextHawk: Exploring Efficient Fine-Grained Perception of Multimodal Large Language Models

Reference 118

Resolution
unresolved
no resolver link, observed 2026-08-10T22:08:09.340902Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:08:09.340902Z digest=sha256:c7cd94be33b4f1ac419bfa143bde5d330ea03321dbec15483113be6f1998fafd

Observation 5f4d8a90-9c17-4568-987c-dd50b5a74b1f · inbound

RedundancyLens: Revealing and Exploiting Visual Token Processing Redundancy for Efficient Decoder-Only MLLMs cites this paper.

RedundancyLens: Revealing and Exploiting Visual Token Processing Redundancy for Efficient Decoder-Only MLLMs TextHawk: Exploring Efficient Fine-Grained Perception of Multimodal Large Language Models

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-09T21:39:21.842719Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T21:39:21.842719Z digest=sha256:fa62afa631862aaa7854d0d82cf9b726abc252e9f418eb2151ebf9bf40bc6889

Observation bab3b555-d38c-47f6-afe0-f98bd8f69ca1 · inbound

Dolphin: Document Image Parsing via Heterogeneous Anchor Prompting cites this paper.

Dolphin: Document Image Parsing via Heterogeneous Anchor Prompting TextHawk: Exploring Efficient Fine-Grained Perception of Multimodal Large Language Models

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-07T15:42:30.992620Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:42:30.992620Z digest=sha256:593deaa1509efe24204bdd5346496b71c7a7cd1d73ffd64acff839b46d5a3de3

Observation bd675666-0726-40a7-ada7-dcd14b2e0ccb · inbound

Native Visual Understanding: Resolving Resolution Dilemmas in Vision-Language Models cites this paper.

Native Visual Understanding: Resolving Resolution Dilemmas in Vision-Language Models TextHawk: Exploring Efficient Fine-Grained Perception of Multimodal Large Language Models

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-07T00:48:33.913826Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:48:33.913826Z digest=sha256:6afb21e4d399882ac94784c01164c363f58d4eee1f505528d59edb3a5e23c16f

Observation da945551-70eb-4509-a555-6964945401cb · inbound

Single-to-mix Modality Alignment with Multimodal Large Language Model for Document Image Machine Translation cites this paper.

Single-to-mix Modality Alignment with Multimodal Large Language Model for Document Image Machine Translation TextHawk: Exploring Efficient Fine-Grained Perception of Multimodal Large Language Models

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-06T18:43:21.184881Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:43:21.184881Z digest=sha256:298c0b1a9601f08fb6b46376b23f8bba1fc14ffeabd276cbd245edef8b3778ac

Observation bf2c4e41-add8-435f-831c-2cee03d3ce58 · inbound

Improving MLLM's Document Image Machine Translation via Synchronously Self-reviewing Its OCR Proficiency cites this paper.

Improving MLLM's Document Image Machine Translation via Synchronously Self-reviewing Its OCR Proficiency TextHawk: Exploring Efficient Fine-Grained Perception of Multimodal Large Language Models

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-06T18:28:13.949209Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:28:13.949209Z digest=sha256:d7d5da43a2221ec223b9a5e234e3a8974024c190808ed5cbf922fec1e871c9a0

Observation 45ded1b3-bba6-442c-9ca4-9370fca55ed5 · inbound

A Survey on MLLM-based Visually Rich Document Understanding: Methods, Challenges, and Emerging Trends cites this paper.

A Survey on MLLM-based Visually Rich Document Understanding: Methods, Challenges, and Emerging Trends TextHawk: Exploring Efficient Fine-Grained Perception of Multimodal Large Language Models

Reference 68

Resolution
verified exact
arxiv_id, observed 2026-05-19T04:42:04.336814Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-05-19T04:38:49.512293Z digest=sha256:d61be6343621c8f31d2198a0f559f095fffc14b0d57f42fe030ddabb66be45a1

Observation d4dd574f-1f6e-4058-a219-a9e9662bea66 · inbound

Cognitive Mismatch in Multimodal Large Language Models for Discrete Symbol Understanding cites this paper.

Cognitive Mismatch in Multimodal Large Language Models for Discrete Symbol Understanding TextHawk: Exploring Efficient Fine-Grained Perception of Multimodal Large Language Models

Reference 27

Resolution
verified exact
arxiv_id, observed 2026-05-15T09:15:21.061230Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-15T09:11:31.870441Z digest=sha256:9132965ba02a911b3339a7c6fcc9ff0ea152fafe8109c70f4709c93d23832253

Observation b5d9ef0d-62c2-4bed-9758-5cc8ab1a50a9 · inbound

From Plausibility to Verifiability: Risk-Controlled Generative OCR with Vision-Language Models cites this paper.

From Plausibility to Verifiability: Risk-Controlled Generative OCR with Vision-Language Models TextHawk: Exploring Efficient Fine-Grained Perception of Multimodal Large Language Models

Reference 45

Resolution
metadata mismatch
arxiv_id, observed 2026-05-15T08:55:19.397669Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-15T08:53:18.268970Z digest=sha256:071fa2ff84b9828407418dcef0f1bf1565526c0b9aac8839529176425bbbb1a1

Observation f2017440-6d04-4e6e-98ab-582e3ae656eb · inbound

ReAlign: Optimizing the Visual Document Retriever with Reasoning-Guided Fine-Grained Alignment cites this paper.

ReAlign: Optimizing the Visual Document Retriever with Reasoning-Guided Fine-Grained Alignment TextHawk: Exploring Efficient Fine-Grained Perception of Multimodal Large Language Models

Reference 95

Resolution
verified exact
arxiv_id, observed 2026-05-11T06:15:58.167006Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-10T17:43:15.630570Z digest=sha256:6ee61de598e510599497e6f34c3c048520d5ada7e0916a08d7333cde66b184cc

Observation c1e3a305-7ffb-4e9a-9c0a-8ffd881c53d2 · inbound

MetaRA: Metamorphic Robustness Assessment for Multimodal Large Language Model-based Visual Question Answering Systems cites this paper.

MetaRA: Metamorphic Robustness Assessment for Multimodal Large Language Model-based Visual Question Answering Systems TextHawk: Exploring Efficient Fine-Grained Perception of Multimodal Large Language Models

Reference 39

Resolution
verified exact
arxiv_id, observed 2026-05-20T07:13:06.655691Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-20T07:08:27.096029Z digest=sha256:fe1ffb02baa120f410f9f0b9440b141ed598162ce442ea96206d99b6323f6541

Observation 619ea8c1-1590-46c9-a539-750845217293 · inbound

Grounded 3D-Aware Spatial Vision-Language Modeling cites this paper.

Grounded 3D-Aware Spatial Vision-Language Modeling TextHawk: Exploring Efficient Fine-Grained Perception of Multimodal Large Language Models

Reference 119

Resolution
verified exact
arxiv_id, observed 2026-06-29T08:13:15.353661Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-06-29T08:08:36.012761Z digest=sha256:f895491deb0e3de993ee01d5c45e48255c1ecd2af2912535e6eb7c037a0b09c8