Pith. sign in

Paper Citation Record · LEDGER

TextHawk2: A Large Vision-Language Model Excels in Bilingual OCR and Grounding with 16x Fewer Tokens

As of 8 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 12 inbound Pith citation observations for arXiv:2410.05261.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2410.05261 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 12 of 12 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 12 of 12 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T15:42:31.079587Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-20T07:13:06.644583Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation a6719a44-8107-49c8-ba88-5c2df4e4c016 · inbound

Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling cites this paper.

Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling TextHawk2: A Large Vision-Language Model Excels in Bilingual OCR and Grounding with 16x Fewer Tokens

Reference 286

Resolution
verified exact
arxiv_id, observed 2026-05-10T13:23:58.272883Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T13:23:57.588851Z digest=sha256:7e71b8edb3226b63e61d0a2b2f64f1f302c88152978df971bc1208ddd9ca1e5f

Observation 51d4c1cb-6033-4345-a32b-c74ed19ae48d · inbound

DeepSeek-VL2: Mixture-of-Experts Vision-Language Models for Advanced Multimodal Understanding cites this paper.

DeepSeek-VL2: Mixture-of-Experts Vision-Language Models for Advanced Multimodal Understanding TextHawk2: A Large Vision-Language Model Excels in Bilingual OCR and Grounding with 16x Fewer Tokens

Reference 104

Resolution
verified exact
arxiv_id, observed 2026-05-11T10:09:26.325640Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-11T10:09:21.542356Z digest=sha256:50d4a250e8b8b21edbbe3ccd9b0482293fbb6f5e039db9857b4ed3494204a6ba

Observation 8e87d110-f4fe-44ac-9097-1b8a3435c0db · inbound

InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models cites this paper.

InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models TextHawk2: A Large Vision-Language Model Excels in Bilingual OCR and Grounding with 16x Fewer Tokens

Reference 141

Resolution
verified exact
arxiv_id, observed 2026-05-10T13:41:08.163440Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T13:41:07.991012Z digest=sha256:b74056c384dc64f595f5ba9d2d07bd9082dafa3ad6b2a694e52f4dc27d2d9a31

Observation d579bcc5-90ce-4206-a9d1-ca44152c45a2 · inbound

Dolphin: Document Image Parsing via Heterogeneous Anchor Prompting cites this paper.

Dolphin: Document Image Parsing via Heterogeneous Anchor Prompting TextHawk2: A Large Vision-Language Model Excels in Bilingual OCR and Grounding with 16x Fewer Tokens

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-07T15:42:31.079587Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:42:31.079587Z digest=sha256:28d625fb4e97bfe4439840303fed5271ed224f934aab1eee708812d081712d56

Observation 73163043-530e-41f2-8ede-891deaf0a7d3 · inbound

OCR-Reasoning Benchmark: Unveiling the True Capabilities of MLLMs in Complex Text-Rich Image Reasoning cites this paper.

OCR-Reasoning Benchmark: Unveiling the True Capabilities of MLLMs in Complex Text-Rich Image Reasoning TextHawk2: A Large Vision-Language Model Excels in Bilingual OCR and Grounding with 16x Fewer Tokens

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-07T14:57:21.429164Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:57:21.429164Z digest=sha256:2429307bda4792f3f788b4add24c370f6c0adab76384e42f8d852743a6c08b4a

Observation 1b75d618-2b7d-4489-bec3-61f553f0e1ad · inbound

A Survey on MLLM-based Visually Rich Document Understanding: Methods, Challenges, and Emerging Trends cites this paper.

A Survey on MLLM-based Visually Rich Document Understanding: Methods, Challenges, and Emerging Trends TextHawk2: A Large Vision-Language Model Excels in Bilingual OCR and Grounding with 16x Fewer Tokens

Reference 69

Resolution
verified exact
arxiv_id, observed 2026-05-19T04:42:04.361196Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-19T04:38:49.512293Z digest=sha256:1e6d008600b367107271987b28b2e6eaa00325ef6c4bbb1ca879d0185866a0ce

Observation 2ddefe12-f8f7-430b-932b-7aec5aeb1566 · inbound

Multi-Agent Interactive Question Generation Framework for Long Document Understanding cites this paper.

Multi-Agent Interactive Question Generation Framework for Long Document Understanding TextHawk2: A Large Vision-Language Model Excels in Bilingual OCR and Grounding with 16x Fewer Tokens

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-06T13:45:18.536837Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:45:18.536837Z digest=sha256:44f114034db9d5b408e8e1ba8b93e8e424ece3dadc92869b70e2e25259ef52a2

Observation 04693b41-5900-4236-b76b-41a5c210de4b · inbound

InternVL3.5: Advancing Open-Source Multimodal Models in Versatility, Reasoning, and Efficiency cites this paper.

InternVL3.5: Advancing Open-Source Multimodal Models in Versatility, Reasoning, and Efficiency TextHawk2: A Large Vision-Language Model Excels in Bilingual OCR and Grounding with 16x Fewer Tokens

Reference 170

Resolution
verified exact
arxiv_id, observed 2026-05-10T11:58:58.933332Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T11:58:58.660564Z digest=sha256:d6142846787cef4139509872a545cf594a0ad5922ac02b99b127be84dac7a274

Observation 81fefa0e-6857-40a4-adf0-e7f5603a0335 · inbound

Q-Mask: Query-driven Causal Masks for Text Anchoring in OCR-Oriented Vision-Language Models cites this paper.

Q-Mask: Query-driven Causal Masks for Text Anchoring in OCR-Oriented Vision-Language Models TextHawk2: A Large Vision-Language Model Excels in Bilingual OCR and Grounding with 16x Fewer Tokens

Reference 41

Resolution
verified exact
arxiv_id, observed 2026-05-13T23:33:26.644288Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-13T23:30:53.449935Z digest=sha256:fda9c7d0e476dfc86f4938be4a018397b4ee2db7fa95f36009143df4e8149544

Observation 19b594b1-1a15-4592-a419-388cb0803ad0 · inbound

UIPress: Bringing Optical Token Compression to UI-to-Code Generation cites this paper.

UIPress: Bringing Optical Token Compression to UI-to-Code Generation TextHawk2: A Large Vision-Language Model Excels in Bilingual OCR and Grounding with 16x Fewer Tokens

Reference 67

Resolution
verified exact
arxiv_id, observed 2026-05-11T07:00:59.204276Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T17:21:32.024105Z digest=sha256:f4ac68c1f2033d6bc249f242faef62241c0472d8ac1cb0d7325a9be83265d8a8

Observation 793c502c-88b8-4fad-ab15-0882408abbc4 · inbound

Visual Preference Optimization with Rubric Rewards cites this paper.

Visual Preference Optimization with Rubric Rewards TextHawk2: A Large Vision-Language Model Excels in Bilingual OCR and Grounding with 16x Fewer Tokens

Reference 57

Resolution
verified exact
arxiv_id, observed 2026-05-11T09:56:00.866790Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T15:45:52.980881Z digest=sha256:45aab3bfceebd91b5e6082fbb544a3e67de8ca480c6eedac97341aaab1c13341

Observation daf5c038-6a87-4bd2-ad6e-2659742359a8 · inbound

MetaRA: Metamorphic Robustness Assessment for Multimodal Large Language Model-based Visual Question Answering Systems cites this paper.

MetaRA: Metamorphic Robustness Assessment for Multimodal Large Language Model-based Visual Question Answering Systems TextHawk2: A Large Vision-Language Model Excels in Bilingual OCR and Grounding with 16x Fewer Tokens

Reference 40

Resolution
verified exact
arxiv_id, observed 2026-05-20T07:13:06.646603Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-20T07:08:27.096029Z digest=sha256:f6b25626ab097ed8759f9188cbd08b4f74bacda550cc2ec598161b21cc013028