Pith. sign in

Paper Citation Record · LEDGER

TextHawk2: A Large Vision-Language Model Excels in Bilingual OCR and Grounding with 16x Fewer Tokens

As of 17 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 19 inbound Pith citation observations for arXiv:2410.05261.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2410.05261 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 19 of 19 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-16T06:30:59.297886+00:00

measured 19 of 19 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-12T15:28:37.529815Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-20T07:13:06.644583Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation fac41a21-f0be-49ee-aa91-a5f5766ced3d · inbound

FocusLLaVA: A Coarse-to-Fine Approach for Efficient and Effective Visual Token Compression cites this paper.

FocusLLaVA: A Coarse-to-Fine Approach for Efficient and Effective Visual Token Compression TextHawk2: A Large Vision-Language Model Excels in Bilingual OCR and Grounding with 16x Fewer Tokens

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-12T15:28:37.529815Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:28:37.529815Z digest=sha256:410620ee09d8453a509eb0cf377d4fb4a82f2a73ea0322c85550367071bcbfdc

Observation a6719a44-8107-49c8-ba88-5c2df4e4c016 · inbound

Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling cites this paper.

Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling TextHawk2: A Large Vision-Language Model Excels in Bilingual OCR and Grounding with 16x Fewer Tokens

Reference 286

Resolution
verified exact
arxiv_id, observed 2026-05-10T13:23:58.272883Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-10T13:23:57.588851Z digest=sha256:2d252edf054af5840b2157831e529fed918e2988f450a57be654360783a25f57

Observation 521f3395-2251-439e-a251-375270dffeed · inbound

Dynamic Cross-Modal Alignment for Robust Semantic Location Prediction cites this paper.

Dynamic Cross-Modal Alignment for Robust Semantic Location Prediction TextHawk2: A Large Vision-Language Model Excels in Bilingual OCR and Grounding with 16x Fewer Tokens

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-11T16:40:38.938311Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T16:40:38.938311Z digest=sha256:df62045488d63a51561c6652cc9dbc93677f6b52be57fbf7c451a8f2d5b34348

Observation 8d7b2d00-ca13-4178-9608-84f2af0f140d · inbound

Selective State Space Memory for Large Vision-Language Models cites this paper.

Selective State Space Memory for Large Vision-Language Models TextHawk2: A Large Vision-Language Model Excels in Bilingual OCR and Grounding with 16x Fewer Tokens

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-11T16:43:15.654699Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T16:43:15.654699Z digest=sha256:f2c046187bfbca5313d61f0d24534436e4e42cf1424e4e572028e9f67ddd375c

Observation 51d4c1cb-6033-4345-a32b-c74ed19ae48d · inbound

DeepSeek-VL2: Mixture-of-Experts Vision-Language Models for Advanced Multimodal Understanding cites this paper.

DeepSeek-VL2: Mixture-of-Experts Vision-Language Models for Advanced Multimodal Understanding TextHawk2: A Large Vision-Language Model Excels in Bilingual OCR and Grounding with 16x Fewer Tokens

Reference 104

Resolution
verified exact
arxiv_id, observed 2026-05-11T10:09:26.325640Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-11T10:09:21.542356Z digest=sha256:99a7a1783e5138f3a5fad868cad66be4c65fa1e1dfe495beffae0b7915100e20

Observation 290f8d76-452e-4d2d-bfc2-4da99cfc6534 · inbound

Optimizing Vision-Language Interactions Through Decoder-Only Models cites this paper.

Optimizing Vision-Language Interactions Through Decoder-Only Models TextHawk2: A Large Vision-Language Model Excels in Bilingual OCR and Grounding with 16x Fewer Tokens

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-11T15:41:07.639261Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:41:07.639261Z digest=sha256:7f1b9b18215be15eff628311dc974ca0095bef1a39da6ba2fa042d02d67eb121

Observation 6280c8fe-0ba1-4b28-a26b-5b53566389d8 · inbound

Temporal Contrastive Learning for Video Temporal Reasoning in Large Vision-Language Models cites this paper.

Temporal Contrastive Learning for Video Temporal Reasoning in Large Vision-Language Models TextHawk2: A Large Vision-Language Model Excels in Bilingual OCR and Grounding with 16x Fewer Tokens

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-11T15:03:10.552828Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:03:10.552828Z digest=sha256:69126f13f8c4cee080cd9e748ead8f172ea0602ca878ee0824cf03765559cc36

Observation 884e615b-1289-4b06-9f1c-15d5311b9041 · inbound

Ocean-OCR: Towards General OCR Application via a Vision-Language Model cites this paper.

Ocean-OCR: Towards General OCR Application via a Vision-Language Model TextHawk2: A Large Vision-Language Model Excels in Bilingual OCR and Grounding with 16x Fewer Tokens

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-10T14:14:54.802004Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T14:14:54.802004Z digest=sha256:59ee3659983d7ee3112057e925b5526a6a59c87e586b3e223659eb35a708f439

Observation 8e87d110-f4fe-44ac-9097-1b8a3435c0db · inbound

InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models cites this paper.

InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models TextHawk2: A Large Vision-Language Model Excels in Bilingual OCR and Grounding with 16x Fewer Tokens

Reference 141

Resolution
verified exact
arxiv_id, observed 2026-05-10T13:41:08.163440Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-10T13:41:07.991012Z digest=sha256:b577eb43f9e46907e756012e7f1c66c275c29e9d8eaaf51acd881af47ad9456e

Observation d579bcc5-90ce-4206-a9d1-ca44152c45a2 · inbound

Dolphin: Document Image Parsing via Heterogeneous Anchor Prompting cites this paper.

Dolphin: Document Image Parsing via Heterogeneous Anchor Prompting TextHawk2: A Large Vision-Language Model Excels in Bilingual OCR and Grounding with 16x Fewer Tokens

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-07T15:42:31.079587Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:42:31.079587Z digest=sha256:b719f3756fe60fbd859651413f159dbbbb50f69afd61eb85757e1998b46acadb

Observation 73163043-530e-41f2-8ede-891deaf0a7d3 · inbound

OCR-Reasoning Benchmark: Unveiling the True Capabilities of MLLMs in Complex Text-Rich Image Reasoning cites this paper.

OCR-Reasoning Benchmark: Unveiling the True Capabilities of MLLMs in Complex Text-Rich Image Reasoning TextHawk2: A Large Vision-Language Model Excels in Bilingual OCR and Grounding with 16x Fewer Tokens

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-07T14:57:21.429164Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:57:21.429164Z digest=sha256:9a9c259d12d90e4a1f1ee29e2bcaf4cd682123ec72c6b767fc0e2a7ec3391f04

Observation 1b75d618-2b7d-4489-bec3-61f553f0e1ad · inbound

A Survey on MLLM-based Visually Rich Document Understanding: Methods, Challenges, and Emerging Trends cites this paper.

A Survey on MLLM-based Visually Rich Document Understanding: Methods, Challenges, and Emerging Trends TextHawk2: A Large Vision-Language Model Excels in Bilingual OCR and Grounding with 16x Fewer Tokens

Reference 69

Resolution
verified exact
arxiv_id, observed 2026-05-19T04:42:04.361196Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-05-19T04:38:49.512293Z digest=sha256:a362e03ccf850d2592034195ce9333b67c0f86df2397ec8082da5030ed7b3398

Observation 2ddefe12-f8f7-430b-932b-7aec5aeb1566 · inbound

Multi-Agent Interactive Question Generation Framework for Long Document Understanding cites this paper.

Multi-Agent Interactive Question Generation Framework for Long Document Understanding TextHawk2: A Large Vision-Language Model Excels in Bilingual OCR and Grounding with 16x Fewer Tokens

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-06T13:45:18.536837Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:45:18.536837Z digest=sha256:ff0c06b6abfb5bdfd208812a0f64c93b893bc78f0ed59c3b4d06b4ed66beea20

Observation 04693b41-5900-4236-b76b-41a5c210de4b · inbound

InternVL3.5: Advancing Open-Source Multimodal Models in Versatility, Reasoning, and Efficiency cites this paper.

InternVL3.5: Advancing Open-Source Multimodal Models in Versatility, Reasoning, and Efficiency TextHawk2: A Large Vision-Language Model Excels in Bilingual OCR and Grounding with 16x Fewer Tokens

Reference 170

Resolution
verified exact
arxiv_id, observed 2026-05-10T11:58:58.933332Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-10T11:58:58.660564Z digest=sha256:0a3df444762a93a4ec67c7e8106430c57e8695ca6c12190edf05cad788fe7d5b

Observation 81fefa0e-6857-40a4-adf0-e7f5603a0335 · inbound

Q-Mask: Query-driven Causal Masks for Text Anchoring in OCR-Oriented Vision-Language Models cites this paper.

Q-Mask: Query-driven Causal Masks for Text Anchoring in OCR-Oriented Vision-Language Models TextHawk2: A Large Vision-Language Model Excels in Bilingual OCR and Grounding with 16x Fewer Tokens

Reference 41

Resolution
verified exact
arxiv_id, observed 2026-05-13T23:33:26.644288Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-13T23:30:53.449935Z digest=sha256:2b0d0ab09b0c356e9e184c0e35bb84eb0259c58290848be9b5fa2af6dea14d35

Observation 19b594b1-1a15-4592-a419-388cb0803ad0 · inbound

UIPress: Bringing Optical Token Compression to UI-to-Code Generation cites this paper.

UIPress: Bringing Optical Token Compression to UI-to-Code Generation TextHawk2: A Large Vision-Language Model Excels in Bilingual OCR and Grounding with 16x Fewer Tokens

Reference 67

Resolution
verified exact
arxiv_id, observed 2026-05-11T07:00:59.204276Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-10T17:21:32.024105Z digest=sha256:bd2bbb7b76f4599b9bb58103b669aa211e934e2ec44b361426698d12ecefbc5f

Observation 793c502c-88b8-4fad-ab15-0882408abbc4 · inbound

Visual Preference Optimization with Rubric Rewards cites this paper.

Visual Preference Optimization with Rubric Rewards TextHawk2: A Large Vision-Language Model Excels in Bilingual OCR and Grounding with 16x Fewer Tokens

Reference 57

Resolution
verified exact
arxiv_id, observed 2026-05-11T09:56:00.866790Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-10T15:45:52.980881Z digest=sha256:c3e0fa749d72048ff9d79aaa84a18274067e214dec2aeec74ecf11f23b8f1491

Observation daf5c038-6a87-4bd2-ad6e-2659742359a8 · inbound

MetaRA: Metamorphic Robustness Assessment for Multimodal Large Language Model-based Visual Question Answering Systems cites this paper.

MetaRA: Metamorphic Robustness Assessment for Multimodal Large Language Model-based Visual Question Answering Systems TextHawk2: A Large Vision-Language Model Excels in Bilingual OCR and Grounding with 16x Fewer Tokens

Reference 40

Resolution
verified exact
arxiv_id, observed 2026-05-20T07:13:06.646603Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-20T07:08:27.096029Z digest=sha256:f0fe48c1077ed586fc3fcd26ea57a176ee92b899207984c52aae94abd83c259d

Observation bf31c70b-3f0c-430e-a618-afbd13d9eefa · inbound

Locating Failure in Multi-Page Visually Rich Document Understanding: An Empirical Attribution cites this paper.

Locating Failure in Multi-Page Visually Rich Document Understanding: An Empirical Attribution TextHawk2: A Large Vision-Language Model Excels in Bilingual OCR and Grounding with 16x Fewer Tokens

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-12T00:43:45.111390Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T00:43:45.111390Z digest=sha256:7840b46bfe40c03d664e2517f192d9179106407822d800e61900b00434c564b2