Pith. sign in

Paper Citation Record · LEDGER

DRISHTIKON: Visual Grounding at Multiple Granularities in Documents

As of 7 August 2026, this Paper Citation Record lists 24 of 24 outbound references and 0 inbound Pith citation observations for arXiv:2506.21316.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.21316 v2

Coverage vector

measured 24 of 24 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T22:34:38.340134Z

measured 24 of 24 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

24 of 24 outbound references displayed

  • verified exact1
  • verified fuzzy15
  • unresolved8
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 08222e7e-2dc9-4fb1-a287-4bf071b47021 · outbound

This paper cites Manmatha.

DRISHTIKON: Visual Grounding at Multiple Granularities in Documents Manmatha

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:34:41.825117Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T22:34:36.769699Z digest=sha256:128b2a4b67e595741144be21a6e1f56f5ed8343eb2b59f68b690c918c422eb6c

Observation 818156ef-effa-486f-b649-57812d72f569 · outbound

This paper cites Paddleocr, awesome multilingual ocr toolkits based on paddlepaddle.

DRISHTIKON: Visual Grounding at Multiple Granularities in Documents Paddleocr, awesome multilingual ocr toolkits based on paddlepaddle

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:34:41.576762Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T22:34:36.835297Z digest=sha256:eda8fc09059cc579f1834fd90284760c183fc66164363b510974edb51cb716f7

Observation 734429a8-cdfb-4d64-a25d-c84fc13fdbfa · outbound

This paper cites Bharatgen unveils patram: India’s pioneer- ing vision-language foundation model for document intelli- gence, 2025.

DRISHTIKON: Visual Grounding at Multiple Granularities in Documents Bharatgen unveils patram: India’s pioneer- ing vision-language foundation model for document intelli- gence, 2025

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:34:41.327537Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T22:34:36.925134Z digest=sha256:2ac72e310745fa2f33467d81eabc7f7a365b14df293e0d75c24aaa116e06273a

Observation 41ddf928-70c8-4b8e-a667-2cb316beb88d · outbound

This paper cites Molmo and pixmo: Open weights and open data for state- of-the-art vision-language models, 2024.

DRISHTIKON: Visual Grounding at Multiple Granularities in Documents Molmo and pixmo: Open weights and open data for state- of-the-art vision-language models, 2024

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:34:41.148905Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T22:34:36.989772Z digest=sha256:e67c9b6b5690b685e7d333a25b86ccb2a9a3bd8aac0f5caa8be5cba3f77c7908

Observation b842fa82-5fc4-4a65-8a4d-9ea92d83c69f · outbound

This paper cites The Llama 3 Herd of Models.

DRISHTIKON: Visual Grounding at Multiple Granularities in Documents The Llama 3 Herd of Models

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-06T22:34:37.070851Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:34:37.070851Z digest=sha256:7fa82c39cbf50d86d9ced2066a04a1e61f70514bb0f9939c4385581522db6d51

Observation d5c9bf1e-a4e8-4fb4-b9f3-a137212ed0c2 · outbound

This paper cites mPLUG-DocOwl2: High-resolution Compressing for OCR-free Multi-page Document Understanding.

DRISHTIKON: Visual Grounding at Multiple Granularities in Documents mPLUG-DocOwl2: High-resolution Compressing for OCR-free Multi-page Document Understanding

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-06T22:34:37.142434Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:34:37.142434Z digest=sha256:ac000ea1d1235372a2646e968be95997e13a8c31c7e39aec36ebcf1ea1124810

Observation f9551178-2d6b-4390-8a77-4060d5b445cb · outbound

This paper cites Layoutlmv3: Pre-training for document ai with uni- fied text and image masking.

DRISHTIKON: Visual Grounding at Multiple Granularities in Documents Layoutlmv3: Pre-training for document ai with uni- fied text and image masking

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:34:40.974738Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T22:34:37.230465Z digest=sha256:2de156934bdbb7f68bd49d19a3b6f12c55ad5394741b96a723583ac310afa038

Observation ff6eef45-1be0-4eb1-a914-190b55fc339d · outbound

This paper cites an unresolved cited work.

DRISHTIKON: Visual Grounding at Multiple Granularities in Documents Unresolved cited work

Reference 8

Resolution
unresolved
raw_fallback, observed 2026-08-06T22:34:40.730665Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T22:34:37.298695Z digest=sha256:c5f6d3930eec6913bd1abee271510f3437bc9d0efa343f0d21e4c1da65a2c40d

Observation a3633cfb-e759-4e92-b1b0-df3cb5bc3d99 · outbound

This paper cites Towards visual text grounding of multimodal large language model.

DRISHTIKON: Visual Grounding at Multiple Granularities in Documents Towards visual text grounding of multimodal large language model

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-06T22:34:37.388852Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:34:37.388852Z digest=sha256:62e7ee05aab8831b5651282ce693351ce829f31c6f40aabd49de981f67c73b2c

Observation f288daed-560b-4aaa-8e87-ae9c25f7473c · outbound

This paper cites Layoutllm: Layout instruction tun- ing with large language models for document understanding.

DRISHTIKON: Visual Grounding at Multiple Granularities in Documents Layoutllm: Layout instruction tun- ing with large language models for document understanding

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:34:40.548635Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T22:34:37.470543Z digest=sha256:bd8028b1e41a66c88df695f01739b277cc8dfaa2b22b01d6b8eed5d695a75a2e

Observation 9d8a63d0-4dfb-4ce0-9cfa-fa3d159a3e8e · outbound

This paper cites Manmatha, and C.

DRISHTIKON: Visual Grounding at Multiple Granularities in Documents Manmatha, and C

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:34:40.370596Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T22:34:37.514020Z digest=sha256:6c8fdeec1725d4cb554decee3ad67f1137c31bf32db748dfb7b41bcca683e464

Observation b8237044-855f-4bfa-9c99-a5fc9f649726 · outbound

This paper cites doctr: Document text recognition.

DRISHTIKON: Visual Grounding at Multiple Granularities in Documents doctr: Document text recognition

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:34:40.168201Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T22:34:37.572451Z digest=sha256:38646d34ed4d644e96ecfc5f0893d85a4cb606fd9104ceec26a09bef3c7599d6

Observation bd85ff38-1840-4c48-99a1-e8d27e393c66 · outbound

This paper cites Nagaraja, Vlad I.

DRISHTIKON: Visual Grounding at Multiple Granularities in Documents Nagaraja, Vlad I

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:34:39.997566Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T22:34:37.633362Z digest=sha256:09a76203fb4d8a5a7a14ca13d2caf18d6b81805cdc1f9adfbcbffb05cf5b5b4f

Observation 7293f1f5-efa8-4bdb-b161-831c66a9849a · outbound

This paper cites Surya: A lightweight document ocr and analysis toolkit.

DRISHTIKON: Visual Grounding at Multiple Granularities in Documents Surya: A lightweight document ocr and analysis toolkit

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:34:39.784297Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T22:34:37.715820Z digest=sha256:071db87be72921c89c98251a90d9c17b8a36c003bbf064d12e46764467d69238

Observation e3ed71f0-a2e2-4a52-9176-52848ac7601f · outbound

This paper cites Going Full-TILT Boogie on Document Understanding with Text-Image-Layout Transformer.

DRISHTIKON: Visual Grounding at Multiple Granularities in Documents Going Full-TILT Boogie on Document Understanding with Text-Image-Layout Transformer

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-06T22:34:37.794185Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:34:37.794185Z digest=sha256:4e411032afaec3d86c1e56b9c461a83d1a385c9758527cf27238ed72626300b5

Observation 03772cf5-7c6c-4ee2-879c-a242821c1315 · outbound

This paper cites Grounding of textual phrases in images by reconstruction.

DRISHTIKON: Visual Grounding at Multiple Granularities in Documents Grounding of textual phrases in images by reconstruction

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:34:39.637526Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T22:34:37.831093Z digest=sha256:78a2319554cbd2d1ef7dba8cb96accb247024c9e0088735b8ad155001da2f607

Observation 6297c7f0-0d49-47ab-ad74-2602afa5bd94 · outbound

This paper cites Unifying vision, text, and layout for universal document processing.

DRISHTIKON: Visual Grounding at Multiple Granularities in Documents Unifying vision, text, and layout for universal document processing

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:34:39.447752Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T22:34:37.916791Z digest=sha256:ca5421530d9ea6aa61f2c8e8860c117e02522411875db701fa9d7ad6bb775ff5

Observation 182a84dd-83fc-4dbc-8716-a71a0b4d09d8 · outbound

This paper cites Gpt-4 technical report.

DRISHTIKON: Visual Grounding at Multiple Granularities in Documents Gpt-4 technical report

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:34:39.249161Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T22:34:38.003951Z digest=sha256:ef400495d2e3c5a073803b84ca2f1b17b244eb1fa34590d231ad6981c55ffd45

Observation 887b9de5-1192-48e0-ac6c-e7577ae81940 · outbound

This paper cites Hierarchical multimodal transformers for Multi-Page DocVQA.

DRISHTIKON: Visual Grounding at Multiple Granularities in Documents Hierarchical multimodal transformers for Multi-Page DocVQA

Reference 19

Resolution
verified exact
local_arxiv, observed 2026-08-06T22:34:38.562815Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T22:34:38.065829Z digest=sha256:b36f19199c29f6ace69976857409fb145e4f0c3c1b3ac3fd0f8746dcb494f969

Observation 53c3c62c-21e8-45b2-908f-e25224daeeaa · outbound

This paper cites Docllm: A layout-aware gener- ative language model for multimodal document understand- ing.

DRISHTIKON: Visual Grounding at Multiple Granularities in Documents Docllm: A layout-aware gener- ative language model for multimodal document understand- ing

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:34:39.037272Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T22:34:38.137910Z digest=sha256:20eada3889d6962e236e82e1f2e6bbfaaf5e2487f22e6f88275ca173a1a2e1ad

Observation 69761ad4-aa54-4e5e-a872-fc183c7c1275 · outbound

This paper cites Qwen2-vl: Enhancing vision-language model’s perception of the world at any resolution, 2024.

DRISHTIKON: Visual Grounding at Multiple Granularities in Documents Qwen2-vl: Enhancing vision-language model’s perception of the world at any resolution, 2024

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:34:38.861845Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T22:34:38.184240Z digest=sha256:1b18f476eff0ad097e1098da8ef8bc8722aabfd3690753eab566c4ef18c3a029

Observation 1703ad8c-9734-4663-ade3-b85c758fef5a · outbound

This paper cites Towards Improving Document Understanding: An Exploration on Text-Grounding via MLLMs.

DRISHTIKON: Visual Grounding at Multiple Granularities in Documents Towards Improving Document Understanding: An Exploration on Text-Grounding via MLLMs

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-06T22:34:38.254340Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:34:38.254340Z digest=sha256:28ec31b52b46ba0d1b598bd0f10a8a6022dbe4406b450831f7ffd1f251bc9385

Observation 2e77bf65-a5ce-412f-873c-d7c3da2d502a · outbound

This paper cites Toward Visual Grounding: A Survey.

DRISHTIKON: Visual Grounding at Multiple Granularities in Documents Toward Visual Grounding: A Survey

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-06T22:34:38.310883Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:34:38.310883Z digest=sha256:8fb8ca5f7f5d742a451c3f2c70666e9b4f70b25a64c69b8cdf1271a14277e492

Observation 83781985-7053-40ae-a698-8e9887e53953 · outbound

This paper cites DOGR: Towards Versatile Visual Document Grounding and Referring.

DRISHTIKON: Visual Grounding at Multiple Granularities in Documents DOGR: Towards Versatile Visual Document Grounding and Referring

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-06T22:34:38.340134Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:34:38.340134Z digest=sha256:52a23943fdbde2772ac7ba6a6992f42c9a364ebfcf4d3c2434c9c15e0affcb87

Pith citing papers

No inbound Pith citation observations are available.