Pith. sign in

Paper Citation Record · LEDGER

When Prompts Become Pixels: Prompt-Region Grounding for Multimodal Reasoning

As of 7 August 2026, this Paper Citation Record lists 17 of 17 outbound references and 0 inbound Pith citation observations for arXiv:2608.04726.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2608.04726 v1

Coverage vector

measured 17 of 17 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T18:02:43.614624Z

measured 17 of 17 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

17 of 17 outbound references displayed

  • verified exact0
  • verified fuzzy1
  • unresolved16
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 7a482709-b5dd-4cb3-a056-9dfd3d727272 · outbound

This paper cites OCRBench v2: An Improved Benchmark for Evaluating Large Multimodal Models on Visual Text Localization and Reasoning.

When Prompts Become Pixels: Prompt-Region Grounding for Multimodal Reasoning OCRBench v2: An Improved Benchmark for Evaluating Large Multimodal Models on Visual Text Localization and Reasoning

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-06T18:02:43.554226Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:02:43.554226Z digest=sha256:c898c541aa6be837907d4f2578746cd405fbd7dc59e75d996c42a426920e3d00

Observation db58b1e2-504e-4452-9dae-b83961accf95 · outbound

This paper cites Li, B.; Zhang, Y.; Guo, D.; Zhang, R.; Li, F.; Zhang, H.; Zhang, K.; Zhang, P.; Li, Y.; Liu, Z.; and Li, C.

When Prompts Become Pixels: Prompt-Region Grounding for Multimodal Reasoning Li, B.; Zhang, Y.; Guo, D.; Zhang, R.; Li, F.; Zhang, H.; Zhang, K.; Zhang, P.; Li, Y.; Liu, Z.; and Li, C

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-06T18:02:43.563621Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:02:43.563621Z digest=sha256:6f1daee54f96e33cf3bdbfbd9b9d98c603dce3150aeb312d63ae7b2efe12fdbe

Observation 9133902b-5ab7-40d2-9911-fdede827cc13 · outbound

This paper cites Lin, H.; Liu, Z.; Zhu, Y.; Qin, C.; Lin, J.; Shang, X.; He, C.; Zhang, W.; and Wu, L.

When Prompts Become Pixels: Prompt-Region Grounding for Multimodal Reasoning Lin, H.; Liu, Z.; Zhu, Y.; Qin, C.; Lin, J.; Shang, X.; He, C.; Zhang, W.; and Wu, L

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-06T18:02:43.567634Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:02:43.567634Z digest=sha256:531a7fffbbb1c8a37afe3c759f6720d6bd6e6327697caa5b035d5757a647ea3f

Observation 073a3ff1-8751-4f9b-9003-9930983f4ffa · outbound

This paper cites arXiv preprint arXiv:2601.21821.

When Prompts Become Pixels: Prompt-Region Grounding for Multimodal Reasoning arXiv preprint arXiv:2601.21821

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-06T18:02:43.572158Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:02:43.572158Z digest=sha256:049716f02bcbf8a1011b1d334c715100d985b1b0ef94cfd2bc433d39bf61fd76

Observation 5082ee0b-8a44-460d-9ebe-fc329006bb18 · outbound

This paper cites VISTA-Bench: Do Vision-Language Models Really Understand Visualized Text as Well as Pure Text?.

When Prompts Become Pixels: Prompt-Region Grounding for Multimodal Reasoning VISTA-Bench: Do Vision-Language Models Really Understand Visualized Text as Well as Pure Text?

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-06T18:02:43.576081Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:02:43.576081Z digest=sha256:fafaa12b73afeabc07d6efd0273dcc5533e67b61be7ebbe21b7a6fd1825ce1b2

Observation e4571943-360e-4046-9080-7be6d136e523 · outbound

This paper cites We-Math 2.0: A Versatile MathBook System for Incentivizing Visual Mathematical Reasoning.

When Prompts Become Pixels: Prompt-Region Grounding for Multimodal Reasoning We-Math 2.0: A Versatile MathBook System for Incentivizing Visual Mathematical Reasoning

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-06T18:02:43.584703Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:02:43.584703Z digest=sha256:8768413e893d2079eebe03c07e3e7f8a6fb148bf197a77b1f285cb1d86ed88a0

Observation 182f00bc-48d2-4ba2-9971-1b52146dab92 · outbound

This paper cites BMMR: A Large-Scale Bilingual Multimodal Multi-Discipline Reasoning Dataset.

When Prompts Become Pixels: Prompt-Region Grounding for Multimodal Reasoning BMMR: A Large-Scale Bilingual Multimodal Multi-Discipline Reasoning Dataset

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-06T18:02:43.598214Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:02:43.598214Z digest=sha256:71b89315a72b90f2bcdd845f49caf732589281bb34225c04affd2c5e5e27e58d

Observation fcdee749-e364-4976-8046-a77cfebff434 · outbound

This paper cites Zheng, C.; Liu, S.; Li, M.; Chen, X.-H.; Yu, B.; Gao, C.; Dang, K.; Liu, Y.; Men, R.; Yang, A.; Zhou, J.; and Lin, J.

When Prompts Become Pixels: Prompt-Region Grounding for Multimodal Reasoning Zheng, C.; Liu, S.; Li, M.; Chen, X.-H.; Yu, B.; Gao, C.; Dang, K.; Liu, Y.; Men, R.; Yang, A.; Zhou, J.; and Lin, J

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-06T18:02:43.602554Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:02:43.602554Z digest=sha256:85408a7d60e124ee310b62197a375ee6c0bdcb2359b068055723ebfd1e837360

Observation 62e59df4-17db-4dd9-b4fb-0c5a31f7e731 · outbound

This paper cites Group Sequence Policy Optimization.

When Prompts Become Pixels: Prompt-Region Grounding for Multimodal Reasoning Group Sequence Policy Optimization

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-06T18:02:43.606491Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:02:43.606491Z digest=sha256:c18931dda9933992724ac9cc20417b0376578a86bee361aeba2aeda44d59bcb1

Observation 393aa8f8-c4fc-49da-9c84-44ecef1a648a · outbound

This paper cites InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models.

When Prompts Become Pixels: Prompt-Region Grounding for Multimodal Reasoning InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-06T18:02:43.610661Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:02:43.610661Z digest=sha256:db8ad9d3484a7260e09ba4d9c3e469b57a16b7b82e6996daf44069254f03041f

Observation fa9de672-90ac-483d-bee0-3e604dfe204a · outbound

This paper cites InInternational Conference on Learning Representations.

When Prompts Become Pixels: Prompt-Region Grounding for Multimodal Reasoning InInternational Conference on Learning Representations

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:02:44.323648Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T18:02:43.614624Z digest=sha256:39d68b97cb5d8c61f9382c63d98f2046056cb0cea1f6e88ab52825cc9d64a12d

Observation 1780b0e9-a7ba-4a10-923c-93daf34ea924 · outbound

This paper cites Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution.

When Prompts Become Pixels: Prompt-Region Grounding for Multimodal Reasoning Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-06T18:02:43.588950Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:02:43.588950Z digest=sha256:576c764df0044ef8a3a89ee35127e93d2e6ed1aa044fbb0abacd121782269d11

Observation f4a02599-d57b-4a32-83b1-4d2d1094bfc0 · outbound

This paper cites Kimi-VL Technical Report.

When Prompts Become Pixels: Prompt-Region Grounding for Multimodal Reasoning Kimi-VL Technical Report

Reference 2020

Resolution
unresolved
no resolver link, observed 2026-08-06T18:02:43.558955Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:02:43.558955Z digest=sha256:e8231768d8c19de749d93ad48c675a060339f1759964cb5cb0029ac2aa987d3f

Observation 7f726666-04ca-4bbe-834b-124876c1cfb1 · outbound

This paper cites MM-Eureka: Exploring the Frontiers of Multimodal Reasoning with Rule-based Reinforcement Learning.

When Prompts Become Pixels: Prompt-Region Grounding for Multimodal Reasoning MM-Eureka: Exploring the Frontiers of Multimodal Reasoning with Rule-based Reinforcement Learning

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-06T18:02:43.580389Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:02:43.580389Z digest=sha256:f7f0864f4adceaae89984a45747593ffefaae8382f593cd4b41ad463bdff24e0

Observation 7b70000d-798d-4848-b40c-80934cbca12a · outbound

This paper cites DeepSeek-VL2: Mixture-of-Experts Vision-Language Models for Advanced Multimodal Understanding.

When Prompts Become Pixels: Prompt-Region Grounding for Multimodal Reasoning DeepSeek-VL2: Mixture-of-Experts Vision-Language Models for Advanced Multimodal Understanding

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-06T18:02:43.593744Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:02:43.593744Z digest=sha256:2ae7dca2da73a4d111e83c76978546de626873016b58dd6dec3e284c33a1566b

Observation 43834796-eb56-4b6b-85e4-d0ba82abc155 · outbound

This paper cites Assran,M.;Duval,Q.;Misra,I.;Bojanowski,P.;Vincent,P.; Rabbat,M.;LeCun,Y.;andBallas,N.2023.

When Prompts Become Pixels: Prompt-Region Grounding for Multimodal Reasoning Assran,M.;Duval,Q.;Misra,I.;Bojanowski,P.;Vincent,P.; Rabbat,M.;LeCun,Y.;andBallas,N.2023

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-06T18:02:43.545297Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:02:43.545297Z digest=sha256:bdaffa97131ed57798ad5d28da69ac10561c2a0fe99fe00191c9429bdc5ede78

Observation 7e1ffec3-0e69-4bff-89f1-4d1b2d5f4022 · outbound

This paper cites arXiv preprint arXiv:2602.09483.

When Prompts Become Pixels: Prompt-Region Grounding for Multimodal Reasoning arXiv preprint arXiv:2602.09483

Reference 2026

Resolution
unresolved
no resolver link, observed 2026-08-06T18:02:43.549873Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:02:43.549873Z digest=sha256:7c2eee7e4e4cc062aaa006e32121ccf519d1cbdd41f6c6dfdbb70940d96b67de

Pith citing papers

No inbound Pith citation observations are available.