Pith. sign in

Paper Citation Record · LEDGER

When Prompts Become Pixels: Prompt-Region Grounding for Multimodal Reasoning

As of 22 August 2026, this Paper Citation Record lists 17 of 17 outbound references and 0 inbound Pith citation observations for arXiv:2608.04726.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2608.04726 v1

Coverage vector

measured 17 of 17 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T18:02:43.614624Z

measured 17 of 17 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-22T06:32:14.747728+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

17 of 17 outbound references displayed

  • verified exact0
  • verified fuzzy1
  • unresolved16
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 7a482709-b5dd-4cb3-a056-9dfd3d727272 · outbound

This paper cites OCRBench v2: An Improved Benchmark for Evaluating Large Multimodal Models on Visual Text Localization and Reasoning.

When Prompts Become Pixels: Prompt-Region Grounding for Multimodal Reasoning OCRBench v2: An Improved Benchmark for Evaluating Large Multimodal Models on Visual Text Localization and Reasoning

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-06T18:02:43.554226Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:02:43.554226Z digest=sha256:360ad8278421823b518da524c7bb4ae21b5af5409a3233f54cedc6c2735b5f8d

Observation db58b1e2-504e-4452-9dae-b83961accf95 · outbound

This paper cites Li, B.; Zhang, Y.; Guo, D.; Zhang, R.; Li, F.; Zhang, H.; Zhang, K.; Zhang, P.; Li, Y.; Liu, Z.; and Li, C.

When Prompts Become Pixels: Prompt-Region Grounding for Multimodal Reasoning Li, B.; Zhang, Y.; Guo, D.; Zhang, R.; Li, F.; Zhang, H.; Zhang, K.; Zhang, P.; Li, Y.; Liu, Z.; and Li, C

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-06T18:02:43.563621Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:02:43.563621Z digest=sha256:b8a5b49ec8675f4aad352028a1dc8e58f6fba674cbb7e6b114d18c5c4ca62bc1

Observation 9133902b-5ab7-40d2-9911-fdede827cc13 · outbound

This paper cites Lin, H.; Liu, Z.; Zhu, Y.; Qin, C.; Lin, J.; Shang, X.; He, C.; Zhang, W.; and Wu, L.

When Prompts Become Pixels: Prompt-Region Grounding for Multimodal Reasoning Lin, H.; Liu, Z.; Zhu, Y.; Qin, C.; Lin, J.; Shang, X.; He, C.; Zhang, W.; and Wu, L

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-06T18:02:43.567634Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:02:43.567634Z digest=sha256:5cd60f3db17f642a3362dd1b7b9c73b3601baafd938d50a6f6f04a60448efab5

Observation 073a3ff1-8751-4f9b-9003-9930983f4ffa · outbound

This paper cites arXiv preprint arXiv:2601.21821.

When Prompts Become Pixels: Prompt-Region Grounding for Multimodal Reasoning arXiv preprint arXiv:2601.21821

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-06T18:02:43.572158Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:02:43.572158Z digest=sha256:9c9b2e68e03baea5a3024934878af7690d117e711aefae2d164569a546fb601b

Observation 5082ee0b-8a44-460d-9ebe-fc329006bb18 · outbound

This paper cites VISTA-Bench: Do Vision-Language Models Really Understand Visualized Text as Well as Pure Text?.

When Prompts Become Pixels: Prompt-Region Grounding for Multimodal Reasoning VISTA-Bench: Do Vision-Language Models Really Understand Visualized Text as Well as Pure Text?

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-06T18:02:43.576081Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:02:43.576081Z digest=sha256:e02aae7476ab8b15d743bcf366502a12191cba4d2be2eb7384ca9cd6671057bc

Observation e4571943-360e-4046-9080-7be6d136e523 · outbound

This paper cites We-Math 2.0: A Versatile MathBook System for Incentivizing Visual Mathematical Reasoning.

When Prompts Become Pixels: Prompt-Region Grounding for Multimodal Reasoning We-Math 2.0: A Versatile MathBook System for Incentivizing Visual Mathematical Reasoning

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-06T18:02:43.584703Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:02:43.584703Z digest=sha256:379c2989f67c94955f3ca4f50a948504277e7a23b3432005412df4174a491fea

Observation 182f00bc-48d2-4ba2-9971-1b52146dab92 · outbound

This paper cites BMMR: A Large-Scale Bilingual Multimodal Multi-Discipline Reasoning Dataset.

When Prompts Become Pixels: Prompt-Region Grounding for Multimodal Reasoning BMMR: A Large-Scale Bilingual Multimodal Multi-Discipline Reasoning Dataset

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-06T18:02:43.598214Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:02:43.598214Z digest=sha256:b121dae99609b09fc4f98c9abb36631df08d13d5704b76122a7c0cff6c441458

Observation fcdee749-e364-4976-8046-a77cfebff434 · outbound

This paper cites Zheng, C.; Liu, S.; Li, M.; Chen, X.-H.; Yu, B.; Gao, C.; Dang, K.; Liu, Y.; Men, R.; Yang, A.; Zhou, J.; and Lin, J.

When Prompts Become Pixels: Prompt-Region Grounding for Multimodal Reasoning Zheng, C.; Liu, S.; Li, M.; Chen, X.-H.; Yu, B.; Gao, C.; Dang, K.; Liu, Y.; Men, R.; Yang, A.; Zhou, J.; and Lin, J

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-06T18:02:43.602554Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:02:43.602554Z digest=sha256:ff28cf93343b33b1dc4c6cefae119622fb5789ca6a483236d6116f62b7d12b4e

Observation 62e59df4-17db-4dd9-b4fb-0c5a31f7e731 · outbound

This paper cites Group Sequence Policy Optimization.

When Prompts Become Pixels: Prompt-Region Grounding for Multimodal Reasoning Group Sequence Policy Optimization

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-06T18:02:43.606491Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:02:43.606491Z digest=sha256:ec739111214c7510dc16c52c2709c791809d394a0168a66b3f6347bcbd5f35b5

Observation 393aa8f8-c4fc-49da-9c84-44ecef1a648a · outbound

This paper cites InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models.

When Prompts Become Pixels: Prompt-Region Grounding for Multimodal Reasoning InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-06T18:02:43.610661Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:02:43.610661Z digest=sha256:782d750d2d6df2a21e72f64861daa13f60579727c9e04a0a2159099acc81c5cc

Observation fa9de672-90ac-483d-bee0-3e604dfe204a · outbound

This paper cites InInternational Conference on Learning Representations.

When Prompts Become Pixels: Prompt-Region Grounding for Multimodal Reasoning InInternational Conference on Learning Representations

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:02:44.323648Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-06T18:02:43.614624Z digest=sha256:35ef1a59daf7d09fa4112b07112e4dde327d50f90f0870504e3212b0dfb9defa

Observation 1780b0e9-a7ba-4a10-923c-93daf34ea924 · outbound

This paper cites Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution.

When Prompts Become Pixels: Prompt-Region Grounding for Multimodal Reasoning Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-06T18:02:43.588950Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:02:43.588950Z digest=sha256:79be0783fb320eea277b52c92b21bb4dda6e7617bcb0ee6522123f4f2fbc7266

Observation f4a02599-d57b-4a32-83b1-4d2d1094bfc0 · outbound

This paper cites Kimi-VL Technical Report.

When Prompts Become Pixels: Prompt-Region Grounding for Multimodal Reasoning Kimi-VL Technical Report

Reference 2020

Resolution
unresolved
no resolver link, observed 2026-08-06T18:02:43.558955Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:02:43.558955Z digest=sha256:e75818cf0a605fd144a7846c1f8c826e7ed4a187a6206599a6201926d427dea5

Observation 7f726666-04ca-4bbe-834b-124876c1cfb1 · outbound

This paper cites MM-Eureka: Exploring the Frontiers of Multimodal Reasoning with Rule-based Reinforcement Learning.

When Prompts Become Pixels: Prompt-Region Grounding for Multimodal Reasoning MM-Eureka: Exploring the Frontiers of Multimodal Reasoning with Rule-based Reinforcement Learning

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-06T18:02:43.580389Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:02:43.580389Z digest=sha256:5bd86cbdc86339356067214b912dc033f78afa98987a7e26cadb8191b83846da

Observation 7b70000d-798d-4848-b40c-80934cbca12a · outbound

This paper cites DeepSeek-VL2: Mixture-of-Experts Vision-Language Models for Advanced Multimodal Understanding.

When Prompts Become Pixels: Prompt-Region Grounding for Multimodal Reasoning DeepSeek-VL2: Mixture-of-Experts Vision-Language Models for Advanced Multimodal Understanding

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-06T18:02:43.593744Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:02:43.593744Z digest=sha256:a447caa22db7d4ba1af1960d4fc51534b3defaadecf80862a56d48e1e1c9bf26

Observation 43834796-eb56-4b6b-85e4-d0ba82abc155 · outbound

This paper cites Assran,M.;Duval,Q.;Misra,I.;Bojanowski,P.;Vincent,P.; Rabbat,M.;LeCun,Y.;andBallas,N.2023.

When Prompts Become Pixels: Prompt-Region Grounding for Multimodal Reasoning Assran,M.;Duval,Q.;Misra,I.;Bojanowski,P.;Vincent,P.; Rabbat,M.;LeCun,Y.;andBallas,N.2023

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-06T18:02:43.545297Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:02:43.545297Z digest=sha256:06ffea47b62ec6e5121529503c60f5719c456edbff450c5136a25dcee192bcf2

Observation 7e1ffec3-0e69-4bff-89f1-4d1b2d5f4022 · outbound

This paper cites arXiv preprint arXiv:2602.09483.

When Prompts Become Pixels: Prompt-Region Grounding for Multimodal Reasoning arXiv preprint arXiv:2602.09483

Reference 2026

Resolution
unresolved
no resolver link, observed 2026-08-06T18:02:43.549873Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:02:43.549873Z digest=sha256:f84d35a52188dba29c97dc89f2412b5b69a8d69137037fd7b8b4546179f9673d

Pith citing papers

No inbound Pith citation observations are available.