Pith. sign in

Paper Citation Record · LEDGER

When Prompts Become Pixels: Prompt-Region Grounding for Multimodal Reasoning

As of 7 August 2026, this Paper Citation Record lists 17 of 17 outbound references and 0 inbound Pith citation observations for arXiv:2608.04726.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2608.04726 v1

Coverage vector

measured 17 of 17 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T18:02:43.614624Z

measured 17 of 17 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

17 of 17 outbound references displayed

  • verified exact0
  • verified fuzzy1
  • unresolved16
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 7a482709-b5dd-4cb3-a056-9dfd3d727272 · outbound

This paper cites OCRBench v2: An Improved Benchmark for Evaluating Large Multimodal Models on Visual Text Localization and Reasoning.

When Prompts Become Pixels: Prompt-Region Grounding for Multimodal Reasoning OCRBench v2: An Improved Benchmark for Evaluating Large Multimodal Models on Visual Text Localization and Reasoning

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-06T18:02:43.554226Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:02:43.554226Z digest=sha256:9f3230ef8f85eb6950684fd3347e70a8199059cdd67692d4872014e54a5dc1c3

Observation db58b1e2-504e-4452-9dae-b83961accf95 · outbound

This paper cites Li, B.; Zhang, Y.; Guo, D.; Zhang, R.; Li, F.; Zhang, H.; Zhang, K.; Zhang, P.; Li, Y.; Liu, Z.; and Li, C.

When Prompts Become Pixels: Prompt-Region Grounding for Multimodal Reasoning Li, B.; Zhang, Y.; Guo, D.; Zhang, R.; Li, F.; Zhang, H.; Zhang, K.; Zhang, P.; Li, Y.; Liu, Z.; and Li, C

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-06T18:02:43.563621Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:02:43.563621Z digest=sha256:3d9c97ea1cf6f158c76ce4a90a4ab23e6ebd56439b5907cd20c743dc9bdbad7a

Observation 9133902b-5ab7-40d2-9911-fdede827cc13 · outbound

This paper cites Lin, H.; Liu, Z.; Zhu, Y.; Qin, C.; Lin, J.; Shang, X.; He, C.; Zhang, W.; and Wu, L.

When Prompts Become Pixels: Prompt-Region Grounding for Multimodal Reasoning Lin, H.; Liu, Z.; Zhu, Y.; Qin, C.; Lin, J.; Shang, X.; He, C.; Zhang, W.; and Wu, L

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-06T18:02:43.567634Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:02:43.567634Z digest=sha256:75b63967518f2d2c0fdff6045db660298401b979a2e9833198237d2105b3ea01

Observation 073a3ff1-8751-4f9b-9003-9930983f4ffa · outbound

This paper cites arXiv preprint arXiv:2601.21821.

When Prompts Become Pixels: Prompt-Region Grounding for Multimodal Reasoning arXiv preprint arXiv:2601.21821

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-06T18:02:43.572158Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:02:43.572158Z digest=sha256:67209b2fd1708cbd6e3a0ebccbfb8eb62bf8de387d77a096e45b036877743916

Observation 5082ee0b-8a44-460d-9ebe-fc329006bb18 · outbound

This paper cites VISTA-Bench: Do Vision-Language Models Really Understand Visualized Text as Well as Pure Text?.

When Prompts Become Pixels: Prompt-Region Grounding for Multimodal Reasoning VISTA-Bench: Do Vision-Language Models Really Understand Visualized Text as Well as Pure Text?

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-06T18:02:43.576081Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:02:43.576081Z digest=sha256:fad592596e4c6bbbd25b3b815499c70a7b72fa0b386d24c0901d04402945ce88

Observation e4571943-360e-4046-9080-7be6d136e523 · outbound

This paper cites We-Math 2.0: A Versatile MathBook System for Incentivizing Visual Mathematical Reasoning.

When Prompts Become Pixels: Prompt-Region Grounding for Multimodal Reasoning We-Math 2.0: A Versatile MathBook System for Incentivizing Visual Mathematical Reasoning

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-06T18:02:43.584703Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:02:43.584703Z digest=sha256:2561cb5408e924c4adb04a10886df008900d74e674cd154d4c3491c7b4060407

Observation 182f00bc-48d2-4ba2-9971-1b52146dab92 · outbound

This paper cites BMMR: A Large-Scale Bilingual Multimodal Multi-Discipline Reasoning Dataset.

When Prompts Become Pixels: Prompt-Region Grounding for Multimodal Reasoning BMMR: A Large-Scale Bilingual Multimodal Multi-Discipline Reasoning Dataset

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-06T18:02:43.598214Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:02:43.598214Z digest=sha256:c58d67315ef968bcc0492cac7ea8aaf9c4feae22459dd18ab0a03888c9119de4

Observation fcdee749-e364-4976-8046-a77cfebff434 · outbound

This paper cites Zheng, C.; Liu, S.; Li, M.; Chen, X.-H.; Yu, B.; Gao, C.; Dang, K.; Liu, Y.; Men, R.; Yang, A.; Zhou, J.; and Lin, J.

When Prompts Become Pixels: Prompt-Region Grounding for Multimodal Reasoning Zheng, C.; Liu, S.; Li, M.; Chen, X.-H.; Yu, B.; Gao, C.; Dang, K.; Liu, Y.; Men, R.; Yang, A.; Zhou, J.; and Lin, J

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-06T18:02:43.602554Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:02:43.602554Z digest=sha256:7fb5be32bb2d1c184d6fb429d69cb8f77cb336bb7e681073b01a49959b179033

Observation 62e59df4-17db-4dd9-b4fb-0c5a31f7e731 · outbound

This paper cites Group Sequence Policy Optimization.

When Prompts Become Pixels: Prompt-Region Grounding for Multimodal Reasoning Group Sequence Policy Optimization

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-06T18:02:43.606491Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:02:43.606491Z digest=sha256:2707a0a029d2c47e33200116684b3da4b6e2512dd5b59571b6f5e2592b6b351d

Observation 393aa8f8-c4fc-49da-9c84-44ecef1a648a · outbound

This paper cites InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models.

When Prompts Become Pixels: Prompt-Region Grounding for Multimodal Reasoning InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-06T18:02:43.610661Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:02:43.610661Z digest=sha256:03ddc5d096e053f3fa7f5c870317a4fae78dfd84a7ceaace3c34c803166699a3

Observation fa9de672-90ac-483d-bee0-3e604dfe204a · outbound

This paper cites InInternational Conference on Learning Representations.

When Prompts Become Pixels: Prompt-Region Grounding for Multimodal Reasoning InInternational Conference on Learning Representations

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:02:44.323648Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T18:02:43.614624Z digest=sha256:cf840d4b5b50340a54eeadf7b2ba54364fa27178fb490a6750af2ccf2492f593

Observation 1780b0e9-a7ba-4a10-923c-93daf34ea924 · outbound

This paper cites Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution.

When Prompts Become Pixels: Prompt-Region Grounding for Multimodal Reasoning Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-06T18:02:43.588950Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:02:43.588950Z digest=sha256:adb52415b5b138b099a0f363ef187d0e46fd1a036653bab1c4b1f37a4554c4bf

Observation f4a02599-d57b-4a32-83b1-4d2d1094bfc0 · outbound

This paper cites Kimi-VL Technical Report.

When Prompts Become Pixels: Prompt-Region Grounding for Multimodal Reasoning Kimi-VL Technical Report

Reference 2020

Resolution
unresolved
no resolver link, observed 2026-08-06T18:02:43.558955Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:02:43.558955Z digest=sha256:da77461eac9a923a44f67cf391dc5c1ba7fd5ba28e0ffea9af1318476e2e7c94

Observation 7f726666-04ca-4bbe-834b-124876c1cfb1 · outbound

This paper cites MM-Eureka: Exploring the Frontiers of Multimodal Reasoning with Rule-based Reinforcement Learning.

When Prompts Become Pixels: Prompt-Region Grounding for Multimodal Reasoning MM-Eureka: Exploring the Frontiers of Multimodal Reasoning with Rule-based Reinforcement Learning

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-06T18:02:43.580389Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:02:43.580389Z digest=sha256:369f936758818e90cadd0214ca0033eb55d85ddbb252d9995bcc7616c4eb3e4c

Observation 7b70000d-798d-4848-b40c-80934cbca12a · outbound

This paper cites DeepSeek-VL2: Mixture-of-Experts Vision-Language Models for Advanced Multimodal Understanding.

When Prompts Become Pixels: Prompt-Region Grounding for Multimodal Reasoning DeepSeek-VL2: Mixture-of-Experts Vision-Language Models for Advanced Multimodal Understanding

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-06T18:02:43.593744Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:02:43.593744Z digest=sha256:8fe647cd17dba93e5a7bb98acfa9b7cbe14542caa59a812756c1680cc24a68eb

Observation 43834796-eb56-4b6b-85e4-d0ba82abc155 · outbound

This paper cites Assran,M.;Duval,Q.;Misra,I.;Bojanowski,P.;Vincent,P.; Rabbat,M.;LeCun,Y.;andBallas,N.2023.

When Prompts Become Pixels: Prompt-Region Grounding for Multimodal Reasoning Assran,M.;Duval,Q.;Misra,I.;Bojanowski,P.;Vincent,P.; Rabbat,M.;LeCun,Y.;andBallas,N.2023

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-06T18:02:43.545297Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:02:43.545297Z digest=sha256:d239e582b909e996f2c23a663f5ca959cf63b1d66cf89610c8addd4727a1e805

Observation 7e1ffec3-0e69-4bff-89f1-4d1b2d5f4022 · outbound

This paper cites arXiv preprint arXiv:2602.09483.

When Prompts Become Pixels: Prompt-Region Grounding for Multimodal Reasoning arXiv preprint arXiv:2602.09483

Reference 2026

Resolution
unresolved
no resolver link, observed 2026-08-06T18:02:43.549873Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:02:43.549873Z digest=sha256:a0136e1980f31dccc8f94c9d8514bb409c86cd9ffe113d104eb0214e65dcbc74

Pith citing papers

No inbound Pith citation observations are available.