Pith. sign in

Paper Citation Record · LEDGER

Boosting MLLM Reasoning with Text-Debiased Hint-GRPO

As of 9 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 9 inbound Pith citation observations for arXiv:2503.23905.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2503.23905 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 9 of 9 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 9 of 9 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T14:49:03.494449Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-07-10T05:16:48.098751Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation efb91d72-2bd2-4ef7-9da4-e9e7d1f49cac · inbound

RePrompt: Reasoning-Augmented Reprompting for Text-to-Image Generation via Reinforcement Learning cites this paper.

RePrompt: Reasoning-Augmented Reprompting for Text-to-Image Generation via Reinforcement Learning Boosting MLLM Reasoning with Text-Debiased Hint-GRPO

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T14:49:03.494449Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:49:03.494449Z digest=sha256:1224e0fab1461f83674485bc916f52a49fad64ffdefa9fd0c35eefeba384c81f

Observation bdc4a137-b272-4730-8576-cbc00ffd0c81 · inbound

Unveiling the Compositional Ability Gap in Vision-Language Reasoning Model cites this paper.

Unveiling the Compositional Ability Gap in Vision-Language Reasoning Model Boosting MLLM Reasoning with Text-Debiased Hint-GRPO

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T14:19:02.964530Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:19:02.964530Z digest=sha256:3189070a6f475d09efb1d9b72ba24401d1b6d4333dfa15367db9a0864f911a65

Observation 3966b4c6-aee4-4b0f-a038-e38f848f33ca · inbound

Reinforcement Learning Meets Large Language Models: A Survey of Advancements and Applications Across the LLM Lifecycle cites this paper.

Reinforcement Learning Meets Large Language Models: A Survey of Advancements and Applications Across the LLM Lifecycle Boosting MLLM Reasoning with Text-Debiased Hint-GRPO

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-04T16:07:30.943890Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T16:07:30.943890Z digest=sha256:609a58d2d4ebd0a4ad4d603d17ed9c73a22c58ef29c0e882d64bd1051fc97668

Observation 0df3c921-6629-48d7-8754-f026660e3197 · inbound

No More Stale Feedback: Co-Evolving Critics for Open-World Agent Learning cites this paper.

No More Stale Feedback: Co-Evolving Critics for Open-World Agent Learning Boosting MLLM Reasoning with Text-Debiased Hint-GRPO

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-05-16T16:03:04.283411Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-16T16:01:48.789986Z digest=sha256:45ec8aa38736e761776c92a9967876e40954d19d52be0d8c8881915f447ab424

Observation 88669c8e-ff9e-460f-a9bc-475363cf071b · inbound

EMO-R3: Reflective Reinforcement Learning for Emotional Reasoning in Multimodal Large Language Models cites this paper.

EMO-R3: Reflective Reinforcement Learning for Emotional Reasoning in Multimodal Large Language Models Boosting MLLM Reasoning with Text-Debiased Hint-GRPO

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-02T20:14:03.983641Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T20:14:03.983641Z digest=sha256:17ae6c4bff1608e97c3de6415b17487e72996b283e5132c1dd34c58da4f48e0e

Observation 290ad136-657d-4a53-8a26-16787c95732e · inbound

DiagramNet: An End-to-End Recognition Framework and Dataset for Non-Standard System-Level Diagrams cites this paper.

DiagramNet: An End-to-End Recognition Framework and Dataset for Non-Standard System-Level Diagrams Boosting MLLM Reasoning with Text-Debiased Hint-GRPO

Reference 8

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T16:51:09.196344Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-09T14:37:35.517734Z digest=sha256:1bd81b84c33a2f80ad133dcff2c91d5d80aece0d492e3f0a86aeeaafcf3de30b

Observation 4fb533e3-d46d-4869-bc49-8d71a1180f5b · inbound

MHPR: Multidimensional Human Perception and Reasoning Benchmark for Large Vision-Languate Models cites this paper.

MHPR: Multidimensional Human Perception and Reasoning Benchmark for Large Vision-Languate Models Boosting MLLM Reasoning with Text-Debiased Hint-GRPO

Reference 4

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T11:01:30.642654Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-08T01:20:54.441367Z digest=sha256:e26e82e9628451eb05363e5b709ef89989063fe803e4ac96d79099f8dc7ed81d

Observation cb77a87f-83cd-4b53-ac29-ce2a0c3b81e7 · inbound

Max Out GRPO Signal: Adaptive Trace Prefix Control for Hard Reasoning Problems cites this paper.

Max Out GRPO Signal: Adaptive Trace Prefix Control for Hard Reasoning Problems Boosting MLLM Reasoning with Text-Debiased Hint-GRPO

Reference 6

Resolution
verified exact
local_arxiv, observed 2026-07-09T03:25:57.810399Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-07-09T03:20:40.229376Z digest=sha256:e13f131f6964e6a10d39fa39996cbce2f7522d36b620ce4e119d4b4e30a22418

Observation acad75f2-78db-4485-91aa-9a6d2cca548a · inbound

Switch-Reasoner: Learn When to Think in Multitask Mixtures via Reinforcement Learning cites this paper.

Switch-Reasoner: Learn When to Think in Multitask Mixtures via Reinforcement Learning Boosting MLLM Reasoning with Text-Debiased Hint-GRPO

Reference 6

Resolution
metadata mismatch
local_arxiv, observed 2026-07-10T05:16:48.100099Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-07-10T05:07:40.106099Z digest=sha256:e29c1ac77efbaf87f33ee84b2a7ce3ea4efdde934307a503608cafd5e49cdc42