Pith. sign in

Paper Citation Record · LEDGER

Multi-Task GRPO: Reliable LLM Reasoning Across Tasks

As of 7 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 7 inbound Pith citation observations for arXiv:2602.05547.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2602.05547 v3

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 7 of 7 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 7 of 7 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-07-03T20:14:46.147637Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T18:30:01.596419Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 344a2ab8-f8f0-43e7-9408-7abc9bd20e44 · inbound

Target Policy Optimization cites this paper.

Target Policy Optimization Multi-Task GRPO: Reliable LLM Reasoning Across Tasks

Reference 12

Resolution
metadata mismatch
arxiv_id, observed 2026-07-28T03:23:25.766972Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T19:21:37.744591Z digest=sha256:4cc4c4288cf6cc5f8fbdac166acc85bdb3f6fa8c21e86ff76c57c94bf2b2883d

Observation ab6fb01a-aee1-4930-a62d-f254786773b2 · inbound

M2A: Synergizing Mathematical and Agentic Reasoning in Large Language Models cites this paper.

M2A: Synergizing Mathematical and Agentic Reasoning in Large Language Models Multi-Task GRPO: Reliable LLM Reasoning Across Tasks

Reference 28

Resolution
verified exact
arxiv_id, observed 2026-07-28T03:23:25.766972Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-12T04:43:50.384980Z digest=sha256:499a5f37167f30ef7a6b8573c911c05af0ddb323a25506c4e6e456f80abb27c9

Observation 0cbc1a8d-a12b-4309-af7d-e50be85a6bf9 · inbound

Harmony in Diversity: Multi-domain Contrastive Policy Optimization for Large Reasoning Models cites this paper.

Harmony in Diversity: Multi-domain Contrastive Policy Optimization for Large Reasoning Models Multi-Task GRPO: Reliable LLM Reasoning Across Tasks

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-07-28T03:23:25.766972Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-29T22:18:10.508034Z digest=sha256:839c97fe23b3ae67f978da6621f5a5cf46519927825fc14c09683747e25fc343

Observation 78273ebe-f991-4a52-ae1b-550e622d8d86 · inbound

CARE-RL: Capability-Aware Reinforcement Learning for Mitigating Cross-Domain Conflicts cites this paper.

CARE-RL: Capability-Aware Reinforcement Learning for Mitigating Cross-Domain Conflicts Multi-Task GRPO: Reliable LLM Reasoning Across Tasks

Reference 39

Resolution
metadata mismatch
arxiv_id, observed 2026-07-28T03:23:25.766972Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-06-28T19:01:14.340754Z digest=sha256:dd4448686a00717d267b3a201960cd21dffb7d5da1414348d7893a1fa1063f2a

Observation 0dfe8bba-30b5-4330-953a-8bdde47a9ed5 · inbound

Transferability for General Reasoning: An Automated Curriculum for Multi-Domain RLVR cites this paper.

Transferability for General Reasoning: An Automated Curriculum for Multi-Domain RLVR Multi-Task GRPO: Reliable LLM Reasoning Across Tasks

Reference 45

Resolution
verified exact
arxiv_id, observed 2026-07-28T03:23:25.766972Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-25T22:47:09.330723Z digest=sha256:782f4ead91d7560e9d3e533c0446d96b5e76e03e5222649c211d3a16fb8fa0a4

Observation d463cd78-c6e6-4721-ba5e-6c78df6acba7 · inbound

Transferability for General Reasoning: An Automated Curriculum for Multi-Domain RLVR cites this paper.

Transferability for General Reasoning: An Automated Curriculum for Multi-Domain RLVR Multi-Task GRPO: Reliable LLM Reasoning Across Tasks

Reference 45

Resolution
verified exact
arxiv_id, observed 2026-07-28T03:23:25.766972Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-30T09:55:19.804589Z digest=sha256:b57d17589ee8c79c8296f2584bdb0db2d82b5755af904e27cf69e8b995226aa0

Observation 9bf43f96-847a-42d8-80b3-2a1b7f22338b · inbound

World Feedback for Clinical Agents: Diagnosing RL in FHIR Environments cites this paper.

World Feedback for Clinical Agents: Diagnosing RL in FHIR Environments Multi-Task GRPO: Reliable LLM Reasoning Across Tasks

Reference 15

Resolution
metadata mismatch
arxiv_id, observed 2026-07-28T03:23:25.766972Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-07-03T20:14:46.147637Z digest=sha256:efa73c99e8f7b684857816e64d1f9b523cedb1d47ffc7079bfcd44aad17a7203