Pith. sign in

Paper Citation Record · LEDGER

Jointly Reinforcing Diversity and Quality in Language Model Generations

As of 8 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 28 inbound Pith citation observations for arXiv:2509.02534.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2509.02534 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 28 of 28 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 28 of 28 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-08T00:52:47.153167Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T21:00:08.446955Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 2c1a1a68-f682-4f72-ab3e-c7d1d3067f65 · inbound

Outcome-based Exploration for LLM Reasoning cites this paper.

Outcome-based Exploration for LLM Reasoning Jointly Reinforcing Diversity and Quality in Language Model Generations

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-04T22:59:14.516795Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T22:59:14.516795Z digest=sha256:6dc8fe4dd95324ff7c8280a598258a677383ee64aa70e979f23516246cf3bfdf

Observation 8c9d354d-9438-4e69-ac92-8514bfa5a166 · inbound

A Survey of Reinforcement Learning for Large Reasoning Models cites this paper.

A Survey of Reinforcement Learning for Large Reasoning Models Jointly Reinforcing Diversity and Quality in Language Model Generations

Reference 282

Resolution
verified exact
arxiv_id, observed 2026-05-18T00:02:24.741526Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-05-18T00:02:24.352947Z digest=sha256:398fe3cfdc354a669501ea8fd95cd8da18d0f0481f7ea31e3c28d09fb1a425c9

Observation 6ec81ccd-1a6d-4a95-94e9-cdf8cca9bd45 · inbound

Task-Dependent Evaluation of LLM Output Homogenization: A Taxonomy-Guided Framework cites this paper.

Task-Dependent Evaluation of LLM Output Homogenization: A Taxonomy-Guided Framework Jointly Reinforcing Diversity and Quality in Language Model Generations

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-05-18T13:36:25.011374Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-05-18T13:34:22.790199Z digest=sha256:dddc2f96b8354b0a522aee4ba0f3880228423bbaca2ef3e5dfcdf536402c2db7

Observation d0ae5256-a762-42b8-908d-a7886cedb03b · inbound

Polychromic Objectives for Reinforcement Learning cites this paper.

Polychromic Objectives for Reinforcement Learning Jointly Reinforcing Diversity and Quality in Language Model Generations

Reference 20

Resolution
verified exact
arxiv_id, observed 2026-05-18T11:56:20.020157Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-18T11:54:29.955833Z digest=sha256:ed4c47f3dc2d24de245042e019f6e640a35cf4435c734541a35bc30411d1c709

Observation 65341c38-834d-40f2-af35-ac56ad9d22fb · inbound

Representation-Based Exploration for Language Models: From Test-Time to Post-Training cites this paper.

Representation-Based Exploration for Language Models: From Test-Time to Post-Training Jointly Reinforcing Diversity and Quality in Language Model Generations

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-04T10:09:02.784967Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T10:09:02.784967Z digest=sha256:8b45a275d881a6a8ac80f87b13541c5dc851b07b688c9d38ac6fafac6bccdcf3

Observation b6666164-4599-4ab3-9a89-f967167c1b11 · inbound

Optimizing Diversity and Quality through Base-Aligned Model Collaboration cites this paper.

Optimizing Diversity and Quality through Base-Aligned Model Collaboration Jointly Reinforcing Diversity and Quality in Language Model Generations

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-03T23:32:54.723467Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T23:32:54.723467Z digest=sha256:d6e9d80eeb35edcad99391393c11c901c686e9add17e06176bb01fe77fcffc6a

Observation 896ca766-b1b1-4f13-9289-8c9efdbe432b · inbound

VOYAGER: A Training Free Approach for Generating Diverse Datasets using LLMs cites this paper.

VOYAGER: A Training Free Approach for Generating Diverse Datasets using LLMs Jointly Reinforcing Diversity and Quality in Language Model Generations

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-05-16T22:31:19.138233Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-16T22:30:06.582263Z digest=sha256:a25e302554b216603c6850a341fb9197fbcddec78db2da01f8c472efe86fa8c8

Observation da651ce3-f68d-40fd-9ddc-7c3d5f1cd224 · inbound

Vocabulary Dropout for Curriculum Diversity in LLM Co-Evolution cites this paper.

Vocabulary Dropout for Curriculum Diversity in LLM Co-Evolution Jointly Reinforcing Diversity and Quality in Language Model Generations

Reference 20

Resolution
verified exact
arxiv_id, observed 2026-05-13T19:28:10.094759Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-13T19:23:10.901441Z digest=sha256:64aec11b107dfef93586c6300c0c64cf63ce31ab8d432b3e8f03b9209664f056

Observation 28e0f42f-f2e9-4dcf-b69b-042d76811b7b · inbound

Vocabulary Dropout for Curriculum Diversity in LLM Co-Evolution cites this paper.

Vocabulary Dropout for Curriculum Diversity in LLM Co-Evolution Jointly Reinforcing Diversity and Quality in Language Model Generations

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-02T16:52:27.782802Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T16:52:27.782802Z digest=sha256:87be89f5f5b863f6ab5f18be9b3de2fae57f9ea8238bb440a7f5a4112a3ec2a1

Observation d85fd09c-20f7-485f-9cac-9bb1d58d1c65 · inbound

Too Correct to Learn: Reinforcement Learning on Saturated Reasoning Data cites this paper.

Too Correct to Learn: Reinforcement Learning on Saturated Reasoning Data Jointly Reinforcing Diversity and Quality in Language Model Generations

Reference 39

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T05:36:01.947611Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-05-10T05:32:23.972335Z digest=sha256:44ac121d930ceb94bc93604c4e25a261e7ba6527c53b375906cf6e181b049fed

Observation 612aa4d4-c586-4d4d-8cf8-5068c75199ca · inbound

Ex Ante Evaluation of AI-Induced Idea Diversity Collapse cites this paper.

Ex Ante Evaluation of AI-Induced Idea Diversity Collapse Jointly Reinforcing Diversity and Quality in Language Model Generations

Reference 30

Resolution
verified exact
arxiv_id, observed 2026-05-11T20:16:09.466150Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-08T09:49:57.122084Z digest=sha256:fe6fe4a71490688d7d0854218779abf02b88d0d6b4da1211a729e588486dd501

Observation cf08cd18-a316-4897-b0b2-8430a1291cd2 · inbound

Breaking $\textit{Winner-Takes-All}$: Cooperative Policy Optimization Improves Diverse LLM Reasoning cites this paper.

Breaking $\textit{Winner-Takes-All}$: Cooperative Policy Optimization Improves Diverse LLM Reasoning Jointly Reinforcing Diversity and Quality in Language Model Generations

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-05-13T02:07:09.130699Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-13T01:57:30.877915Z digest=sha256:a708cdf15e284f80c11ca7bd2ae351b32413c0efcc4c0eef902955a26568ec3c

Observation 72698e9c-aeeb-4466-9574-642f650d7f34 · inbound

Breaking $\textit{Winner-Takes-All}$: Cooperative Policy Optimization Improves Diverse LLM Reasoning cites this paper.

Breaking $\textit{Winner-Takes-All}$: Cooperative Policy Optimization Improves Diverse LLM Reasoning Jointly Reinforcing Diversity and Quality in Language Model Generations

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-05-20T22:49:10.604622Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-20T22:45:56.857810Z digest=sha256:3870ebb669d88e4029f0935437afe387f30867419bb2495be3957518570b3970

Observation 15894360-a27a-4a05-b28b-472d67f53cca · inbound

Weak-to-Strong Elicitation via Mismatched Wrong Drafts cites this paper.

Weak-to-Strong Elicitation via Mismatched Wrong Drafts Jointly Reinforcing Diversity and Quality in Language Model Generations

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-05-20T14:33:21.355499Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-20T14:32:04.778738Z digest=sha256:0a30626d82a5eb454846e737b9f167fdc00581a9c59243736e0a456d28d0754b

Observation 3cf3dda4-60e3-4df8-81f0-b0ef0e37317f · inbound

Weak-to-Strong Elicitation via Mismatched Wrong Drafts cites this paper.

Weak-to-Strong Elicitation via Mismatched Wrong Drafts Jointly Reinforcing Diversity and Quality in Language Model Generations

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-06-30T19:45:01.300944Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-30T19:25:47.710148Z digest=sha256:985fad58c5f43f4fa4ea8b495f782227625e724ece0779f2cba83a21a0d951e0

Observation 46694769-7e5f-4653-8de0-3befa6f7c54a · inbound

Cast a Wider Net: Coordinated Pass@K Policy Optimization for Code Reasoning cites this paper.

Cast a Wider Net: Coordinated Pass@K Policy Optimization for Code Reasoning Jointly Reinforcing Diversity and Quality in Language Model Generations

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-06-29T18:43:50.675745Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-29T18:38:12.066692Z digest=sha256:f3ee47d234035e4c2ad3aa4fbc9c275907e8bcace0e8195ea5e8b6b71d83fcf4

Observation c9ad50ca-f78f-4e37-93aa-d8f48db953a6 · inbound

Cast a Wider Net: Coordinated Pass@K Policy Optimization for Code Reasoning cites this paper.

Cast a Wider Net: Coordinated Pass@K Policy Optimization for Code Reasoning Jointly Reinforcing Diversity and Quality in Language Model Generations

Reference 2

Resolution
unresolved
no resolver link, observed 2026-07-12T15:56:23.031807Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T15:56:23.031807Z digest=sha256:7ce164924ffadfe9287c9830ee40c7ba56e3abc7a9e6f227cb5a31b19dee01b6

Observation f194d0ca-825e-4046-9468-86dfb47678d5 · inbound

DEI: Diversity in Evolutionary Inference for Quality-Diversity Search cites this paper.

DEI: Diversity in Evolutionary Inference for Quality-Diversity Search Jointly Reinforcing Diversity and Quality in Language Model Generations

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-06-29T18:43:50.481352Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-29T18:41:53.436686Z digest=sha256:a8516d548195b61105dc4a792850d85ecea0aea38f92aec271e7de83f473c48c

Observation eed184b3-4317-4ac5-a57b-07e48e8baf58 · inbound

Using Reward Uncertainty to Induce Diverse Behaviour in Reinforcement Learning cites this paper.

Using Reward Uncertainty to Induce Diverse Behaviour in Reinforcement Learning Jointly Reinforcing Diversity and Quality in Language Model Generations

Reference 12

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T02:26:26.270305Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-06-28T10:59:44.092482Z digest=sha256:5497f39126b25512166a01a2c5469c97f8734d89a0eb5efdbefe5dc154fe5c13

Observation 566591fd-f894-43e7-9ac0-cc2a9b78f151 · inbound

On Advantage Estimates for Max@K Policy Gradients cites this paper.

On Advantage Estimates for Max@K Policy Gradients Jointly Reinforcing Diversity and Quality in Language Model Generations

Reference 29

Resolution
verified exact
arxiv_id, observed 2026-07-02T12:06:56.451211Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-28T02:21:57.143016Z digest=sha256:c68a794aac7b6099864ff5fadfe0ab208ae9b87af6a7e4ea3e2356546e008c09

Observation 3670faa7-ac27-4631-b212-00732dcaccf4 · inbound

OrderGrad: Optimizing Beyond the Mean with Order-Statistic Policy Gradient Estimation cites this paper.

OrderGrad: Optimizing Beyond the Mean with Order-Statistic Policy Gradient Estimation Jointly Reinforcing Diversity and Quality in Language Model Generations

Reference 53

Resolution
verified exact
arxiv_id, observed 2026-07-02T12:16:56.970665Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-28T02:17:30.974692Z digest=sha256:261133d960f0a70a5d11a1362b540f7fa5427aed6cdf2b404cc3e0bc71cb935d

Observation 6805ecb0-bf8f-47bc-a768-415d5b40026e · inbound

Teacher-Free Self-Training Amplifies but Does Not Compound: A Pass@$K$ Crossover on a Free-Verifier Domain cites this paper.

Teacher-Free Self-Training Amplifies but Does Not Compound: A Pass@$K$ Crossover on a Free-Verifier Domain Jointly Reinforcing Diversity and Quality in Language Model Generations

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-07-02T16:47:09.146501Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-27T22:26:50.223393Z digest=sha256:242b2b909f7d1ee8267a874eee4af3c3b3608b0bd8d4e3f3510713ae736b676e

Observation ea897559-3dcd-4a69-8fa7-99089ab3a7ce · inbound

Reasoning or Memorization? Direction-Aware Diversity Exploration in LLM Reinforcement Learning cites this paper.

Reasoning or Memorization? Direction-Aware Diversity Exploration in LLM Reinforcement Learning Jointly Reinforcing Diversity and Quality in Language Model Generations

Reference 11

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T04:47:38.182193Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-06-27T13:39:42.620290Z digest=sha256:a3535167f01ebf522f40c01a5148b61d10a321e94eec3428fbd9b41429da83e8

Observation e20bb264-d25d-4544-b387-b6cfc569cf3c · inbound

On-Policy Self-Distillation with Sampled Demonstrations Reduces Output Diversity cites this paper.

On-Policy Self-Distillation with Sampled Demonstrations Reduces Output Diversity Jointly Reinforcing Diversity and Quality in Language Model Generations

Reference 19

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T21:00:08.450601Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-06-25T19:23:56.452083Z digest=sha256:470fdf42c6bd07e90208e6f592e7573cf01cc22a0de557764be6724e10407873

Observation c72d8999-34a7-4bdd-afce-9b19c257c954 · inbound

Decomposer: Learning to Decompile Symbolic Music to Programs cites this paper.

Decomposer: Learning to Decompile Symbolic Music to Programs Jointly Reinforcing Diversity and Quality in Language Model Generations

Reference 2

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T17:58:46.838735Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-07-03T17:51:10.756155Z digest=sha256:c2bbcccc2c1f4dc66f24686a7f2a917effd6c1fbd249d66d7d77f3b781b49334

Observation 0f44fba6-c473-437f-9b6f-25c2c03e8c57 · inbound

The One-Word Census: Answer-Choice Conformity Across 44 Language Models cites this paper.

The One-Word Census: Answer-Choice Conformity Across 44 Language Models Jointly Reinforcing Diversity and Quality in Language Model Generations

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-02T06:23:47.839824Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T06:23:47.839824Z digest=sha256:b4bace6aea4029f2ac777cbc291961d4531338077d06a5c4e70ad45d1fd32dff

Observation 096ae7fa-3420-4382-aa65-f5575fbcf5b0 · inbound

When Correct Solutions Repeat: Rarity-Aware Credit Redistribution for GRPO cites this paper.

When Correct Solutions Repeat: Rarity-Aware Credit Redistribution for GRPO Jointly Reinforcing Diversity and Quality in Language Model Generations

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-05T18:43:16.055506Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T18:43:16.055506Z digest=sha256:5ee222cc589d0b29f80506ed8028008868d0f6930223d93ef0ce020df9b9f358

Observation 14741cdb-ebe5-4946-8e2a-c9f1ee7b1038 · inbound

When Correct Solutions Repeat: Rarity-Aware Credit Redistribution for GRPO cites this paper.

When Correct Solutions Repeat: Rarity-Aware Credit Redistribution for GRPO Jointly Reinforcing Diversity and Quality in Language Model Generations

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-08T00:52:47.153167Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T00:52:47.153167Z digest=sha256:430f3244209e2add4b6a566df1433cf7b917ee8aab247addf874daad36250d29