Pith. sign in

Paper Citation Record · LEDGER

On the Generalization Gap in Self-Evolving Language Model Reasoning

As of 4 August 2026, this Paper Citation Record lists 46 of 46 outbound references and 0 inbound Pith citation observations for arXiv:2606.01075.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2606.01075 v2

Coverage vector

measured 46 of 46 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-07-11T11:50:26.030339Z

measured 46 of 46 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-04T06:34:03.388597+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

46 of 46 outbound references displayed

  • verified exact33
  • verified fuzzy0
  • unresolved4
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch9

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 9b2654f6-84c6-40e9-af8f-46ba3b22ae61 · outbound

This paper cites Miranda, Hanyang Zhao, Mohammad Rifqi Farhansyah, Garry Kuwanto, Derry Wijaya, and Genta Indra Winata.

On the Generalization Gap in Self-Evolving Language Model Reasoning Miranda, Hanyang Zhao, Mohammad Rifqi Farhansyah, Garry Kuwanto, Derry Wijaya, and Genta Indra Winata

Reference 1

Resolution
metadata mismatch
arxiv_id, observed 2026-06-28T17:22:24.933942Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-28T17:18:37.671369Z digest=sha256:10d1f120990e2127179b28b8f35f22af28249292076d5931c477988c093d2308

Observation 8ef754e8-b454-4eca-96e8-e60de265ad14 · outbound

This paper cites Annotating the annotators: Analysis, insightsandmodellingfromanannotationcampaignonpersuasiontechniques detection.

On the Generalization Gap in Self-Evolving Language Model Reasoning Annotating the annotators: Analysis, insightsandmodellingfromanannotationcampaignonpersuasiontechniques detection

Reference 2

Resolution
unresolved
no resolver link, observed 2026-06-28T17:18:37.671369Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T17:18:37.671369Z digest=sha256:a5a5a96c870be4b2cf5e08e806f318dc5dbb8c73bee5a96dc9b867951d03b042

Observation 79d82125-d43a-4c5f-a8f4-f8e9ea9cc3dc · outbound

This paper cites Avrim Blum, Daniel Hsu, Cyrus Rashtchian, and Donya Saless.

On the Generalization Gap in Self-Evolving Language Model Reasoning Avrim Blum, Daniel Hsu, Cyrus Rashtchian, and Donya Saless

Reference 3

Resolution
unresolved
no resolver link, observed 2026-06-28T17:18:37.671369Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T17:18:37.671369Z digest=sha256:7a880181bde49bd828df8f05293ae1e7595cd4df00c57670da865bf827058a08

Observation da42ab51-9ed2-4ed8-bb2c-146d8f3330f0 · outbound

This paper cites Training Verifiers to Solve Math Word Problems.

On the Generalization Gap in Self-Evolving Language Model Reasoning Training Verifiers to Solve Math Word Problems

Reference 4

Resolution
verified exact
local_arxiv, observed 2026-06-28T17:22:24.862173Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-28T17:18:37.671369Z digest=sha256:340a47b14c4cc80b3bf5e35d4e8f6aa3e89071de10c4abf430704cb3881ff863

Observation ac52c175-ccf9-44d8-8445-67eac2bf696f · outbound

This paper cites Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities.

On the Generalization Gap in Self-Evolving Language Model Reasoning Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities

Reference 5

Resolution
verified exact
local_arxiv, observed 2026-06-28T17:22:24.884454Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-28T17:18:37.671369Z digest=sha256:97fa74e2ad5fa2001f60776bed53cb994cfbe4590d5ca496a8f82803f0b1d01e

Observation 22d9bd67-12de-4616-86ab-0d46057438c9 · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

On the Generalization Gap in Self-Evolving Language Model Reasoning DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 6

Resolution
metadata mismatch
local_arxiv, observed 2026-06-28T17:22:24.870434Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-28T17:18:37.671369Z digest=sha256:f57b2d5ec6bc22e65779dd87d4fea10fbb08fc161d70af884b23ebb89880dbd8

Observation 39980ace-226b-4322-8de7-cf919e314565 · outbound

This paper cites On Designing Effective RL Reward at Training Time for LLM Reasoning.

On the Generalization Gap in Self-Evolving Language Model Reasoning On Designing Effective RL Reward at Training Time for LLM Reasoning

Reference 7

Resolution
metadata mismatch
arxiv_id, observed 2026-06-28T17:22:24.859815Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-28T17:18:37.671369Z digest=sha256:6f191da882d0bde709bf380e4a122408c9657282c101d7ed3cbbed35cdeef599

Observation 2cb1455c-c6ab-4c98-8adc-ecf9c6831f92 · outbound

This paper cites Gemma 3 Technical Report.

On the Generalization Gap in Self-Evolving Language Model Reasoning Gemma 3 Technical Report

Reference 9

Resolution
metadata mismatch
local_arxiv, observed 2026-06-28T17:22:24.924563Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-28T17:18:37.671369Z digest=sha256:d7910ef46224a69240020afef52de9f05dbdf88052bdde028268d76dfe6926a4

Observation b111a94c-994f-4f75-8016-ab24774a1a39 · outbound

This paper cites OpenThoughts: Data Recipes for Reasoning Models.

On the Generalization Gap in Self-Evolving Language Model Reasoning OpenThoughts: Data Recipes for Reasoning Models

Reference 10

Resolution
verified exact
local_arxiv, observed 2026-06-28T17:22:24.870793Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-28T17:18:37.671369Z digest=sha256:b6962c7daadbb31df2fcd08b92d49245ae79e1b2a113ba19715d5ef05f73aa5c

Observation 536da892-ee22-4e77-a2d9-3ba2cea9676c · outbound

This paper cites Rubrics as Rewards: Reinforcement Learning Beyond Verifiable Domains.

On the Generalization Gap in Self-Evolving Language Model Reasoning Rubrics as Rewards: Reinforcement Learning Beyond Verifiable Domains

Reference 11

Resolution
verified exact
local_arxiv, observed 2026-06-28T17:22:24.856776Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-28T17:18:37.671369Z digest=sha256:4a4eff87088146d976e9d29f22242f11512024c7001580e74c1420795079d456

Observation 9a0c1eb6-0ce2-442e-9a96-8f627c9a2a06 · outbound

This paper cites Measuring Mathematical Problem Solving With the MATH Dataset.

On the Generalization Gap in Self-Evolving Language Model Reasoning Measuring Mathematical Problem Solving With the MATH Dataset

Reference 12

Resolution
verified exact
local_arxiv, observed 2026-06-28T17:22:24.904800Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-28T17:18:37.671369Z digest=sha256:488284f7bfa9607a7696d5d56b289c2c24304fa6c4820d12a6ee6b691fba3783

Observation 73f72c26-913e-4e00-9097-f07e002fca28 · outbound

This paper cites Self-Improvement in Language Models: The Sharpening Mechanism.

On the Generalization Gap in Self-Evolving Language Model Reasoning Self-Improvement in Language Models: The Sharpening Mechanism

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-06-28T17:22:24.862765Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-28T17:18:37.671369Z digest=sha256:484bbd453cda23c819ca816d1670be762772bf41508ff9736b936055873ed8ce

Observation 5aa575b6-57f0-4a85-9993-76f83eeccc51 · outbound

This paper cites Is Best-of-N the Best of Them? Coverage, Scaling, and Optimality in Inference-Time Alignment.

On the Generalization Gap in Self-Evolving Language Model Reasoning Is Best-of-N the Best of Them? Coverage, Scaling, and Optimality in Inference-Time Alignment

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-06-28T17:22:24.893889Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-28T17:18:37.671369Z digest=sha256:a5df7b29e875ef06aa051913b2e3c3de86673e3b0be2d4ebb9274f91d39ffe11

Observation d418c805-f767-42cb-8fc5-743144cb575d · outbound

This paper cites Compute as Teacher: Turning Inference Compute Into Reference-Free Supervision.

On the Generalization Gap in Self-Evolving Language Model Reasoning Compute as Teacher: Turning Inference Compute Into Reference-Free Supervision

Reference 15

Resolution
verified exact
local_arxiv, observed 2026-06-28T17:22:24.904073Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-28T17:18:37.671369Z digest=sha256:227f5dc375851c7eca234562c8e1bcf85a7ac01e88958b92bd22dc9dd969783a

Observation 0d110353-efde-4ecc-9ad3-290604fad18d · outbound

This paper cites LLMs Could Autonomously Learn Without External Supervision.

On the Generalization Gap in Self-Evolving Language Model Reasoning LLMs Could Autonomously Learn Without External Supervision

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-06-28T17:22:24.896495Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-28T17:18:37.671369Z digest=sha256:95d6a0dea1e1909eeec182f7b4681f92d56770ca95cca4813918ecc45f07f1bf

Observation 5b2fbd29-b6ea-48d7-b0b2-79804be7b781 · outbound

This paper cites Demystifying synthetic data in llm pre-training: A systematic study of scaling laws, benefits, and pitfalls.

On the Generalization Gap in Self-Evolving Language Model Reasoning Demystifying synthetic data in llm pre-training: A systematic study of scaling laws, benefits, and pitfalls

Reference 17

Resolution
unresolved
no resolver link, observed 2026-06-28T17:18:37.671369Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T17:18:37.671369Z digest=sha256:f072e7052a3355abe0c4c65f179483864345c09e9b686b13795e6f57117bb112

Observation 9ccf7c5f-5c9c-4925-9c20-6319af8bde5b · outbound

This paper cites Search-based correction of reasoning chains for language models.arXiv preprint arXiv:2505.11824,.

On the Generalization Gap in Self-Evolving Language Model Reasoning Search-based correction of reasoning chains for language models.arXiv preprint arXiv:2505.11824,

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-06-28T17:22:24.901017Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-28T17:18:37.671369Z digest=sha256:5ca60cd943fbc5234a8f01e1773cd6796bb0c8c15a8ab6714c7666f8aa808b3e

Observation 2767187f-e6d4-44ec-8358-2a2a6d1a939f · outbound

This paper cites Language self-play for data-free training.

On the Generalization Gap in Self-Evolving Language Model Reasoning Language self-play for data-free training

Reference 19

Resolution
verified exact
arxiv_id, observed 2026-06-28T17:22:24.834128Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-28T17:18:37.671369Z digest=sha256:ce5cb190ae2b1d88ea4a8d34e9aab1b4ab6dbe36f08ff5b34583a26329fb2e1e

Observation 30caf84e-ec89-42ce-b53f-7c8efb986e0a · outbound

This paper cites Training Language Models to Self-Correct via Reinforcement Learning.

On the Generalization Gap in Self-Evolving Language Model Reasoning Training Language Models to Self-Correct via Reinforcement Learning

Reference 20

Resolution
verified exact
local_arxiv, observed 2026-06-28T17:22:24.915769Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-28T17:18:37.671369Z digest=sha256:a3f4ac7dfbc966ab0464bb5dbfe901f3daed02ceada60cc269c981e53adb7e11

Observation ec45cdcf-ec91-4ab9-b5cb-639a6397113f · outbound

This paper cites Tulu 3: Pushing Frontiers in Open Language Model Post-Training.

On the Generalization Gap in Self-Evolving Language Model Reasoning Tulu 3: Pushing Frontiers in Open Language Model Post-Training

Reference 21

Resolution
verified exact
local_arxiv, observed 2026-06-28T17:22:24.909914Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-28T17:18:37.671369Z digest=sha256:9efdb5b73589c599110272fc43bb8d9e254db1ad896f66ac3b15220674b21f7f

Observation aec1a5f7-acac-4f2c-8072-27ecff59d401 · outbound

This paper cites ReVISE: Learning to Refine at Test-Time via Intrinsic Self-Verification.

On the Generalization Gap in Self-Evolving Language Model Reasoning ReVISE: Learning to Refine at Test-Time via Intrinsic Self-Verification

Reference 22

Resolution
verified exact
arxiv_id, observed 2026-06-28T17:22:24.906664Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-28T17:18:37.671369Z digest=sha256:19a00334c58310089433e7e510b78680b7696e40827ef3a22a0807d71a1dcadf

Observation 071efbf6-4c99-49ef-9fb6-a0e5c6ca72cb · outbound

This paper cites Multi-Agent Verification: Scaling Test-Time Compute with Multiple Verifiers.

On the Generalization Gap in Self-Evolving Language Model Reasoning Multi-Agent Verification: Scaling Test-Time Compute with Multiple Verifiers

Reference 23

Resolution
metadata mismatch
arxiv_id, observed 2026-06-28T17:22:24.899116Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-28T17:18:37.671369Z digest=sha256:f849aa04295a17bc8dd55f413b3fceb1e4c65e4d528033dea74b16fdfedcf835

Observation 38a7524c-f242-49d8-bcce-59accbd5e717 · outbound

This paper cites Self-improving vlm judges without human annotations.arXiv preprint arXiv:2512.05145,.

On the Generalization Gap in Self-Evolving Language Model Reasoning Self-improving vlm judges without human annotations.arXiv preprint arXiv:2512.05145,

Reference 24

Resolution
verified exact
arxiv_id, observed 2026-06-28T17:22:24.845015Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-28T17:18:37.671369Z digest=sha256:7fcc7f6b3b5ff9755386a69b37b77dbdf5042b84f1d80db3c18abef1d0c2e7a0

Observation 818e8dfd-a375-4be8-b94c-906bad8d0842 · outbound

This paper cites ProRL: Prolonged Reinforcement Learning Expands Reasoning Boundaries in Large Language Models.

On the Generalization Gap in Self-Evolving Language Model Reasoning ProRL: Prolonged Reinforcement Learning Expands Reasoning Boundaries in Large Language Models

Reference 25

Resolution
verified exact
local_arxiv, observed 2026-06-28T17:22:24.820455Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-28T17:18:37.671369Z digest=sha256:d7ae4f3c313ba6f56b335bf9725faf2d3211cf4a341d0bf9c08cb9192a6d3b9c

Observation 03258f5a-a419-4f52-962e-1e01d6164398 · outbound

This paper cites Dynamic Prompt Learning via Policy Gradient for Semi-structured Mathematical Reasoning.

On the Generalization Gap in Self-Evolving Language Model Reasoning Dynamic Prompt Learning via Policy Gradient for Semi-structured Mathematical Reasoning

Reference 26

Resolution
verified exact
arxiv_id, observed 2026-06-28T17:22:24.934086Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-28T17:18:37.671369Z digest=sha256:adfae14ec16619fbfd3df75dd0724936dc393f2f474fe72bfa5b5563fbced8e3

Observation 14c942b6-4f89-4d4b-b320-f6db9d609586 · outbound

This paper cites The “problem” of human label variation: On ground truth in data, modeling and evaluation.

On the Generalization Gap in Self-Evolving Language Model Reasoning The “problem” of human label variation: On ground truth in data, modeling and evaluation

Reference 27

Resolution
unresolved
no resolver link, observed 2026-06-28T17:18:37.671369Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T17:18:37.671369Z digest=sha256:32963fc901e4c451c68cf3184051bae5238728facf1b74bf033c35cfa6d98845

Observation 3281a1a2-3759-4e31-a6de-dc46c4f75653 · outbound

This paper cites The answer is (X)\.

On the Generalization Gap in Self-Evolving Language Model Reasoning The answer is (X)\

Reference 28

Resolution
metadata mismatch
arxiv_id, observed 2026-06-28T17:22:24.823056Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-28T17:18:37.671369Z digest=sha256:b339b16df2f4186d43e2e94ffc35c1c34a453b5eb66887f644ea27ade2839d01

Observation 959bee1d-0987-4be4-adfc-c3212bae2d98 · outbound

This paper cites Scaling Test-Time Compute Without Verification or RL is Suboptimal.

On the Generalization Gap in Self-Evolving Language Model Reasoning Scaling Test-Time Compute Without Verification or RL is Suboptimal

Reference 29

Resolution
verified exact
arxiv_id, observed 2026-06-28T17:22:24.924302Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-28T17:18:37.671369Z digest=sha256:8a2129349e38fdb454e6db5d9ca29b5c588b9ffde5706db5c76b90a578a25892

Observation f03008c7-d87b-4c6a-af33-17a8242f9f23 · outbound

This paper cites Can large reasoning models self-train?.

On the Generalization Gap in Self-Evolving Language Model Reasoning Can large reasoning models self-train?

Reference 30

Resolution
verified exact
arxiv_id, observed 2026-06-28T17:22:24.846010Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-28T17:18:37.671369Z digest=sha256:73317dad655e124ab3fa1a90ef33ef0b75fdf93a2f58e79b171d710bfaa88a7d

Observation 03f90375-e681-4bc1-915a-4f83010bb428 · outbound

This paper cites Spurious Rewards: Rethinking Training Signals in RLVR.

On the Generalization Gap in Self-Evolving Language Model Reasoning Spurious Rewards: Rethinking Training Signals in RLVR

Reference 31

Resolution
verified exact
local_arxiv, observed 2026-06-28T17:22:24.873180Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-28T17:18:37.671369Z digest=sha256:96244e14eda112df1dc1a6119c66b53625ecd4166ae84528867238d8290c349e

Observation a301ee11-2939-4f69-8dc7-da88abacbfca · outbound

This paper cites ResearchRubrics : A benchmark of prompts and rubrics for evaluating deep research agents.

On the Generalization Gap in Self-Evolving Language Model Reasoning ResearchRubrics : A benchmark of prompts and rubrics for evaluating deep research agents

Reference 32

Resolution
verified exact
arxiv_id, observed 2026-06-28T17:22:24.880044Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-28T17:18:37.671369Z digest=sha256:a0ca451d9a54f8068d1fb335fbf90aba9dd5a9a05fd223120b7a51d93808568a

Observation d65c1290-9258-4ad1-accd-ea2af3017a66 · outbound

This paper cites Satori: Reinforcement Learning with Chain-of-Action-Thought Enhances LLM Reasoning via Autoregressive Search.

On the Generalization Gap in Self-Evolving Language Model Reasoning Satori: Reinforcement Learning with Chain-of-Action-Thought Enhances LLM Reasoning via Autoregressive Search

Reference 33

Resolution
verified exact
arxiv_id, observed 2026-06-28T17:22:24.890131Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-28T17:18:37.671369Z digest=sha256:e9f22d7579f677e0f0233e603fa867c1b27fe56366286d21d428c8c2fad62706

Observation f5591126-1c6c-4e21-b113-78511853ae51 · outbound

This paper cites Mind the Gap: Examining the Self-Improvement Capabilities of Large Language Models.

On the Generalization Gap in Self-Evolving Language Model Reasoning Mind the Gap: Examining the Self-Improvement Capabilities of Large Language Models

Reference 34

Resolution
verified exact
arxiv_id, observed 2026-06-28T17:22:24.928244Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-28T17:18:37.671369Z digest=sha256:9a03851e548f59e387e3da5e8e2f5e1fa2fdd3c4294bad36af0bd778f5c90a0a

Observation bf023d46-7256-4c25-ba92-143c70e8fad8 · outbound

This paper cites arXiv preprint arXiv:2507.00075 , year=.

On the Generalization Gap in Self-Evolving Language Model Reasoning arXiv preprint arXiv:2507.00075 , year=

Reference 35

Resolution
metadata mismatch
arxiv_id, observed 2026-06-28T17:22:24.892779Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-28T17:18:37.671369Z digest=sha256:4deba2ed25e57e714709fc015cf92162e9e2fb2aeab4703086bfbc584413a286

Observation 56b0ad6b-cd6a-47cc-877f-7df9ddc22f11 · outbound

This paper cites Reinforcement Learning for Reasoning in Large Language Models with One Training Example.

On the Generalization Gap in Self-Evolving Language Model Reasoning Reinforcement Learning for Reasoning in Large Language Models with One Training Example

Reference 36

Resolution
metadata mismatch
local_arxiv, observed 2026-06-28T17:22:24.111531Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-07-11T11:50:26.030339Z digest=sha256:c638a34a87ab252dcaec4ba70b3e0c76abb5b09537bc722a56877f0715723428

Observation 5f9d1c4b-b980-4951-bed9-2d7cad5ab78c · outbound

This paper cites CREAM: Consistency Regularized Self-Rewarding Language Models.

On the Generalization Gap in Self-Evolving Language Model Reasoning CREAM: Consistency Regularized Self-Rewarding Language Models

Reference 37

Resolution
verified exact
arxiv_id, observed 2026-06-28T17:22:24.883038Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-28T17:18:37.671369Z digest=sha256:788d3bfd200208448a9e21aedf8e600971105b9d2b6b131a20c47ffac910bdd8

Observation 2d63038d-9174-40a4-9294-783aa9b9a061 · outbound

This paper cites The invisible leash: Why rlvr may or may not escape its origin.

On the Generalization Gap in Self-Evolving Language Model Reasoning The invisible leash: Why rlvr may or may not escape its origin

Reference 38

Resolution
verified exact
arxiv_id, observed 2026-06-28T17:22:24.891235Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-28T17:18:37.671369Z digest=sha256:69963db37c89d5b3e4516fa1bfd999de7db971ad7df474675749afd131404944

Observation 205eac9a-7f3f-473f-9b7e-ba3b271e323d · outbound

This paper cites On Memorization of Large Language Models in Logical Reasoning.

On the Generalization Gap in Self-Evolving Language Model Reasoning On Memorization of Large Language Models in Logical Reasoning

Reference 39

Resolution
verified exact
arxiv_id, observed 2026-06-28T17:22:24.888547Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-28T17:18:37.671369Z digest=sha256:d8cb7208aa98877cd0d644ffb7d6912331b6a3c5d10204f645126126279601e7

Observation 17319cef-dde8-4c2c-9d86-28417351a0c9 · outbound

This paper cites Qwen3 Technical Report.

On the Generalization Gap in Self-Evolving Language Model Reasoning Qwen3 Technical Report

Reference 40

Resolution
verified exact
local_arxiv, observed 2026-06-28T17:22:24.875497Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-28T17:18:37.671369Z digest=sha256:a0cce30a25a18089cbacfcef58ffacc1a49727a802dc28040acc91a3e8ee77b6

Observation e315e6d2-d6b1-4853-858b-f75e8fbbf44e · outbound

This paper cites CoT-Self-Instruct: Building high-quality synthetic prompts for reasoning and non-reasoning tasks.

On the Generalization Gap in Self-Evolving Language Model Reasoning CoT-Self-Instruct: Building high-quality synthetic prompts for reasoning and non-reasoning tasks

Reference 41

Resolution
verified exact
arxiv_id, observed 2026-06-28T17:22:24.918750Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-28T17:18:37.671369Z digest=sha256:83db2f8c8528ddffe942f62f0d8c0cb13ad284ec435b3fb7af225e5d23fe4199

Observation f0e7d95e-baf8-4494-b6b0-69750ca5a0be · outbound

This paper cites Wisdom of the Crowd: Reinforcement Learning from Coevolutionary Collective Feedback.

On the Generalization Gap in Self-Evolving Language Model Reasoning Wisdom of the Crowd: Reinforcement Learning from Coevolutionary Collective Feedback

Reference 42

Resolution
verified exact
arxiv_id, observed 2026-06-28T17:22:24.921535Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-28T17:18:37.671369Z digest=sha256:33494a9873cbff7db50d91888d8cf5affff45d97994d2b628d5a70549c8e3a7a

Observation 20b0da13-b25d-40f3-be3d-edf1d8366c7d · outbound

This paper cites Learning to Discover at Test Time.

On the Generalization Gap in Self-Evolving Language Model Reasoning Learning to Discover at Test Time

Reference 43

Resolution
verified exact
local_arxiv, observed 2026-06-28T17:22:24.881981Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-28T17:18:37.671369Z digest=sha256:2c3aa6852d9565e295dbdb12e2979ff44568a5e8c95fdecab38eea6ab953bcec

Observation 9ca7a1b8-1fae-4c19-a619-8726fa10704e · outbound

This paper cites arXiv preprint arXiv:2502.05605 , year =.

On the Generalization Gap in Self-Evolving Language Model Reasoning arXiv preprint arXiv:2502.05605 , year =

Reference 44

Resolution
verified exact
arxiv_id, observed 2026-06-28T17:22:24.921961Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-28T17:18:37.671369Z digest=sha256:1de1d7b956b5aa82dcb7c1cee953d232dd063190bb31f21c1cf264b80ddf070d

Observation c26257f0-6d10-4a9e-ba02-094f8a5eb183 · outbound

This paper cites arXiv preprint arXiv:2601.05280 , doi=.

On the Generalization Gap in Self-Evolving Language Model Reasoning arXiv preprint arXiv:2601.05280 , doi=

Reference 45

Resolution
metadata mismatch
arxiv_id, observed 2026-06-28T17:22:24.907486Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-28T17:18:37.671369Z digest=sha256:c50b57f6faf26f8b280d2fe6716990856fdcb980ddc88c71ca02170ea3a761f2

Observation 24bdac98-83ff-46a8-911f-ab6accbe959b · outbound

This paper cites Right Question is Already Half the Answer: Fully Unsupervised LLM Reasoning Incentivization.

On the Generalization Gap in Self-Evolving Language Model Reasoning Right Question is Already Half the Answer: Fully Unsupervised LLM Reasoning Incentivization

Reference 46

Resolution
verified exact
arxiv_id, observed 2026-06-28T17:22:24.919029Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-28T17:18:37.671369Z digest=sha256:7aa465a0b1a7a5d26600541130fe53adb8334c9ad91de3c70c7248ff63fd517a

Observation 44d1bcc3-e190-4656-bd72-f09fac76d353 · outbound

This paper cites TTRL: Test-Time Reinforcement Learning.

On the Generalization Gap in Self-Evolving Language Model Reasoning TTRL: Test-Time Reinforcement Learning

Reference 47

Resolution
verified exact
local_arxiv, observed 2026-06-28T17:22:24.927167Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-28T17:18:37.671369Z digest=sha256:4cdb7abc6f657e6298beafd8727eca909e9f2fa99e4fbbf8db194a2f3f0cb106

Pith citing papers

No inbound Pith citation observations are available.