Pith. sign in

Paper Citation Record · LEDGER

SHARP: Synthesizing High-quality Aligned Reasoning Problems for Large Reasoning Models Reinforcement Learning

As of 7 August 2026, this Paper Citation Record lists 20 of 20 outbound references and 2 inbound Pith citation observations for arXiv:2505.14147.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.14147 v3

Coverage vector

measured 20 of 20 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T15:42:31.015003Z

measured 22 of 22 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-05-10T03:23:18.770963Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-11T19:26:08.912385Z

Reference resolution

20 of 20 outbound references displayed

  • verified exact0
  • verified fuzzy17
  • unresolved3
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 7f04fcfc-1a55-4958-b8f5-7d006ea8383f · outbound

This paper cites problem":.

SHARP: Synthesizing High-quality Aligned Reasoning Problems for Large Reasoning Models Reinforcement Learning problem":

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:42:34.927985Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T15:42:28.412837Z digest=sha256:644381e5d8c1cc8b7802fb4b0ea8429373117a02d30515b89277416e4024e6f7

Observation 7e7078b4-d55b-4f4c-9e0f-e43d3db70aa3 · outbound

This paper cites Limitations.

SHARP: Synthesizing High-quality Aligned Reasoning Problems for Large Reasoning Models Reinforcement Learning Limitations

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:42:34.375996Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T15:42:28.891053Z digest=sha256:4b91ef4fb086b726a9b9730b4beeea27566615b2c8086012fb9149b761a257e1

Observation 82be9593-8397-4f25-9b7d-1f4f7fbe994f · outbound

This paper cites It does not appear to present new theoretical results in the form of theorems or mathematical proofs that would require a separate section for assumptions and proofs.

SHARP: Synthesizing High-quality Aligned Reasoning Problems for Large Reasoning Models Reinforcement Learning It does not appear to present new theoretical results in the form of theorems or mathematical proofs that would require a separate section for assumptions and proofs

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:42:34.207184Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T15:42:29.082213Z digest=sha256:399afaa4c09b414896c329183d5d83da1431717b568c46ea4b7a09578917aa73

Observation 42f43e75-b3ec-4589-a7e5-c2204aa64a7a · outbound

This paper cites Three-Tier Category.

SHARP: Synthesizing High-quality Aligned Reasoning Problems for Large Reasoning Models Reinforcement Learning Three-Tier Category

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:42:34.773530Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T15:42:28.599117Z digest=sha256:7f95c4d5569ff9fd4d15c97085ab9172435f6f5447a977047488440747763bf8

Observation 87a2bf42-4d00-4ba9-96d4-5867e6d946d3 · outbound

This paper cites It claimsSHARP encompasses self-alignment principles and a three-phase framework.

SHARP: Synthesizing High-quality Aligned Reasoning Problems for Large Reasoning Models Reinforcement Learning It claimsSHARP encompasses self-alignment principles and a three-phase framework

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:42:34.599356Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T15:42:28.766306Z digest=sha256:1f727ba97a88657ad5417cf289412fec456381ad4a70e3b42a7765b5d48c8688

Observation 5f74e292-5fc9-494b-999a-d9294be82bc0 · outbound

This paper cites It names the models used for comparison.

SHARP: Synthesizing High-quality Aligned Reasoning Problems for Large Reasoning Models Reinforcement Learning It names the models used for comparison

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:42:33.684352Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T15:42:29.552774Z digest=sha256:b4510520def8566ac529a53ee3802f6dbcd11e394e9e8e4a01b4537ca37c525e

Observation 6fbe80fa-9e49-4f4b-ac47-f11c2c22c2c2 · outbound

This paper cites If needed, we will include them in the camera-ready version.

SHARP: Synthesizing High-quality Aligned Reasoning Problems for Large Reasoning Models Reinforcement Learning If needed, we will include them in the camera-ready version

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:42:33.486092Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T15:42:29.669218Z digest=sha256:58f37ea47959fdb3bb275adc108441088bee27aaba284258bc313542ccada4ca

Observation 85901396-447b-4289-8f6f-c06ad66e6f1e · outbound

This paper cites It specifies the comparison models used for distillation and RL Zero training.

SHARP: Synthesizing High-quality Aligned Reasoning Problems for Large Reasoning Models Reinforcement Learning It specifies the comparison models used for distillation and RL Zero training

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:42:34.038605Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T15:42:29.300544Z digest=sha256:b458113e6a8da96a62554c328df15480185434a232d5fe776a154fec52d14ba1

Observation 4b4724b5-d8b0-4381-85e2-289156866be5 · outbound

This paper cites Moreover, we will open-source all necessary codes and related data for industry use during the review period.

SHARP: Synthesizing High-quality Aligned Reasoning Problems for Large Reasoning Models Reinforcement Learning Moreover, we will open-source all necessary codes and related data for industry use during the review period

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:42:33.854206Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T15:42:29.451461Z digest=sha256:0254f01bd30818dfcb876d410e2980dbe0f8fa5090e218ecae192af965239f2b

Observation 6450b544-6edd-49bf-8f5f-b3fc71aa8bcf · outbound

This paper cites It discusses the potential to push LRM performance closer to expert-level proficiency and superintelligence in STEM.

SHARP: Synthesizing High-quality Aligned Reasoning Problems for Large Reasoning Models Reinforcement Learning It discusses the potential to push LRM performance closer to expert-level proficiency and superintelligence in STEM

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:42:32.984653Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T15:42:30.090733Z digest=sha256:c88724d4a956e44a22027b29b1097e4733855301c6d9eeb73d55e6e41daed6ea

Observation 90292439-1e14-45bc-ab71-c3eaad2a752d · outbound

This paper cites The generated data consists of STEM problems.

SHARP: Synthesizing High-quality Aligned Reasoning Problems for Large Reasoning Models Reinforcement Learning The generated data consists of STEM problems

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:42:32.812818Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T15:42:30.195549Z digest=sha256:e91a110b1ad57086978fedd780305af430b1e542987ca4041d52f516dec11fa3

Observation 9a7968a1-7220-4c17-aac7-99556ee41ed2 · outbound

This paper cites Guidelines: • The answer NA means that the paper does not include experiments.

SHARP: Synthesizing High-quality Aligned Reasoning Problems for Large Reasoning Models Reinforcement Learning Guidelines: • The answer NA means that the paper does not include experiments

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:42:33.313116Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T15:42:29.800723Z digest=sha256:b097964d2a998fb9e8e99e78a3529ffba7836eb71953b8cab434ad3427fd92a9

Observation 8410c1cb-ea12-4ee0-b787-2ffa178fc9e3 · outbound

This paper cites It does not involve human subjects or obviously ethically sensitive applications, and we assume it conforms to the NeurIPS Code of Ethics.

SHARP: Synthesizing High-quality Aligned Reasoning Problems for Large Reasoning Models Reinforcement Learning It does not involve human subjects or obviously ethically sensitive applications, and we assume it conforms to the NeurIPS Code of Ethics

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:42:33.134018Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T15:42:29.913205Z digest=sha256:cf011a65a45b85c9e2d8a86408c9180c9bf0df00087bc96375c4b37cc55806ea

Observation 7badaaff-7eaf-45a3-bea3-177752fb8903 · outbound

This paper cites The process involves using LLMs to generate and verify problems, and then training other LLMs.

SHARP: Synthesizing High-quality Aligned Reasoning Problems for Large Reasoning Models Reinforcement Learning The process involves using LLMs to generate and verify problems, and then training other LLMs

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:42:32.063873Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T15:42:30.685344Z digest=sha256:0d1177141c626761b078a447a6f5a9445d047b6fdcddaf5686c02554a80d4d13

Observation de1bf953-6522-4f59-a4c1-1c85b5ea7919 · outbound

This paper cites Guidelines: • The answer NA means that the paper does not involve crowdsourcing nor research with human subjects.

SHARP: Synthesizing High-quality Aligned Reasoning Problems for Large Reasoning Models Reinforcement Learning Guidelines: • The answer NA means that the paper does not involve crowdsourcing nor research with human subjects

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:42:31.629140Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T15:42:30.852235Z digest=sha256:e235653c963c418d4f25ffbfdc6e235924126cb35ab0e21908fd5b6cb8adffb9

Observation 21a3a1f5-4a7a-4d1b-8a06-dc5aea7878c9 · outbound

This paper cites However, the specific licenses and terms of use for these assets are not explicitly mentioned in the paper text or the appendix.

SHARP: Synthesizing High-quality Aligned Reasoning Problems for Large Reasoning Models Reinforcement Learning However, the specific licenses and terms of use for these assets are not explicitly mentioned in the paper text or the appendix

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:42:32.659233Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T15:42:30.306739Z digest=sha256:3c08f9451975e418ff59ee5debfc11d9fd0d94ca13d9110ad5d50154b755413c

Observation cd1591e1-98d0-4078-bf50-35168ecc0b5b · outbound

This paper cites an unresolved cited work.

SHARP: Synthesizing High-quality Aligned Reasoning Problems for Large Reasoning Models Reinforcement Learning Unresolved cited work

Reference 17

Resolution
unresolved
raw_fallback, observed 2026-08-07T15:42:32.401501Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T15:42:30.476026Z digest=sha256:119806a7a0da0f600861258e437f5e75297d6a0c68e86cfa34b04eb6d4726a63

Observation bcdc7789-64c7-4d50-b5c4-b4a7d1c95f2f · outbound

This paper cites We implement SHARP by leveraging a state-of-the-art LRM to infer and verify challenging STEM questions.

SHARP: Synthesizing High-quality Aligned Reasoning Problems for Large Reasoning Models Reinforcement Learning We implement SHARP by leveraging a state-of-the-art LRM to infer and verify challenging STEM questions

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:42:31.289456Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T15:42:31.015003Z digest=sha256:8a937eed96626cb4d0c4ed21a28fd94b75b8b46828018c058799cc0ec7f7a0c7

Observation c203d76b-99d5-4d09-8387-726927cae348 · outbound

This paper cites Scaling Synthetic Data Creation with 1,000,000,000 Personas.

SHARP: Synthesizing High-quality Aligned Reasoning Problems for Large Reasoning Models Reinforcement Learning Scaling Synthetic Data Creation with 1,000,000,000 Personas

Reference 757

Resolution
unresolved
no resolver link, observed 2026-08-07T15:42:28.214460Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:42:28.214460Z digest=sha256:e23858b1db54c010d6da5097f0d77dbde2e1003f6a97dd1c4cf8a9a67b1901b9

Observation 24136f9c-ed0a-4f17-8bc5-15a2777b1ccc · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

SHARP: Synthesizing High-quality Aligned Reasoning Problems for Large Reasoning Models Reinforcement Learning DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-07T15:42:28.103899Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:42:28.103899Z digest=sha256:88d8fb27fb6b4ff04f7dcd8776e86f6bf8fa02b16f07b8660e90b1c12e72b4ca

Pith citing papers

Observation 774862cf-cf25-499f-88a4-57cfcdb9c030 · inbound

Fine-Tuning Small Reasoning Models for Quantum Field Theory cites this paper.

Fine-Tuning Small Reasoning Models for Quantum Field Theory SHARP: Synthesizing High-quality Aligned Reasoning Problems for Large Reasoning Models Reinforcement Learning

Reference 30

Resolution
verified exact
arxiv_id, observed 2026-05-11T12:31:06.168237Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T03:23:18.770963Z digest=sha256:38e3a75de6da8837a25763b0a1c3300b63d02fda0c2f0108e5bc99ba33185c72

Observation a48c0be4-6b0c-4213-8155-262274143fc4 · inbound

Verifier-Backed Hard Problem Generation for Mathematical Reasoning cites this paper.

Verifier-Backed Hard Problem Generation for Mathematical Reasoning SHARP: Synthesizing High-quality Aligned Reasoning Problems for Large Reasoning Models Reinforcement Learning

Reference 28

Resolution
verified exact
arxiv_id, observed 2026-05-11T19:26:08.915509Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-08T11:59:44.850797Z digest=sha256:6b4dfc518595b32ffc4b088e2d413ed290aa80a0178180ed67e476e9f054c3d0