Pith. sign in

Paper Citation Record · LEDGER

Advancing LLM Reasoning Generalists with Preference Trees

As of 9 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 22 inbound Pith citation observations for arXiv:2404.02078.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2404.02078 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 22 of 22 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 22 of 22 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-08T12:17:46.033521Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-03T09:47:59.922599Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 751ebca6-97a9-4d89-845a-c0bddda875f0 · inbound

Skywork-Reward: Bag of Tricks for Reward Modeling in LLMs cites this paper.

Skywork-Reward: Bag of Tricks for Reward Modeling in LLMs Advancing LLM Reasoning Generalists with Preference Trees

Reference 25

Resolution
verified exact
arxiv_id, observed 2026-05-17T16:18:01.688186Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-17T16:18:01.560780Z digest=sha256:0620a63312ff9b31cad8618e4407fa99d9290bbb7898c5ce4959906ef54fb00b

Observation fe2c7e9f-6eeb-4093-b116-adbf38238b95 · inbound

Training Software Engineering Agents and Verifiers with SWE-Gym cites this paper.

Training Software Engineering Agents and Verifiers with SWE-Gym Advancing LLM Reasoning Generalists with Preference Trees

Reference 9

Resolution
metadata mismatch
arxiv_id, observed 2026-05-18T05:20:40.532699Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-18T05:20:40.483057Z digest=sha256:2ab986bff7ce2f8d7976e48ecdbe93748083d61749f941b95683bc8f35035ce0

Observation bbb07ed4-5c56-43c8-9149-8d4d764f1977 · inbound

SmolLM2: When Smol Goes Big -- Data-Centric Training of a Small Language Model cites this paper.

SmolLM2: When Smol Goes Big -- Data-Centric Training of a Small Language Model Advancing LLM Reasoning Generalists with Preference Trees

Reference 249

Resolution
metadata mismatch
arxiv_id, observed 2026-05-13T17:30:03.070601Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-05-13T17:30:02.803757Z digest=sha256:8e8665b23dd22004f96ec49e52f04a033ae9f1473a56f97c6e3c97ef2c96be3a

Observation 1403d4b0-07a5-4db9-9c9d-63e497ddc699 · inbound

DPO-Shift: Shifting the Distribution of Direct Preference Optimization cites this paper.

DPO-Shift: Shifting the Distribution of Direct Preference Optimization Advancing LLM Reasoning Generalists with Preference Trees

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-08T12:17:46.033521Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T12:17:46.033521Z digest=sha256:2f2be06dc0fa6eab081dd5ee3db823ab6721c865eaf3c7696a5ff4dbe24d0ec1

Observation 165f613f-8b50-463b-87c7-d8edfa19b879 · inbound

Measuring Diversity in Synthetic Datasets cites this paper.

Measuring Diversity in Synthetic Datasets Advancing LLM Reasoning Generalists with Preference Trees

Reference 88

Resolution
unresolved
no resolver link, observed 2026-08-08T04:54:50.972531Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T04:54:50.972531Z digest=sha256:28cd3f2daf1de5a02923f1c245711461e25e3f2c3c030a3b21785000eaad3a72

Observation 97156b06-e60e-412f-bc12-c2c7951d7809 · inbound

Video-R1: Reinforcing Video Reasoning in MLLMs cites this paper.

Video-R1: Reinforcing Video Reasoning in MLLMs Advancing LLM Reasoning Generalists with Preference Trees

Reference 42

Resolution
verified exact
arxiv_id, observed 2026-05-12T09:43:00.303113Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-12T09:43:00.208065Z digest=sha256:2fd3f9fbb419da115bba7242af66c3e1a67ef0122a904ed230dcdf1a3d0543de

Observation 5bab72b9-5519-4cfb-8f12-b78c64860908 · inbound

On the Effect of Negative Gradient in Group Relative Deep Reinforcement Optimization cites this paper.

On the Effect of Negative Gradient in Group Relative Deep Reinforcement Optimization Advancing LLM Reasoning Generalists with Preference Trees

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T14:29:33.431718Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:29:33.431718Z digest=sha256:5c598761ef2e0d423118c35846c9b78e1d0d0f14d8b1f60ce3055e36753966fc

Observation a0780e28-bddc-4838-9500-e3b1140c06a7 · inbound

Towards Reliable, Uncertainty-Aware Alignment cites this paper.

Towards Reliable, Uncertainty-Aware Alignment Advancing LLM Reasoning Generalists with Preference Trees

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-06T15:40:17.967384Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T15:40:17.967384Z digest=sha256:ee0137cb4f4c0c561bba9c852a8d60cef9d202238cdbe2f47b8bf8a43e2ed7a5

Observation 7053f5d5-dea0-40bc-9503-8717edc46fbd · inbound

SC2Arena and StarEvolve: Benchmark and Self-Improvement Framework for LLMs in Complex Decision-Making Tasks cites this paper.

SC2Arena and StarEvolve: Benchmark and Self-Improvement Framework for LLMs in Complex Decision-Making Tasks Advancing LLM Reasoning Generalists with Preference Trees

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-05T20:30:30.162000Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T20:30:30.162000Z digest=sha256:6f7faa1bee3617b1b0a531d7b7e52bb3abff7f638bc3d046753719542c5e6597

Observation de9fd7fd-02f4-48a6-90e0-a0ae86389acd · inbound

Anchoring Refusal Direction: Mitigating Safety Risks in Tuning via Projection Constraint cites this paper.

Anchoring Refusal Direction: Mitigating Safety Risks in Tuning via Projection Constraint Advancing LLM Reasoning Generalists with Preference Trees

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-04T23:10:21.344039Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T23:10:21.344039Z digest=sha256:ce1eb88b57618611a61414cbb7e00f7cd192da0876cc3787333d1aeb285ead0d

Observation 66f4bc15-23dc-4d73-aba2-5d4f2540df54 · inbound

MoGU V2: Toward a Higher Pareto Frontier Between Model Usability and Security cites this paper.

MoGU V2: Toward a Higher Pareto Frontier Between Model Usability and Security Advancing LLM Reasoning Generalists with Preference Trees

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-04T23:09:41.578345Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T23:09:41.578345Z digest=sha256:a5b18d0363c0275bbd2661e410f7f4ad16b6922c5419a19b7761f562b1d6703a

Observation b03affbf-8748-48a8-b01b-0021ec3c3e58 · inbound

PaTaRM: Bridging Pairwise and Pointwise Signals via Preference-Aware Task-Adaptive Reward Modeling cites this paper.

PaTaRM: Bridging Pairwise and Pointwise Signals via Preference-Aware Task-Adaptive Reward Modeling Advancing LLM Reasoning Generalists with Preference Trees

Reference 22

Resolution
metadata mismatch
arxiv_id, observed 2026-05-18T03:20:49.262191Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-18T03:15:57.744706Z digest=sha256:2a4e73c88e0d0f06b7f9d5acef6c9744f0a9a3eaffe976f2e48b5a7ca5ab3570

Observation c93977e0-94bb-4c22-9f53-e65977d49b3d · inbound

CompliBench: Benchmarking LLM Judges for Compliance Violation Detection in Dialogue Systems cites this paper.

CompliBench: Benchmarking LLM Judges for Compliance Violation Detection in Dialogue Systems Advancing LLM Reasoning Generalists with Preference Trees

Reference 21

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T11:26:02.393606Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-10T14:56:00.449776Z digest=sha256:3a86ff61caa1172c0fa40fbb7ffb83b63a6a4ae827a1680b8b36802a362a77f2

Observation af10dbf6-ac5e-4141-b505-c0e394271f2b · inbound

Breaking the Reward Barrier: Accelerating Tree-of-Thought Reasoning via Speculative Exploration cites this paper.

Breaking the Reward Barrier: Accelerating Tree-of-Thought Reasoning via Speculative Exploration Advancing LLM Reasoning Generalists with Preference Trees

Reference 66

Resolution
verified exact
arxiv_id, observed 2026-05-12T06:51:30.256032Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-12T03:51:52.375703Z digest=sha256:c6b05f7af7eed243c1b6068ddf26b9c7a9b807e7737745915dd735da0bcf788e

Observation afbbf399-40d4-47e1-8cc6-6f4a444e99ce · inbound

Breaking the Reward Barrier: Accelerating Tree-of-Thought Reasoning via Speculative Exploration cites this paper.

Breaking the Reward Barrier: Accelerating Tree-of-Thought Reasoning via Speculative Exploration Advancing LLM Reasoning Generalists with Preference Trees

Reference 66

Resolution
verified exact
arxiv_id, observed 2026-05-15T05:15:03.285300Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-15T05:11:32.053440Z digest=sha256:bc81be104628a5e6121dc7b9d8d7034cbc8e4425c926c9cb89ed2068fe83247f

Observation c197a95d-51d7-4132-993a-48c8947a0adb · inbound

Holder Policy Optimisation cites this paper.

Holder Policy Optimisation Advancing LLM Reasoning Generalists with Preference Trees

Reference 32

Resolution
verified exact
arxiv_id, observed 2026-05-13T06:12:22.771833Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-05-13T06:08:28.855671Z digest=sha256:2f760311d037b8ad7a4aa889da26f27244f3425d2388cc6a404cf86666347950

Observation cdafa392-70f2-41ba-b055-d46aef9d1d2e · inbound

Holder Policy Optimisation cites this paper.

Holder Policy Optimisation Advancing LLM Reasoning Generalists with Preference Trees

Reference 32

Resolution
verified exact
arxiv_id, observed 2026-05-22T10:01:23.055182Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-05-22T10:00:58.600743Z digest=sha256:5c0d6553e7f4a3464e34dabb3d83ed700f6f3d03d7da9a1cc8326516a0fb23ce

Observation 0d72bc2b-36a2-4861-a354-1bc2b1e43cf5 · inbound

Selective-Advantage Entropy-Adaptive Horizon GRPO: Asymmetric Token-Level Discounting for Efficient Reinforcement Learning of Language Models cites this paper.

Selective-Advantage Entropy-Adaptive Horizon GRPO: Asymmetric Token-Level Discounting for Efficient Reinforcement Learning of Language Models Advancing LLM Reasoning Generalists with Preference Trees

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-07-02T07:26:46.066822Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-28T06:56:01.601701Z digest=sha256:a496eb4e545b99cf5ad65313a3953922d3261b3a34ddc7aa9dbf89644637fcbe

Observation 0a1dc478-abb7-42a0-b754-eaed8449afc6 · inbound

Representation-Aware Advantage Estimation: Your Reward Model Provides More Than A Scalar Output cites this paper.

Representation-Aware Advantage Estimation: Your Reward Model Provides More Than A Scalar Output Advancing LLM Reasoning Generalists with Preference Trees

Reference 27

Resolution
verified exact
arxiv_id, observed 2026-07-03T04:47:38.255909Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-06-27T13:39:17.701196Z digest=sha256:95357319a7d45f7050d196ed3bda5d215102229ef305e117aeb8ce12da5166c8

Observation 35613fd2-39ca-4567-9df6-1dd09537c8c0 · inbound

Architecture-Aware Reinforcement Learning Makes Sliding-Window Attention Competitive in Math Reasoning cites this paper.

Architecture-Aware Reinforcement Learning Makes Sliding-Window Attention Competitive in Math Reasoning Advancing LLM Reasoning Generalists with Preference Trees

Reference 66

Resolution
verified exact
arxiv_id, observed 2026-07-03T09:47:59.923933Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-06-27T10:18:54.163862Z digest=sha256:c6c28b604ed4bfa2f70daf431cc1da1212015adcd1566524e2e5b917c984f68c

Observation cf321595-60bd-4b5a-a8f3-66f03fa2802c · inbound

Multi-Turn On-Policy Distillation with Prefix Replay cites this paper.

Multi-Turn On-Policy Distillation with Prefix Replay Advancing LLM Reasoning Generalists with Preference Trees

Reference 67

Resolution
unresolved
no resolver link, observed 2026-07-11T13:53:36.775836Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-11T13:53:36.775836Z digest=sha256:3df6b751ab6d3c92b6b00ad90b1408ea8fe7c4188c5317cdfb27af11871c4de5

Observation 350e3b83-e95c-497e-96fa-f4c9f3db9a24 · inbound

Multi-Turn On-Policy Distillation with Prefix Replay cites this paper.

Multi-Turn On-Policy Distillation with Prefix Replay Advancing LLM Reasoning Generalists with Preference Trees

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-02T08:40:39.023948Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T08:40:39.023948Z digest=sha256:22fc84c796c3ca8f612fff945f15d12c6f3302a8406723b11cc9dd347e051f60