Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
Paper Citation Record · LEDGER
As of 9 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 22 inbound Pith citation observations for arXiv:2404.02078.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-08T12:17:46.033521Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-07-03T09:47:59.922599Z
0 of 0 outbound references displayed
External citation measurements
No source-named external measurement is stored.
No outbound reference observations are available for this paper version.
Observation 751ebca6-97a9-4d89-845a-c0bddda875f0 · inbound
Skywork-Reward: Bag of Tricks for Reward Modeling in LLMs Advancing LLM Reasoning Generalists with Preference Trees
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation fe2c7e9f-6eeb-4093-b116-adbf38238b95 · inbound
Training Software Engineering Agents and Verifiers with SWE-Gym Advancing LLM Reasoning Generalists with Preference Trees
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation bbb07ed4-5c56-43c8-9149-8d4d764f1977 · inbound
SmolLM2: When Smol Goes Big -- Data-Centric Training of a Small Language Model Advancing LLM Reasoning Generalists with Preference Trees
Reference 249
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 1403d4b0-07a5-4db9-9c9d-63e497ddc699 · inbound
DPO-Shift: Shifting the Distribution of Direct Preference Optimization Advancing LLM Reasoning Generalists with Preference Trees
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 165f613f-8b50-463b-87c7-d8edfa19b879 · inbound
Measuring Diversity in Synthetic Datasets Advancing LLM Reasoning Generalists with Preference Trees
Reference 88
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 97156b06-e60e-412f-bc12-c2c7951d7809 · inbound
Video-R1: Reinforcing Video Reasoning in MLLMs Advancing LLM Reasoning Generalists with Preference Trees
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 5bab72b9-5519-4cfb-8f12-b78c64860908 · inbound
On the Effect of Negative Gradient in Group Relative Deep Reinforcement Optimization Advancing LLM Reasoning Generalists with Preference Trees
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a0780e28-bddc-4838-9500-e3b1140c06a7 · inbound
Towards Reliable, Uncertainty-Aware Alignment Advancing LLM Reasoning Generalists with Preference Trees
Reference 54
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7053f5d5-dea0-40bc-9503-8717edc46fbd · inbound
SC2Arena and StarEvolve: Benchmark and Self-Improvement Framework for LLMs in Complex Decision-Making Tasks Advancing LLM Reasoning Generalists with Preference Trees
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation de9fd7fd-02f4-48a6-90e0-a0ae86389acd · inbound
Anchoring Refusal Direction: Mitigating Safety Risks in Tuning via Projection Constraint Advancing LLM Reasoning Generalists with Preference Trees
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 66f4bc15-23dc-4d73-aba2-5d4f2540df54 · inbound
MoGU V2: Toward a Higher Pareto Frontier Between Model Usability and Security Advancing LLM Reasoning Generalists with Preference Trees
Reference 66
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b03affbf-8748-48a8-b01b-0021ec3c3e58 · inbound
PaTaRM: Bridging Pairwise and Pointwise Signals via Preference-Aware Task-Adaptive Reward Modeling Advancing LLM Reasoning Generalists with Preference Trees
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation c93977e0-94bb-4c22-9f53-e65977d49b3d · inbound
CompliBench: Benchmarking LLM Judges for Compliance Violation Detection in Dialogue Systems Advancing LLM Reasoning Generalists with Preference Trees
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation af10dbf6-ac5e-4141-b505-c0e394271f2b · inbound
Breaking the Reward Barrier: Accelerating Tree-of-Thought Reasoning via Speculative Exploration Advancing LLM Reasoning Generalists with Preference Trees
Reference 66
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation afbbf399-40d4-47e1-8cc6-6f4a444e99ce · inbound
Breaking the Reward Barrier: Accelerating Tree-of-Thought Reasoning via Speculative Exploration Advancing LLM Reasoning Generalists with Preference Trees
Reference 66
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation c197a95d-51d7-4132-993a-48c8947a0adb · inbound
Holder Policy Optimisation Advancing LLM Reasoning Generalists with Preference Trees
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation cdafa392-70f2-41ba-b055-d46aef9d1d2e · inbound
Holder Policy Optimisation Advancing LLM Reasoning Generalists with Preference Trees
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 0d72bc2b-36a2-4861-a354-1bc2b1e43cf5 · inbound
Selective-Advantage Entropy-Adaptive Horizon GRPO: Asymmetric Token-Level Discounting for Efficient Reinforcement Learning of Language Models Advancing LLM Reasoning Generalists with Preference Trees
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 0a1dc478-abb7-42a0-b754-eaed8449afc6 · inbound
Representation-Aware Advantage Estimation: Your Reward Model Provides More Than A Scalar Output Advancing LLM Reasoning Generalists with Preference Trees
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 35613fd2-39ca-4567-9df6-1dd09537c8c0 · inbound
Architecture-Aware Reinforcement Learning Makes Sliding-Window Attention Competitive in Math Reasoning Advancing LLM Reasoning Generalists with Preference Trees
Reference 66
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation cf321595-60bd-4b5a-a8f3-66f03fa2802c · inbound
Multi-Turn On-Policy Distillation with Prefix Replay Advancing LLM Reasoning Generalists with Preference Trees
Reference 67
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 350e3b83-e95c-497e-96fa-f4c9f3db9a24 · inbound
Multi-Turn On-Policy Distillation with Prefix Replay Advancing LLM Reasoning Generalists with Preference Trees
Reference 68
Source-reported events for the cited work
Unavailable: canonical work link unavailable.