Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
Paper Citation Record · LEDGER
As of 22 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 41 inbound Pith citation observations for arXiv:2404.02078.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-22T06:32:14.747728+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-16T00:35:19.143448Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-07-03T09:47:59.922599Z
0 of 0 outbound references displayed
External citation measurements
No source-named external measurement is stored.
No outbound reference observations are available for this paper version.
Observation 751ebca6-97a9-4d89-845a-c0bddda875f0 · inbound
Skywork-Reward: Bag of Tricks for Reward Modeling in LLMs Advancing LLM Reasoning Generalists with Preference Trees
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 0b4d4772-3d26-4719-94d0-a63a35b42b18 · inbound
Towards Adaptive Mechanism Activation in Language Agent Advancing LLM Reasoning Generalists with Preference Trees
Reference 61
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation acca620a-0b77-43c2-9e4b-578ac7f043d1 · inbound
Large Language Models for Scholarly Ontology Generation: An Extensive Analysis in the Engineering Field Advancing LLM Reasoning Generalists with Preference Trees
Reference 68
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6099d9b7-f6b1-4abf-b6d6-bd8e7a675af1 · inbound
JuStRank: Benchmarking LLM Judges for System Ranking Advancing LLM Reasoning Generalists with Preference Trees
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 99b3e0ae-f37d-4861-986f-cdd089439c71 · inbound
Cal-DPO: Calibrated Direct Preference Optimization for Language Model Alignment Advancing LLM Reasoning Generalists with Preference Trees
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f85fa636-c19f-47b0-ae73-25e54172b216 · inbound
Progressive Multimodal Reasoning via Active Retrieval Advancing LLM Reasoning Generalists with Preference Trees
Reference 117
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cd143d7e-9622-438d-a359-28b3a2e8a701 · inbound
AceMath: Advancing Frontier Math Reasoning with Post-Training and Reward Modeling Advancing LLM Reasoning Generalists with Preference Trees
Reference 74
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fe2c7e9f-6eeb-4093-b116-adbf38238b95 · inbound
Training Software Engineering Agents and Verifiers with SWE-Gym Advancing LLM Reasoning Generalists with Preference Trees
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation af2455f1-6d70-40e0-a5f2-6f8daa8f7209 · inbound
LLM-Virus: Evolutionary Jailbreak Attack on Large Language Models Advancing LLM Reasoning Generalists with Preference Trees
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a1ed69d1-9d74-414b-9184-f965fb0de936 · inbound
AlphaPO: Reward Shape Matters for LLM Alignment Advancing LLM Reasoning Generalists with Preference Trees
Reference 46
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2dfc7745-fffc-455d-94a0-c9c1e4ca79cd · inbound
VidChain: Chain-of-Tasks with Metric-based Direct Preference Optimization for Dense Video Captioning Advancing LLM Reasoning Generalists with Preference Trees
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4a54c6a8-65d1-4f45-b9b5-e850d0fd7f98 · inbound
VideoWorld: Exploring Knowledge Learning from Unlabeled Videos Advancing LLM Reasoning Generalists with Preference Trees
Reference 74
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3a07252d-3349-4ce5-94f7-05ec9388a8c1 · inbound
Improving Influence-based Instruction Tuning Data Selection for Balanced Learning of Diverse Capabilities Advancing LLM Reasoning Generalists with Preference Trees
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 96d4f58c-83e4-4ec7-a4e3-d052d399ecbb · inbound
PairJudge RM: Perform Best-of-N Sampling with Knockout Tournament Advancing LLM Reasoning Generalists with Preference Trees
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ae83feb8-fc69-41a7-a818-32a886b08617 · inbound
Eagle 2: Building Post-Training Data Strategies from Scratch for Frontier Vision-Language Models Advancing LLM Reasoning Generalists with Preference Trees
Reference 171
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bbb07ed4-5c56-43c8-9149-8d4d764f1977 · inbound
SmolLM2: When Smol Goes Big -- Data-Centric Training of a Small Language Model Advancing LLM Reasoning Generalists with Preference Trees
Reference 249
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 1403d4b0-07a5-4db9-9c9d-63e497ddc699 · inbound
DPO-Shift: Shifting the Distribution of Direct Preference Optimization Advancing LLM Reasoning Generalists with Preference Trees
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 165f613f-8b50-463b-87c7-d8edfa19b879 · inbound
Measuring Diversity in Synthetic Datasets Advancing LLM Reasoning Generalists with Preference Trees
Reference 88
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 97156b06-e60e-412f-bc12-c2c7951d7809 · inbound
Video-R1: Reinforcing Video Reasoning in MLLMs Advancing LLM Reasoning Generalists with Preference Trees
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation b527fa5f-5611-4f88-811b-e9b95c6ec9f1 · inbound
InfoPO: On Mutual Information Maximization for Large Language Model Alignment Advancing LLM Reasoning Generalists with Preference Trees
Reference 57
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0e536b2c-b0e1-49c3-a74d-b593f5ba1926 · inbound
GE-Chat: A Graph Enhanced RAG Framework for Evidential Response Generation of LLMs Advancing LLM Reasoning Generalists with Preference Trees
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f2527d3e-6789-464e-b0c6-c7f67fa7a817 · inbound
SEED-GRPO: Semantic Entropy Enhanced GRPO for Uncertainty-Aware Policy Optimization Advancing LLM Reasoning Generalists with Preference Trees
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 551757cb-7d42-4d87-9df2-a53eae1d4dcc · inbound
Rethinking Reward Model Evaluation Through the Lens of Reward Overoptimization Advancing LLM Reasoning Generalists with Preference Trees
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5bab72b9-5519-4cfb-8f12-b78c64860908 · inbound
On the Effect of Negative Gradient in Group Relative Deep Reinforcement Optimization Advancing LLM Reasoning Generalists with Preference Trees
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a0780e28-bddc-4838-9500-e3b1140c06a7 · inbound
Towards Reliable, Uncertainty-Aware Alignment Advancing LLM Reasoning Generalists with Preference Trees
Reference 54
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7053f5d5-dea0-40bc-9503-8717edc46fbd · inbound
SC2Arena and StarEvolve: Benchmark and Self-Improvement Framework for LLMs in Complex Decision-Making Tasks Advancing LLM Reasoning Generalists with Preference Trees
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation de9fd7fd-02f4-48a6-90e0-a0ae86389acd · inbound
Anchoring Refusal Direction: Mitigating Safety Risks in Tuning via Projection Constraint Advancing LLM Reasoning Generalists with Preference Trees
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 66f4bc15-23dc-4d73-aba2-5d4f2540df54 · inbound
MoGU V2: Toward a Higher Pareto Frontier Between Model Usability and Security Advancing LLM Reasoning Generalists with Preference Trees
Reference 66
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b03affbf-8748-48a8-b01b-0021ec3c3e58 · inbound
PaTaRM: Bridging Pairwise and Pointwise Signals via Preference-Aware Task-Adaptive Reward Modeling Advancing LLM Reasoning Generalists with Preference Trees
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation c93977e0-94bb-4c22-9f53-e65977d49b3d · inbound
CompliBench: Benchmarking LLM Judges for Compliance Violation Detection in Dialogue Systems Advancing LLM Reasoning Generalists with Preference Trees
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation af10dbf6-ac5e-4141-b505-c0e394271f2b · inbound
Breaking the Reward Barrier: Accelerating Tree-of-Thought Reasoning via Speculative Exploration Advancing LLM Reasoning Generalists with Preference Trees
Reference 66
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation afbbf399-40d4-47e1-8cc6-6f4a444e99ce · inbound
Breaking the Reward Barrier: Accelerating Tree-of-Thought Reasoning via Speculative Exploration Advancing LLM Reasoning Generalists with Preference Trees
Reference 66
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation c197a95d-51d7-4132-993a-48c8947a0adb · inbound
Holder Policy Optimisation Advancing LLM Reasoning Generalists with Preference Trees
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation cdafa392-70f2-41ba-b055-d46aef9d1d2e · inbound
Holder Policy Optimisation Advancing LLM Reasoning Generalists with Preference Trees
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 0d72bc2b-36a2-4861-a354-1bc2b1e43cf5 · inbound
Selective-Advantage Entropy-Adaptive Horizon GRPO: Asymmetric Token-Level Discounting for Efficient Reinforcement Learning of Language Models Advancing LLM Reasoning Generalists with Preference Trees
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 0a1dc478-abb7-42a0-b754-eaed8449afc6 · inbound
Representation-Aware Advantage Estimation: Your Reward Model Provides More Than A Scalar Output Advancing LLM Reasoning Generalists with Preference Trees
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 35613fd2-39ca-4567-9df6-1dd09537c8c0 · inbound
Architecture-Aware Reinforcement Learning Makes Sliding-Window Attention Competitive in Math Reasoning Advancing LLM Reasoning Generalists with Preference Trees
Reference 66
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation cf321595-60bd-4b5a-a8f3-66f03fa2802c · inbound
Multi-Turn On-Policy Distillation with Prefix Replay Advancing LLM Reasoning Generalists with Preference Trees
Reference 67
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 350e3b83-e95c-497e-96fa-f4c9f3db9a24 · inbound
Multi-Turn On-Policy Distillation with Prefix Replay Advancing LLM Reasoning Generalists with Preference Trees
Reference 68
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bd8c6ada-e011-45c3-ac60-530c11e2b87b · inbound
REOPD: Reliability-Adaptive Reward Extrapolation for On-Policy Distillation Advancing LLM Reasoning Generalists with Preference Trees
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b7b9d316-19c1-4044-9d84-7a9764830712 · inbound
Preference Tree Optimization: Enhancing Goal-Oriented Dialogue with Look-Ahead Simulations Advancing LLM Reasoning Generalists with Preference Trees
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.