Pith. sign in

Paper Citation Record · LEDGER

Advancing LLM Reasoning Generalists with Preference Trees

As of 22 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 41 inbound Pith citation observations for arXiv:2404.02078.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2404.02078 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 41 of 41 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-22T06:32:14.747728+00:00

measured 41 of 41 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-16T00:35:19.143448Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-03T09:47:59.922599Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 751ebca6-97a9-4d89-845a-c0bddda875f0 · inbound

Skywork-Reward: Bag of Tricks for Reward Modeling in LLMs cites this paper.

Skywork-Reward: Bag of Tricks for Reward Modeling in LLMs Advancing LLM Reasoning Generalists with Preference Trees

Reference 25

Resolution
verified exact
arxiv_id, observed 2026-05-17T16:18:01.688186Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-17T16:18:01.560780Z digest=sha256:7dfb2887d67edf20394af4ed7f1cbbc5594542afa8a07b388c8ca356db6b659f

Observation 0b4d4772-3d26-4719-94d0-a63a35b42b18 · inbound

Towards Adaptive Mechanism Activation in Language Agent cites this paper.

Towards Adaptive Mechanism Activation in Language Agent Advancing LLM Reasoning Generalists with Preference Trees

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-12T05:09:47.983171Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T05:09:47.983171Z digest=sha256:cf5020879ef658bcf50ec6a37d08bebc14d861d463ee711524aca86fed7a2fdf

Observation acca620a-0b77-43c2-9e4b-578ac7f043d1 · inbound

Large Language Models for Scholarly Ontology Generation: An Extensive Analysis in the Engineering Field cites this paper.

Large Language Models for Scholarly Ontology Generation: An Extensive Analysis in the Engineering Field Advancing LLM Reasoning Generalists with Preference Trees

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-11T18:05:58.788948Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T18:05:58.788948Z digest=sha256:01a7d45f7c3c2c1b4634e9090c9f9af90ca898842a0f18ab16480050e65b56f7

Observation 6099d9b7-f6b1-4abf-b6d6-bd8e7a675af1 · inbound

JuStRank: Benchmarking LLM Judges for System Ranking cites this paper.

JuStRank: Benchmarking LLM Judges for System Ranking Advancing LLM Reasoning Generalists with Preference Trees

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-11T16:59:49.654539Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T16:59:49.654539Z digest=sha256:210a5b8313ff2387a83ba598aaf589de67196e15db584540a10cef318d4d24ee

Observation 99b3e0ae-f37d-4861-986f-cdd089439c71 · inbound

Cal-DPO: Calibrated Direct Preference Optimization for Language Model Alignment cites this paper.

Cal-DPO: Calibrated Direct Preference Optimization for Language Model Alignment Advancing LLM Reasoning Generalists with Preference Trees

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-11T12:16:31.827786Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:16:31.827786Z digest=sha256:4374454ca0124eb294ad665a7666d85d58d00f0cdfa142e03041a1d309fe87ce

Observation f85fa636-c19f-47b0-ae73-25e54172b216 · inbound

Progressive Multimodal Reasoning via Active Retrieval cites this paper.

Progressive Multimodal Reasoning via Active Retrieval Advancing LLM Reasoning Generalists with Preference Trees

Reference 117

Resolution
unresolved
no resolver link, observed 2026-08-11T11:55:10.604222Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:55:10.604222Z digest=sha256:c8755a648165aa39b22fe4f58f6ee7c044ca1595b24e2228a0e07d52a52f3582

Observation cd143d7e-9622-438d-a359-28b3a2e8a701 · inbound

AceMath: Advancing Frontier Math Reasoning with Post-Training and Reward Modeling cites this paper.

AceMath: Advancing Frontier Math Reasoning with Post-Training and Reward Modeling Advancing LLM Reasoning Generalists with Preference Trees

Reference 74

Resolution
unresolved
no resolver link, observed 2026-08-11T11:43:24.886660Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T11:43:24.886660Z digest=sha256:e5990dd9e9690be92c9fc72e7c7cb5d9bfcf17d76afffa0c7ce776e3b3c5b8f5

Observation fe2c7e9f-6eeb-4093-b116-adbf38238b95 · inbound

Training Software Engineering Agents and Verifiers with SWE-Gym cites this paper.

Training Software Engineering Agents and Verifiers with SWE-Gym Advancing LLM Reasoning Generalists with Preference Trees

Reference 9

Resolution
metadata mismatch
arxiv_id, observed 2026-05-18T05:20:40.532699Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-18T05:20:40.483057Z digest=sha256:b819af4dd4fde4c6890013ab56a8095fa563a5c92567c9b04d401b21a4b91762

Observation af2455f1-6d70-40e0-a5f2-6f8daa8f7209 · inbound

LLM-Virus: Evolutionary Jailbreak Attack on Large Language Models cites this paper.

LLM-Virus: Evolutionary Jailbreak Attack on Large Language Models Advancing LLM Reasoning Generalists with Preference Trees

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-10T23:40:38.766040Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:40:38.766040Z digest=sha256:9281cefdfee4e44df0c7a78acae470ed4f49ef87b21f78f150abb72d88734033

Observation a1ed69d1-9d74-414b-9184-f965fb0de936 · inbound

AlphaPO: Reward Shape Matters for LLM Alignment cites this paper.

AlphaPO: Reward Shape Matters for LLM Alignment Advancing LLM Reasoning Generalists with Preference Trees

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-10T21:51:09.374753Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T21:51:09.374753Z digest=sha256:70bdf7a1245f13ad8691987f5d73c714393506c19f974c36341eac0f6a4a722f

Observation 2dfc7745-fffc-455d-94a0-c9c1e4ca79cd · inbound

VidChain: Chain-of-Tasks with Metric-based Direct Preference Optimization for Dense Video Captioning cites this paper.

VidChain: Chain-of-Tasks with Metric-based Direct Preference Optimization for Dense Video Captioning Advancing LLM Reasoning Generalists with Preference Trees

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-10T20:54:13.330788Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T20:54:13.330788Z digest=sha256:c4a8c728327159dbe3e9cddbc9285f1581ee5ea0c7bc2bfa3098971c14dfc677

Observation 4a54c6a8-65d1-4f45-b9b5-e850d0fd7f98 · inbound

VideoWorld: Exploring Knowledge Learning from Unlabeled Videos cites this paper.

VideoWorld: Exploring Knowledge Learning from Unlabeled Videos Advancing LLM Reasoning Generalists with Preference Trees

Reference 74

Resolution
unresolved
no resolver link, observed 2026-08-10T19:49:20.138864Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T19:49:20.138864Z digest=sha256:e57b1021ee81d593fb165d68fef25ca9d56f2514d1f646bb53c9618102682b45

Observation 3a07252d-3349-4ce5-94f7-05ec9388a8c1 · inbound

Improving Influence-based Instruction Tuning Data Selection for Balanced Learning of Diverse Capabilities cites this paper.

Improving Influence-based Instruction Tuning Data Selection for Balanced Learning of Diverse Capabilities Advancing LLM Reasoning Generalists with Preference Trees

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-10T17:34:52.529423Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T17:34:52.529423Z digest=sha256:ed7bf8480f66b8405d2421bb95caca7c0edfa36be8ad4ad01861b6288995d00b

Observation 96d4f58c-83e4-4ec7-a4e3-d052d399ecbb · inbound

PairJudge RM: Perform Best-of-N Sampling with Knockout Tournament cites this paper.

PairJudge RM: Perform Best-of-N Sampling with Knockout Tournament Advancing LLM Reasoning Generalists with Preference Trees

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-10T16:39:31.274704Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T16:39:31.274704Z digest=sha256:8f037ed43408540330c58809e7a48a041a6827818f66f95ae2874e29744557e4

Observation ae83feb8-fc69-41a7-a818-32a886b08617 · inbound

Eagle 2: Building Post-Training Data Strategies from Scratch for Frontier Vision-Language Models cites this paper.

Eagle 2: Building Post-Training Data Strategies from Scratch for Frontier Vision-Language Models Advancing LLM Reasoning Generalists with Preference Trees

Reference 171

Resolution
unresolved
no resolver link, observed 2026-08-10T18:04:34.732726Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T18:04:34.732726Z digest=sha256:a25e6da78c2a7a1fa613bfd0ea2dba697c97c651ed1baee861ffc599fcabd5cb

Observation bbb07ed4-5c56-43c8-9149-8d4d764f1977 · inbound

SmolLM2: When Smol Goes Big -- Data-Centric Training of a Small Language Model cites this paper.

SmolLM2: When Smol Goes Big -- Data-Centric Training of a Small Language Model Advancing LLM Reasoning Generalists with Preference Trees

Reference 249

Resolution
metadata mismatch
arxiv_id, observed 2026-05-13T17:30:03.070601Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-05-13T17:30:02.803757Z digest=sha256:1d31780b43d634669142850c5e9e001f432bd0615143ccfd6fa47c43c0c0e30c

Observation 1403d4b0-07a5-4db9-9c9d-63e497ddc699 · inbound

DPO-Shift: Shifting the Distribution of Direct Preference Optimization cites this paper.

DPO-Shift: Shifting the Distribution of Direct Preference Optimization Advancing LLM Reasoning Generalists with Preference Trees

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-08T12:17:46.033521Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T12:17:46.033521Z digest=sha256:b2538218ae59b649dab2812a892cc8bc6f9f36955014bca571d874735b061c55

Observation 165f613f-8b50-463b-87c7-d8edfa19b879 · inbound

Measuring Diversity in Synthetic Datasets cites this paper.

Measuring Diversity in Synthetic Datasets Advancing LLM Reasoning Generalists with Preference Trees

Reference 88

Resolution
unresolved
no resolver link, observed 2026-08-08T04:54:50.972531Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T04:54:50.972531Z digest=sha256:b0b301c52ba01bdd18a6f0e32f2a1c04dd03dd4f7cd5bef0bc5a8758168aefd1

Observation 97156b06-e60e-412f-bc12-c2c7951d7809 · inbound

Video-R1: Reinforcing Video Reasoning in MLLMs cites this paper.

Video-R1: Reinforcing Video Reasoning in MLLMs Advancing LLM Reasoning Generalists with Preference Trees

Reference 42

Resolution
verified exact
arxiv_id, observed 2026-05-12T09:43:00.303113Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-12T09:43:00.208065Z digest=sha256:9e788903978b6180242c956df8636484a44b4fa923fb9357c489bf3b38bc0d2a

Observation b527fa5f-5611-4f88-811b-e9b95c6ec9f1 · inbound

InfoPO: On Mutual Information Maximization for Large Language Model Alignment cites this paper.

InfoPO: On Mutual Information Maximization for Large Language Model Alignment Advancing LLM Reasoning Generalists with Preference Trees

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-15T22:01:30.142617Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T22:01:30.142617Z digest=sha256:d0c4be08e146dbdff8a5a21908258975033060fc6c508e219340e52432f5280b

Observation 0e536b2c-b0e1-49c3-a74d-b593f5ba1926 · inbound

GE-Chat: A Graph Enhanced RAG Framework for Evidential Response Generation of LLMs cites this paper.

GE-Chat: A Graph Enhanced RAG Framework for Evidential Response Generation of LLMs Advancing LLM Reasoning Generalists with Preference Trees

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-15T21:18:38.578510Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:18:38.578510Z digest=sha256:59545c9fab2301b72a6a6fd9f344d1a2f59eb0a2ed4796c671a09595ad60c1ff

Observation f2527d3e-6789-464e-b0c6-c7f67fa7a817 · inbound

SEED-GRPO: Semantic Entropy Enhanced GRPO for Uncertainty-Aware Policy Optimization cites this paper.

SEED-GRPO: Semantic Entropy Enhanced GRPO for Uncertainty-Aware Policy Optimization Advancing LLM Reasoning Generalists with Preference Trees

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-15T20:40:13.147339Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:40:13.147339Z digest=sha256:151d2b21a4487c3c899679fb767f624d00b2a104eea9ffae98edf0d71f3e790f

Observation 551757cb-7d42-4d87-9df2-a53eae1d4dcc · inbound

Rethinking Reward Model Evaluation Through the Lens of Reward Overoptimization cites this paper.

Rethinking Reward Model Evaluation Through the Lens of Reward Overoptimization Advancing LLM Reasoning Generalists with Preference Trees

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-15T20:34:32.707744Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:34:32.707744Z digest=sha256:bcff8aff8bee414d833e8e84427282f7c5a1a11a3b0ffa29c3779021a1498821

Observation 5bab72b9-5519-4cfb-8f12-b78c64860908 · inbound

On the Effect of Negative Gradient in Group Relative Deep Reinforcement Optimization cites this paper.

On the Effect of Negative Gradient in Group Relative Deep Reinforcement Optimization Advancing LLM Reasoning Generalists with Preference Trees

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T14:29:33.431718Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:29:33.431718Z digest=sha256:91e5c88e768d1836a5f20b0b99baf1e91b254006d1e8cad4cb4851df21a4cddd

Observation a0780e28-bddc-4838-9500-e3b1140c06a7 · inbound

Towards Reliable, Uncertainty-Aware Alignment cites this paper.

Towards Reliable, Uncertainty-Aware Alignment Advancing LLM Reasoning Generalists with Preference Trees

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-06T15:40:17.967384Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T15:40:17.967384Z digest=sha256:adb8f55cc0096a7a8112d77a946e0f8e0da6f840dfd05c59043419d1e9b5be6e

Observation 7053f5d5-dea0-40bc-9503-8717edc46fbd · inbound

SC2Arena and StarEvolve: Benchmark and Self-Improvement Framework for LLMs in Complex Decision-Making Tasks cites this paper.

SC2Arena and StarEvolve: Benchmark and Self-Improvement Framework for LLMs in Complex Decision-Making Tasks Advancing LLM Reasoning Generalists with Preference Trees

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-05T20:30:30.162000Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T20:30:30.162000Z digest=sha256:7c50255508eba7ea8b20ca9adb0bee4a14dc979cfacc5b91929f59120a69e2ef

Observation de9fd7fd-02f4-48a6-90e0-a0ae86389acd · inbound

Anchoring Refusal Direction: Mitigating Safety Risks in Tuning via Projection Constraint cites this paper.

Anchoring Refusal Direction: Mitigating Safety Risks in Tuning via Projection Constraint Advancing LLM Reasoning Generalists with Preference Trees

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-04T23:10:21.344039Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T23:10:21.344039Z digest=sha256:1ce09ecefc8b5c470afc52508e0be1aff2d60f3f456db0c9d2caafc01e976c2e

Observation 66f4bc15-23dc-4d73-aba2-5d4f2540df54 · inbound

MoGU V2: Toward a Higher Pareto Frontier Between Model Usability and Security cites this paper.

MoGU V2: Toward a Higher Pareto Frontier Between Model Usability and Security Advancing LLM Reasoning Generalists with Preference Trees

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-04T23:09:41.578345Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T23:09:41.578345Z digest=sha256:c6b38e26cf796e07a51355d885aa5d63969f375cc51966649ffd9b2f266d1617

Observation b03affbf-8748-48a8-b01b-0021ec3c3e58 · inbound

PaTaRM: Bridging Pairwise and Pointwise Signals via Preference-Aware Task-Adaptive Reward Modeling cites this paper.

PaTaRM: Bridging Pairwise and Pointwise Signals via Preference-Aware Task-Adaptive Reward Modeling Advancing LLM Reasoning Generalists with Preference Trees

Reference 22

Resolution
metadata mismatch
arxiv_id, observed 2026-05-18T03:20:49.262191Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-18T03:15:57.744706Z digest=sha256:8a539b8200f99a3fac2cbdcc0c4bd361edc69414c03e1143dd1b3b9f0cc490b3

Observation c93977e0-94bb-4c22-9f53-e65977d49b3d · inbound

CompliBench: Benchmarking LLM Judges for Compliance Violation Detection in Dialogue Systems cites this paper.

CompliBench: Benchmarking LLM Judges for Compliance Violation Detection in Dialogue Systems Advancing LLM Reasoning Generalists with Preference Trees

Reference 21

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T11:26:02.393606Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-10T14:56:00.449776Z digest=sha256:fd2886416359d4bd14ebe6dcc87eda3567d6484c9e2b82629e423d205bc35fb4

Observation af10dbf6-ac5e-4141-b505-c0e394271f2b · inbound

Breaking the Reward Barrier: Accelerating Tree-of-Thought Reasoning via Speculative Exploration cites this paper.

Breaking the Reward Barrier: Accelerating Tree-of-Thought Reasoning via Speculative Exploration Advancing LLM Reasoning Generalists with Preference Trees

Reference 66

Resolution
verified exact
arxiv_id, observed 2026-05-12T06:51:30.256032Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-12T03:51:52.375703Z digest=sha256:a9744d8fcfd46d3f85cbd6e75ccf0c00b8b145f16edc9e8b40a25f2cd5b28e14

Observation afbbf399-40d4-47e1-8cc6-6f4a444e99ce · inbound

Breaking the Reward Barrier: Accelerating Tree-of-Thought Reasoning via Speculative Exploration cites this paper.

Breaking the Reward Barrier: Accelerating Tree-of-Thought Reasoning via Speculative Exploration Advancing LLM Reasoning Generalists with Preference Trees

Reference 66

Resolution
verified exact
arxiv_id, observed 2026-05-15T05:15:03.285300Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-15T05:11:32.053440Z digest=sha256:6b1e9e37351b10a8fb9e93568cdbd622b548d439ad328b820320f35bac77cfb6

Observation c197a95d-51d7-4132-993a-48c8947a0adb · inbound

Holder Policy Optimisation cites this paper.

Holder Policy Optimisation Advancing LLM Reasoning Generalists with Preference Trees

Reference 32

Resolution
verified exact
arxiv_id, observed 2026-05-13T06:12:22.771833Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-05-13T06:08:28.855671Z digest=sha256:6dfee3dae96f73bcf9e0c11e3b561dbfb382f4397e70ba980adc45cd95ea31b4

Observation cdafa392-70f2-41ba-b055-d46aef9d1d2e · inbound

Holder Policy Optimisation cites this paper.

Holder Policy Optimisation Advancing LLM Reasoning Generalists with Preference Trees

Reference 32

Resolution
verified exact
arxiv_id, observed 2026-05-22T10:01:23.055182Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-05-22T10:00:58.600743Z digest=sha256:9ad84a98f4d1f7d82170e4eede66dfcebe4b2d5b4889cbfc8987ad56c476a779

Observation 0d72bc2b-36a2-4861-a354-1bc2b1e43cf5 · inbound

Selective-Advantage Entropy-Adaptive Horizon GRPO: Asymmetric Token-Level Discounting for Efficient Reinforcement Learning of Language Models cites this paper.

Selective-Advantage Entropy-Adaptive Horizon GRPO: Asymmetric Token-Level Discounting for Efficient Reinforcement Learning of Language Models Advancing LLM Reasoning Generalists with Preference Trees

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-07-02T07:26:46.066822Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-06-28T06:56:01.601701Z digest=sha256:6526c31d607651d8e61804f7822cfca8028f7ba9f18c1456ab50cb5649edb06a

Observation 0a1dc478-abb7-42a0-b754-eaed8449afc6 · inbound

Representation-Aware Advantage Estimation: Your Reward Model Provides More Than A Scalar Output cites this paper.

Representation-Aware Advantage Estimation: Your Reward Model Provides More Than A Scalar Output Advancing LLM Reasoning Generalists with Preference Trees

Reference 27

Resolution
verified exact
arxiv_id, observed 2026-07-03T04:47:38.255909Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-06-27T13:39:17.701196Z digest=sha256:ccd7e56ffac039fef2da8056d12d79bc5cf359237724a5ca237b648796981da9

Observation 35613fd2-39ca-4567-9df6-1dd09537c8c0 · inbound

Architecture-Aware Reinforcement Learning Makes Sliding-Window Attention Competitive in Math Reasoning cites this paper.

Architecture-Aware Reinforcement Learning Makes Sliding-Window Attention Competitive in Math Reasoning Advancing LLM Reasoning Generalists with Preference Trees

Reference 66

Resolution
verified exact
arxiv_id, observed 2026-07-03T09:47:59.923933Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-06-27T10:18:54.163862Z digest=sha256:203e583fb6bcab18160aad294a2f0f5acfa90b23382b6b8f869acfed495b0a66

Observation cf321595-60bd-4b5a-a8f3-66f03fa2802c · inbound

Multi-Turn On-Policy Distillation with Prefix Replay cites this paper.

Multi-Turn On-Policy Distillation with Prefix Replay Advancing LLM Reasoning Generalists with Preference Trees

Reference 67

Resolution
unresolved
no resolver link, observed 2026-07-11T13:53:36.775836Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-11T13:53:36.775836Z digest=sha256:7e40a765bdaadb6c100feeaf73f577db9fe6ac77e3063999f389d5d671a20db1

Observation 350e3b83-e95c-497e-96fa-f4c9f3db9a24 · inbound

Multi-Turn On-Policy Distillation with Prefix Replay cites this paper.

Multi-Turn On-Policy Distillation with Prefix Replay Advancing LLM Reasoning Generalists with Preference Trees

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-02T08:40:39.023948Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T08:40:39.023948Z digest=sha256:3452f1e0f4e9e8e512197ac0510cb4f376747abab81d06fa4a5d3fc2d3d00b70

Observation bd8c6ada-e011-45c3-ac60-530c11e2b87b · inbound

REOPD: Reliability-Adaptive Reward Extrapolation for On-Policy Distillation cites this paper.

REOPD: Reliability-Adaptive Reward Extrapolation for On-Policy Distillation Advancing LLM Reasoning Generalists with Preference Trees

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-16T00:35:19.143448Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:35:19.143448Z digest=sha256:7e297a4a3dd0f5e2bbd24c74d918de5fbab8dc3eb19a6cb398ed02c4d74a8d79

Observation b7b9d316-19c1-4044-9d84-7a9764830712 · inbound

Preference Tree Optimization: Enhancing Goal-Oriented Dialogue with Look-Ahead Simulations cites this paper.

Preference Tree Optimization: Enhancing Goal-Oriented Dialogue with Look-Ahead Simulations Advancing LLM Reasoning Generalists with Preference Trees

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-16T00:21:59.263922Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T00:21:59.263922Z digest=sha256:5bcd9c901edc3a0f2b93106c580ed451b16e86d9074f3f534b3cff225ae0dabe