Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
Paper Citation Record · LEDGER
As of 14 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 31 inbound Pith citation observations for arXiv:2406.12845.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-14T06:32:32.682623+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-12T20:31:53.198146Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-07-04T13:49:52.398618Z
0 of 0 outbound references displayed
External citation measurements
No source-named external measurement is stored.
No outbound reference observations are available for this paper version.
Observation 65ddd036-d75f-47de-afe4-99952039e0d8 · inbound
Adaptive Decoding via Latent Preference Optimization Interpretable Preferences via Multi-Objective Reward Modeling and Mixture-of-Experts
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6586546e-2a46-48dc-8144-e1e528e25074 · inbound
Reward Modeling with Ordinal Feedback: Wisdom of the Crowd Interpretable Preferences via Multi-Objective Reward Modeling and Mixture-of-Experts
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 179382e3-09a0-4fee-93bb-aa3b7174946c · inbound
Interpreting Language Reward Models via Contrastive Explanations Interpretable Preferences via Multi-Objective Reward Modeling and Mixture-of-Experts
Reference 2017
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ae6ea1ec-a58e-4b27-bf42-0d54fce4fbca · inbound
T-REG: Preference Optimization with Token-Level Reward Regularization Interpretable Preferences via Multi-Objective Reward Modeling and Mixture-of-Experts
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4bbd1de7-89bb-4245-845b-a3ce80ee060e · inbound
Boosting LLM via Learning from Data Iteratively and Selectively Interpretable Preferences via Multi-Objective Reward Modeling and Mixture-of-Experts
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 76dd9384-d843-481b-9c36-4ba97b5fbc21 · inbound
CoT-based Synthesizer: Enhancing LLM Performance through Answer Synthesis Interpretable Preferences via Multi-Objective Reward Modeling and Mixture-of-Experts
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ff6c558d-bff6-401c-8d5d-ec1eeff439c6 · inbound
AlphaPO: Reward Shape Matters for LLM Alignment Interpretable Preferences via Multi-Objective Reward Modeling and Mixture-of-Experts
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e7092d19-35bc-4643-861b-ceaa4b504d6c · inbound
InternLM-XComposer2.5-Reward: A Simple Yet Effective Multi-Modal Reward Model Interpretable Preferences via Multi-Objective Reward Modeling and Mixture-of-Experts
Reference 89
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 589cc8e1-0d44-499e-99c3-9300818f6f27 · inbound
Data-adaptive Safety Rules for Training Reward Models Interpretable Preferences via Multi-Objective Reward Modeling and Mixture-of-Experts
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 59b02dbf-8291-4708-9a20-58cfbc4b8699 · inbound
Mixture of Experts (MoE): A Big Data Perspective Interpretable Preferences via Multi-Objective Reward Modeling and Mixture-of-Experts
Reference 164
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b1bdb19a-0c51-4eb1-9ac9-686627d8924a · inbound
Atla Selene Mini: A General Purpose Evaluation Model Interpretable Preferences via Multi-Objective Reward Modeling and Mixture-of-Experts
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6b6b7f60-0c5e-4956-a978-5c3389a3891c · inbound
R.I.P.: Better Models by Survival of the Fittest Prompts Interpretable Preferences via Multi-Objective Reward Modeling and Mixture-of-Experts
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 338590b3-d24c-4faf-bf56-9d251942cf9c · inbound
MTSA: Multi-turn Safety Alignment for LLMs through Multi-round Red-teaming Interpretable Preferences via Multi-Objective Reward Modeling and Mixture-of-Experts
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7323d348-b0fc-4847-9203-2a850d839a04 · inbound
MOSLIM:Align with diverse preferences in prompts through reward classification Interpretable Preferences via Multi-Objective Reward Modeling and Mixture-of-Experts
Reference 2020
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4be40304-6593-490b-9633-e738ab1320a2 · inbound
Accelerating RLHF Training with Reward Variance Increase Interpretable Preferences via Multi-Objective Reward Modeling and Mixture-of-Experts
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 60e0b91e-3536-448c-ad4b-885071fab307 · inbound
RewardBench 2: Advancing Reward Model Evaluation Interpretable Preferences via Multi-Objective Reward Modeling and Mixture-of-Experts
Reference 57
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 875c2e27-052c-4e7f-bbbd-d4b8646455d2 · inbound
Well Begun is Half Done: Low-resource Preference Alignment by Weak-to-Strong Decoding Interpretable Preferences via Multi-Objective Reward Modeling and Mixture-of-Experts
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e82a1853-8285-4461-8166-fcdc7998f07b · inbound
DuaShepherd: Integrating Stepwise Correctness and Potential Rewards for Mathematical Reasoning Interpretable Preferences via Multi-Objective Reward Modeling and Mixture-of-Experts
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cceac6b1-0107-4437-8266-c51db615d615 · inbound
BEST-Route: Adaptive LLM Routing with Test-Time Optimal Compute Interpretable Preferences via Multi-Objective Reward Modeling and Mixture-of-Experts
Reference 49
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 849056ab-8732-4605-8bed-aa8161486bed · inbound
BMMR: A Large-Scale Bilingual Multimodal Multi-Discipline Reasoning Dataset Interpretable Preferences via Multi-Objective Reward Modeling and Mixture-of-Experts
Reference 53
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b06ce073-c8ad-4ce4-bb51-450be7245415 · inbound
T-POP: Test-Time Personalization with Online Preference Feedback Interpretable Preferences via Multi-Objective Reward Modeling and Mixture-of-Experts
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5daa66ef-e109-4667-939c-3f89ab5c0f62 · inbound
Simultaneous Multi-objective Alignment Across Verifiable and Non-verifiable Rewards Interpretable Preferences via Multi-Objective Reward Modeling and Mixture-of-Experts
Reference 58
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 73b6f605-1e24-4871-bb6d-233259723df4 · inbound
On the Overscaling Curse of Parallel Thinking: System Efficacy Contradicts Sample Efficiency Interpretable Preferences via Multi-Objective Reward Modeling and Mixture-of-Experts
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 78c93807-1306-4273-b3fb-328262e3caf1 · inbound
When Errors Can Be Beneficial: A Categorization of Imperfect Rewards for Policy Gradient Interpretable Preferences via Multi-Objective Reward Modeling and Mixture-of-Experts
Reference 88
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation aa5406f5-3998-43dc-a48e-2dcc17ae23a3 · inbound
OpsLLM: Construction of Large Language Model for Software Operations with Multi-stage Learning Interpretable Preferences via Multi-Objective Reward Modeling and Mixture-of-Experts
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation ef04ec64-c2e6-464d-91e8-d7ff1db14216 · inbound
OpsLLM: Construction of Large Language Model for Software Operations with Multi-stage Learning Interpretable Preferences via Multi-Objective Reward Modeling and Mixture-of-Experts
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 684ed106-daed-4f8e-9b11-24a1aced04e7 · inbound
OpsLLM: Construction of Large Language Model for Software Operations with Multi-stage Learning Interpretable Preferences via Multi-Objective Reward Modeling and Mixture-of-Experts
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 881263e6-a814-4432-ba11-50152e4a4d62 · inbound
Explaining and Breaking the Safety-Helpfulness Ceiling via Preference Dimensional Expansion Interpretable Preferences via Multi-Objective Reward Modeling and Mixture-of-Experts
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation afc53f3c-b12b-409c-94b1-edb69dece7a5 · inbound
Explaining and Breaking the Safety-Helpfulness Ceiling via Preference Dimensional Expansion Interpretable Preferences via Multi-Objective Reward Modeling and Mixture-of-Experts
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation aae0c89a-f6f3-440e-8496-99ba88e4727b · inbound
GEOALIGN: Geometric Rollout Curation for Robust LLM Reinforcement Learning Interpretable Preferences via Multi-Objective Reward Modeling and Mixture-of-Experts
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 863bc362-198e-4315-9319-54c2924a9a3c · inbound
Chain-of-Models: Cross-Model Auditing for Bias-Robust LLM Judges Interpretable Preferences via Multi-Objective Reward Modeling and Mixture-of-Experts
Reference 64
Source-reported events for the cited work
Unavailable: canonical work link unavailable.