Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
Paper Citation Record · LEDGER
As of 10 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 5 inbound Pith citation observations for arXiv:2406.07295.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-09T12:11:21.962879Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-05-21T06:44:00.890441Z
0 of 0 outbound references displayed
External citation measurements
No source-named external measurement is stored.
No outbound reference observations are available for this paper version.
Observation e37066ae-2d7d-40b8-8b8e-00bec3edcc6f · inbound
e-SimFT: Alignment of Generative Models with Simulation Feedback for Pareto-Front Design Exploration Multi-objective Reinforcement learning from AI Feedback
Reference 2020
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 07800d77-66de-4df0-9f86-f2063f4be016 · inbound
Pareto-Optimal Offline Reinforcement Learning via Smooth Tchebysheff Scalarization Multi-objective Reinforcement learning from AI Feedback
Reference 89
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation e39d376c-1f67-42ca-82c6-8f88f92173b4 · inbound
Not Every Rubric Teaches Equally: Policy-Aware Rubric Rewards for RLVR Multi-objective Reinforcement learning from AI Feedback
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation fee51cc4-a833-44f0-a300-61d4c998fe69 · inbound
SURF: Steering the Scalarization Weight to Uniformly Traverse the Pareto Front Multi-objective Reinforcement learning from AI Feedback
Reference 96
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 53e6144e-33e8-454b-b059-9793794acb39 · inbound
Don't Mix Rewards, Mix Policies: Policy Decomposition and Optimization for Multi-Reward RL Multi-objective Reinforcement learning from AI Feedback
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.