Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-06T16:17:35.918503Z
Paper Citation Record · LEDGER
As of 8 August 2026, this Paper Citation Record lists 55 of 55 outbound references and 0 inbound Pith citation observations for arXiv:2507.14066.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-06T16:17:35.918503Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
55 of 55 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 9fc1a82f-f4fb-4255-be81-4d7da518b296 · outbound
Preference-based Multi-Objective Reinforcement Learning A survey on modeling and optimizing multi-objective systems,
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation ae06a308-6755-455b-a074-e2e0c3686134 · outbound
Preference-based Multi-Objective Reinforcement Learning Active learning with fairness-aware clustering for fair classification considering multiple sensitive attributes,
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 0c8692a9-07b7-45d4-b321-303a35428fbd · outbound
Preference-based Multi-Objective Reinforcement Learning Constrained ordinal opti- mization—a feasibility model based approach,
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation c696ac33-1444-4b15-8189-5e10c27456d2 · outbound
Preference-based Multi-Objective Reinforcement Learning Reliability/cost-based multi-objective pareto optimal design of stand- alone wind/pv/fc generation microgrid system,
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 2d48f134-df5b-413c-a3aa-0c3ee75c410e · outbound
Preference-based Multi-Objective Reinforcement Learning Toward personalized decision making for autonomous vehicles: a constrained multi-objective reinforcement learning tech- nique,
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 716eb782-3d6c-4a6c-8ed9-1be6e67e00ff · outbound
Preference-based Multi-Objective Reinforcement Learning B-pref: Benchmarking preference-based reinforcement learning,
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 9e9b1cfd-c2d6-4a14-9fc6-cbd0d5bc2d51 · outbound
Preference-based Multi-Objective Reinforcement Learning Self- supervised online reward shaping in sparse-reward environments,
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 13ceb81f-37c5-41c2-83c9-72311c454326 · outbound
Preference-based Multi-Objective Reinforcement Learning Reward shaping for knowledge-based multi-objective multi-agent reinforcement learning,
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation a2583ef5-ca33-4e10-82ce-64ef063c2802 · outbound
Preference-based Multi-Objective Reinforcement Learning Pebble: Feedback-efficient interactive reinforcement learning via relabeling experience and unsuper- vised pre-training,
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 4537b1db-9c38-4270-ae4b-71a51d79da40 · outbound
Preference-based Multi-Objective Reinforcement Learning Reinforcement learning and the reward engineering princi- ple,
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 24ca2cd7-5064-40f5-99ea-12b31bbaef82 · outbound
Preference-based Multi-Objective Reinforcement Learning Deep reinforcement learning from human preferences,
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b220a6de-e529-44a7-a136-6f85cbaf7a7f · outbound
Preference-based Multi-Objective Reinforcement Learning A bayesian approach for policy learning from trajectory preference queries,
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 5e73303c-0400-4634-b2d6-680945ac21e7 · outbound
Preference-based Multi-Objective Reinforcement Learning Unresolved cited work
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 59e50602-f0f1-4bc3-8a65-75e1f0d8e058 · outbound
Preference-based Multi-Objective Reinforcement Learning Playing Atari with Deep Reinforcement Learning
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 45c39390-16cf-419b-8bc3-0cf91b5b9b7f · outbound
Preference-based Multi-Objective Reinforcement Learning E-mapp: Efficient multi-agent reinforcement learning with parallel program guidance,
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 188d3ba9-859c-4067-93ac-7c801054fb7e · outbound
Preference-based Multi-Objective Reinforcement Learning Mastering the game of go without human knowledge,
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 191fe72d-0e3f-4e18-92d0-ca2e6acdd8a8 · outbound
Preference-based Multi-Objective Reinforcement Learning Simplify twin crane scheduling in railway yard by spatial task assignment,
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation b5bb24a4-fef6-4148-a288-c2e0654dcf78 · outbound
Preference-based Multi-Objective Reinforcement Learning Large-scale data center cooling control via sample-efficient reinforcement learning,
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation a82b67cb-516f-4951-92b1-4de16f88d1de · outbound
Preference-based Multi-Objective Reinforcement Learning An efficient real- time railway container yard management method based on partial de- coupling,
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 97ff2a81-6704-405c-af47-15839da3abfc · outbound
Preference-based Multi-Objective Reinforcement Learning Integrating mechanism and data: Rein- forcement learning based on multi-fidelity model for data center cooling control,
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 1dffd4ba-39e9-4fea-b1e1-816f02af9399 · outbound
Preference-based Multi-Objective Reinforcement Learning Incentive-oriented power-carbon emissions trading-tradable green certificate integrated market mecha- nisms using multi-agent deep reinforcement learning,
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 7786b8cc-d84f-4ec2-a684-b758ac9015f7 · outbound
Preference-based Multi-Objective Reinforcement Learning OpenAI Gym
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7be9d931-503a-43e0-9f1a-065a9f2119ae · outbound
Preference-based Multi-Objective Reinforcement Learning Exploration by random network distillation,
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 88d8419b-af3a-4b5b-8697-f17d167b40be · outbound
Preference-based Multi-Objective Reinforcement Learning Direct preference optimization: Your language model is secretly a reward model,
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2af1f740-84f9-45f4-8b03-a71c0e055626 · outbound
Preference-based Multi-Objective Reinforcement Learning Reward learning from human preferences and demonstrations in atari,
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 86d955f6-956d-453c-b71f-bf247c652a92 · outbound
Preference-based Multi-Objective Reinforcement Learning S-EPOA: Overcoming the Indistinguishability of Segments with Skill-Driven Preference-Based Reinforcement Learning
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 8e1175aa-1ece-41da-a025-3a324564b00f · outbound
Preference-based Multi-Objective Reinforcement Learning Surf: Semi- supervised reward learning with data augmentation for feedback-efficient preference-based reinforcement learning,
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation b70ba05f-268c-45fb-9cf8-5e797f7c0d36 · outbound
Preference-based Multi-Objective Reinforcement Learning Few-shot preference learning for human- in-the-loop rl,
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation ad4b84f9-c650-4b4c-a6a4-c2e3eba99b98 · outbound
Preference-based Multi-Objective Reinforcement Learning Learning to summarize with human feedback,
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 468c6574-e772-4935-be5d-7550d1d20647 · outbound
Preference-based Multi-Objective Reinforcement Learning A toolkit for reliable benchmarking and research in multi-objective reinforcement learning,
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 112be415-f344-4a58-af0f-03915d8a1d6e · outbound
Preference-based Multi-Objective Reinforcement Learning A practical guide to multi-objective reinforcement learning and planning,
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 420d363e-a9f4-47b2-bd3e-6ef4f759e91c · outbound
Preference-based Multi-Objective Reinforcement Learning Human-in-the-Loop Policy Optimization for Preference-Based Multi-Objective Reinforcement Learning
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 2712168f-f393-4094-ab35-cbfe5a78983a · outbound
Preference-based Multi-Objective Reinforcement Learning A generalized algorithm for multi-objective reinforcement learning and policy adaptation,
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 9eb1a771-cbb8-4150-a54b-76732f7c49cf · outbound
Preference-based Multi-Objective Reinforcement Learning Multi-objective rein- forcement learning for the expected utility of the return,
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation e893cdc1-7bfb-4641-8956-8dd0749bdc23 · outbound
Preference-based Multi-Objective Reinforcement Learning Prediction- guided multi-objective reinforcement learning for continuous robot con- trol,
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a2dca11c-79cc-4b92-a95e-a824b1f95d7a · outbound
Preference-based Multi-Objective Reinforcement Learning Pareto conditioned net- works,
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 1386df91-dd9e-4e2f-b578-ea283fac0c47 · outbound
Preference-based Multi-Objective Reinforcement Learning Sample-efficient multi-objective learning via generalized pol- icy improvement prioritization,
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 08479a28-5f7d-426b-b471-a1c561ec2fc0 · outbound
Preference-based Multi-Objective Reinforcement Learning Q-learning,
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8c28b02e-16f2-441c-8153-7ae6a43bdfd5 · outbound
Preference-based Multi-Objective Reinforcement Learning Unresolved cited work
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation cdb7ec45-fc02-495e-9814-7e508caa2089 · outbound
Preference-based Multi-Objective Reinforcement Learning Convergence of q-learning: A simple proof,
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4dfa9a84-855f-4b7c-b202-e87470ba748c · outbound
Preference-based Multi-Objective Reinforcement Learning Decentralized multi-agent reinforcement learning: An off-policy method,
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 7a7e2417-c332-43fb-bf14-0778f084600e · outbound
Preference-based Multi-Objective Reinforcement Learning An ocba-based method for efficient sample collection in reinforcement learning,
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 3ddde7a4-ef4e-4b4b-beb2-67bb2f31f995 · outbound
Preference-based Multi-Objective Reinforcement Learning Rank analysis of incomplete block designs: I. the method of paired comparisons,
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation efa1052b-c7ae-4221-9dfe-a7afba883828 · outbound
Preference-based Multi-Objective Reinforcement Learning Preference-based multi-objective reinforcement learning with explicit reward modeling,
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation fb1ad97a-90c7-48d6-a16e-e32ffce1f627 · outbound
Preference-based Multi-Objective Reinforcement Learning Clarify: Contrastive preference reinforcement learning for untangling ambigu- ous queries,
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 01e14327-69ca-4bd8-9cc6-d519c38dbc89 · outbound
Preference-based Multi-Objective Reinforcement Learning Zitzler,Evolutionary algorithms for multiobjective optimization: Methods and applications
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 5470b709-8978-47df-bc91-e5981e8a03b3 · outbound
Preference-based Multi-Objective Reinforcement Learning Query-policy mis- alignment in preference-based reinforcement learning,
Reference 47
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 04d09981-ffca-4a2a-bd60-afbd07ba336f · outbound
Preference-based Multi-Objective Reinforcement Learning Empirical evaluation methods for multiobjective reinforcement learning algorithms,
Reference 48
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 5c6d56b8-ce60-4bb6-bd21-65d694dd9faf · outbound
Preference-based Multi-Objective Reinforcement Learning Learning all optimal policies with multiple criteria,
Reference 49
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation c6af2e34-a117-40a6-89e9-3164030e5849 · outbound
Preference-based Multi-Objective Reinforcement Learning An environment for autonomous driving decision-making,
Reference 50
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 99c5bb15-f318-4c62-bf1a-764392ecd9d8 · outbound
Preference-based Multi-Objective Reinforcement Learning Congested traffic states in empirical observations and microscopic simulations,
Reference 51
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dc9de4e9-cb9f-4b63-b0de-a2ba264dd6e8 · outbound
Preference-based Multi-Objective Reinforcement Learning General lane-changing model mobil for car-following models,
Reference 52
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c6181032-1c83-4bd9-8cd5-c8e95428db13 · outbound
Preference-based Multi-Objective Reinforcement Learning Implementing deep reinforcement learning (drl)-based driving styles for non-player vehicles,
Reference 53
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation c10ce3c5-a198-4c91-8f6b-450aeb1c1485 · outbound
Preference-based Multi-Objective Reinforcement Learning Enhancing Autonomous Vehicle Training with Language Model Integration and Critical Scenario Generation
Reference 54
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2ab81b5d-f8d5-4169-aca5-ff1f79acf333 · outbound
Preference-based Multi-Objective Reinforcement Learning Listwise reward estimation for offline preference-based reinforcement learning,
Reference 55
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
No inbound Pith citation observations are available.