Pith. sign in

Paper Citation Record · LEDGER

Preference-based Multi-Objective Reinforcement Learning

As of 8 August 2026, this Paper Citation Record lists 55 of 55 outbound references and 0 inbound Pith citation observations for arXiv:2507.14066.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.14066 v1

Coverage vector

measured 55 of 55 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T16:17:35.918503Z

measured 55 of 55 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

55 of 55 outbound references displayed

  • verified exact2
  • verified fuzzy37
  • unresolved16
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 9fc1a82f-f4fb-4255-be81-4d7da518b296 · outbound

This paper cites A survey on modeling and optimizing multi-objective systems,.

Preference-based Multi-Objective Reinforcement Learning A survey on modeling and optimizing multi-objective systems,

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:17:36.328027Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T16:17:32.960135Z digest=sha256:d595036a8996a9bb156177a05afaeae2299e943237e97af1780ecd26af669c23

Observation ae06a308-6755-455b-a074-e2e0c3686134 · outbound

This paper cites Active learning with fairness-aware clustering for fair classification considering multiple sensitive attributes,.

Preference-based Multi-Objective Reinforcement Learning Active learning with fairness-aware clustering for fair classification considering multiple sensitive attributes,

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:17:36.320638Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T16:17:33.086379Z digest=sha256:6c63a2785b719cbd2b3d93143376637d2b6bd6a27dd4d4fcc800ae1ce566962f

Observation 0c8692a9-07b7-45d4-b321-303a35428fbd · outbound

This paper cites Constrained ordinal opti- mization—a feasibility model based approach,.

Preference-based Multi-Objective Reinforcement Learning Constrained ordinal opti- mization—a feasibility model based approach,

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:17:36.313585Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T16:17:33.158951Z digest=sha256:55df1987120d291660e6089ad7f135461edcc7716b1f60c875943283552e388a

Observation c696ac33-1444-4b15-8189-5e10c27456d2 · outbound

This paper cites Reliability/cost-based multi-objective pareto optimal design of stand- alone wind/pv/fc generation microgrid system,.

Preference-based Multi-Objective Reinforcement Learning Reliability/cost-based multi-objective pareto optimal design of stand- alone wind/pv/fc generation microgrid system,

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:17:36.306194Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T16:17:33.319524Z digest=sha256:3b6a7575ccaa35aaf880311ebe6e3a8a62905637d1264a548008fc700bd58af9

Observation 2d48f134-df5b-413c-a3aa-0c3ee75c410e · outbound

This paper cites Toward personalized decision making for autonomous vehicles: a constrained multi-objective reinforcement learning tech- nique,.

Preference-based Multi-Objective Reinforcement Learning Toward personalized decision making for autonomous vehicles: a constrained multi-objective reinforcement learning tech- nique,

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:17:36.298042Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T16:17:33.469179Z digest=sha256:1f153ec294500d519e29d9cbbb8ddd4dc1b86ae18daaf0a409d86f6b51ade290

Observation 716eb782-3d6c-4a6c-8ed9-1be6e67e00ff · outbound

This paper cites B-pref: Benchmarking preference-based reinforcement learning,.

Preference-based Multi-Objective Reinforcement Learning B-pref: Benchmarking preference-based reinforcement learning,

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:17:36.290058Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T16:17:33.685131Z digest=sha256:543e6951e761d6cd25894b978f8879fc14d0f6ff099708b3760dae58e65e4abb

Observation 9e9b1cfd-c2d6-4a14-9fc6-cbd0d5bc2d51 · outbound

This paper cites Self- supervised online reward shaping in sparse-reward environments,.

Preference-based Multi-Objective Reinforcement Learning Self- supervised online reward shaping in sparse-reward environments,

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:17:36.282576Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T16:17:33.851165Z digest=sha256:7c7da4847fbf6c873d2fb2c010fa11e107c97f6254cf995167ea3c6fc39fd0c8

Observation 13ceb81f-37c5-41c2-83c9-72311c454326 · outbound

This paper cites Reward shaping for knowledge-based multi-objective multi-agent reinforcement learning,.

Preference-based Multi-Objective Reinforcement Learning Reward shaping for knowledge-based multi-objective multi-agent reinforcement learning,

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:17:36.275241Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T16:17:34.081232Z digest=sha256:ffe510b368c496971d53c1628e766766ac88874815d21515045c7c1578c6aa62

Observation a2583ef5-ca33-4e10-82ce-64ef063c2802 · outbound

This paper cites Pebble: Feedback-efficient interactive reinforcement learning via relabeling experience and unsuper- vised pre-training,.

Preference-based Multi-Objective Reinforcement Learning Pebble: Feedback-efficient interactive reinforcement learning via relabeling experience and unsuper- vised pre-training,

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:17:36.267833Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T16:17:34.340535Z digest=sha256:7e481267b9b000bda05e4a3df2c59b00612c86a47a6b64a0e87aeefc5c7ba787

Observation 4537b1db-9c38-4270-ae4b-71a51d79da40 · outbound

This paper cites Reinforcement learning and the reward engineering princi- ple,.

Preference-based Multi-Objective Reinforcement Learning Reinforcement learning and the reward engineering princi- ple,

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:17:36.260407Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T16:17:34.526104Z digest=sha256:7a65eb9a2664f2d95d9fd907eac47c989fa8321f81f83b0d0c6a66c338698b3d

Observation 24ca2cd7-5064-40f5-99ea-12b31bbaef82 · outbound

This paper cites Deep reinforcement learning from human preferences,.

Preference-based Multi-Objective Reinforcement Learning Deep reinforcement learning from human preferences,

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-06T16:17:34.529527Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:17:34.529527Z digest=sha256:3432fe208c4819f28c03ed379290575a9b01afa7153f927f9dc35dd533dc51a4

Observation b220a6de-e529-44a7-a136-6f85cbaf7a7f · outbound

This paper cites A bayesian approach for policy learning from trajectory preference queries,.

Preference-based Multi-Objective Reinforcement Learning A bayesian approach for policy learning from trajectory preference queries,

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:17:36.247729Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T16:17:34.569067Z digest=sha256:be6347ace244066179697e2e31e9c72b47e510fe71400d1249685caaf870e205

Observation 5e73303c-0400-4634-b2d6-680945ac21e7 · outbound

This paper cites an unresolved cited work.

Preference-based Multi-Objective Reinforcement Learning Unresolved cited work

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-06T16:17:34.641984Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:17:34.641984Z digest=sha256:1652d906fe2496be561cb515ef010b4a440df2856ec41c8cd9baf4efd60519b0

Observation 59e50602-f0f1-4bc3-8a65-75e1f0d8e058 · outbound

This paper cites Playing Atari with Deep Reinforcement Learning.

Preference-based Multi-Objective Reinforcement Learning Playing Atari with Deep Reinforcement Learning

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-06T16:17:34.803156Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:17:34.803156Z digest=sha256:049e54b164c653fec7d5b294a2114c245e6c3e3bb164601bdeae92ac95d91728

Observation 45c39390-16cf-419b-8bc3-0cf91b5b9b7f · outbound

This paper cites E-mapp: Efficient multi-agent reinforcement learning with parallel program guidance,.

Preference-based Multi-Objective Reinforcement Learning E-mapp: Efficient multi-agent reinforcement learning with parallel program guidance,

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:17:36.235748Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T16:17:34.908704Z digest=sha256:a35ca6cfeb1e89795426296517fc9e21d4039d837181291ead3c4a44f11407c3

Observation 188d3ba9-859c-4067-93ac-7c801054fb7e · outbound

This paper cites Mastering the game of go without human knowledge,.

Preference-based Multi-Objective Reinforcement Learning Mastering the game of go without human knowledge,

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-06T16:17:34.990348Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:17:34.990348Z digest=sha256:72b9ee4a803598c8a293441a53b7a49797b03700a554eee9fae1962dd3f75b65

Observation 191fe72d-0e3f-4e18-92d0-ca2e6acdd8a8 · outbound

This paper cites Simplify twin crane scheduling in railway yard by spatial task assignment,.

Preference-based Multi-Objective Reinforcement Learning Simplify twin crane scheduling in railway yard by spatial task assignment,

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:17:36.223250Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T16:17:35.073240Z digest=sha256:89f6e9ac31de84b63c4556cc49ac0bb9f640a6a9a78222677c963af8604789b0

Observation b5bb24a4-fef6-4148-a288-c2e0654dcf78 · outbound

This paper cites Large-scale data center cooling control via sample-efficient reinforcement learning,.

Preference-based Multi-Objective Reinforcement Learning Large-scale data center cooling control via sample-efficient reinforcement learning,

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:17:36.215576Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T16:17:35.159987Z digest=sha256:aaaa52bff598d275c88e89ebb173ac901cdda80e7a9eff269213864019f9bf14

Observation a82b67cb-516f-4951-92b1-4de16f88d1de · outbound

This paper cites An efficient real- time railway container yard management method based on partial de- coupling,.

Preference-based Multi-Objective Reinforcement Learning An efficient real- time railway container yard management method based on partial de- coupling,

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:17:36.208046Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T16:17:35.240063Z digest=sha256:731c6fa1dd83315cc73f9482c786921349e347e23b5a4f381ba0ccb4e5816c1a

Observation 97ff2a81-6704-405c-af47-15839da3abfc · outbound

This paper cites Integrating mechanism and data: Rein- forcement learning based on multi-fidelity model for data center cooling control,.

Preference-based Multi-Objective Reinforcement Learning Integrating mechanism and data: Rein- forcement learning based on multi-fidelity model for data center cooling control,

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:17:36.200746Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T16:17:35.322733Z digest=sha256:ba61ff325c848dd7dfbfe3b55c8512a1d7ddfcc3a83958001e10b1d930040275

Observation 1dffd4ba-39e9-4fea-b1e1-816f02af9399 · outbound

This paper cites Incentive-oriented power-carbon emissions trading-tradable green certificate integrated market mecha- nisms using multi-agent deep reinforcement learning,.

Preference-based Multi-Objective Reinforcement Learning Incentive-oriented power-carbon emissions trading-tradable green certificate integrated market mecha- nisms using multi-agent deep reinforcement learning,

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:17:36.192649Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T16:17:35.382374Z digest=sha256:8e0ee47de271fe134bba9e02fb8137634ad74666c741e9053e17cfffce2c44ba

Observation 7786b8cc-d84f-4ec2-a684-b758ac9015f7 · outbound

This paper cites OpenAI Gym.

Preference-based Multi-Objective Reinforcement Learning OpenAI Gym

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-06T16:17:35.450783Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:17:35.450783Z digest=sha256:59fff2cd2cd9bb7e3753cf5e14c2f31f1f40c312f60aad4ab4bd102825952c54

Observation 7be9d931-503a-43e0-9f1a-065a9f2119ae · outbound

This paper cites Exploration by random network distillation,.

Preference-based Multi-Objective Reinforcement Learning Exploration by random network distillation,

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:17:36.184804Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T16:17:35.511221Z digest=sha256:2cb080f30ad7b197524200d8a3bc8e247f60f1a0c20812fc75f97e6608a5dc88

Observation 88d8419b-af3a-4b5b-8697-f17d167b40be · outbound

This paper cites Direct preference optimization: Your language model is secretly a reward model,.

Preference-based Multi-Objective Reinforcement Learning Direct preference optimization: Your language model is secretly a reward model,

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-06T16:17:35.584119Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:17:35.584119Z digest=sha256:e990d4383f571ae08da23afd9dc0802715c6afa955fd42631de0151b7a7d3fc3

Observation 2af1f740-84f9-45f4-8b03-a71c0e055626 · outbound

This paper cites Reward learning from human preferences and demonstrations in atari,.

Preference-based Multi-Objective Reinforcement Learning Reward learning from human preferences and demonstrations in atari,

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:17:36.173015Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T16:17:35.668501Z digest=sha256:d2fff95affb0fc8f2fb2bd7794f2c31595bd7d01383a16b2f8a0881a7317f3dd

Observation 86d955f6-956d-453c-b71f-bf247c652a92 · outbound

This paper cites S-EPOA: Overcoming the Indistinguishability of Segments with Skill-Driven Preference-Based Reinforcement Learning.

Preference-based Multi-Objective Reinforcement Learning S-EPOA: Overcoming the Indistinguishability of Segments with Skill-Driven Preference-Based Reinforcement Learning

Reference 26

Resolution
verified exact
local_arxiv, observed 2026-08-06T16:17:35.967349Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T16:17:35.749954Z digest=sha256:4704114ee4761ac63a38ecddb45892f6d059cc33731402cb3996362c79abd985

Observation 8e1175aa-1ece-41da-a025-3a324564b00f · outbound

This paper cites Surf: Semi- supervised reward learning with data augmentation for feedback-efficient preference-based reinforcement learning,.

Preference-based Multi-Objective Reinforcement Learning Surf: Semi- supervised reward learning with data augmentation for feedback-efficient preference-based reinforcement learning,

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:17:36.165275Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T16:17:35.810695Z digest=sha256:4df46bc1b58b7585d578534e73233682c7b7677bdb92d4606272c67802a7c97a

Observation b70ba05f-268c-45fb-9cf8-5e797f7c0d36 · outbound

This paper cites Few-shot preference learning for human- in-the-loop rl,.

Preference-based Multi-Objective Reinforcement Learning Few-shot preference learning for human- in-the-loop rl,

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:17:36.156940Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T16:17:35.840299Z digest=sha256:7005e6a90cb668d9d52ee7f6b56931bf286cd1722a95f2db4722ff51e0a03ae2

Observation ad4b84f9-c650-4b4c-a6a4-c2e3eba99b98 · outbound

This paper cites Learning to summarize with human feedback,.

Preference-based Multi-Objective Reinforcement Learning Learning to summarize with human feedback,

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-06T16:17:35.844045Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:17:35.844045Z digest=sha256:bfc42d4f55eaaddab46d6bfe87e20b7141280c0444665f4c1519d2348c7e4511

Observation 468c6574-e772-4935-be5d-7550d1d20647 · outbound

This paper cites A toolkit for reliable benchmarking and research in multi-objective reinforcement learning,.

Preference-based Multi-Objective Reinforcement Learning A toolkit for reliable benchmarking and research in multi-objective reinforcement learning,

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:17:36.145259Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T16:17:35.848219Z digest=sha256:37fb72cf3bbedaedbaf4a5990fda55226e8290304e328801b89e55109667de9c

Observation 112be415-f344-4a58-af0f-03915d8a1d6e · outbound

This paper cites A practical guide to multi-objective reinforcement learning and planning,.

Preference-based Multi-Objective Reinforcement Learning A practical guide to multi-objective reinforcement learning and planning,

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:17:36.137836Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T16:17:35.851254Z digest=sha256:7a56978aa65c6b59d740b4684efe963129bb88e2958194c6770c95353fcc7c6c

Observation 420d363e-a9f4-47b2-bd3e-6ef4f759e91c · outbound

This paper cites Human-in-the-Loop Policy Optimization for Preference-Based Multi-Objective Reinforcement Learning.

Preference-based Multi-Objective Reinforcement Learning Human-in-the-Loop Policy Optimization for Preference-Based Multi-Objective Reinforcement Learning

Reference 32

Resolution
verified exact
local_arxiv, observed 2026-08-06T16:17:35.956243Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T16:17:35.855636Z digest=sha256:1894bef621df9f44423f4786dd2377b01a9edca1c7cdef9dc814c24bfb2c2dba

Observation 2712168f-f393-4094-ab35-cbfe5a78983a · outbound

This paper cites A generalized algorithm for multi-objective reinforcement learning and policy adaptation,.

Preference-based Multi-Objective Reinforcement Learning A generalized algorithm for multi-objective reinforcement learning and policy adaptation,

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:17:36.130297Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T16:17:35.859598Z digest=sha256:78a309116fbf07cfda37a27910d3181e01c9c8b1d5ccbce050159b77a1368311

Observation 9eb1a771-cbb8-4150-a54b-76732f7c49cf · outbound

This paper cites Multi-objective rein- forcement learning for the expected utility of the return,.

Preference-based Multi-Objective Reinforcement Learning Multi-objective rein- forcement learning for the expected utility of the return,

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:17:36.122647Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T16:17:35.863547Z digest=sha256:4cb1f9b1607419d9457868ebd7d0e7c430459d79e07686c543ddbf50783f17ec

Observation e893cdc1-7bfb-4641-8956-8dd0749bdc23 · outbound

This paper cites Prediction- guided multi-objective reinforcement learning for continuous robot con- trol,.

Preference-based Multi-Objective Reinforcement Learning Prediction- guided multi-objective reinforcement learning for continuous robot con- trol,

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-06T16:17:35.865902Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:17:35.865902Z digest=sha256:a47a82a385b3fdf595e0d2f84acebbef7989908cd88dcc8814c7c918a0a1ed6b

Observation a2dca11c-79cc-4b92-a95e-a824b1f95d7a · outbound

This paper cites Pareto conditioned net- works,.

Preference-based Multi-Objective Reinforcement Learning Pareto conditioned net- works,

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:17:36.109162Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T16:17:35.869150Z digest=sha256:b4ff0470159fcd74280d553f7bf684a621f70d6d5c226cc7af1de5450e72ff64

Observation 1386df91-dd9e-4e2f-b578-ea283fac0c47 · outbound

This paper cites Sample-efficient multi-objective learning via generalized pol- icy improvement prioritization,.

Preference-based Multi-Objective Reinforcement Learning Sample-efficient multi-objective learning via generalized pol- icy improvement prioritization,

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:17:36.101150Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T16:17:35.871852Z digest=sha256:714c0be360a39cbb065d02646ebd6204ddd9615ac26c65cd20458f0fa9e43655

Observation 08479a28-5f7d-426b-b471-a1c561ec2fc0 · outbound

This paper cites Q-learning,.

Preference-based Multi-Objective Reinforcement Learning Q-learning,

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-06T16:17:35.874650Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:17:35.874650Z digest=sha256:f018710e9dc24f8cc7e63c8aa836b2c65a60da726d3ce7d12daaec49b47a7885

Observation 8c28b02e-16f2-441c-8153-7ae6a43bdfd5 · outbound

This paper cites an unresolved cited work.

Preference-based Multi-Objective Reinforcement Learning Unresolved cited work

Reference 39

Resolution
unresolved
raw_fallback, observed 2026-08-06T16:17:36.089262Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T16:17:35.876997Z digest=sha256:8d20fbaf290b1340d944966bf685f344cdff2a112ce9bf53d0614622c089c225

Observation cdb7ec45-fc02-495e-9814-7e508caa2089 · outbound

This paper cites Convergence of q-learning: A simple proof,.

Preference-based Multi-Objective Reinforcement Learning Convergence of q-learning: A simple proof,

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-06T16:17:35.879279Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:17:35.879279Z digest=sha256:8000e18fc3e9262993889060fd0e7d5cc1bb1b64bda54283a1c85b885a951844

Observation 4dfa9a84-855f-4b7c-b202-e87470ba748c · outbound

This paper cites Decentralized multi-agent reinforcement learning: An off-policy method,.

Preference-based Multi-Objective Reinforcement Learning Decentralized multi-agent reinforcement learning: An off-policy method,

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:17:36.076420Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T16:17:35.881717Z digest=sha256:f64e349206bfaa3dd698dbb33181f5bf1558f6505883448fa0a5b7694c6f5c42

Observation 7a7e2417-c332-43fb-bf14-0778f084600e · outbound

This paper cites An ocba-based method for efficient sample collection in reinforcement learning,.

Preference-based Multi-Objective Reinforcement Learning An ocba-based method for efficient sample collection in reinforcement learning,

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:17:36.068792Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T16:17:35.884101Z digest=sha256:c6c79e991b70d159544c78c11fae4e248af0bce32ee18cfb927a3150cf041c11

Observation 3ddde7a4-ef4e-4b4b-beb2-67bb2f31f995 · outbound

This paper cites Rank analysis of incomplete block designs: I. the method of paired comparisons,.

Preference-based Multi-Objective Reinforcement Learning Rank analysis of incomplete block designs: I. the method of paired comparisons,

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-06T16:17:35.886959Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:17:35.886959Z digest=sha256:cc56c3029b93202f1a1353d9709fdb8cdbdaae6adfbeacd38aab3f8fb2cd3ad7

Observation efa1052b-c7ae-4221-9dfe-a7afba883828 · outbound

This paper cites Preference-based multi-objective reinforcement learning with explicit reward modeling,.

Preference-based Multi-Objective Reinforcement Learning Preference-based multi-objective reinforcement learning with explicit reward modeling,

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:17:36.057027Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T16:17:35.889314Z digest=sha256:e1a82ac4001115e06e14ac30ebb6506b3b8fe21dd1471f3b17ac3e222b7a6b6a

Observation fb1ad97a-90c7-48d6-a16e-e32ffce1f627 · outbound

This paper cites Clarify: Contrastive preference reinforcement learning for untangling ambigu- ous queries,.

Preference-based Multi-Objective Reinforcement Learning Clarify: Contrastive preference reinforcement learning for untangling ambigu- ous queries,

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:17:36.048623Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T16:17:35.891737Z digest=sha256:2b15d2278c47a196b83ff93ddb5f74c944f05d98b7bb655f5cc6993bca68cc72

Observation 01e14327-69ca-4bd8-9cc6-d519c38dbc89 · outbound

This paper cites Zitzler,Evolutionary algorithms for multiobjective optimization: Methods and applications.

Preference-based Multi-Objective Reinforcement Learning Zitzler,Evolutionary algorithms for multiobjective optimization: Methods and applications

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:17:36.040417Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T16:17:35.894131Z digest=sha256:6a860d647baf5d9b74e804806fff5638bb94ba5d6fe4aa450f20250e8fa1add0

Observation 5470b709-8978-47df-bc91-e5981e8a03b3 · outbound

This paper cites Query-policy mis- alignment in preference-based reinforcement learning,.

Preference-based Multi-Objective Reinforcement Learning Query-policy mis- alignment in preference-based reinforcement learning,

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:17:36.032810Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T16:17:35.896952Z digest=sha256:3974e8075dbbad4acadf918dd7a931cafab2902ba948563996349b98d7625993

Observation 04d09981-ffca-4a2a-bd60-afbd07ba336f · outbound

This paper cites Empirical evaluation methods for multiobjective reinforcement learning algorithms,.

Preference-based Multi-Objective Reinforcement Learning Empirical evaluation methods for multiobjective reinforcement learning algorithms,

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:17:36.024987Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T16:17:35.899502Z digest=sha256:cdd6558f915fe2487d44c2750d1c79997ac77bebbd67819b3691d449fd99715f

Observation 5c6d56b8-ce60-4bb6-bd21-65d694dd9faf · outbound

This paper cites Learning all optimal policies with multiple criteria,.

Preference-based Multi-Objective Reinforcement Learning Learning all optimal policies with multiple criteria,

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:17:36.017157Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T16:17:35.902317Z digest=sha256:0b6437a8aa7dfc688dd2ef6d04d0d346ca851acf8881d83faddfa999c042959b

Observation c6af2e34-a117-40a6-89e9-3164030e5849 · outbound

This paper cites An environment for autonomous driving decision-making,.

Preference-based Multi-Objective Reinforcement Learning An environment for autonomous driving decision-making,

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-06T16:17:35.905119Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:17:35.905119Z digest=sha256:a72c22039291e9671a4eaa025251f8f5d8b485761d32dd61170463b898b0afee

Observation 99c5bb15-f318-4c62-bf1a-764392ecd9d8 · outbound

This paper cites Congested traffic states in empirical observations and microscopic simulations,.

Preference-based Multi-Objective Reinforcement Learning Congested traffic states in empirical observations and microscopic simulations,

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-06T16:17:35.907272Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:17:35.907272Z digest=sha256:91e4933b477ba75817b876e17b108e65bacf6cd116e765c73a4ed13e01e64496

Observation dc9de4e9-cb9f-4b63-b0de-a2ba264dd6e8 · outbound

This paper cites General lane-changing model mobil for car-following models,.

Preference-based Multi-Objective Reinforcement Learning General lane-changing model mobil for car-following models,

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-06T16:17:35.909933Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:17:35.909933Z digest=sha256:cdf0be87e10f2ceb4ee3600d86f8ec1a2f89b5791cb412791dcd6975fac5d701

Observation c6181032-1c83-4bd9-8cd5-c8e95428db13 · outbound

This paper cites Implementing deep reinforcement learning (drl)-based driving styles for non-player vehicles,.

Preference-based Multi-Objective Reinforcement Learning Implementing deep reinforcement learning (drl)-based driving styles for non-player vehicles,

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:17:35.996133Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T16:17:35.912851Z digest=sha256:993cf3bf6d5137ddc50f450f7b820d0282f805f6a20fbc5f640920437a9bcef6

Observation c10ce3c5-a198-4c91-8f6b-450aeb1c1485 · outbound

This paper cites Enhancing Autonomous Vehicle Training with Language Model Integration and Critical Scenario Generation.

Preference-based Multi-Objective Reinforcement Learning Enhancing Autonomous Vehicle Training with Language Model Integration and Critical Scenario Generation

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-06T16:17:35.915292Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:17:35.915292Z digest=sha256:b1f1e61d5b639eadf0d987fc2d7a93cce6097e3e5ea6af68798a4a81f16f92e3

Observation 2ab81b5d-f8d5-4169-aca5-ff1f79acf333 · outbound

This paper cites Listwise reward estimation for offline preference-based reinforcement learning,.

Preference-based Multi-Objective Reinforcement Learning Listwise reward estimation for offline preference-based reinforcement learning,

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:17:35.988790Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T16:17:35.918503Z digest=sha256:e4188b413ccc4bceaabb2bf9669f79d7c66f4d528888211a3b9dc85e30cf3236

Pith citing papers

No inbound Pith citation observations are available.