Pith. sign in

Paper Citation Record · LEDGER

LEMUR: Learning to Align with Multi-Objective Reinforcement Learning from Preference Feedback

As of 21 August 2026, this Paper Citation Record lists 100 of 297 outbound references and 0 inbound Pith citation observations for arXiv:2607.29559.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2607.29559 v1

Coverage vector

measured 100 of 297 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-03T04:39:31.713492Z

measured 100 of 100 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-20T06:33:59.587034+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

100 of 297 outbound references displayed

  • verified exact14
  • verified fuzzy0
  • unresolved81
  • parse uncertain3
  • malformed identifier0
  • metadata mismatch2

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 5bee8ff9-5c58-4934-b338-19479b3eb5f2 · outbound

This paper cites Structure and Interpretation of Computer Programs.

LEMUR: Learning to Align with Multi-Objective Reinforcement Learning from Preference Feedback Structure and Interpretation of Computer Programs

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-03T04:39:30.589269Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T04:39:30.589269Z digest=sha256:2cb07d2a1881a850f77755e10c980deb1d07fa92515247f03021aaeac364ce83

Observation 9b30f37c-da2c-481c-bfa0-4e9f6ecc15f4 · outbound

This paper cites Visual Information Extraction with Lixto.

LEMUR: Learning to Align with Multi-Objective Reinforcement Learning from Preference Feedback Visual Information Extraction with Lixto

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-03T04:39:30.593231Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T04:39:30.593231Z digest=sha256:617aa59706e107de62d8c591a422b83fedcaa4e0633a6b1060b10fa3f5a72065

Observation f188bd46-0683-490f-8df6-c71ba6d0e7b3 · outbound

This paper cites Brachman and James G.

LEMUR: Learning to Align with Multi-Objective Reinforcement Learning from Preference Feedback Brachman and James G

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-03T04:39:30.597128Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T04:39:30.597128Z digest=sha256:17b0900d7cf2f3a0e82d64e90a63856ac6d088f8d9b56897d16e317f9e504282

Observation a4b422d4-937d-47da-9294-400f8b6dd384 · outbound

This paper cites Complexity results for nonmonotonic logics.

LEMUR: Learning to Align with Multi-Objective Reinforcement Learning from Preference Feedback Complexity results for nonmonotonic logics

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-03T04:39:30.601312Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T04:39:30.601312Z digest=sha256:d8fe52f86eb2fa5bc2d3f91b2530f0ae81e9ba71204d8aafd7610cd75f5d892b

Observation 1791ea39-2bd3-4a9a-845a-4e9d3e1396b7 · outbound

This paper cites International conference on machine learning , pages=.

LEMUR: Learning to Align with Multi-Objective Reinforcement Learning from Preference Feedback International conference on machine learning , pages=

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-03T04:39:30.604828Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T04:39:30.604828Z digest=sha256:6f0c8b763755c1a66bb6c811dc6c153c86bd9a493bca0355a1331e3888e7f5af

Observation a2b936d4-1aac-468b-ad57-d588e4de3995 · outbound

This paper cites Hypertree Decompositions and Tractable Queries.

LEMUR: Learning to Align with Multi-Objective Reinforcement Learning from Preference Feedback Hypertree Decompositions and Tractable Queries

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-03T04:39:30.608263Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T04:39:30.608263Z digest=sha256:21b65bfdc5a5cbb60b2a436e5882169665aa2f70d9f268db1b74040ed8b4a71f

Observation 874628da-b357-4d2b-b021-20f4df54aff4 · outbound

This paper cites Levesque.

LEMUR: Learning to Align with Multi-Objective Reinforcement Learning from Preference Feedback Levesque

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-03T04:39:30.611943Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T04:39:30.611943Z digest=sha256:bd7a49c059660932f53b7fa84e2311bf2845ce296f47f195a6554b1f42979a85

Observation 718dda53-a904-439c-8cec-b0e522001c5b · outbound

This paper cites Levesque.

LEMUR: Learning to Align with Multi-Objective Reinforcement Learning from Preference Feedback Levesque

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-03T04:39:30.615468Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T04:39:30.615468Z digest=sha256:a0852944ce1071f71ed5cb93ad3eb97538af6746d075ae4c63ddad1573982b74

Observation 3f5d6fc0-6329-4d30-ac5b-612dbc22a701 · outbound

This paper cites On the compilability and expressive power of propositional planning formalisms.

LEMUR: Learning to Align with Multi-Objective Reinforcement Learning from Preference Feedback On the compilability and expressive power of propositional planning formalisms

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-03T04:39:30.619045Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T04:39:30.619045Z digest=sha256:f4c794c74e7d6d50cc8429c91c6a7781f8dce0cc1ff8e24f1eadd1c47b6b74fc

Observation 5c012139-7e05-4380-8366-8b1ecf6f9a68 · outbound

This paper cites an unresolved cited work.

LEMUR: Learning to Align with Multi-Objective Reinforcement Learning from Preference Feedback Unresolved cited work

Reference 10

Resolution
parse uncertain
no resolver link, observed 2026-08-03T04:39:30.622204Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T04:39:30.622204Z digest=sha256:afae4f66462356713115f360ee6fe77d662e162077ea0101b62fa837dd989b3d

Observation 6857fca7-2dc5-41bd-ac20-fec6fefca1b9 · outbound

This paper cites The Knowledge Engineering Review , volume =.

LEMUR: Learning to Align with Multi-Objective Reinforcement Learning from Preference Feedback The Knowledge Engineering Review , volume =

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-03T04:39:30.625671Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T04:39:30.625671Z digest=sha256:c58f2508749d88f53fbf23f4cf345b8d993667ce23b375a2f02ec1852ef7c503

Observation 5145e282-dee6-4f7a-aaf6-794fbe7f9b2d · outbound

This paper cites Artificial Intelligence , volume =.

LEMUR: Learning to Align with Multi-Objective Reinforcement Learning from Preference Feedback Artificial Intelligence , volume =

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-03T04:39:30.629782Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T04:39:30.629782Z digest=sha256:bf0c3b020e02043ccd877496b0bbad00eed2a6a7fe66b901c04c229cdbb0e192

Observation 28349626-1ec0-4800-bc66-552eef11895c · outbound

This paper cites Logics of programs: axiomatics and descriptive power.

LEMUR: Learning to Align with Multi-Objective Reinforcement Learning from Preference Feedback Logics of programs: axiomatics and descriptive power

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-03T04:39:30.633251Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T04:39:30.633251Z digest=sha256:97b34eaf5182c1a5fab2e54505b0932249ada8f452cc9e5cef5d2fc17e36f793

Observation 00802faa-6e1c-4d96-ab0b-ac437ee5a7a2 · outbound

This paper cites Clarkson.

LEMUR: Learning to Align with Multi-Objective Reinforcement Learning from Preference Feedback Clarkson

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-03T04:39:30.637314Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T04:39:30.637314Z digest=sha256:7da0d3e37aa1e4062e38a964e13ce91b814241d8517bfc2701acec8f15f340e0

Observation 9a2c6029-d150-4ed9-b5f9-13ef503b7f71 · outbound

This paper cites A More Perfect Union.

LEMUR: Learning to Align with Multi-Objective Reinforcement Learning from Preference Feedback A More Perfect Union

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-03T04:39:30.640616Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T04:39:30.640616Z digest=sha256:c97c84a875a9ff28dbe90ab6fcf5662ccff1e25ddcacfa86eaa876977d4d9a68

Observation fa38e6a1-b867-410e-96e5-aaf08eedfcb9 · outbound

This paper cites The fountain of youth.

LEMUR: Learning to Align with Multi-Objective Reinforcement Learning from Preference Feedback The fountain of youth

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-03T04:39:30.644592Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T04:39:30.644592Z digest=sha256:3a75959f6bb64612b3acbd4a7d81f40f6888c9ffa1cdbbe5754e844e567d6d00

Observation 94a46ff4-018c-4ca9-8cfc-eb50925f6c5c · outbound

This paper cites an unresolved cited work.

LEMUR: Learning to Align with Multi-Objective Reinforcement Learning from Preference Feedback Unresolved cited work

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-03T04:39:30.648325Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T04:39:30.648325Z digest=sha256:c5807115ca2e7f4d4da62fbf7dd4f0a9e744e14f75e71b9a324704bab9adcc25

Observation 7b35fef3-5325-4b20-855b-3503a655c8b2 · outbound

This paper cites Proceedings of the 20th International Colloquium on Automata, Languages and Programming , series =.

LEMUR: Learning to Align with Multi-Objective Reinforcement Learning from Preference Feedback Proceedings of the 20th International Colloquium on Automata, Languages and Programming , series =

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-03T04:39:30.651978Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T04:39:30.651978Z digest=sha256:5af7d74f81c5bc0f37b1a22d99e55c7872657e2d39353790d7f5ba5a2c663fb7

Observation e0696e14-2f17-45cd-9113-515c1ccab64e · outbound

This paper cites an unresolved cited work.

LEMUR: Learning to Align with Multi-Objective Reinforcement Learning from Preference Feedback Unresolved cited work

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-03T04:39:30.655272Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T04:39:30.655272Z digest=sha256:25f3d6cbd836fa07cb5c50d242f4df84cc6ab7e4ba7a50ee3da782ad55bac028

Observation b0ca4f76-3181-44dc-9806-784ed28f8ec7 · outbound

This paper cites Anisi , title =.

LEMUR: Learning to Align with Multi-Objective Reinforcement Learning from Preference Feedback Anisi , title =

Reference 20

Resolution
parse uncertain
no resolver link, observed 2026-08-03T04:39:30.658830Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T04:39:30.658830Z digest=sha256:a574fc997d9a719ed584ffb174290b23f6d5826e64deb9f1524b7528fbcb3ea5

Observation e8027626-9377-4990-a60d-30ef340c0f38 · outbound

This paper cites Journal of Artificial Intelligence Research , author =.

LEMUR: Learning to Align with Multi-Objective Reinforcement Learning from Preference Feedback Journal of Artificial Intelligence Research , author =

Reference 21

Resolution
verified exact
doi, observed 2026-08-03T04:44:22.845500Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-03T04:39:30.663083Z digest=sha256:a65d362f0c0acedf0d3d11d9607cceef7e31a086a381b12bf22ac885970f1c5b

Observation 6cf07bab-36b6-42d1-aaa7-7cf783553107 · outbound

This paper cites an unresolved cited work.

LEMUR: Learning to Align with Multi-Objective Reinforcement Learning from Preference Feedback Unresolved cited work

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-03T04:39:30.667119Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T04:39:30.667119Z digest=sha256:27eaa2c4c1dbfccc34e68b940cda128e1034576403ecc2e122a2b3c5ada0e03d

Observation 8f9dd01d-2f57-4723-82e6-86a18ceda1b4 · outbound

This paper cites Reinforcement Learning , author =.

LEMUR: Learning to Align with Multi-Objective Reinforcement Learning from Preference Feedback Reinforcement Learning , author =

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-03T04:39:30.671077Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T04:39:30.671077Z digest=sha256:eacdb0cca65d28f9473f455b42a6e07ad08c8c747066cef617b69d7ba78156d8

Observation 5295577c-5a5c-4fd0-a3fd-26657c5e5361 · outbound

This paper cites an unresolved cited work.

LEMUR: Learning to Align with Multi-Objective Reinforcement Learning from Preference Feedback Unresolved cited work

Reference 24

Resolution
parse uncertain
no resolver link, observed 2026-08-03T04:39:30.674550Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T04:39:30.674550Z digest=sha256:17de4712b0168bdea940bbe0f40df653955bd04abf9e09cd63a6c0b34836b23e

Observation 98c2686b-c84a-4893-9817-5c9a2595a88e · outbound

This paper cites Human-Aligned Skill Discovery: Balancing Behaviour Exploration and Alignment.

LEMUR: Learning to Align with Multi-Objective Reinforcement Learning from Preference Feedback Human-Aligned Skill Discovery: Balancing Behaviour Exploration and Alignment

Reference 25

Resolution
verified exact
local_arxiv, observed 2026-08-03T04:44:22.697127Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-03T04:39:30.677776Z digest=sha256:21eccdeb7413e00a7aeb4e4f1733c7ae04451624911440847ec367158f59a7a3

Observation 6b68dfec-b62d-4cd2-a103-c5f50a2c0958 · outbound

This paper cites Divide and Conquer: Provably Unveiling the Pareto Front with Multi-Objective Reinforcement Learning.

LEMUR: Learning to Align with Multi-Objective Reinforcement Learning from Preference Feedback Divide and Conquer: Provably Unveiling the Pareto Front with Multi-Objective Reinforcement Learning

Reference 26

Resolution
metadata mismatch
local_arxiv, observed 2026-08-03T04:44:22.607261Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-03T04:39:30.681667Z digest=sha256:0d4cec9edcabc00e3030b1644254a738f99f324b7f6de7b0a4ed242a2ca63f0c

Observation bbc5ba6b-ab98-4543-aa82-cbf85e72c605 · outbound

This paper cites Fairness in Preference-based Reinforcement Learning.

LEMUR: Learning to Align with Multi-Objective Reinforcement Learning from Preference Feedback Fairness in Preference-based Reinforcement Learning

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-03T04:39:30.685468Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T04:39:30.685468Z digest=sha256:2820ab0dfadeffa9283744be2c336706a478432cf95e2c9fd1785c5b9c803b1f

Observation 7311e445-1e27-4c5f-ba38-dddb4f2487ce · outbound

This paper cites Personalized Soups: Personalized Large Language Model Alignment via Post-hoc Parameter Merging.

LEMUR: Learning to Align with Multi-Objective Reinforcement Learning from Preference Feedback Personalized Soups: Personalized Large Language Model Alignment via Post-hoc Parameter Merging

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-03T04:39:30.689529Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T04:39:30.689529Z digest=sha256:8b6b5d7205a39789c15c120eb04401405bd30820a7eba7a10966adbc98a2b1ea

Observation e4467589-f136-4aa9-88bd-8a31f7fa049c · outbound

This paper cites Advances in Neural Information Processing Systems , volume=.

LEMUR: Learning to Align with Multi-Objective Reinforcement Learning from Preference Feedback Advances in Neural Information Processing Systems , volume=

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-03T04:39:30.693991Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T04:39:30.693991Z digest=sha256:e17007953eb477bbae81ac3e3f5d6113b2f21cd7e301e3d79505771baea302d3

Observation 9295a79a-5eeb-4002-8212-51dcb5f837c8 · outbound

This paper cites Promptable Behaviors: Personalizing Multi-Objective Rewards from Human Preferences.

LEMUR: Learning to Align with Multi-Objective Reinforcement Learning from Preference Feedback Promptable Behaviors: Personalizing Multi-Objective Rewards from Human Preferences

Reference 30

Resolution
verified exact
local_arxiv, observed 2026-08-03T04:44:22.467309Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-03T04:39:30.697817Z digest=sha256:b69149a9c4dd527f0b21cb23d3c14e1d3a43905bf184a76dd67b28b687d74ea0

Observation 14452a8c-e7fb-4342-8515-23dc6816b271 · outbound

This paper cites Proceedings of the 41st International Conference on Machine Learning , pages =.

LEMUR: Learning to Align with Multi-Objective Reinforcement Learning from Preference Feedback Proceedings of the 41st International Conference on Machine Learning , pages =

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-03T04:39:30.702262Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T04:39:30.702262Z digest=sha256:9c3499a313fe9f5f30b23b727261cc0f849bdf06894d38635a0ae3ebcea2faa5

Observation 996524c1-8261-4f57-9c33-36780dec384a · outbound

This paper cites Fine-tuning language models to find agreement among humans with diverse preferences.

LEMUR: Learning to Align with Multi-Objective Reinforcement Learning from Preference Feedback Fine-tuning language models to find agreement among humans with diverse preferences

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-03T04:39:30.705616Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T04:39:30.705616Z digest=sha256:5a901647d3340e6e8be3e1556479d259a4596c25bcf9f2cf7c5da6046e0f07e3

Observation 3bd86fb5-4b35-472d-a65a-1b88e422cdaa · outbound

This paper cites Adaptive Alignment: Dynamic Preference Adjustments via Multi-Objective Reinforcement Learning for Pluralistic AI.

LEMUR: Learning to Align with Multi-Objective Reinforcement Learning from Preference Feedback Adaptive Alignment: Dynamic Preference Adjustments via Multi-Objective Reinforcement Learning for Pluralistic AI

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-03T04:39:30.709335Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T04:39:30.709335Z digest=sha256:606bd474f8159e82be70cbb002318dc8e96208642fc4232769d1ecfa79a353ed

Observation b756a39e-1aa6-4116-b690-3c3671133fb9 · outbound

This paper cites an unresolved cited work.

LEMUR: Learning to Align with Multi-Objective Reinforcement Learning from Preference Feedback Unresolved cited work

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-03T04:39:30.713064Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T04:39:30.713064Z digest=sha256:6d0a5187f3b87cd9bfbda88755c9959579a6c2be1e18c953db0a8b6dcff92edb

Observation 9a6c6a0f-78f6-4c38-84fa-8fdfeed44829 · outbound

This paper cites 2024 , pages =.

LEMUR: Learning to Align with Multi-Objective Reinforcement Learning from Preference Feedback 2024 , pages =

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-03T04:39:30.716496Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T04:39:30.716496Z digest=sha256:a1a273d1a1c7ac9adcc0b4882553fc7b0c16bfbaed7eb4c763730677373f7d50

Observation 3e0eed1b-71b7-41d5-b112-ebdecbe953e8 · outbound

This paper cites an unresolved cited work.

LEMUR: Learning to Align with Multi-Objective Reinforcement Learning from Preference Feedback Unresolved cited work

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-03T04:39:30.719807Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T04:39:30.719807Z digest=sha256:ff72ad1ebe450e4c1b5e32026ea6bedb195d0770ff10a5f588e112c0893edec7

Observation 8994e582-dfc2-45a5-bb47-dd15029abe7b · outbound

This paper cites and Yang, Diyi and Vosoughi, Soroush , month = oct, year =.

LEMUR: Learning to Align with Multi-Objective Reinforcement Learning from Preference Feedback and Yang, Diyi and Vosoughi, Soroush , month = oct, year =

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-03T04:39:30.723081Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T04:39:30.723081Z digest=sha256:ce9e3942c08b138528094c663d7d40071a4d757b69b5835572513662e700c051

Observation 0dc6aa27-f5cb-4155-ad93-1d27d700f73e · outbound

This paper cites an unresolved cited work.

LEMUR: Learning to Align with Multi-Objective Reinforcement Learning from Preference Feedback Unresolved cited work

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-03T04:39:30.727111Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T04:39:30.727111Z digest=sha256:374505460e3cc2f39eefea171a6744aaba1e23581a2cd748502af3b4adc50008

Observation 9f051224-659f-489d-b2b7-27447a6814e0 · outbound

This paper cites and Hassenzahl, Marc , month = jul, year =.

LEMUR: Learning to Align with Multi-Objective Reinforcement Learning from Preference Feedback and Hassenzahl, Marc , month = jul, year =

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-03T04:39:30.730473Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T04:39:30.730473Z digest=sha256:fae5e1f02fb28ec6dd148cb80b902cb5b82eae20f87cff6582d97b8b10d9b2c9

Observation 766bef44-b4e9-43b8-849b-f56bf7b7ddce · outbound

This paper cites IEEE Access , author =.

LEMUR: Learning to Align with Multi-Objective Reinforcement Learning from Preference Feedback IEEE Access , author =

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-03T04:39:30.733862Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T04:39:30.733862Z digest=sha256:f49634f1115b041c33684c2ba5288263a18b80c61940b25c42dc8d1e224f7dfe

Observation d63e0ac1-be9c-4eed-8225-f0d62014ec83 · outbound

This paper cites Advice to.

LEMUR: Learning to Align with Multi-Objective Reinforcement Learning from Preference Feedback Advice to

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-03T04:39:30.737965Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T04:39:30.737965Z digest=sha256:8ad617e584ce1b93e0acc0af37b80c93f0b859856c5cc39d3149d41c591e5e23

Observation 763a2c3b-bcc3-4a93-b075-69576f7a0dc9 · outbound

This paper cites Applied Intelligence , author =.

LEMUR: Learning to Align with Multi-Objective Reinforcement Learning from Preference Feedback Applied Intelligence , author =

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-03T04:39:30.741891Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T04:39:30.741891Z digest=sha256:4ae9eaaf902a341603e8eeb0c0905f19f6a256bf3d9d9811f94bf506633fa033

Observation 0eb6dbf2-6d28-47ca-992f-176b78bfe290 · outbound

This paper cites Frontiers in Computer Science , author =.

LEMUR: Learning to Align with Multi-Objective Reinforcement Learning from Preference Feedback Frontiers in Computer Science , author =

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-03T04:39:30.745384Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T04:39:30.745384Z digest=sha256:12906f3319bddcc2a66bda5855e91d4297a7b2b82007b647d3dbb083a4858eff

Observation 8f962405-36d3-43db-af86-be6858457f23 · outbound

This paper cites International Journal of Social Robotics , author =.

LEMUR: Learning to Align with Multi-Objective Reinforcement Learning from Preference Feedback International Journal of Social Robotics , author =

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-03T04:39:30.749576Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T04:39:30.749576Z digest=sha256:89e8ead73f159e7552890805456cff3f306f11a2232119697065011d5aee773d

Observation c7e68047-3b0c-4162-a5c2-ee5573d5470f · outbound

This paper cites Analysis of.

LEMUR: Learning to Align with Multi-Objective Reinforcement Learning from Preference Feedback Analysis of

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-03T04:39:30.756277Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T04:39:30.756277Z digest=sha256:815f3446de4d6e2ebe9529844a1a45237836c34e9d4092f700ee02f0513e6236

Observation 08808c1b-9707-47da-bdc9-b8130d669f69 · outbound

This paper cites Multimodal Technologies and Interaction , author =.

LEMUR: Learning to Align with Multi-Objective Reinforcement Learning from Preference Feedback Multimodal Technologies and Interaction , author =

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-03T04:39:30.759538Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T04:39:30.759538Z digest=sha256:d11b2e143bee1228e3f824d431faffae61d8a29ab306da100621ca681f75d07f

Observation cb74602b-fb75-4018-a1c1-72a8ffa45802 · outbound

This paper cites Trends in Cognitive Sciences , author =.

LEMUR: Learning to Align with Multi-Objective Reinforcement Learning from Preference Feedback Trends in Cognitive Sciences , author =

Reference 48

Resolution
verified exact
doi, observed 2026-08-03T04:44:22.332375Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-03T04:39:30.762969Z digest=sha256:fa8a4102b31857c54e296f490f2e4277b2078b2cb064efaac5e1432c38a21f65

Observation c759fdb9-7b23-49b2-ba9f-5eaa6cf3527d · outbound

This paper cites iScience , author =.

LEMUR: Learning to Align with Multi-Objective Reinforcement Learning from Preference Feedback iScience , author =

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-03T04:39:30.767177Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T04:39:30.767177Z digest=sha256:79582649f276152956d42bda183233131ec1ef54e2c6f5fa2031981e37f5d9c8

Observation 50facedf-7392-42dc-814d-530e0afdf03e · outbound

This paper cites Journal of Artificial Intelligence Research , author =.

LEMUR: Learning to Align with Multi-Objective Reinforcement Learning from Preference Feedback Journal of Artificial Intelligence Research , author =

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-03T04:39:30.770465Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T04:39:30.770465Z digest=sha256:ab5905ed39b58b8dfc88a46af807a6914389d22ec9d15f68dabf5341c864cf2e

Observation 63418822-c77a-42ba-a1df-f2bd9c35d25e · outbound

This paper cites Transformers are.

LEMUR: Learning to Align with Multi-Objective Reinforcement Learning from Preference Feedback Transformers are

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-03T04:39:30.774322Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T04:39:30.774322Z digest=sha256:33488bbeab13a6d292bb4eba2fbc58dafdf6728c4e5a95f6034df26c0889176a

Observation a2e3a93c-482e-4e7c-a128-769964af40e0 · outbound

This paper cites Evaluating Agents using Social Choice Theory.

LEMUR: Learning to Align with Multi-Objective Reinforcement Learning from Preference Feedback Evaluating Agents using Social Choice Theory

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-03T04:39:30.777817Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T04:39:30.777817Z digest=sha256:2bfc020cb6b086fc2a86a78d0b5d07ff420a0d7f91017ac4441cc65f73e8c59b

Observation 013cc6e2-fc5f-43b4-b00a-c9d1ca894a43 · outbound

This paper cites Developing, Evaluating and Scaling Learning Agents in Multi-Agent Environments.

LEMUR: Learning to Align with Multi-Objective Reinforcement Learning from Preference Feedback Developing, Evaluating and Scaling Learning Agents in Multi-Agent Environments

Reference 53

Resolution
metadata mismatch
local_arxiv, observed 2026-08-03T04:44:22.209340Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-03T04:39:30.781448Z digest=sha256:09bd3e6371ece49253e043e380a8e052b1218711bc5d5d757b8b1bede88de291

Observation 82229c54-26b0-4761-9832-e67686709713 · outbound

This paper cites Autonomous Agents and Multi-Agent Systems , author =.

LEMUR: Learning to Align with Multi-Objective Reinforcement Learning from Preference Feedback Autonomous Agents and Multi-Agent Systems , author =

Reference 54

Resolution
verified exact
doi, observed 2026-08-03T04:44:22.123359Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-03T04:39:30.785712Z digest=sha256:e0a351e29786114a5ce964ecbdeeb8732c8cf84afd68e710d7dc5e483b74e321

Observation 1d8f45c5-c3be-4bd4-aee9-03c779c17589 · outbound

This paper cites Journal of Artificial Intelligence Research , author =.

LEMUR: Learning to Align with Multi-Objective Reinforcement Learning from Preference Feedback Journal of Artificial Intelligence Research , author =

Reference 55

Resolution
verified exact
doi, observed 2026-08-03T04:44:21.966221Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-03T04:39:30.789290Z digest=sha256:a18a1d2b3697bcf031df9d071862a48620450d518d0e2a8be1a0e7031ec0a200

Observation a9e4b278-bcf0-49c4-b9de-dde5656552a6 · outbound

This paper cites Journal of Artificial Intelligence Research , author =.

LEMUR: Learning to Align with Multi-Objective Reinforcement Learning from Preference Feedback Journal of Artificial Intelligence Research , author =

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-03T04:39:30.793340Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T04:39:30.793340Z digest=sha256:0f30157a7f0dfa1f365d39e53a99b570e5e6f4dbb3db177f129a9be7d890c15a

Observation 31449dda-d39a-46d5-b32e-a85bc33eebba · outbound

This paper cites Neural Architecture Search: Insights from 1000 Papers.

LEMUR: Learning to Align with Multi-Objective Reinforcement Learning from Preference Feedback Neural Architecture Search: Insights from 1000 Papers

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-03T04:39:30.797264Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T04:39:30.797264Z digest=sha256:52ca093d7cb186a482dea8e548d317f513d612548e3fea6e758632400b55ae73

Observation b5f807f2-f37b-4114-ab1a-3798a692f2aa · outbound

This paper cites and Barto, Andrew , year =.

LEMUR: Learning to Align with Multi-Objective Reinforcement Learning from Preference Feedback and Barto, Andrew , year =

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-03T04:39:30.800852Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T04:39:30.800852Z digest=sha256:58d744ceceeafaff6e07b9dce4aed9099be82a745dbc3845e1d52e8f8a354a64

Observation 19a06b44-d4d5-48c1-8b96-26123e292310 · outbound

This paper cites and Barto, Andrew G.

LEMUR: Learning to Align with Multi-Objective Reinforcement Learning from Preference Feedback and Barto, Andrew G

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-03T04:39:30.803950Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T04:39:30.803950Z digest=sha256:a8b850e076c9a514a3c820e44e214f0f97c9a048e07dd7a343c5b0681aed3117

Observation 9894208b-911b-4749-8fe2-e5180db8bb08 · outbound

This paper cites Royal Society Open Science , author =.

LEMUR: Learning to Align with Multi-Objective Reinforcement Learning from Preference Feedback Royal Society Open Science , author =

Reference 61

Resolution
verified exact
doi, observed 2026-08-03T04:44:21.806905Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-03T04:39:30.813042Z digest=sha256:e8add18000d602ccb4e7df38119054e7c911c185b681c797909b8e46b96c2fea

Observation 21b12909-a796-4a19-a9a3-5e000d6edab9 · outbound

This paper cites A Review of Cooperation in Multi-agent Learning.

LEMUR: Learning to Align with Multi-Objective Reinforcement Learning from Preference Feedback A Review of Cooperation in Multi-agent Learning

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-03T04:39:30.816292Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T04:39:30.816292Z digest=sha256:3473a87a39aafe42334d9305429b2791795516e7bb9f15ea3884c2b986866742

Observation 63bd049d-e006-4ea7-84f5-e410c8be53f5 · outbound

This paper cites , month = oct, year =.

LEMUR: Learning to Align with Multi-Objective Reinforcement Learning from Preference Feedback , month = oct, year =

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-03T04:39:30.820337Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T04:39:30.820337Z digest=sha256:35f06173d1411d242f7d9a9b993601987168630b16316bb30c8c8faf75b78dd0

Observation 6498bee1-836a-4139-ad8a-e68c6e9b103c · outbound

This paper cites Mediated.

LEMUR: Learning to Align with Multi-Objective Reinforcement Learning from Preference Feedback Mediated

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-03T04:39:30.823576Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T04:39:30.823576Z digest=sha256:cf334fcd644d9a8cf161290c4f14f7af2769f9840e75633f9a3c2ec7a4f73204

Observation 5c2cf02a-971c-473f-8edf-f60a5a12e900 · outbound

This paper cites Autonomous Agents and Multi-Agent Systems , author =.

LEMUR: Learning to Align with Multi-Objective Reinforcement Learning from Preference Feedback Autonomous Agents and Multi-Agent Systems , author =

Reference 65

Resolution
verified exact
doi, observed 2026-08-03T04:44:21.678380Z

Source-reported events for the cited work

correction dated 2024-06-07. Source: crossref record 10.1007/s10458-024-09654-9->10.1007/s10458-024-09649-6:correction, observed 2026-07-11T03:00:36.460696+00:00. This notice travels one citation hop only.

source=arxiv_source observed=2026-08-03T04:39:30.827182Z digest=sha256:a014e4fc3b36363aae8d427174af7a8e27f96205dd795acb41b2f98da5a4bc56

Observation 9817d800-dc4a-4ad2-8ebc-132e96241b54 · outbound

This paper cites Cognition , author =.

LEMUR: Learning to Align with Multi-Objective Reinforcement Learning from Preference Feedback Cognition , author =

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-03T04:39:30.830453Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T04:39:30.830453Z digest=sha256:8f6ce0da0a2d22638cd80e56f6b76a03890fdd23ac69f7dd06593a026e0dab65

Observation f2eee9dd-060e-4613-97fd-9fdbe0974ca5 · outbound

This paper cites and Everett, Richard and Weidinger, Laura and Isaac, William S.

LEMUR: Learning to Align with Multi-Objective Reinforcement Learning from Preference Feedback and Everett, Richard and Weidinger, Laura and Isaac, William S

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-03T04:39:30.833868Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T04:39:30.833868Z digest=sha256:2e1ae4193fd6f9c5c1e78938bd7baef2309a0aa614a3fa92059196b588f32e8b

Observation 60ba10fd-b2c5-4b05-b302-883f468d5c08 · outbound

This paper cites Cooperative.

LEMUR: Learning to Align with Multi-Objective Reinforcement Learning from Preference Feedback Cooperative

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-03T04:39:30.838430Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T04:39:30.838430Z digest=sha256:e7a42018573bcf782cb71d9fb283507811ca382201144bdc5c0ad14bd26b0ba0

Observation 3ee80319-2cc3-4ffc-9cbd-c32069438a0a · outbound

This paper cites IEEE Robotics and Automation Letters , author =.

LEMUR: Learning to Align with Multi-Objective Reinforcement Learning from Preference Feedback IEEE Robotics and Automation Letters , author =

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-03T04:39:30.846473Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T04:39:30.846473Z digest=sha256:f3c3b477a0f05ba942927aa2842297767450eb18042abd23ce959536976fdb6f

Observation 2141b031-3153-4c47-9e74-12baf8384c26 · outbound

This paper cites Current Robotics Reports , author =.

LEMUR: Learning to Align with Multi-Objective Reinforcement Learning from Preference Feedback Current Robotics Reports , author =

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-03T04:39:30.861005Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T04:39:30.861005Z digest=sha256:d4f5032b9e78f08e049c6535e379dcbdeeda11dc0804211fe27021f21bc9411f

Observation 944e8b32-f415-499c-ab9b-dde16231b2a5 · outbound

This paper cites Social Cognitive and Affective Neuroscience , author =.

LEMUR: Learning to Align with Multi-Objective Reinforcement Learning from Preference Feedback Social Cognitive and Affective Neuroscience , author =

Reference 71

Resolution
verified exact
doi, observed 2026-08-03T04:44:21.512071Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-03T04:39:30.877323Z digest=sha256:1397759d98187aa07a2543d4242fa8b63f6be6dc4a016d7041d604d6d0d22852

Observation 0a0f1814-5932-41ea-8f1e-0aee02a71c29 · outbound

This paper cites Trends in Cognitive Sciences , author =.

LEMUR: Learning to Align with Multi-Objective Reinforcement Learning from Preference Feedback Trends in Cognitive Sciences , author =

Reference 72

Resolution
verified exact
doi, observed 2026-08-03T04:44:21.381159Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-03T04:39:30.893139Z digest=sha256:be4c01492ff010523422c3a06c208aac4307a339a81f93dcffb0fe3efd3367ac

Observation a6865620-a6b8-4435-8907-a69076c020d0 · outbound

This paper cites Computational Intelligence and Neuroscience , author =.

LEMUR: Learning to Align with Multi-Objective Reinforcement Learning from Preference Feedback Computational Intelligence and Neuroscience , author =

Reference 73

Resolution
verified exact
doi, observed 2026-08-03T04:44:21.205263Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-03T04:39:30.905584Z digest=sha256:69f76ebd9ae3d314b951d9b0672f192617b4ab11ada58d5bbeb4b05d9dd27277

Observation 55af0887-b944-430a-82f4-8788f7724e23 · outbound

This paper cites Adapting a.

LEMUR: Learning to Align with Multi-Objective Reinforcement Learning from Preference Feedback Adapting a

Reference 74

Resolution
unresolved
no resolver link, observed 2026-08-03T04:39:30.915570Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T04:39:30.915570Z digest=sha256:203e3cd65beeaa6ba1b675375038a82aeec523c6967262cc6d9aac9eb8574706

Observation bc90fbd5-be26-4c4f-8889-a23294a9334e · outbound

This paper cites IEEE Transactions on Affective Computing , author =.

LEMUR: Learning to Align with Multi-Objective Reinforcement Learning from Preference Feedback IEEE Transactions on Affective Computing , author =

Reference 75

Resolution
unresolved
no resolver link, observed 2026-08-03T04:39:30.931146Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T04:39:30.931146Z digest=sha256:030ffbb74ec76847d517aba5939a7b88be9bba02e1050d247ee8ccaefdd20ae4

Observation d1fd4bd7-c2b8-46bc-b200-c7178fe8802a · outbound

This paper cites an unresolved cited work.

LEMUR: Learning to Align with Multi-Objective Reinforcement Learning from Preference Feedback Unresolved cited work

Reference 76

Resolution
unresolved
no resolver link, observed 2026-08-03T04:39:30.948606Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T04:39:30.948606Z digest=sha256:8f93af0ace677f9a65e004e885b7a861f9daaf66763d0a6cdcbe39cdf7acc775

Observation e3abd1ed-5039-4116-9bfd-306a0b9298a2 · outbound

This paper cites Metin and Yemez, Yucel , month = aug, year =.

LEMUR: Learning to Align with Multi-Objective Reinforcement Learning from Preference Feedback Metin and Yemez, Yucel , month = aug, year =

Reference 77

Resolution
unresolved
no resolver link, observed 2026-08-03T04:39:30.960528Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T04:39:30.960528Z digest=sha256:1bbd72ae1aca0e2c76764c53273792c913283a10024a461604192570d2b59670

Observation 26e52079-457f-4d75-b913-1c032f5ac3ab · outbound

This paper cites Metin and Yemez, Yucel , month = aug, year =.

LEMUR: Learning to Align with Multi-Objective Reinforcement Learning from Preference Feedback Metin and Yemez, Yucel , month = aug, year =

Reference 78

Resolution
unresolved
no resolver link, observed 2026-08-03T04:39:30.972625Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T04:39:30.972625Z digest=sha256:bf8f3edd8a6678cb4b098e094513a2c7f384d3a26f70417999ecb0a8e4b88c85

Observation 97dbda7e-315a-49f6-8356-7dd4b3812788 · outbound

This paper cites an unresolved cited work.

LEMUR: Learning to Align with Multi-Objective Reinforcement Learning from Preference Feedback Unresolved cited work

Reference 79

Resolution
unresolved
no resolver link, observed 2026-08-03T04:39:30.984006Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T04:39:30.984006Z digest=sha256:abc022d3e2dc59795c1392e02ba07ca8e32cdef3437dad20facfce4a77224dda

Observation 9e32b8c5-426e-40cc-91f6-084daaee667c · outbound

This paper cites Human-to-.

LEMUR: Learning to Align with Multi-Objective Reinforcement Learning from Preference Feedback Human-to-

Reference 80

Resolution
unresolved
no resolver link, observed 2026-08-03T04:39:31.060609Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T04:39:31.060609Z digest=sha256:2b792be0dbc4ac29234226603310c2da9d38a4dae00107245918cb1cbf57f4be

Observation eee5a56b-dfad-48f7-a4fd-52b13a7b474e · outbound

This paper cites IEEE Robotics and Automation Letters , author =.

LEMUR: Learning to Align with Multi-Objective Reinforcement Learning from Preference Feedback IEEE Robotics and Automation Letters , author =

Reference 81

Resolution
unresolved
no resolver link, observed 2026-08-03T04:39:31.177488Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T04:39:31.177488Z digest=sha256:5ca7e2c46909e0d8c38b0666861442b7f7e874fc5fa9cdedba7ad4564d6d02f0

Observation 80d2053f-f062-430e-813e-56aed56d01c7 · outbound

This paper cites Efficient.

LEMUR: Learning to Align with Multi-Objective Reinforcement Learning from Preference Feedback Efficient

Reference 82

Resolution
unresolved
no resolver link, observed 2026-08-03T04:39:31.289728Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T04:39:31.289728Z digest=sha256:cbe7736c4bcc31dd26211f7f18f7b865760b5bf79b35c021c6d5dba09fa5c76c

Observation d2362b2b-3b77-45c1-96b3-14db60f9121c · outbound

This paper cites Modeling.

LEMUR: Learning to Align with Multi-Objective Reinforcement Learning from Preference Feedback Modeling

Reference 83

Resolution
unresolved
no resolver link, observed 2026-08-03T04:39:31.385547Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T04:39:31.385547Z digest=sha256:05e1e2ae6750b1bf0151c80db137c9a56f257355f0548fb461d53159913cf3cf

Observation 398d1d3a-0f5b-4192-8465-e7616806cbf7 · outbound

This paper cites Machine Learning , author =.

LEMUR: Learning to Align with Multi-Objective Reinforcement Learning from Preference Feedback Machine Learning , author =

Reference 84

Resolution
verified exact
doi, observed 2026-08-03T04:44:21.076143Z

Source-reported events for the cited work

correction dated 2022-12-28. Source: crossref record 10.1007/s10994-022-06298-2->10.1007/s10994-022-06273-x:correction, observed 2026-07-11T02:56:29.207212+00:00. This notice travels one citation hop only.

source=arxiv_source observed=2026-08-03T04:39:31.481959Z digest=sha256:ade7e494c404e8f90c654f0fc7f01f093088f99789ce04161841853b64f1f931

Observation aac3b22b-e924-4590-a1fa-6a6bdd1207d3 · outbound

This paper cites an unresolved cited work.

LEMUR: Learning to Align with Multi-Objective Reinforcement Learning from Preference Feedback Unresolved cited work

Reference 85

Resolution
unresolved
no resolver link, observed 2026-08-03T04:39:31.547361Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T04:39:31.547361Z digest=sha256:72e9dec9fcf1fee0d2a51aa17ea39c8223bd591994db3c80ca28e3595796fdea

Observation dfc7ef3b-3178-45c8-924a-2f082fa68de0 · outbound

This paper cites Emotional.

LEMUR: Learning to Align with Multi-Objective Reinforcement Learning from Preference Feedback Emotional

Reference 86

Resolution
unresolved
no resolver link, observed 2026-08-03T04:39:31.664294Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T04:39:31.664294Z digest=sha256:3dee3cd62af2e2a67266885a78162e9551b389951e66148002a3d19351427e6f

Observation 5f42a059-110c-4e68-a3ef-896b353f0183 · outbound

This paper cites and Islam, Usman and Willis, Richard and Sunehag, Peter , month = dec, year =.

LEMUR: Learning to Align with Multi-Objective Reinforcement Learning from Preference Feedback and Islam, Usman and Willis, Richard and Sunehag, Peter , month = dec, year =

Reference 87

Resolution
unresolved
no resolver link, observed 2026-08-03T04:39:31.667664Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T04:39:31.667664Z digest=sha256:f133e38051e20a4f3e504a6e2d840bd03b508b0fac28fe2f6522bfb21bbedefa

Observation 58d7e322-5396-48c8-a347-3319934c83fd · outbound

This paper cites and Gemp, Ian and McWilliams, Brian and Duéñez-Guzmán, Edgar A.

LEMUR: Learning to Align with Multi-Objective Reinforcement Learning from Preference Feedback and Gemp, Ian and McWilliams, Brian and Duéñez-Guzmán, Edgar A

Reference 88

Resolution
unresolved
no resolver link, observed 2026-08-03T04:39:31.670843Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T04:39:31.670843Z digest=sha256:3a04092459d0ef603b91008841ea3488e8a29d4ddc47a8cc665c5aa8b0522de3

Observation 6c916df2-d1b0-4243-9d6e-f782a21519ef · outbound

This paper cites an unresolved cited work.

LEMUR: Learning to Align with Multi-Objective Reinforcement Learning from Preference Feedback Unresolved cited work

Reference 89

Resolution
unresolved
no resolver link, observed 2026-08-03T04:39:31.674019Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T04:39:31.674019Z digest=sha256:7236675d7a73a35ba1f80751d46c72c54aafb0f993db47e61d7bcdd97a5ad7ac

Observation 10f9f5b4-8f97-4e32-b704-623d55bcffed · outbound

This paper cites an unresolved cited work.

LEMUR: Learning to Align with Multi-Objective Reinforcement Learning from Preference Feedback Unresolved cited work

Reference 90

Resolution
unresolved
no resolver link, observed 2026-08-03T04:39:31.677004Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T04:39:31.677004Z digest=sha256:4bf08e8795635b2b3800cbbd9ca486a579111ab75abc4f3801c16ae99493422f

Observation 57260140-fee5-4958-bad2-e8a50ee4028e · outbound

This paper cites an unresolved cited work.

LEMUR: Learning to Align with Multi-Objective Reinforcement Learning from Preference Feedback Unresolved cited work

Reference 91

Resolution
unresolved
no resolver link, observed 2026-08-03T04:39:31.680086Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T04:39:31.680086Z digest=sha256:d175d11060f22374258678cf870cbcf81e9291a7f5e7090c3184bffea7d03d2b

Observation 480e41c0-4ef8-4e12-acaa-0ace24f19522 · outbound

This paper cites Autonomous Agents and Multi-Agent Systems , author =.

LEMUR: Learning to Align with Multi-Objective Reinforcement Learning from Preference Feedback Autonomous Agents and Multi-Agent Systems , author =

Reference 92

Resolution
unresolved
no resolver link, observed 2026-08-03T04:39:31.683227Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T04:39:31.683227Z digest=sha256:2ed71b2030faf7047ece938b0d8fe90d77050af32a3ae927ef8cca9d0c350f37

Observation a5a6ac41-2cdf-45aa-890d-506b7065f227 · outbound

This paper cites Bradley and Sadigh, Dorsa , month = apr, year =.

LEMUR: Learning to Align with Multi-Objective Reinforcement Learning from Preference Feedback Bradley and Sadigh, Dorsa , month = apr, year =

Reference 93

Resolution
unresolved
no resolver link, observed 2026-08-03T04:39:31.686200Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T04:39:31.686200Z digest=sha256:0456a7ec5f9b3fd176e4164187ec45e3674d4a9542560ce147c2e713cfa6a831

Observation 9ffd6a80-e9ba-40a8-8103-9a3c3a09fd14 · outbound

This paper cites and Fidler, Sanja and Torralba, Antonio , month = may, year =.

LEMUR: Learning to Align with Multi-Objective Reinforcement Learning from Preference Feedback and Fidler, Sanja and Torralba, Antonio , month = may, year =

Reference 94

Resolution
unresolved
no resolver link, observed 2026-08-03T04:39:31.689071Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T04:39:31.689071Z digest=sha256:0a24faf60a45dac7509f0a3a0ee8a075fd213a877a44fb0b85643b866a9e7a15

Observation 99fc0100-028f-4dc4-b98e-931f6000a011 · outbound

This paper cites Promptable.

LEMUR: Learning to Align with Multi-Objective Reinforcement Learning from Preference Feedback Promptable

Reference 95

Resolution
unresolved
no resolver link, observed 2026-08-03T04:39:31.691892Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T04:39:31.691892Z digest=sha256:6c3ba20fa90455b13717f623546ea2cf7d053a517f47fc30b4796bf47e35cffb

Observation 86f42e63-cc98-44bb-a265-3c4581a50234 · outbound

This paper cites an unresolved cited work.

LEMUR: Learning to Align with Multi-Objective Reinforcement Learning from Preference Feedback Unresolved cited work

Reference 96

Resolution
unresolved
no resolver link, observed 2026-08-03T04:39:31.694946Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T04:39:31.694946Z digest=sha256:fb5b9783e11b993428c0573e3d3ae77ed39d710675ed71ed85a34d5bf200b9c1

Observation 55496758-f7d2-46bc-a06f-36b1c5706af4 · outbound

This paper cites an unresolved cited work.

LEMUR: Learning to Align with Multi-Objective Reinforcement Learning from Preference Feedback Unresolved cited work

Reference 97

Resolution
unresolved
no resolver link, observed 2026-08-03T04:39:31.697972Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T04:39:31.697972Z digest=sha256:15bf179f1a92a321afaa163eeef7ff3cb79923ede58d92073223eec1987c50c9

Observation 4a0f04da-3b3b-41a8-b155-4ce8e945408b · outbound

This paper cites an unresolved cited work.

LEMUR: Learning to Align with Multi-Objective Reinforcement Learning from Preference Feedback Unresolved cited work

Reference 98

Resolution
unresolved
no resolver link, observed 2026-08-03T04:39:31.701073Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T04:39:31.701073Z digest=sha256:5a9f8a4dcd5eac8679d0adf1b450079b0c33e01cc4181f4c48d4b27f64546df9

Observation a3936488-a1cb-4a90-912f-d7bf327a38e6 · outbound

This paper cites an unresolved cited work.

LEMUR: Learning to Align with Multi-Objective Reinforcement Learning from Preference Feedback Unresolved cited work

Reference 99

Resolution
unresolved
no resolver link, observed 2026-08-03T04:39:31.704035Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T04:39:31.704035Z digest=sha256:34bd440db72ab7a957c22c005373953c42a69cad7920cfc5119d74ffe68fb0c8

Observation 8e997296-835a-4343-b4d5-2f0a78e07893 · outbound

This paper cites Cognitive Systems Research , author =.

LEMUR: Learning to Align with Multi-Objective Reinforcement Learning from Preference Feedback Cognitive Systems Research , author =

Reference 100

Resolution
verified exact
doi, observed 2026-08-03T04:44:20.901824Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-03T04:39:31.707129Z digest=sha256:2978787c8dcb3a25beab1a837486787a217740d43885ef94540462189b0de8af

Observation 0f55f6cf-6310-4e71-bc2c-05897c0bd980 · outbound

This paper cites Artificial Intelligence Review , author =.

LEMUR: Learning to Align with Multi-Objective Reinforcement Learning from Preference Feedback Artificial Intelligence Review , author =

Reference 101

Resolution
unresolved
no resolver link, observed 2026-08-03T04:39:31.710251Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T04:39:31.710251Z digest=sha256:30bc5dcf9721c1fbfaab60d00fa8cd8a3a8787612f7881c1ae7f95225d08c1dc

Observation 0f746c6a-e146-45f4-8560-67a2c894362d · outbound

This paper cites ACM Transactions on Intelligent Systems and Technology , author =.

LEMUR: Learning to Align with Multi-Objective Reinforcement Learning from Preference Feedback ACM Transactions on Intelligent Systems and Technology , author =

Reference 102

Resolution
verified exact
doi, observed 2026-08-03T04:44:20.819365Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-03T04:39:31.713492Z digest=sha256:09afcbd819be34b13834b530d1ce99b2dfdc5cac240bf8927fb772650b2f76f1

Pith citing papers

No inbound Pith citation observations are available.