Pith. sign in

Paper Citation Record · LEDGER

CLARIFY: Contrastive Preference Reinforcement Learning for Untangling Ambiguous Queries

As of 7 August 2026, this Paper Citation Record lists 40 of 40 outbound references and 0 inbound Pith citation observations for arXiv:2506.00388.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.00388 v3

Coverage vector

measured 40 of 40 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T12:13:01.213190Z

measured 40 of 40 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

40 of 40 outbound references displayed

  • verified exact2
  • verified fuzzy9
  • unresolved29
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 250a93c1-e861-440f-b914-2164a8bd24a1 · outbound

This paper cites G., Dabney, W., and Munos, R.

CLARIFY: Contrastive Preference Reinforcement Learning for Untangling Ambiguous Queries G., Dabney, W., and Munos, R

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T12:12:57.074379Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:12:57.074379Z digest=sha256:adc2a8804081fc5abc67f5b79221e573e245d0fcb6404c880c3ad68f778f5850

Observation 54d7213e-edb2-437b-95bc-fa5e10e26b29 · outbound

This paper cites G., Candido, S., Castro, P.

CLARIFY: Contrastive Preference Reinforcement Learning for Untangling Ambiguous Queries G., Candido, S., Castro, P

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:13:03.592044Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T12:12:57.147263Z digest=sha256:cf2505a07fc3dcd7107f148a26a32b423a965d13ad3a82e20c9aa357dc950335

Observation e04203ac-c923-4752-a3f6-a23e54f4bb58 · outbound

This paper cites Learning to distinguish: shared perceptual features and discrimination practice tune behavioural pattern separation.

CLARIFY: Contrastive Preference Reinforcement Learning for Untangling Ambiguous Queries Learning to distinguish: shared perceptual features and discrimination practice tune behavioural pattern separation

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:13:03.418448Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T12:12:57.238918Z digest=sha256:ff3419db415d0dc1e38cfb241e877beaf294844903b5ef1c0abfbca4370a2d47

Observation 1a505b6a-ca0d-44e2-aba9-e49c090cce6d · outbound

This paper cites Active preference-based gaussian process regression for reward learning.

CLARIFY: Contrastive Preference Reinforcement Learning for Untangling Ambiguous Queries Active preference-based gaussian process regression for reward learning

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:13:03.268428Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T12:12:57.330118Z digest=sha256:543f7a0751349b6f1a216044cbd17ff0f9370ec0ab0d29fa0056c9c3a8fb585f

Observation e479e1df-6574-48d8-a2e8-b097915e16e3 · outbound

This paper cites an unresolved cited work.

CLARIFY: Contrastive Preference Reinforcement Learning for Untangling Ambiguous Queries Unresolved cited work

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T12:12:57.354573Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:12:57.354573Z digest=sha256:f319be1d741ea0553d4caca0db5a17ce582569755e9f3059ab227c69a289f011

Observation 1bd4b567-0992-476d-aa06-2b0964c20d66 · outbound

This paper cites and He, K.

CLARIFY: Contrastive Preference Reinforcement Learning for Untangling Ambiguous Queries and He, K

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T12:12:57.439208Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:12:57.439208Z digest=sha256:a2c8fc04c62c3cc8359b97903b28df0b6120b9fc2c6989e6b733b400f80da04d

Observation d7ec4a76-b6b0-47f6-ae4f-ccec4cdd5be6 · outbound

This paper cites RIME: Robust Preference-based Reinforcement Learning with Noisy Preferences.

CLARIFY: Contrastive Preference Reinforcement Learning for Untangling Ambiguous Queries RIME: Robust Preference-based Reinforcement Learning with Noisy Preferences

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T12:12:57.502734Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:12:57.502734Z digest=sha256:8601ea974c19acc0ff8141860935d49ab3e63f08b2539aa4f59dc27c53ccfd63

Observation c9248b55-c092-436d-9845-69455364ba71 · outbound

This paper cites Listwise Reward Estimation for Offline Preference-based Reinforcement Learning.

CLARIFY: Contrastive Preference Reinforcement Learning for Untangling Ambiguous Queries Listwise Reward Estimation for Offline Preference-based Reinforcement Learning

Reference 8

Resolution
verified exact
local_arxiv, observed 2026-08-07T12:13:01.753321Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T12:12:57.547010Z digest=sha256:550620ce8ca85f21fcce413df416ad7ecad2b384238ac9aa3505e2adbcdd1583

Observation 89a3358b-236d-486f-80f2-80e68e37d0a8 · outbound

This paper cites Learning a similarity metric discriminatively, with application to face verification.

CLARIFY: Contrastive Preference Reinforcement Learning for Untangling Ambiguous Queries Learning a similarity metric discriminatively, with application to face verification

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T12:12:57.626849Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:12:57.626849Z digest=sha256:ae4a2fce8119bf33b544d68132decc66828f45fdfb575516ed9d1e04f0cecc8f

Observation 43919007-ec05-4297-9da6-36c5f86c37bd · outbound

This paper cites F., Leike, J., Brown, T., Martic, M., Legg, S., and Amodei, D.

CLARIFY: Contrastive Preference Reinforcement Learning for Untangling Ambiguous Queries F., Leike, J., Brown, T., Martic, M., Legg, S., and Amodei, D

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T12:12:57.702927Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:12:57.702927Z digest=sha256:7c56db244688a9b609452c95c445d1d307fb847f013370b1f25d90ae333717aa

Observation 1a34d678-3436-4a3a-a4fe-65d1b7b3e432 · outbound

This paper cites D4RL: Datasets for Deep Data-Driven Reinforcement Learning.

CLARIFY: Contrastive Preference Reinforcement Learning for Untangling Ambiguous Queries D4RL: Datasets for Deep Data-Driven Reinforcement Learning

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T12:12:57.766405Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:12:57.766405Z digest=sha256:1a49264dd0975acfebde97fc50de85b552d8a56aedc809ee9cc5fd2d31664217

Observation abc8b36e-fc44-4deb-ad0b-d09b727f5b1d · outbound

This paper cites Generalized Decision Transformer for Offline Hindsight Information Matching.

CLARIFY: Contrastive Preference Reinforcement Learning for Untangling Ambiguous Queries Generalized Decision Transformer for Offline Hindsight Information Matching

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T12:12:57.859163Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:12:57.859163Z digest=sha256:eff4fa16f22c8d5f74918eb318d59d652ae2ba6459fd7d36043bfc7be02c7473

Observation b45fc730-8cc9-4621-9f44-ee2ae67fa062 · outbound

This paper cites an unresolved cited work.

CLARIFY: Contrastive Preference Reinforcement Learning for Untangling Ambiguous Queries Unresolved cited work

Reference 13

Resolution
unresolved
raw_fallback, observed 2026-08-07T12:13:03.000325Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T12:12:57.997015Z digest=sha256:dc0a3edf8fac98369801e81c5a5772475c8b1518bf8786a8c4df38d218ad10a9

Observation 1179e094-1d9b-4423-bdc4-75bde49d1d28 · outbound

This paper cites Hindsight Preference Learning for Offline Preference-based Reinforcement Learning.

CLARIFY: Contrastive Preference Reinforcement Learning for Untangling Ambiguous Queries Hindsight Preference Learning for Offline Preference-based Reinforcement Learning

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T12:12:58.048180Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:12:58.048180Z digest=sha256:a127398540a708b058710db9eda98a105377ecbcd06153804dc28da8b174315e

Observation b71a3d6c-8600-487c-8108-c3687755c31a · outbound

This paper cites and Sadigh, D.

CLARIFY: Contrastive Preference Reinforcement Learning for Untangling Ambiguous Queries and Sadigh, D

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:13:02.802723Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T12:12:58.112002Z digest=sha256:a336558bd6e0d95fe48bdd3efb7690cff5267b8cecb221063bacb6a0fa800b93

Observation 144323eb-6a86-404c-afb8-5cdfbfdcd49f · outbound

This paper cites Reward learning from human preferences and demonstrations in atari.

CLARIFY: Contrastive Preference Reinforcement Learning for Untangling Ambiguous Queries Reward learning from human preferences and demonstrations in atari

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T12:12:58.174843Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:12:58.174843Z digest=sha256:973585786549dd298180df68b9bde363780dfaadb7e9fa0b3e75c6aca854279b

Observation 0378777a-f6aa-451c-8e86-d440816a9297 · outbound

This paper cites Tuning in to how neurons distinguish between stimuli.

CLARIFY: Contrastive Preference Reinforcement Learning for Untangling Ambiguous Queries Tuning in to how neurons distinguish between stimuli

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:13:02.486917Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T12:12:58.256962Z digest=sha256:4751ba828a3d141d0f8a0a45166b1465d517919d8db8a67307b72eda03ccd4e8

Observation 99d4ff69-fea4-443c-83dc-edd97301a97c · outbound

This paper cites Episodic novelty through temporal distance.

CLARIFY: Contrastive Preference Reinforcement Learning for Untangling Ambiguous Queries Episodic novelty through temporal distance

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:13:02.287093Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T12:12:58.309735Z digest=sha256:bbc05869238a1730fcf0abcb48f481c83d0a6f61e29e1d4f53e77ce93d2a64f6

Observation 769b3f9d-210d-4324-9761-b8f37d1b5367 · outbound

This paper cites Transferring policy of deep reinforcement learning from simulation to reality for robotics.

CLARIFY: Contrastive Preference Reinforcement Learning for Untangling Ambiguous Queries Transferring policy of deep reinforcement learning from simulation to reality for robotics

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:13:02.157844Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T12:12:58.354488Z digest=sha256:62255d83cc42e8cfa4c487c114abe1611466148b954dc92d7d0c06a7148e31e1

Observation e33ead37-5bc7-4215-84bf-5ac5a8506a86 · outbound

This paper cites Scalable deep reinforcement learning for vision-based robotic manipulation.

CLARIFY: Contrastive Preference Reinforcement Learning for Untangling Ambiguous Queries Scalable deep reinforcement learning for vision-based robotic manipulation

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T12:12:58.463092Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:12:58.463092Z digest=sha256:46964287e2c5319c1d5fddc870d9f0bf878d6e1a9cea4146c9e5b6be526fb914

Observation 39893ad9-687e-45ad-949f-f43fe432dbf1 · outbound

This paper cites Beyond Reward: Offline Preference-guided Policy Optimization.

CLARIFY: Contrastive Preference Reinforcement Learning for Untangling Ambiguous Queries Beyond Reward: Offline Preference-guided Policy Optimization

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T12:12:58.566244Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:12:58.566244Z digest=sha256:7a1f5e93afc3632f9300130d8c9b1156343f572cd62cabc47b1ccee92bc59a6e

Observation f1aa0140-78bb-4265-974a-256780365eb6 · outbound

This paper cites Preference transformer: Modeling human preferences using transformers for rl.

CLARIFY: Contrastive Preference Reinforcement Learning for Untangling Ambiguous Queries Preference transformer: Modeling human preferences using transformers for rl

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:13:02.066230Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T12:12:58.705934Z digest=sha256:f88d02e5bd86858ca022816ce7b5a7742a9fe1c0985caa9cbd8e328f563571d7

Observation 573e15d2-3710-4a67-9edd-851b0dc63bef · outbound

This paper cites Offline Reinforcement Learning with Implicit Q-Learning.

CLARIFY: Contrastive Preference Reinforcement Learning for Untangling Ambiguous Queries Offline Reinforcement Learning with Implicit Q-Learning

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T12:12:58.782816Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:12:58.782816Z digest=sha256:6a8b17b3010171faaeda3cc2cffde5b9fd4c9fbdefce44ef278b49f6b6650e62

Observation 90389d52-fd13-4079-84f8-66984933191a · outbound

This paper cites Curl: Contrastive unsupervised representations for reinforcement learning.

CLARIFY: Contrastive Preference Reinforcement Learning for Untangling Ambiguous Queries Curl: Contrastive unsupervised representations for reinforcement learning

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T12:12:58.924071Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:12:58.924071Z digest=sha256:a8cd9d701de764df7579a2a120102e4c84d0e4ff465ae3219773de19ebce3b67

Observation 8190da58-22ab-43de-a3f7-162654f53e83 · outbound

This paper cites PEBBLE: Feedback-Efficient Interactive Reinforcement Learning via Relabeling Experience and Unsupervised Pre-training.

CLARIFY: Contrastive Preference Reinforcement Learning for Untangling Ambiguous Queries PEBBLE: Feedback-Efficient Interactive Reinforcement Learning via Relabeling Experience and Unsupervised Pre-training

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T12:12:59.030007Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:12:59.030007Z digest=sha256:03abee6c2b43971b5690a319a3758abcb9c36bd8ee17bb8bddaa2642651c4d96

Observation 6c8e3ae6-9185-403b-869a-96c023284f05 · outbound

This paper cites B-Pref: Benchmarking Preference-Based Reinforcement Learning.

CLARIFY: Contrastive Preference Reinforcement Learning for Untangling Ambiguous Queries B-Pref: Benchmarking Preference-Based Reinforcement Learning

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T12:12:59.151835Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:12:59.151835Z digest=sha256:bf7f7cfb6bcc851e16c69030b91aaa1875d5ea0be0e69d37c72a5e579dd95a0b

Observation f3ab3b8f-f2e9-4a75-93fe-7b11dd15946e · outbound

This paper cites Survival instinct in offline reinforcement learning.

CLARIFY: Contrastive Preference Reinforcement Learning for Untangling Ambiguous Queries Survival instinct in offline reinforcement learning

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:13:01.970434Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T12:12:59.288060Z digest=sha256:9c0839da233bfb0cc18e7f4d65839aa6058cd4233fbb589d6dbecb03decd70bf

Observation 1387e94c-3fdf-408f-b021-b70d45fbb5b9 · outbound

This paper cites Reward Uncertainty for Exploration in Preference-based Reinforcement Learning.

CLARIFY: Contrastive Preference Reinforcement Learning for Untangling Ambiguous Queries Reward Uncertainty for Exploration in Preference-based Reinforcement Learning

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T12:12:59.440374Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:12:59.440374Z digest=sha256:2605f0227436903adcab1ab4f0b290e556405e608df4a64a44fadf89b652c4d3

Observation 8e0ac5bf-e130-47d6-adab-5a92f2968f88 · outbound

This paper cites A., Veness, J., Bellemare, M.

CLARIFY: Contrastive Preference Reinforcement Learning for Untangling Ambiguous Queries A., Veness, J., Bellemare, M

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T12:12:59.633355Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:12:59.633355Z digest=sha256:3096c6e409677e8634a56cf23df42ca2806a37d6fe108757b7963a1a504356bc

Observation f89b50c1-8f96-4fc7-b1c9-87a7b865bfcf · outbound

This paper cites S-EPOA: Overcoming the Indistinguishability of Segments with Skill-Driven Preference-Based Reinforcement Learning.

CLARIFY: Contrastive Preference Reinforcement Learning for Untangling Ambiguous Queries S-EPOA: Overcoming the Indistinguishability of Segments with Skill-Driven Preference-Based Reinforcement Learning

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T12:12:59.815667Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:12:59.815667Z digest=sha256:6e46fba2da97cef9a34f22b3b81701af6603f796a20c623d7405b5fccde9665a

Observation faf88960-dd49-4059-a571-3472806c415b · outbound

This paper cites Learning Temporal Distances: Contrastive Successor Features Can Provide a Metric Structure for Decision-Making.

CLARIFY: Contrastive Preference Reinforcement Learning for Untangling Ambiguous Queries Learning Temporal Distances: Contrastive Successor Features Can Provide a Metric Structure for Decision-Making

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T12:12:59.969023Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:12:59.969023Z digest=sha256:86d42c12b773d2357ece30adab104f217c5daf34bd25fe2dc6430698308c66d3

Observation c95aa10b-79be-4370-a553-f3c490942db1 · outbound

This paper cites Representation Learning with Contrastive Predictive Coding.

CLARIFY: Contrastive Preference Reinforcement Learning for Untangling Ambiguous Queries Representation Learning with Contrastive Predictive Coding

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T12:13:00.091060Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:13:00.091060Z digest=sha256:e300ad8ebbddea32c5ba22914f821458fb79f528201095118512649a8fe02238

Observation 4c46c2c8-2b1f-4eb0-a9a4-b39f94a9886a · outbound

This paper cites Training language models to follow instructions with human feedback.

CLARIFY: Contrastive Preference Reinforcement Learning for Untangling Ambiguous Queries Training language models to follow instructions with human feedback

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T12:13:00.201111Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:13:00.201111Z digest=sha256:8706f62e935206317f78df7842be8b7216bebcbef5b106ea52df9a0b4dade079

Observation cba3728a-2163-4aaa-8eee-1e95afce2404 · outbound

This paper cites SURF: Semi-supervised Reward Learning with Data Augmentation for Feedback-efficient Preference-based Reinforcement Learning.

CLARIFY: Contrastive Preference Reinforcement Learning for Untangling Ambiguous Queries SURF: Semi-supervised Reward Learning with Data Augmentation for Feedback-efficient Preference-based Reinforcement Learning

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T12:13:00.332041Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:13:00.332041Z digest=sha256:195c28e3a9f5e28fc28051b7165ef615c884bef1e0ba5d21c77e3bab442992dd

Observation 2d234ffd-ef7e-4178-98fd-518cc1164d1e · outbound

This paper cites Trust Region Policy Optimization.

CLARIFY: Contrastive Preference Reinforcement Learning for Untangling Ambiguous Queries Trust Region Policy Optimization

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T12:13:00.476431Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:13:00.476431Z digest=sha256:027d6097935429bb6c2440190495635f42d7ef9a4661f32ce4b77358d5725a33

Observation 3931ac7b-a868-4eec-9af1-bf714c641b6c · outbound

This paper cites Benchmarks and Algorithms for Offline Preference-Based Reward Learning.

CLARIFY: Contrastive Preference Reinforcement Learning for Untangling Ambiguous Queries Benchmarks and Algorithms for Offline Preference-Based Reward Learning

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T12:13:00.672841Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:13:00.672841Z digest=sha256:308ec1e7400a75d9453a1c317f22febfa58f6095dd98327eadfb1f547e6af482

Observation 17e588d6-466d-4205-938a-469714d91c97 · outbound

This paper cites DeepMind Control Suite.

CLARIFY: Contrastive Preference Reinforcement Learning for Untangling Ambiguous Queries DeepMind Control Suite

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-07T12:13:00.774609Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:13:00.774609Z digest=sha256:9f6f9910161749450d312215d861d1e66df4b1ca880d53146c04d7638f62be7e

Observation 38bed9a6-7783-490b-a08f-063e354bd693 · outbound

This paper cites Reinforcement Learning from Diverse Human Preferences.

CLARIFY: Contrastive Preference Reinforcement Learning for Untangling Ambiguous Queries Reinforcement Learning from Diverse Human Preferences

Reference 38

Resolution
verified exact
local_arxiv, observed 2026-08-07T12:13:01.391953Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T12:13:00.884525Z digest=sha256:df6eec15219041c896191ad36e84a3c53360bc0cf37c4f16f8f343c9e18f2f36

Observation 6cb82dbc-6008-47ce-8d91-09beca49acdc · outbound

This paper cites Meta-world: A benchmark and evaluation for multi-task and meta reinforcement learning.

CLARIFY: Contrastive Preference Reinforcement Learning for Untangling Ambiguous Queries Meta-world: A benchmark and evaluation for multi-task and meta reinforcement learning

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-07T12:13:01.053360Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:13:01.053360Z digest=sha256:e40b85c5b62fafb9de8099b2c88129d0af28c7cf4268206d4eaec02876482634

Observation f5f1a679-ecbf-42ec-8793-0f69ebc1ae7a · outbound

This paper cites and Lu, Z.

CLARIFY: Contrastive Preference Reinforcement Learning for Untangling Ambiguous Queries and Lu, Z

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-07T12:13:01.213190Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:13:01.213190Z digest=sha256:c8015e0f38353264d93182178d413380626ffa8e2ba7dd49355c44381166955e

Pith citing papers

No inbound Pith citation observations are available.