Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-15T20:11:50.133035Z
Paper Citation Record · LEDGER
As of 19 August 2026, this Paper Citation Record lists 74 of 74 outbound references and 5 inbound Pith citation observations for arXiv:2506.12822.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-15T20:11:50.133035Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-19T06:32:44.657259+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-06T10:13:19.750837Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-07-04T09:59:44.448427Z
74 of 74 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 68ee3853-d796-4007-b9af-4e41c84ac977 · outbound
Enhancing Rating-Based Reinforcement Learning to Effectively Leverage Feedback from Large Vision-Language Models write newline
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f9489836-0ef5-42a7-b7e3-5561118e5122 · outbound
Enhancing Rating-Based Reinforcement Learning to Effectively Leverage Feedback from Large Vision-Language Models Deep Reinforcement Learning from Policy-Dependent Human Feedback
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 65d87bfe-81af-4a10-a255-9fe155452616 · outbound
Enhancing Rating-Based Reinforcement Learning to Effectively Leverage Feedback from Large Vision-Language Models Constitutional AI: Harmlessness from AI Feedback
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dce72d38-18ba-417f-b6dc-5f4f7fdd2b76 · outbound
Enhancing Rating-Based Reinforcement Learning to Effectively Leverage Feedback from Large Vision-Language Models AutoBench-V: Can Large Vision-Language Models Benchmark Themselves?
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a8affdad-f919-43ec-a38c-de0c8d33911c · outbound
Enhancing Rating-Based Reinforcement Learning to Effectively Leverage Feedback from Large Vision-Language Models Vision-Language Models as a Source of Rewards
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2f01d4f4-886d-4ae6-af93-5f83a825f8a0 · outbound
Enhancing Rating-Based Reinforcement Learning to Effectively Leverage Feedback from Large Vision-Language Models G., Candido, S., Castro, P
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 2e3019e8-f9cf-41f4-972b-636ed83a6da7 · outbound
Enhancing Rating-Based Reinforcement Learning to Effectively Leverage Feedback from Large Vision-Language Models K., Scheurer, J., Rando, J., Freedman, R., Korbak, T., Lindner, D., Freire, P., et al
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 860e48b7-e2c2-4796-922c-af17bf59873c · outbound
Enhancing Rating-Based Reinforcement Learning to Effectively Leverage Feedback from Large Vision-Language Models Towards human-level bimanual dexterous manipulation with reinforcement learning
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 0660c414-b87b-48e7-934a-bcacd23dec3a · outbound
Enhancing Rating-Based Reinforcement Learning to Effectively Leverage Feedback from Large Vision-Language Models F., Leike, J., Brown, T., Martic, M., Legg, S., and Amodei, D
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 58359cb2-5f24-4f9b-bb9f-37e620a6e465 · outbound
Enhancing Rating-Based Reinforcement Learning to Effectively Leverage Feedback from Large Vision-Language Models Can foundation models perform zero-shot task specification for robot manipulation? In Learning for Dynamics and Control Conference (L4DC), 2022
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation b4dae81b-de05-4d3f-bbc8-4c1f9d69ceb9 · outbound
Enhancing Rating-Based Reinforcement Learning to Effectively Leverage Feedback from Large Vision-Language Models K., Joty, S., Li, B., and Bing, L
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 9287fbae-e1a6-4cb9-9258-fc38125a756c · outbound
Enhancing Rating-Based Reinforcement Learning to Effectively Leverage Feedback from Large Vision-Language Models Minedojo: Building open-ended embodied agents with internet-scale knowledge
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation ee3c3b87-86ad-403d-8cb1-a56d8f250120 · outbound
Enhancing Rating-Based Reinforcement Learning to Effectively Leverage Feedback from Large Vision-Language Models Guided cost learning: Deep inverse optimal control via policy optimization
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation d7061f3f-87ab-42b9-902a-69ec7285a0ca · outbound
Enhancing Rating-Based Reinforcement Learning to Effectively Leverage Feedback from Large Vision-Language Models Learning robust rewards with adversarial inverse reinforcement learning
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation cd8c370e-555d-4377-9cff-e77261b78781 · outbound
Enhancing Rating-Based Reinforcement Learning to Effectively Leverage Feedback from Large Vision-Language Models Unresolved cited work
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 96291906-b98e-4abe-83df-0f85346a9661 · outbound
Enhancing Rating-Based Reinforcement Learning to Effectively Leverage Feedback from Large Vision-Language Models Chatgpt outperforms crowd workers for text-annotation tasks
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation cba57434-9900-4784-99cb-471828cd9122 · outbound
Enhancing Rating-Based Reinforcement Learning to Effectively Leverage Feedback from Large Vision-Language Models B., and Kambhampati, S
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation e0aff4bf-415a-4e57-b994-a914491521ae · outbound
Enhancing Rating-Based Reinforcement Learning to Effectively Leverage Feedback from Large Vision-Language Models Soft actor-critic: Off-policy maximum entropy deep reinforcement learning with a stochastic actor
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b367b4f9-25ee-4e4e-87b4-8cffe889b291 · outbound
Enhancing Rating-Based Reinforcement Learning to Effectively Leverage Feedback from Large Vision-Language Models Autoreward: Closed-loop reward design with large language models for autonomous driving
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation b29f3da4-f1ce-4fa2-b596-371ae53ded88 · outbound
Enhancing Rating-Based Reinforcement Learning to Effectively Leverage Feedback from Large Vision-Language Models Deep residual learning for image recognition
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 56e91edb-ca8a-4af1-ab82-53145a5911d7 · outbound
Enhancing Rating-Based Reinforcement Learning to Effectively Leverage Feedback from Large Vision-Language Models Unresolved cited work
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 64db9b0d-ac07-469d-8703-f57432542c2e · outbound
Enhancing Rating-Based Reinforcement Learning to Effectively Leverage Feedback from Large Vision-Language Models and Ermon, S
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 7665c5da-b4a6-4c52-8913-6ed6d4587f0b · outbound
Enhancing Rating-Based Reinforcement Learning to Effectively Leverage Feedback from Large Vision-Language Models Reward learning from human preferences and demonstrations in atari
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 99c935c1-5c6d-45e7-bbdb-109e99b889bc · outbound
Enhancing Rating-Based Reinforcement Learning to Effectively Leverage Feedback from Large Vision-Language Models Reinforcement learning friendly vision-language model for minecraft
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 1e42208a-231a-402e-bcae-58c2a232a244 · outbound
Enhancing Rating-Based Reinforcement Learning to Effectively Leverage Feedback from Large Vision-Language Models Scalable deep reinforcement learning for vision-based robotic manipulation
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation a739b6b5-dc1e-4991-bc8b-cbb1d9b92bf1 · outbound
Enhancing Rating-Based Reinforcement Learning to Effectively Leverage Feedback from Large Vision-Language Models Champion-level drone racing using deep reinforcement learning
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 46f564a4-a23e-4fbd-a944-409edba38e42 · outbound
Enhancing Rating-Based Reinforcement Learning to Effectively Leverage Feedback from Large Vision-Language Models Motif: Intrinsic Motivation from Artificial Intelligence Feedback
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 32cf2daa-3a5f-413f-961b-8853995d9ca0 · outbound
Enhancing Rating-Based Reinforcement Learning to Effectively Leverage Feedback from Large Vision-Language Models Unresolved cited work
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 9e467907-f43b-4ea8-8428-c96654c184ec · outbound
Enhancing Rating-Based Reinforcement Learning to Effectively Leverage Feedback from Large Vision-Language Models Offline reinforcement learning with implicit q-learning
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation b6705ab9-4b1c-48ed-a82d-f860aa98633c · outbound
Enhancing Rating-Based Reinforcement Learning to Effectively Leverage Feedback from Large Vision-Language Models M., Bullard, K., and Sadigh, D
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 78b54101-5bfe-455c-b8be-a59f587bc846 · outbound
Enhancing Rating-Based Reinforcement Learning to Effectively Leverage Feedback from Large Vision-Language Models R., Bishop, C., Hall, E., Carbune, V., Rastogi, A., et al
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation ffeee6bf-655e-42ed-a489-0e9af6233fc0 · outbound
Enhancing Rating-Based Reinforcement Learning to Effectively Leverage Feedback from Large Vision-Language Models PEBBLE: Feedback-Efficient Interactive Reinforcement Learning via Relabeling Experience and Unsupervised Pre-training
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5013d33a-a277-4ec7-98f6-b4739f917d35 · outbound
Enhancing Rating-Based Reinforcement Learning to Effectively Leverage Feedback from Large Vision-Language Models B-pref: Benchmarking preference-based reinforcement learning
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 37b4b192-0527-42df-b159-8b5db85c7832 · outbound
Enhancing Rating-Based Reinforcement Learning to Effectively Leverage Feedback from Large Vision-Language Models Scalable agent alignment via reward modeling: a research direction
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ac8d4815-f657-43a6-b599-22ea01fb1a56 · outbound
Enhancing Rating-Based Reinforcement Learning to Effectively Leverage Feedback from Large Vision-Language Models Nonlinear inverse reinforcement learning with gaussian processes
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 4aa0930a-324b-4efe-922d-c1deadfc2893 · outbound
Enhancing Rating-Based Reinforcement Learning to Effectively Leverage Feedback from Large Vision-Language Models Silkie: Preference Distillation for Large Visual Language Models
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 88383d7e-6338-4a20-8ef8-969b5ba4b932 · outbound
Enhancing Rating-Based Reinforcement Learning to Effectively Leverage Feedback from Large Vision-Language Models M., Stepputtis, S., Campbell, J., and Sycara, K
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 9fe1f92b-d999-4cec-8231-4b9433cb29cb · outbound
Enhancing Rating-Based Reinforcement Learning to Effectively Leverage Feedback from Large Vision-Language Models J., Kumar, V., Zhang, A., Bastani, O., and Jayaraman, D
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 22492e8d-ccf7-48ed-bee9-3f42daa534ad · outbound
Enhancing Rating-Based Reinforcement Learning to Effectively Leverage Feedback from Large Vision-Language Models J., Liang, W., Wang, G., Huang, D.-A., Bastani, O., Jayaraman, D., Zhu, Y., Fan, L., and Anandkumar, A
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 286effa6-3e25-47d1-994a-81615bd9a976 · outbound
Enhancing Rating-Based Reinforcement Learning to Effectively Leverage Feedback from Large Vision-Language Models L., Muresan, S., Squire, S., Tellex, S., Arumugam, D., and Yang, L
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation a0be76ed-85e9-421c-983b-f99d72739127 · outbound
Enhancing Rating-Based Reinforcement Learning to Effectively Leverage Feedback from Large Vision-Language Models K., Loftin, R., Peng, B., Wang, G., Roberts, D
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 506d9324-f27b-40fb-bcf9-8b2ba9bfef7a · outbound
Enhancing Rating-Based Reinforcement Learning to Effectively Leverage Feedback from Large Vision-Language Models Zero-shot reward specification via grounded natural language
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation ad59621a-3199-4a47-a542-f36ce6db24e7 · outbound
Enhancing Rating-Based Reinforcement Learning to Effectively Leverage Feedback from Large Vision-Language Models R3m: A universal visual representation for robot manipulation
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 9a240ec4-d408-4b05-9466-77bf85cc2931 · outbound
Enhancing Rating-Based Reinforcement Learning to Effectively Leverage Feedback from Large Vision-Language Models T., Burch, N., Anthony, T., et al
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 7cc25b73-e6f9-474c-8e1b-9bbb409c0048 · outbound
Enhancing Rating-Based Reinforcement Learning to Effectively Leverage Feedback from Large Vision-Language Models W., Hallacy, C., Ramesh, A., Goh, G., Agarwal, S., Sastry, G., Askell, A., Mishkin, P., Clark, J., et al
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 79f6e373-a879-4232-a29c-ab44d0342d36 · outbound
Enhancing Rating-Based Reinforcement Learning to Effectively Leverage Feedback from Large Vision-Language Models Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context
Reference 46
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a4ec7725-1129-417a-b37c-6c56cc759b92 · outbound
Enhancing Rating-Based Reinforcement Learning to Effectively Leverage Feedback from Large Vision-Language Models Vision-language models are zero-shot reward models for reinforcement learning
Reference 47
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 152e4a5f-c165-4bbf-b6ea-99fa8a4866d3 · outbound
Enhancing Rating-Based Reinforcement Learning to Effectively Leverage Feedback from Large Vision-Language Models Alfred: A benchmark for interpreting grounded instructions for everyday tasks
Reference 48
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 47344c4e-5fef-4ef5-8a8b-101afd9f329e · outbound
Enhancing Rating-Based Reinforcement Learning to Effectively Leverage Feedback from Large Vision-Language Models Mastering the game of go without human knowledge
Reference 49
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation fb2e8de6-62c5-4a73-b20b-371a725d6b4d · outbound
Enhancing Rating-Based Reinforcement Learning to Effectively Leverage Feedback from Large Vision-Language Models Defining and characterizing reward gaming
Reference 50
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation bc7cffd5-1760-4a80-b29c-ed8ab87adccf · outbound
Enhancing Rating-Based Reinforcement Learning to Effectively Leverage Feedback from Large Vision-Language Models Roboclip: One demonstration is enough to learn robot policies
Reference 51
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation bf5652b2-7b96-4a8d-b8e8-691262f6813c · outbound
Enhancing Rating-Based Reinforcement Learning to Effectively Leverage Feedback from Large Vision-Language Models Unresolved cited work
Reference 52
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 31100d27-f70c-4fa8-adb3-8340e9ff7d8a · outbound
Enhancing Rating-Based Reinforcement Learning to Effectively Leverage Feedback from Large Vision-Language Models BC-IRL: Learning Generalizable Reward Functions from Demonstrations
Reference 53
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation b82791c4-b971-4804-91ea-9bb40540b6c1 · outbound
Enhancing Rating-Based Reinforcement Learning to Effectively Leverage Feedback from Large Vision-Language Models N., Klissarov, M., Precup, D., Yang, S., and Anand, A
Reference 54
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 3501a304-cb98-4603-9b4d-5673e051e416 · outbound
Enhancing Rating-Based Reinforcement Learning to Effectively Leverage Feedback from Large Vision-Language Models M., Mathieu, M., Dudzik, A., Chung, J., Choi, D
Reference 55
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 92a42e9a-0cbe-4c9a-bf1e-62977c6c1554 · outbound
Enhancing Rating-Based Reinforcement Learning to Effectively Leverage Feedback from Large Vision-Language Models Prefclm: Enhancing preference-based reinforcement learning with crowdsourced large language models
Reference 56
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 03bfdcc6-f154-466d-99a4-dde754301ab3 · outbound
Enhancing Rating-Based Reinforcement Learning to Effectively Leverage Feedback from Large Vision-Language Models Rl-vlm-f: Reinforcement learning from vision language foundation model feedback
Reference 57
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation e0471fd0-7f57-4b2c-af2a-4befcde083e3 · outbound
Enhancing Rating-Based Reinforcement Learning to Effectively Leverage Feedback from Large Vision-Language Models Robogen: Towards unleashing infinite data for automated robot learning via generative simulation
Reference 58
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 0edab5db-26c1-4525-876f-eed1a70c5a89 · outbound
Enhancing Rating-Based Reinforcement Learning to Effectively Leverage Feedback from Large Vision-Language Models Deep tamer: Interactive agent shaping in high-dimensional state spaces
Reference 59
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 3f135604-048a-4d78-9985-8e7d0b55331c · outbound
Enhancing Rating-Based Reinforcement Learning to Effectively Leverage Feedback from Large Vision-Language Models To Smooth or Not? When Label Smoothing Meets Noisy Labels
Reference 60
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f4722e6c-bd5d-4f97-be78-d59bd0e4c1b4 · outbound
Enhancing Rating-Based Reinforcement Learning to Effectively Leverage Feedback from Large Vision-Language Models J., Waytowich, N., and Cao, Y
Reference 61
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 0454f6b4-69e9-4380-8b8d-e7e885f68c85 · outbound
Enhancing Rating-Based Reinforcement Learning to Effectively Leverage Feedback from Large Vision-Language Models A survey of preference-based reinforcement learning methods
Reference 62
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 8cbf753c-9b97-412e-aba5-7d66e7933de1 · outbound
Enhancing Rating-Based Reinforcement Learning to Effectively Leverage Feedback from Large Vision-Language Models Maximum Entropy Deep Inverse Reinforcement Learning
Reference 63
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1f3c16bc-3931-4810-a9e1-23dff2cc7c20 · outbound
Enhancing Rating-Based Reinforcement Learning to Effectively Leverage Feedback from Large Vision-Language Models Rethinking spatiotemporal feature learning: Speed-accuracy trade-offs in video classification
Reference 64
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 6c125c0c-8ac2-4e67-8025-1f79f779982a · outbound
Enhancing Rating-Based Reinforcement Learning to Effectively Leverage Feedback from Large Vision-Language Models H., Liu, Y., Luo, Q., Zhong, V., Yang, Y., and Yu, T
Reference 65
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 8dfe1053-c665-45b2-97a7-4efdf4604d4b · outbound
Enhancing Rating-Based Reinforcement Learning to Effectively Leverage Feedback from Large Vision-Language Models S., Hasegawa-Johnson, M
Reference 66
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 76f92c68-e6b3-4c0e-9d94-6299c2633139 · outbound
Enhancing Rating-Based Reinforcement Learning to Effectively Leverage Feedback from Large Vision-Language Models Meta-world: A benchmark and evaluation for multi-task and meta reinforcement learning
Reference 67
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation fd1e9ae3-03b0-441f-993c-9d2c3098e3ff · outbound
Enhancing Rating-Based Reinforcement Learning to Effectively Leverage Feedback from Large Vision-Language Models Rlaif-v: Aligning mllms through open-source ai feedback for super gpt-4v trustworthiness
Reference 68
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 50d0fdfb-a02a-45a9-8138-a4f0d9b62ad1 · outbound
Enhancing Rating-Based Reinforcement Learning to Effectively Leverage Feedback from Large Vision-Language Models Uni-rlhf: Universal platform and benchmark suite for reinforcement learning with diverse human feedback
Reference 69
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 8843c7e3-8d65-444d-a5d2-8a30acfaeeab · outbound
Enhancing Rating-Based Reinforcement Learning to Effectively Leverage Feedback from Large Vision-Language Models Unresolved cited work
Reference 70
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 3283974b-49e9-405c-a254-8fabfc8cf2ac · outbound
Enhancing Rating-Based Reinforcement Learning to Effectively Leverage Feedback from Large Vision-Language Models Vision-language models for vision tasks: A survey
Reference 71
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation fd38fa2a-c157-4b9f-9580-a45c191aad7c · outbound
Enhancing Rating-Based Reinforcement Learning to Effectively Leverage Feedback from Large Vision-Language Models Siren's Song in the AI Ocean: A Survey on Hallucination in Large Language Models
Reference 72
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 15e08b51-856f-419a-bd78-e2a3c9a54d08 · outbound
Enhancing Rating-Based Reinforcement Learning to Effectively Leverage Feedback from Large Vision-Language Models Unresolved cited work
Reference 73
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 1c7435ff-d998-4abb-88c6-544dba34c2c7 · outbound
Reference 74
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 8674e6ba-bf86-4071-b0a2-26340077f810 · inbound
Occlusion-robust Stylization for Drawing-based 3D Animation Enhancing Rating-Based Reinforcement Learning to Effectively Leverage Feedback from Large Vision-Language Models
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5b704c8c-fc4a-4656-9aa9-60b8aa3f330d · inbound
Self-Rewarding Vision-Language Model via Reasoning Decomposition Enhancing Rating-Based Reinforcement Learning to Effectively Leverage Feedback from Large Vision-Language Models
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 0e9d3458-f541-44cc-bb1f-2ae14dbf8d0c · inbound
SOLE-R1: Video-Language Reasoning as the Sole Reward for On-Robot Reinforcement Learning Enhancing Rating-Based Reinforcement Learning to Effectively Leverage Feedback from Large Vision-Language Models
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 97bc0aea-a88f-49e0-9bdf-0e44fc281504 · inbound
Learning Process Rewards via Success Visitation Matching for Efficient RL Enhancing Rating-Based Reinforcement Learning to Effectively Leverage Feedback from Large Vision-Language Models
Reference 51
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 38e1e6d2-6dc9-4c28-bcfd-3ae2414ab0c7 · inbound
Towards General Language-Conditioned Latent Safety Filters Enhancing Rating-Based Reinforcement Learning to Effectively Leverage Feedback from Large Vision-Language Models
Reference 213
Source-reported events for the cited work
Unavailable: canonical work link unavailable.