Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-04T12:43:27.228812Z
Paper Citation Record · LEDGER
As of 13 August 2026, this Paper Citation Record lists 17 of 17 outbound references and 0 inbound Pith citation observations for arXiv:2510.02561.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-04T12:43:27.228812Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-13T06:32:02.005865+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
17 of 17 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 7cf321b0-87e6-485b-a24d-3d6a4a98476e · outbound
Oracle-RLAIF: An Improved Fine-Tuning Framework for Multi-modal Video Models using Reinforcement Learning from Ranking Feedback RLAIF vs. RLHF: Scaling Reinforcement Learning from Human Feedback with AI Feedback
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d511aa32-e20f-437d-a84d-cbd0157179fe · outbound
Oracle-RLAIF: An Improved Fine-Tuning Framework for Multi-modal Video Models using Reinforcement Learning from Ranking Feedback Video-llava: Learning united visual representation by alignment before projection
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 43a8ced4-544f-4856-a4f3-1850d05a8341 · outbound
Oracle-RLAIF: An Improved Fine-Tuning Framework for Multi-modal Video Models using Reinforcement Learning from Ranking Feedback doi: 10.18653/v1/2024.emnlp-main.342
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4b5a927f-3563-448b-b4f1-cbf3e29cf05d · outbound
Oracle-RLAIF: An Improved Fine-Tuning Framework for Multi-modal Video Models using Reinforcement Learning from Ranking Feedback BT-Adapter: Video Conversation is Feasible Without Video Instruction Tuning
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e96f5fa5-9c47-4287-b1a2-a73a8c623e35 · outbound
Oracle-RLAIF: An Improved Fine-Tuning Framework for Multi-modal Video Models using Reinforcement Learning from Ranking Feedback Direct preference optimization: Your language model is secretly a reward model
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3d8ab01b-b7fc-47b3-9052-01d855daeaee · outbound
Oracle-RLAIF: An Improved Fine-Tuning Framework for Multi-modal Video Models using Reinforcement Learning from Ranking Feedback Proximal Policy Optimization Algorithms
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6eda4fb1-a3c0-42f7-bc1d-8bed18076076 · outbound
Oracle-RLAIF: An Improved Fine-Tuning Framework for Multi-modal Video Models using Reinforcement Learning from Ranking Feedback LLaMA: Open and Efficient Foundation Language Models
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 116573cf-f3c5-45d8-8ec7-f40e46dac953 · outbound
Oracle-RLAIF: An Improved Fine-Tuning Framework for Multi-modal Video Models using Reinforcement Learning from Ranking Feedback Raft: Reward-ranked fine-tuning for generative foundation model alignment
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 597da5fc-a21c-406d-8a06-12c1a5a55e55 · outbound
Oracle-RLAIF: An Improved Fine-Tuning Framework for Multi-modal Video Models using Reinforcement Learning from Ranking Feedback LLaMA-Adapter: Efficient Fine-tuning of Language Models with Zero-init Attention
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fbb0c4d1-fcf3-4de5-9125-aea7c65c6655 · outbound
Oracle-RLAIF: An Improved Fine-Tuning Framework for Multi-modal Video Models using Reinforcement Learning from Ranking Feedback Unresolved cited work
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d9ea132b-402b-4138-b2d4-e438e1c1daa5 · outbound
Oracle-RLAIF: An Improved Fine-Tuning Framework for Multi-modal Video Models using Reinforcement Learning from Ranking Feedback DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
Reference 2017
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d856d017-90c0-44e8-8ee0-6df666939542 · outbound
Oracle-RLAIF: An Improved Fine-Tuning Framework for Multi-modal Video Models using Reinforcement Learning from Ranking Feedback PandaGPT: One Model To Instruction-Follow Them All
Reference 2020
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8200a115-dbf0-4055-87ca-a8a71ac058d1 · outbound
Oracle-RLAIF: An Improved Fine-Tuning Framework for Multi-modal Video Models using Reinforcement Learning from Ranking Feedback The accuracy paradox in rlhf: When better reward models don’t yield better language models
Reference 2021
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e13194ef-23b6-422f-b0dd-408c34e93c02 · outbound
Oracle-RLAIF: An Improved Fine-Tuning Framework for Multi-modal Video Models using Reinforcement Learning from Ranking Feedback Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback
Reference 2022
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 546abc83-3175-495f-8eb4-18acdfcc4ff9 · outbound
Oracle-RLAIF: An Improved Fine-Tuning Framework for Multi-modal Video Models using Reinforcement Learning from Ranking Feedback VideoChat: Chat-Centric Video Understanding
Reference 2023
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b41a9b51-1ea5-442c-9df0-2eea39130c92 · outbound
Oracle-RLAIF: An Improved Fine-Tuning Framework for Multi-modal Video Models using Reinforcement Learning from Ranking Feedback DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
Reference 2024
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cc54c339-f4b6-4761-841c-302c20bcdb77 · outbound
Oracle-RLAIF: An Improved Fine-Tuning Framework for Multi-modal Video Models using Reinforcement Learning from Ranking Feedback Video-ChatGPT: Towards Detailed Video Understanding via Large Vision and Language Models
Reference 2025
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
No inbound Pith citation observations are available.