Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-14T04:40:11.975172Z
Paper Citation Record · LEDGER
As of 14 August 2026, this Paper Citation Record lists 52 of 52 outbound references and 0 inbound Pith citation observations for arXiv:2608.08491.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-14T04:40:11.975172Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-14T06:32:32.682623+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
52 of 52 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation e0d22af6-881a-40a9-ad56-cfcaa581314c · outbound
TrustRoboReward: Preference-Ordered Isotonic Score Editing for Multi-Paradigm Robot Reward Models Eureka: Human-level reward design via coding large language models
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation e5ab305d-2dae-49b4-8184-b6bda825693b · outbound
TrustRoboReward: Preference-Ordered Isotonic Score Editing for Multi-Paradigm Robot Reward Models Serl: A software suite for sample- efficient robotic reinforcement learning
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 62a73c9e-3392-40cf-a88e-febe6f022191 · outbound
TrustRoboReward: Preference-Ordered Isotonic Score Editing for Multi-Paradigm Robot Reward Models Visual instruction tuning.Advances in neural information processing systems, 36:34892–34916, 2023
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2e38b2f5-f651-47ad-923c-a4f7d76c52ea · outbound
TrustRoboReward: Preference-Ordered Isotonic Score Editing for Multi-Paradigm Robot Reward Models Improving vision-language-action model with online reinforcement learning
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 769055bf-55da-4452-90ad-e1cce2da1523 · outbound
TrustRoboReward: Preference-Ordered Isotonic Score Editing for Multi-Paradigm Robot Reward Models Robot-r1: Reinforcement learning for enhanced embodied reasoning in robotics
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 9ca47e5f-3f36-4b14-8605-de65615983c1 · outbound
TrustRoboReward: Preference-Ordered Isotonic Score Editing for Multi-Paradigm Robot Reward Models Reinforcement learning with foundation priors: Let embodied agent efficiently learn on its own
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 30d88430-1e91-41cd-822f-e86ab0170b9f · outbound
TrustRoboReward: Preference-Ordered Isotonic Score Editing for Multi-Paradigm Robot Reward Models Self-improving embodied foundation models
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation d9f4bd21-72a3-4fa7-9909-efd9b3646c86 · outbound
TrustRoboReward: Preference-Ordered Isotonic Score Editing for Multi-Paradigm Robot Reward Models Deep reinforcement learning from human preferences.Advances in neural information processing systems, 30, 2017
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ee2fe596-3c97-4093-a37c-1e5df48ade05 · outbound
TrustRoboReward: Preference-Ordered Isotonic Score Editing for Multi-Paradigm Robot Reward Models Wainwright, Pamela Mishkin, Chong Zhang, Sandhini Agarwal, Katarina Slama, Alex Ray, et al
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 6d2e1e5f-42ba-49b0-9dfc-961bd938fc13 · outbound
TrustRoboReward: Preference-Ordered Isotonic Score Editing for Multi-Paradigm Robot Reward Models Roboreward: General-purpose vision-language reward models for robotics.arXiv preprint arXiv:2601.00675, 2026
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c6ced9ac-b537-47b1-b6e7-e1ea237941f2 · outbound
TrustRoboReward: Preference-Ordered Isotonic Score Editing for Multi-Paradigm Robot Reward Models DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5e30c658-ab9d-462b-9014-8f3557ebefcc · outbound
TrustRoboReward: Preference-Ordered Isotonic Score Editing for Multi-Paradigm Robot Reward Models Deepseek-r1: Incentivizing reasoning capability in llms via reinforcement learning.arXiv e-prints, pages arXiv–2501, 2025
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation c4470bf2-8fcc-4120-8bdb-1f8db33533c1 · outbound
TrustRoboReward: Preference-Ordered Isotonic Score Editing for Multi-Paradigm Robot Reward Models rstar-math: Small llms can master math reasoning with self-evolved deep thinking
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 1cddd270-4b8e-4d7d-9528-9d61ca4b08e2 · outbound
TrustRoboReward: Preference-Ordered Isotonic Score Editing for Multi-Paradigm Robot Reward Models Training software engineering agents and verifiers with swe-gym
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 8a491032-55db-4aa9-8f3f-b0b755e49a0d · outbound
TrustRoboReward: Preference-Ordered Isotonic Score Editing for Multi-Paradigm Robot Reward Models Swe-bench: Can language models resolve real-world github issues? In The twelfth international conference on learning representations, 2023
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation aab7bcd4-4428-4309-883a-25b4b258a4de · outbound
TrustRoboReward: Preference-Ordered Isotonic Score Editing for Multi-Paradigm Robot Reward Models Swe-agent: Agent-computer interfaces enable automated software engineering.Advances in Neural Information Processing Systems, 37:50528–50652, 2024
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0f46069f-7c1b-4654-84df-8d5951222a0d · outbound
TrustRoboReward: Preference-Ordered Isotonic Score Editing for Multi-Paradigm Robot Reward Models OpenHands: An Open Platform for AI Software Developers as Generalist Agents
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cd4ec506-4b13-428a-944a-5586424a8fdd · outbound
TrustRoboReward: Preference-Ordered Isotonic Score Editing for Multi-Paradigm Robot Reward Models GPT-4 Technical Report
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0244c97d-84c3-426d-b3c0-5cac0892ee80 · outbound
TrustRoboReward: Preference-Ordered Isotonic Score Editing for Multi-Paradigm Robot Reward Models Internvl: Scaling up vision foundation models and aligning for generic visual-linguistic tasks
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cbd3c701-2e96-4d91-a28f-00fa08833227 · outbound
TrustRoboReward: Preference-Ordered Isotonic Score Editing for Multi-Paradigm Robot Reward Models Qwen3-VL Technical Report
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 49f5f19b-2114-465f-b9f8-0c06d8428758 · outbound
TrustRoboReward: Preference-Ordered Isotonic Score Editing for Multi-Paradigm Robot Reward Models Manning, and Chelsea Finn
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b0ec75f3-7a02-4059-91b8-766d49da58c8 · outbound
TrustRoboReward: Preference-Ordered Isotonic Score Editing for Multi-Paradigm Robot Reward Models Unresolved cited work
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9bae433f-b067-4f58-be23-2d5f145d7305 · outbound
TrustRoboReward: Preference-Ordered Isotonic Score Editing for Multi-Paradigm Robot Reward Models RLAIF vs. RLHF: Scaling Reinforcement Learning from Human Feedback with AI Feedback
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 24d777e0-76c9-4ab5-85bd-eae9eff9b2a4 · outbound
TrustRoboReward: Preference-Ordered Isotonic Score Editing for Multi-Paradigm Robot Reward Models Alpacaeval: An automatic evaluator of instruction-following models, 2023
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6fb073d5-c527-4045-a412-cf0a26b973ef · outbound
TrustRoboReward: Preference-Ordered Isotonic Score Editing for Multi-Paradigm Robot Reward Models Chatbot Arena: An Open Platform for Evaluating LLMs by Human Preference
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2c9bbe76-1df1-4aef-ae18-83594b368b6d · outbound
TrustRoboReward: Preference-Ordered Isotonic Score Editing for Multi-Paradigm Robot Reward Models Trustjudge: Inconsistencies of LLM-as-a-judge and how to alleviate them
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 95d835a8-80d7-4953-a92f-9aba5429a9ae · outbound
TrustRoboReward: Preference-Ordered Isotonic Score Editing for Multi-Paradigm Robot Reward Models OpenAI GPT-5 System Card
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3d3d6085-8313-4b13-a5d8-55ee0111d7e2 · outbound
TrustRoboReward: Preference-Ordered Isotonic Score Editing for Multi-Paradigm Robot Reward Models Pai-bench: A comprehensive benchmark for physical ai.arXiv preprint arXiv:2512.01989, 2025
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 567042e8-a546-4d6a-9c96-f9deddd1aa5c · outbound
TrustRoboReward: Preference-Ordered Isotonic Score Editing for Multi-Paradigm Robot Reward Models Llava-onevision: Easy visual task transfer.Transactions on Machine Learning Research
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation aa3b390b-492d-4fd5-967a-aeddbd515426 · outbound
TrustRoboReward: Preference-Ordered Isotonic Score Editing for Multi-Paradigm Robot Reward Models Xing, Hao Zhang, Joseph E
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1a132033-3cb5-484d-8f74-39ce15a92cb4 · outbound
TrustRoboReward: Preference-Ordered Isotonic Score Editing for Multi-Paradigm Robot Reward Models Prometheus-vision: Vision-language model as a judge for fine-grained evaluation
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0358f445-0b76-4706-ad16-d7655157e7d1 · outbound
TrustRoboReward: Preference-Ordered Isotonic Score Editing for Multi-Paradigm Robot Reward Models MLLM-as-a-Judge: Assessing multimodal LLM-as-a-Judge with vision-language benchmark
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 2c624792-f540-4523-b288-24a52e8861ae · outbound
TrustRoboReward: Preference-Ordered Isotonic Score Editing for Multi-Paradigm Robot Reward Models Llava-critic: Learning to evaluate multimodal models
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e25d5e47-2cdd-4167-bc22-087b125ba711 · outbound
TrustRoboReward: Preference-Ordered Isotonic Score Editing for Multi-Paradigm Robot Reward Models Generative Reward Models
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation be7c3ea8-06e3-45d4-a8a0-e9b0d15f2db7 · outbound
TrustRoboReward: Preference-Ordered Isotonic Score Editing for Multi-Paradigm Robot Reward Models CompassJudger-1: All-in-one Judge Model Helps Model Evaluation and Evolution
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 82c510b3-bbc4-4023-8563-b5d3aeec15a6 · outbound
TrustRoboReward: Preference-Ordered Isotonic Score Editing for Multi-Paradigm Robot Reward Models Vision- language models are zero-shot reward models for reinforcement learning
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation d15df90e-4803-44d6-92a7-80e66e763c95 · outbound
TrustRoboReward: Preference-Ordered Isotonic Score Editing for Multi-Paradigm Robot Reward Models Rl-vlm-f: reinforcement learning from vision language foundation model feedback
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 0884635c-e168-4f7e-8c97-f4efbcac098c · outbound
TrustRoboReward: Preference-Ordered Isotonic Score Editing for Multi-Paradigm Robot Reward Models Vision-Language Models as a Source of Rewards
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e6b023bb-0e25-414a-9ac6-145d4b8800f1 · outbound
TrustRoboReward: Preference-Ordered Isotonic Score Editing for Multi-Paradigm Robot Reward Models VisionReward: Fine-grained multi-dimensional human preference learning for image and video generation
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 2ed783aa-f9e8-42c0-9c80-2013255c69df · outbound
TrustRoboReward: Preference-Ordered Isotonic Score Editing for Multi-Paradigm Robot Reward Models Improving video generation with human feedback
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation aa38349e-b5d5-4685-97c6-03eaab10ba6b · outbound
TrustRoboReward: Preference-Ordered Isotonic Score Editing for Multi-Paradigm Robot Reward Models RewardBench: Evaluating Reward Models for Language Modeling
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9be46e54-4553-46b7-9318-4825b8aaabbb · outbound
TrustRoboReward: Preference-Ordered Isotonic Score Editing for Multi-Paradigm Robot Reward Models Multimodal RewardBench: Holistic Evaluation of Reward Models for Vision Language Models
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 04a0d552-7e46-4fe0-8ec3-db08307b25d3 · outbound
TrustRoboReward: Preference-Ordered Isotonic Score Editing for Multi-Paradigm Robot Reward Models Vl-rewardbench: a challenging benchmark for vision-language generative reward models
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e0a2c348-bf0c-4fa1-a269-d6549e76e10b · outbound
TrustRoboReward: Preference-Ordered Isotonic Score Editing for Multi-Paradigm Robot Reward Models Simpo: Simple preference optimization with a reference-free reward.Advances in Neural Information Processing Systems, 37:124198–124235, 2024
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 33621821-7d61-40bb-ab70-e5d5e7c53fb5 · outbound
TrustRoboReward: Preference-Ordered Isotonic Score Editing for Multi-Paradigm Robot Reward Models Aligning with human judgement: The role of pairwise preference in large language model evaluators
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 85df1f86-a21a-437e-95d8-6f8232afdfb2 · outbound
TrustRoboReward: Preference-Ordered Isotonic Score Editing for Multi-Paradigm Robot Reward Models Constitutional AI: Harmlessness from AI Feedback
Reference 46
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1d27a4f9-312c-4712-b3c9-eb31e566a8ad · outbound
TrustRoboReward: Preference-Ordered Isotonic Score Editing for Multi-Paradigm Robot Reward Models Meta-rewarding language models: Self-improving alignment with llm-as-a-meta-judge
Reference 47
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1f9e4c15-1c26-4f84-8480-61c12bb86da4 · outbound
TrustRoboReward: Preference-Ordered Isotonic Score Editing for Multi-Paradigm Robot Reward Models An empirical distribution function for sampling with incomplete information.The annals of mathematical statistics, pages 641–647, 1955
Reference 48
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 6daa94cf-372b-4674-9224-44b2b1c4959a · outbound
TrustRoboReward: Preference-Ordered Isotonic Score Editing for Multi-Paradigm Robot Reward Models Finetuned language models are zero-shot learners
Reference 49
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5c570ab1-c0c9-4ba6-85ed-58cc0c1e2242 · outbound
TrustRoboReward: Preference-Ordered Isotonic Score Editing for Multi-Paradigm Robot Reward Models Swift: a scalable lightweight infrastructure for fine-tuning
Reference 50
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 49b5de5b-6a5f-423a-bddb-39e2d77891dc · outbound
TrustRoboReward: Preference-Ordered Isotonic Score Editing for Multi-Paradigm Robot Reward Models Hybridflow: A flexible and efficient rlhf framework
Reference 51
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5231abc1-8303-491c-bc12-84ef2fe769e3 · outbound
TrustRoboReward: Preference-Ordered Isotonic Score Editing for Multi-Paradigm Robot Reward Models World Simulation with Video Foundation Models for Physical AI
Reference 52
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
No inbound Pith citation observations are available.