Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-12T12:58:31.026570Z
Paper Citation Record · LEDGER
As of 21 August 2026, this Paper Citation Record lists 69 of 69 outbound references and 13 inbound Pith citation observations for arXiv:2411.16646.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-12T12:58:31.026570Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-21T06:32:19.484+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-16T06:04:14.616880Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-07-01T12:25:43.079840Z
69 of 69 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation ffe186cc-3c5b-4c67-a231-575c198eb217 · outbound
Self-Generated Critiques Boost Reward Modeling for Language Models write newline
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation be74fe89-6d27-403d-90e7-8c9b03a9fafc · outbound
Self-Generated Critiques Boost Reward Modeling for Language Models GPT-4 Technical Report
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fdc67657-8d22-4276-bdcf-5dccef2e053b · outbound
Self-Generated Critiques Boost Reward Modeling for Language Models Nemotron-4 340B Technical Report
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2da54106-6976-473d-9860-52191c7b82fd · outbound
Self-Generated Critiques Boost Reward Modeling for Language Models On-policy distillation of language models: Learning from self-generated mistakes
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 7692ca6f-1e13-4ddc-b1df-f6598b6d6efd · outbound
Self-Generated Critiques Boost Reward Modeling for Language Models Critique-out-Loud Reward Models
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5495ec06-07b3-4dc0-b3cf-22ea57457e7e · outbound
Self-Generated Critiques Boost Reward Modeling for Language Models LLMs instead of Human Judges? A Large Scale Empirical Study across 20 NLP Evaluation Tasks
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 10718ebf-7c38-4deb-ab5a-493340d7fa5e · outbound
Self-Generated Critiques Boost Reward Modeling for Language Models Rank analysis of incomplete block designs: I
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d881a421-a2f7-4499-aefa-9bc0ce2211a7 · outbound
Self-Generated Critiques Boost Reward Modeling for Language Models ODIN : Disentangled reward mitigates hacking in RLHF
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 0ed77cfe-6af7-476f-9be6-77b567048768 · outbound
Self-Generated Critiques Boost Reward Modeling for Language Models Training Verifiers to Solve Math Word Problems
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f9439820-ed0a-43d8-948e-53aee259aec3 · outbound
Self-Generated Critiques Boost Reward Modeling for Language Models Reward model ensembles help mitigate overoptimization
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation ebdcdafe-36ad-42ac-abc8-873f127fe448 · outbound
Self-Generated Critiques Boost Reward Modeling for Language Models ULTRAFEEDBACK : Boosting language models with scaled AI feedback
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 99a51931-811e-4536-8d67-03eec6ed9b91 · outbound
Self-Generated Critiques Boost Reward Modeling for Language Models Safe RLHF : Safe reinforcement learning from human feedback
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 1df86ae6-5bf9-477a-ab99-5f9324186052 · outbound
Self-Generated Critiques Boost Reward Modeling for Language Models The Llama 3 Herd of Models
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bdd7b0af-553f-4e15-b4fe-a755fbbfd4b8 · outbound
Self-Generated Critiques Boost Reward Modeling for Language Models Alpacafarm: A simulation framework for methods that learn from human feedback
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation ab384553-124d-4260-bf41-f028c093f2b9 · outbound
Self-Generated Critiques Boost Reward Modeling for Language Models Understanding dataset difficulty with V -usable information
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation d7c37332-c054-4d80-8180-4cab9d7947a4 · outbound
Self-Generated Critiques Boost Reward Modeling for Language Models Reinforced Self-Training (ReST) for Language Modeling
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b937e7ea-ae3b-41b2-839d-4f7d69ec706d · outbound
Self-Generated Critiques Boost Reward Modeling for Language Models Measuring mathematical problem solving with the MATH dataset
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 9167072b-59c6-403d-bb6a-f0beff333e73 · outbound
Self-Generated Critiques Boost Reward Modeling for Language Models Large language models are reasoning teachers
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation e3d3cfd2-2f93-47f0-a5a9-4c76d2050892 · outbound
Self-Generated Critiques Boost Reward Modeling for Language Models GPT-4o System Card
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation df515ede-6881-4dc0-883c-228befcbe788 · outbound
Self-Generated Critiques Boost Reward Modeling for Language Models Unpacking DPO and PPO: Disentangling Best Practices for Learning from Preference Feedback
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ccfdbcb1-a093-42bb-882b-351bde9e5176 · outbound
Self-Generated Critiques Boost Reward Modeling for Language Models Prometheus: Inducing fine-grained evaluation capability in language models
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 4952cc09-d83e-42c5-8f0c-60f732d50cf1 · outbound
Self-Generated Critiques Boost Reward Modeling for Language Models Prometheus 2: An Open Source Language Model Specialized in Evaluating Other Language Models
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2b1a55d5-82df-4bee-a113-6fffd0b28b8c · outbound
Self-Generated Critiques Boost Reward Modeling for Language Models Adam: A Method for Stochastic Optimization
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6826f4ad-82ca-4a33-befc-a9094d6f0b12 · outbound
Self-Generated Critiques Boost Reward Modeling for Language Models RewardBench: Evaluating Reward Models for Language Modeling
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ca8cdb4b-64be-4003-824b-f6a7f005b014 · outbound
Self-Generated Critiques Boost Reward Modeling for Language Models RLAIF vs
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation ead8f73b-a27e-49a0-a4c1-4f6efa6ad336 · outbound
Self-Generated Critiques Boost Reward Modeling for Language Models Generative judge for evaluating alignment
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 8b2bad33-2fdb-4f3d-a422-e48cd5c4495b · outbound
Self-Generated Critiques Boost Reward Modeling for Language Models Self-alignment with instruction backtranslation
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 1ebf980e-f036-43f8-9ae6-cddbb62346d3 · outbound
Self-Generated Critiques Boost Reward Modeling for Language Models Alpacaeval: An automatic evaluator of instruction-following models, 2023
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0914b2ba-0d33-4926-badd-596a3b042408 · outbound
Self-Generated Critiques Boost Reward Modeling for Language Models C ritic B ench: Benchmarking LLM s for critique-correct reasoning
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 23ba04f1-cdbe-4813-9a5e-c25a3f1e13c7 · outbound
Self-Generated Critiques Boost Reward Modeling for Language Models RRM: Robust Reward Model Training Mitigates Reward Hacking
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 487d4a13-6e52-441d-a9e3-11fe3b6751c3 · outbound
Self-Generated Critiques Boost Reward Modeling for Language Models Self-refine: Iterative refinement with self-feedback
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 9aed85a1-f604-4f0f-b5ab-0e65a84fceda · outbound
Self-Generated Critiques Boost Reward Modeling for Language Models Generative Reward Models
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6bba815f-51b3-421d-a58c-56362845ad24 · outbound
Self-Generated Critiques Boost Reward Modeling for Language Models LLM Critics Help Catch LLM Bugs
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4ddc2213-a8b0-4b6e-9aea-33b298231266 · outbound
Self-Generated Critiques Boost Reward Modeling for Language Models Introducing ChatGPT , 2022
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 5194ba0e-c31a-44f4-85d5-75a67c2e0cf9 · outbound
Self-Generated Critiques Boost Reward Modeling for Language Models Chatterji, Faisal Ladhak, and Tatsunori Hashimoto
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation dff5371a-f534-47f2-8100-5053755232a0 · outbound
Self-Generated Critiques Boost Reward Modeling for Language Models Training language models to follow instructions with human feedback
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f83f81e4-53fb-4b58-bc24-dc87258681b0 · outbound
Self-Generated Critiques Boost Reward Modeling for Language Models West-of-N: Synthetic Preferences for Self-Improving Reward Models
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fe18b1f5-4420-43cc-9fa1-796a0d13cab2 · outbound
Self-Generated Critiques Boost Reward Modeling for Language Models Iterative reasoning preference optimization
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 3a3ae537-c3f5-4657-af8a-3492ada82995 · outbound
Self-Generated Critiques Boost Reward Modeling for Language Models Self-Consistency Preference Optimization
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5a344517-4d59-4bf7-acb0-701819fa96b1 · outbound
Self-Generated Critiques Boost Reward Modeling for Language Models Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3bae4a46-f07d-4305-adf9-30f3e4002ff6 · outbound
Self-Generated Critiques Boost Reward Modeling for Language Models Self-critiquing models for assisting human evaluators
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f3241364-a4d7-4f55-9ece-e712e0410b41 · outbound
Self-Generated Critiques Boost Reward Modeling for Language Models BOND: Aligning LLMs with Best-of-N Distillation
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2073dceb-9bcd-4492-bdee-b4b33894bd74 · outbound
Self-Generated Critiques Boost Reward Modeling for Language Models Boosting Reward Model with Preference-Conditional Multi-Aspect Synthetic Data Generation
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 01e1fd14-dc31-425c-8d7e-69470379eebb · outbound
Self-Generated Critiques Boost Reward Modeling for Language Models The trickle-down impact of reward inconsistency on RLHF
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 34899596-ca26-4c55-96fb-253fc2dd3352 · outbound
Self-Generated Critiques Boost Reward Modeling for Language Models A Long Way to Go: Investigating Length Correlations in RLHF
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d9e3cab9-e24f-474d-a7aa-1fdb5f9518e3 · outbound
Self-Generated Critiques Boost Reward Modeling for Language Models Unresolved cited work
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 3646890f-9ca2-4e2c-a6b3-a73bdd7e84b4 · outbound
Self-Generated Critiques Boost Reward Modeling for Language Models Learning to summarize with human feedback
Reference 47
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f3ddc1eb-74cf-4886-95f6-612dd368a724 · outbound
Self-Generated Critiques Boost Reward Modeling for Language Models Large Language Models are Inconsistent and Biased Evaluators
Reference 48
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 03829fd2-a518-4681-a276-726bdbe383ed · outbound
Self-Generated Critiques Boost Reward Modeling for Language Models SALMON : Self-alignment with instructable reward models
Reference 49
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 87774883-6e0d-499c-b615-1a7ab9dd2eae · outbound
Self-Generated Critiques Boost Reward Modeling for Language Models Llama 2: Open Foundation and Fine-Tuned Chat Models
Reference 50
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a75f3008-98fd-41c2-9d5e-f1395e529a24 · outbound
Self-Generated Critiques Boost Reward Modeling for Language Models Interpretable preferences via multi-objective reward modeling and mixture-of-experts
Reference 51
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 68f3bf2c-00b1-44bc-a2b4-bf32d29cbcd0 · outbound
Self-Generated Critiques Boost Reward Modeling for Language Models Self-Taught Evaluators
Reference 52
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation db5b1d9e-0179-4254-8595-3c1d171a551a · outbound
Self-Generated Critiques Boost Reward Modeling for Language Models Chi, Sharan Narang, Aakanksha Chowdhery, and Denny Zhou
Reference 53
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation ae03238b-bd98-4a39-b8f1-4f0487bc72da · outbound
Self-Generated Critiques Boost Reward Modeling for Language Models HelpSteer2-Preference: Complementing Ratings with Preferences
Reference 54
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5277e0aa-39f4-4e3d-9cdb-0541f6c20deb · outbound
Self-Generated Critiques Boost Reward Modeling for Language Models HelpSteer2: Open-source dataset for training top-performing reward models
Reference 55
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 462bf89c-54f7-4454-8d6b-7376217d64f3 · outbound
Self-Generated Critiques Boost Reward Modeling for Language Models H elp S teer: Multi-attribute helpfulness dataset for S teer LM
Reference 56
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 3be7e336-c88e-4252-8529-764dd791fde2 · outbound
Self-Generated Critiques Boost Reward Modeling for Language Models Meta-Rewarding Language Models: Self-Improving Alignment with LLM-as-a-Meta-Judge
Reference 57
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c0a009cb-092a-4642-9c3f-d7f0457bafc0 · outbound
Self-Generated Critiques Boost Reward Modeling for Language Models Smith, Mari Ostendorf, and Hannaneh Hajishirzi
Reference 58
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation fe73cde1-93d5-4685-850c-a737b1a660a5 · outbound
Self-Generated Critiques Boost Reward Modeling for Language Models WizardLM: Empowering large pre-trained language models to follow complex instructions
Reference 59
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 92f1ac4a-af66-4441-a1e6-c82f7e85207d · outbound
Self-Generated Critiques Boost Reward Modeling for Language Models The Perfect Blend: Redefining RLHF with Mixture of Judges
Reference 60
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6d441469-538a-448d-97c9-7de403a71d73 · outbound
Self-Generated Critiques Boost Reward Modeling for Language Models Predicting text preference via structured comparative reasoning
Reference 61
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation c085e414-9992-48f3-a9ba-74ff99df1479 · outbound
Self-Generated Critiques Boost Reward Modeling for Language Models Improving Reward Models with Synthetic Critiques
Reference 62
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0bf6c7bb-daa5-409e-8f91-038ad55d8786 · outbound
Self-Generated Critiques Boost Reward Modeling for Language Models Self-rewarding language models
Reference 63
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 37055a91-22f2-4359-b46f-04cdc0d4c6e6 · outbound
Self-Generated Critiques Boost Reward Modeling for Language Models ST ar: Bootstrapping reasoning with reasoning
Reference 64
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation bdbc6414-9384-429c-921d-2bcaafe35f5f · outbound
Self-Generated Critiques Boost Reward Modeling for Language Models Evaluating large language models at evaluating instruction following
Reference 65
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 87930707-cc45-4c6a-b59d-14439ba24ee1 · outbound
Self-Generated Critiques Boost Reward Modeling for Language Models Generative Verifiers: Reward Modeling as Next-Token Prediction
Reference 66
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cd504e4e-fc1e-4e42-864f-f59ca8389635 · outbound
Self-Generated Critiques Boost Reward Modeling for Language Models Gonzalez, and Ion Stoica
Reference 67
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation c5263d52-72f4-4547-965e-7db28e34b226 · outbound
Self-Generated Critiques Boost Reward Modeling for Language Models Law of the Weakest Link: Cross Capabilities of Large Language Models
Reference 68
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9b5d11c8-4154-4cfe-b795-fb9bb11cc8c9 · outbound
Self-Generated Critiques Boost Reward Modeling for Language Models Iterative data smoothing: Mitigating reward overfitting and overoptimization in RLHF
Reference 69
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 4d8777be-64db-4555-bd5a-c50a6fe447bf · inbound
In Context Learning and Reasoning for Symbolic Regression with Large Language Models Self-Generated Critiques Boost Reward Modeling for Language Models
Reference 95
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 11b227f3-7a53-48b0-a279-df007f79eb27 · inbound
AceMath: Advancing Frontier Math Reasoning with Post-Training and Reward Modeling Self-Generated Critiques Boost Reward Modeling for Language Models
Reference 73
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 199020c2-2a2a-4000-8eb8-0509df32500a · inbound
Do NOT Think That Much for 2+3=? On the Overthinking of o1-Like LLMs Self-Generated Critiques Boost Reward Modeling for Language Models
Reference 278
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 860b6c67-1087-476e-b6cd-f5f4d8e47482 · inbound
CoT-based Synthesizer: Enhancing LLM Performance through Answer Synthesis Self-Generated Critiques Boost Reward Modeling for Language Models
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2bf10488-a08d-454a-a18f-e96d366e2c8e · inbound
LongDPO: Unlock Better Long-form Generation Abilities for LLMs via Critique-augmented Stepwise Information Self-Generated Critiques Boost Reward Modeling for Language Models
Reference 52
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d4229abe-225d-4a77-996c-286beb3372a5 · inbound
MM-RLHF: The Next Step Forward in Multimodal LLM Alignment Self-Generated Critiques Boost Reward Modeling for Language Models
Reference 74
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e96f158f-e359-4e4f-bac7-52b63dee4c60 · inbound
SPC: Evolving Self-Play Critic via Adversarial Games for LLM Reasoning Self-Generated Critiques Boost Reward Modeling for Language Models
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 13be2a22-8da4-46b1-9941-d6d1fb329489 · inbound
J1: Exploring Simple Test-Time Scaling for LLM-as-a-Judge Self-Generated Critiques Boost Reward Modeling for Language Models
Reference 62
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation faf6a9ce-0c08-4f47-8de0-a77a1dec0b40 · inbound
Think-RM: Enabling Long-Horizon Reasoning in Generative Reward Models Self-Generated Critiques Boost Reward Modeling for Language Models
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a978274c-3e04-4b41-9561-e696d770c2d0 · inbound
RewardAnything: Generalizable Principle-Following Reward Models Self-Generated Critiques Boost Reward Modeling for Language Models
Reference 96
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 422ef1bf-506e-4adc-bdba-de513b822b03 · inbound
Alignment and Safety in Large Language Models: Safety Mechanisms, Training Paradigms, and Emerging Challenges Self-Generated Critiques Boost Reward Modeling for Language Models
Reference 282
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 58c58590-3bd8-48e6-b25b-a0b1448d2c79 · inbound
Building a Precise Video Language with Human-AI Oversight Self-Generated Critiques Boost Reward Modeling for Language Models
Reference 83
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 0dfaad6f-0025-45e8-99fb-15e8ced842ba · inbound
Test-Time Verification for Text-to-SQL via Outcome Reward Models Self-Generated Critiques Boost Reward Modeling for Language Models
Reference 83
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.