Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-06T17:48:07.210826Z
Paper Citation Record · LEDGER
As of 7 August 2026, this Paper Citation Record lists 49 of 49 outbound references and 0 inbound Pith citation observations for arXiv:2507.09973.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-06T17:48:07.210826Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
49 of 49 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation bb79d5da-4c7a-40e6-b4d7-0ed61c769ad0 · outbound
Tiny Reward Models write newline
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5a83696c-c906-427c-90cf-c33df3bad179 · outbound
Tiny Reward Models Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9fc9e5de-1c04-4e9a-8616-6d0672d9b7d1 · outbound
Tiny Reward Models PPL-MCTS: Constrained Textual Generation Through Discriminator-Guided MCTS Decoding
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 86b77f26-9ca0-4d57-8d65-8bcf86e0ee1e · outbound
Tiny Reward Models Rm-r1: Reward modeling as reasoning, 2025
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8b4b6dfb-eff9-41c9-a1a2-99a9b1cd50b8 · outbound
Tiny Reward Models B., Martic, M., Legg, S., and Amodei, D
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 5027851b-45cc-41f4-9282-970a71d218c4 · outbound
Tiny Reward Models Scaling Instruction-Finetuned Language Models
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f86633c8-6b4d-4265-b4bc-f7dbf11bcbe6 · outbound
Tiny Reward Models It's All in The [MASK]: Simple Instruction-Tuning Enables BERT-like Masked Language Models As Generative Classifiers
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation c0138e18-c753-41df-a667-1b813b9a507e · outbound
Tiny Reward Models BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7ffd1b75-1817-4897-87f9-9b0a4fe3efaa · outbound
Tiny Reward Models TinyStories: How Small Can Language Models Be and Still Speak Coherent English?
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f98dd840-a992-4e95-b9f3-932fd9983cac · outbound
Tiny Reward Models L., and Vigliocco, G
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 17de84cb-7b24-49a6-ae43-762183081e98 · outbound
Tiny Reward Models Scaling Laws for Reward Model Overoptimization
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 07f91a81-d6f9-4dca-9acd-b5e61a593cd5 · outbound
Tiny Reward Models Making pre-trained language models better few-shot learners
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 02458464-c5a9-4475-8514-9a64609c1b63 · outbound
Tiny Reward Models Towards Data-Efficient Language Models: A Child-Inspired Approach to Language Learning
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 80c8c4af-44f5-4f4f-8a0c-f6eb2c026317 · outbound
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c11d3822-6db0-494b-a4f7-02e261dfed68 · outbound
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a15d9dde-7f19-4637-a4f5-ecb0c0c58b8e · outbound
Tiny Reward Models Does RLHF Scale? Exploring the Impacts From Data, Model, and Method
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 841364ec-0573-4c03-a26a-f6ec90b140b8 · outbound
Tiny Reward Models Universal Language Model Fine-tuning for Text Classification
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation db601eee-593a-4c61-a1d8-3d38dfbc423c · outbound
Tiny Reward Models LoRA: Low-Rank Adaptation of Large Language Models
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 42fe3188-5383-423d-b63f-299496c0b3a7 · outbound
Tiny Reward Models PKU-SafeRLHF: Towards Multi-Level Safety Alignment for LLMs with Human Preference
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 53dde41c-fc6c-4f35-97a7-5e515f38a3fa · outbound
Tiny Reward Models OpenAssistant Conversations -- Democratizing Large Language Model Alignment
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 459455e7-a116-4b41-a7fc-2cb1cfac5cc5 · outbound
Tiny Reward Models RewardBench: Evaluating Reward Models for Language Modeling
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8156c2eb-62bd-4796-9ebe-ca1460e7ab41 · outbound
Tiny Reward Models What Would Elsa Do? Freezing Layers During Transformer Fine-Tuning
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8e9201f9-ecad-4092-8e74-1c7c14d616c4 · outbound
Tiny Reward Models BART: Denoising Sequence-to-Sequence Pre-training for Natural Language Generation, Translation, and Comprehension
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5a54112a-3525-44c0-9a96-08ef27a9f9dc · outbound
Tiny Reward Models Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 34cc45f9-ca75-409f-9df3-578986314c80 · outbound
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e99e856f-6e33-4b1c-817e-67625d4bcaca · outbound
Tiny Reward Models Skywork-Reward: Bag of Tricks for Reward Modeling in LLMs
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8c8ae2a7-6cf2-4a81-9f5d-2416827e6c61 · outbound
Tiny Reward Models DoRA: Weight-Decomposed Low-Rank Adaptation
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e8578aff-60bd-4b87-a2f6-fbdf9c4f6b1c · outbound
Tiny Reward Models Routing to the Expert: Efficient Reward-guided Ensemble of Large Language Models
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7cb9e172-efc3-4fa9-932a-ea5822a66474 · outbound
Tiny Reward Models Improve Mathematical Reasoning in Language Models by Automated Process Supervision
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f02cbb16-bc9d-4043-8623-29554abc7501 · outbound
Tiny Reward Models WebGPT: Browser-assisted question-answering with human feedback
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3279575f-3339-45b6-9782-4f22b9fa39ce · outbound
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 57c0428d-05be-40a2-bc31-aae00aa506bc · outbound
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6a490881-e7e0-4992-9191-c5b2b4fd2089 · outbound
Tiny Reward Models Training language models to follow instructions with human feedback
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7a8c7288-d584-4bb4-95e9-f3b0ba4fb327 · outbound
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation af90b18e-5bee-4183-85a5-62a94e9b2624 · outbound
Tiny Reward Models Exploring the Limits of Transfer Learning with a Unified Text-to-Text Transformer
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c96758dd-57fc-40a8-9c3f-1eaab651fd41 · outbound
Tiny Reward Models WARM: On the Benefits of Weight Averaged Reward Models
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e265842c-7e61-40cc-afc5-6409dd85f99f · outbound
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 55c7ffa2-a916-4d5a-9f42-37c62c029a48 · outbound
Tiny Reward Models Exploiting Cloze Questions for Few Shot Text Classification and Natural Language Inference
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation abb3c454-889b-487e-9e31-95e36072bc82 · outbound
Tiny Reward Models It's Not Just Size That Matters: Small Language Models Are Also Few-Shot Learners
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2237a1c5-ba21-4367-b8f6-dc294cd7fbbc · outbound
Tiny Reward Models The Trickle-down Impact of Reward (In-)consistency on RLHF
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 986bf3e7-c475-4689-8ffb-a0a7508a53d1 · outbound
Tiny Reward Models Layer by Layer: Uncovering Hidden Representations in Language Models
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 60319909-5b82-4069-bfa0-023280b205c8 · outbound
Tiny Reward Models D., and Su, W
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation faba13a6-cc57-4e9c-9c9b-8f8b81993a6d · outbound
Tiny Reward Models Large GPT-like Models are Bad Babies: A Closer Look at the Relationship between Linguistic Competence and Psycholinguistic Measures
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation d63fe4c4-efae-4003-a284-2e58a6e4d6a0 · outbound
Tiny Reward Models How to Fine-Tune BERT for Text Classification?
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5be4873b-40e3-45a8-896b-91a648f1633d · outbound
Tiny Reward Models HelpSteer2: Open-source dataset for training top-performing reward models
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 79311ea0-1930-4be7-87c8-9a9fc00b6d25 · outbound
Tiny Reward Models Smarter, Better, Faster, Longer: A Modern Bidirectional Encoder for Fast, Memory Efficient, and Long Context Finetuning and Inference
Reference 46
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation de32fa4b-0030-450f-a9c6-4897c78b760c · outbound
Tiny Reward Models Finetuned Language Models Are Zero-Shot Learners
Reference 47
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a71856fb-f940-4e8f-8098-d0b4794b6b94 · outbound
Tiny Reward Models MetaMetrics: Calibrating Metrics For Generation Tasks Using Human Preferences
Reference 48
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b5ece352-57bb-4918-bb9e-91cd6961c748 · outbound
Tiny Reward Models How transferable are features in deep neural networks?
Reference 49
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
No inbound Pith citation observations are available.