Pith. sign in

Paper Citation Record · LEDGER

Tiny Reward Models

As of 9 August 2026, this Paper Citation Record lists 49 of 49 outbound references and 0 inbound Pith citation observations for arXiv:2507.09973.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.09973 v1

Coverage vector

measured 49 of 49 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T17:48:07.210826Z

measured 49 of 49 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

49 of 49 outbound references displayed

  • verified exact4
  • verified fuzzy2
  • unresolved43
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation bb79d5da-4c7a-40e6-b4d7-0ed61c769ad0 · outbound

This paper cites write newline.

Tiny Reward Models write newline

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-06T17:48:07.089065Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:48:07.089065Z digest=sha256:71c57c526a89726f51a58bce6176bf59b4b42bd4f4aea048e77396f2b7842294

Observation 5a83696c-c906-427c-90cf-c33df3bad179 · outbound

This paper cites Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback.

Tiny Reward Models Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-06T17:48:07.092692Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:48:07.092692Z digest=sha256:7ec6514d9cc52a6af7bbbd3c8128779258d0fdd1277aed0847430b5c85ea4d22

Observation 9fc9e5de-1c04-4e9a-8616-6d0672d9b7d1 · outbound

This paper cites PPL-MCTS: Constrained Textual Generation Through Discriminator-Guided MCTS Decoding.

Tiny Reward Models PPL-MCTS: Constrained Textual Generation Through Discriminator-Guided MCTS Decoding

Reference 3

Resolution
verified exact
local_arxiv, observed 2026-08-06T17:48:07.695107Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-06T17:48:07.095711Z digest=sha256:2b981876889fb475aff11fefb100fb60458e9e1a5339901ffa18e8f829c95b40

Observation 86b77f26-9ca0-4d57-8d65-8bcf86e0ee1e · outbound

This paper cites Rm-r1: Reward modeling as reasoning, 2025.

Tiny Reward Models Rm-r1: Reward modeling as reasoning, 2025

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-06T17:48:07.098573Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:48:07.098573Z digest=sha256:27a99b03599b27fa458160a9ca58a0660372078a192d5c3ff966bf770ff9f298

Observation 8b4b6dfb-eff9-41c9-a1a2-99a9b1cd50b8 · outbound

This paper cites B., Martic, M., Legg, S., and Amodei, D.

Tiny Reward Models B., Martic, M., Legg, S., and Amodei, D

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:48:07.718376Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-06T17:48:07.101195Z digest=sha256:a2a3c4b670bcc5c9fd62fe8a28f0021142ba73964541baedfa6350c1309d64ad

Observation 5027851b-45cc-41f4-9282-970a71d218c4 · outbound

This paper cites Scaling Instruction-Finetuned Language Models.

Tiny Reward Models Scaling Instruction-Finetuned Language Models

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-06T17:48:07.103616Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:48:07.103616Z digest=sha256:f998cd26caf14c7cbb8d66213471541b01de29518989e4318e1ecc53a7141a86

Observation f86633c8-6b4d-4265-b4bc-f7dbf11bcbe6 · outbound

This paper cites It's All in The [MASK]: Simple Instruction-Tuning Enables BERT-like Masked Language Models As Generative Classifiers.

Tiny Reward Models It's All in The [MASK]: Simple Instruction-Tuning Enables BERT-like Masked Language Models As Generative Classifiers

Reference 7

Resolution
verified exact
local_arxiv, observed 2026-08-06T17:48:07.607922Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-06T17:48:07.106359Z digest=sha256:a12403a0005865938fcdc25840a68eef2995845b443831d568921cfdf3611da9

Observation c0138e18-c753-41df-a667-1b813b9a507e · outbound

This paper cites BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding.

Tiny Reward Models BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-06T17:48:07.109354Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:48:07.109354Z digest=sha256:c42c1bbbf13c41d4f5fedc4b9db54e3449398d9591e09337c20307737e7d68d4

Observation 7ffd1b75-1817-4897-87f9-9b0a4fe3efaa · outbound

This paper cites TinyStories: How Small Can Language Models Be and Still Speak Coherent English?.

Tiny Reward Models TinyStories: How Small Can Language Models Be and Still Speak Coherent English?

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-06T17:48:07.111658Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:48:07.111658Z digest=sha256:76784e09730102da7d8960af19fecdf5743f00271b34b815de240ab57f4b2475

Observation f98dd840-a992-4e95-b9f3-932fd9983cac · outbound

This paper cites L., and Vigliocco, G.

Tiny Reward Models L., and Vigliocco, G

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:48:07.710863Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-06T17:48:07.114627Z digest=sha256:6c54b99d468f1f09ac267fd82618d9d3c2c95c65cd8b13a6f611982987444056

Observation 17de84cb-7b24-49a6-ae43-762183081e98 · outbound

This paper cites Scaling Laws for Reward Model Overoptimization.

Tiny Reward Models Scaling Laws for Reward Model Overoptimization

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-06T17:48:07.117097Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:48:07.117097Z digest=sha256:a5b93e6d387b7aca55b961bbd440345a8d091b76929116a49c7c3effb2cdbf50

Observation 07f91a81-d6f9-4dca-9acd-b5e61a593cd5 · outbound

This paper cites Making pre-trained language models better few-shot learners.

Tiny Reward Models Making pre-trained language models better few-shot learners

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-06T17:48:07.119680Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:48:07.119680Z digest=sha256:8ed48b5beece2721d404b1665320fc223f41740888c2d155f8478461efbd580c

Observation 02458464-c5a9-4475-8514-9a64609c1b63 · outbound

This paper cites Towards Data-Efficient Language Models: A Child-Inspired Approach to Language Learning.

Tiny Reward Models Towards Data-Efficient Language Models: A Child-Inspired Approach to Language Learning

Reference 13

Resolution
verified exact
local_arxiv, observed 2026-08-06T17:48:07.575213Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-06T17:48:07.122038Z digest=sha256:894124f3cb403a810827844dface30d771f0ffc7fd6f5b21cc30135976380381

Observation 80c8c4af-44f5-4f4f-8a0c-f6eb2c026317 · outbound

This paper cites The Llama 3 Herd of Models.

Tiny Reward Models The Llama 3 Herd of Models

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-06T17:48:07.124633Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:48:07.124633Z digest=sha256:c2867823cec6456851030c9d4ac7b70f326864b39bf0b4d7b7a19bd28af1fb14

Observation c11d3822-6db0-494b-a4f7-02e261dfed68 · outbound

This paper cites Textbooks Are All You Need.

Tiny Reward Models Textbooks Are All You Need

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-06T17:48:07.127115Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:48:07.127115Z digest=sha256:ac6bad8811f8851e12d0df9704f4a26020f04c8b4342a004631b716e1f15a1e2

Observation a15d9dde-7f19-4637-a4f5-ecb0c0c58b8e · outbound

This paper cites Does RLHF Scale? Exploring the Impacts From Data, Model, and Method.

Tiny Reward Models Does RLHF Scale? Exploring the Impacts From Data, Model, and Method

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-06T17:48:07.129543Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:48:07.129543Z digest=sha256:3479436e8493e405d169b5401bf4418d5693c4cf95024e65f8bfbd74675b7a25

Observation 841364ec-0573-4c03-a26a-f6ec90b140b8 · outbound

This paper cites Universal Language Model Fine-tuning for Text Classification.

Tiny Reward Models Universal Language Model Fine-tuning for Text Classification

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-06T17:48:07.131920Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:48:07.131920Z digest=sha256:ca1cec9cc454fba828447daa79242111d7033b6b991158e9ad5ed1c55ce25a8a

Observation db601eee-593a-4c61-a1d8-3d38dfbc423c · outbound

This paper cites LoRA: Low-Rank Adaptation of Large Language Models.

Tiny Reward Models LoRA: Low-Rank Adaptation of Large Language Models

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-06T17:48:07.134619Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:48:07.134619Z digest=sha256:7d422e4457f4c6c836d3c11eb9bb265869d853ec8ed223822b696964076f1018

Observation 42fe3188-5383-423d-b63f-299496c0b3a7 · outbound

This paper cites PKU-SafeRLHF: Towards Multi-Level Safety Alignment for LLMs with Human Preference.

Tiny Reward Models PKU-SafeRLHF: Towards Multi-Level Safety Alignment for LLMs with Human Preference

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-06T17:48:07.137146Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:48:07.137146Z digest=sha256:357a4c2c9d1ee27a0757790e1739a3f4e0582f9dc56bc2e41213eab52d946147

Observation 53dde41c-fc6c-4f35-97a7-5e515f38a3fa · outbound

This paper cites OpenAssistant Conversations -- Democratizing Large Language Model Alignment.

Tiny Reward Models OpenAssistant Conversations -- Democratizing Large Language Model Alignment

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-06T17:48:07.139475Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:48:07.139475Z digest=sha256:5f8b8c2a52fa3daea23db2b01a7d07047a9c38746ac94238d4e13cc77ef9c35b

Observation 459455e7-a116-4b41-a7fc-2cb1cfac5cc5 · outbound

This paper cites RewardBench: Evaluating Reward Models for Language Modeling.

Tiny Reward Models RewardBench: Evaluating Reward Models for Language Modeling

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-06T17:48:07.141889Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:48:07.141889Z digest=sha256:862e37dd54a2116f90589e17bd8662c7744db37b7549a4e6880e097214ab5d2a

Observation 8156c2eb-62bd-4796-9ebe-ca1460e7ab41 · outbound

This paper cites What Would Elsa Do? Freezing Layers During Transformer Fine-Tuning.

Tiny Reward Models What Would Elsa Do? Freezing Layers During Transformer Fine-Tuning

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-06T17:48:07.144511Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:48:07.144511Z digest=sha256:cc823ebcb3138ce88b5594db4d41ed3d72e4ba8e4b510a5df1ee83c36376871d

Observation 8e9201f9-ecad-4092-8e74-1c7c14d616c4 · outbound

This paper cites BART: Denoising Sequence-to-Sequence Pre-training for Natural Language Generation, Translation, and Comprehension.

Tiny Reward Models BART: Denoising Sequence-to-Sequence Pre-training for Natural Language Generation, Translation, and Comprehension

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-06T17:48:07.146902Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:48:07.146902Z digest=sha256:ea09eeee654fd91a54baa0f557da395de39cf420d6ced6a5c003c8688e4b8174

Observation 5a54112a-3525-44c0-9a96-08ef27a9f9dc · outbound

This paper cites Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks.

Tiny Reward Models Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-06T17:48:07.149469Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:48:07.149469Z digest=sha256:7c9dc20e832862cc5910d063a0a57fb64e5411f3fb48a39b5573de7d927236c3

Observation 34cc45f9-ca75-409f-9df3-578986314c80 · outbound

This paper cites Let's Verify Step by Step.

Tiny Reward Models Let's Verify Step by Step

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-06T17:48:07.152003Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:48:07.152003Z digest=sha256:e9ef7e00ee551d6be67b71994792e5b15599aa9a26bcc84baff586c3ad8f43b5

Observation e99e856f-6e33-4b1c-817e-67625d4bcaca · outbound

This paper cites Skywork-Reward: Bag of Tricks for Reward Modeling in LLMs.

Tiny Reward Models Skywork-Reward: Bag of Tricks for Reward Modeling in LLMs

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-06T17:48:07.154579Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:48:07.154579Z digest=sha256:c36d6ac84f0b4f266161e9c057b936e4c6b1c30e488b00e1159e6d1f94ee47a2

Observation 8c8ae2a7-6cf2-4a81-9f5d-2416827e6c61 · outbound

This paper cites DoRA: Weight-Decomposed Low-Rank Adaptation.

Tiny Reward Models DoRA: Weight-Decomposed Low-Rank Adaptation

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-06T17:48:07.157051Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:48:07.157051Z digest=sha256:3b849f0fda033c5a5eb5c06306e1413d7a3959119077a232034b77cf1cd65150

Observation e8578aff-60bd-4b87-a2f6-fbdf9c4f6b1c · outbound

This paper cites Routing to the Expert: Efficient Reward-guided Ensemble of Large Language Models.

Tiny Reward Models Routing to the Expert: Efficient Reward-guided Ensemble of Large Language Models

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-06T17:48:07.159727Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:48:07.159727Z digest=sha256:f20562fe40420244b2147913ad3ce83606c4138f1009c095fc574aa5e9bb6987

Observation 7cb9e172-efc3-4fa9-932a-ea5822a66474 · outbound

This paper cites Improve Mathematical Reasoning in Language Models by Automated Process Supervision.

Tiny Reward Models Improve Mathematical Reasoning in Language Models by Automated Process Supervision

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-06T17:48:07.162023Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:48:07.162023Z digest=sha256:de2f079e4cdca57b81e1cdae0ab683d0750522b12c7d1a8a9c99ad4ccacf69ad

Observation f02cbb16-bc9d-4043-8623-29554abc7501 · outbound

This paper cites WebGPT: Browser-assisted question-answering with human feedback.

Tiny Reward Models WebGPT: Browser-assisted question-answering with human feedback

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-06T17:48:07.164185Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:48:07.164185Z digest=sha256:a1e11f8dd6368c9b87803384a9a43daa1c5be16748695e845fb7cad86e73cde3

Observation 3279575f-3339-45b6-9782-4f22b9fa39ce · outbound

This paper cites Passage Re-ranking with BERT.

Tiny Reward Models Passage Re-ranking with BERT

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-06T17:48:07.166590Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:48:07.166590Z digest=sha256:a5d715d431ab328908e69d54f73bd1321480b6be6f355cce2e8653692ed5ba5f

Observation 57c0428d-05be-40a2-bc31-aae00aa506bc · outbound

This paper cites Nemotron-4 340B Technical Report.

Tiny Reward Models Nemotron-4 340B Technical Report

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-06T17:48:07.168990Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:48:07.168990Z digest=sha256:bb1085eff79cf27a6f163bd98b1efc1a1cc21b7ae36e5a3fa1fed7258c467d9b

Observation 6a490881-e7e0-4992-9191-c5b2b4fd2089 · outbound

This paper cites Training language models to follow instructions with human feedback.

Tiny Reward Models Training language models to follow instructions with human feedback

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-06T17:48:07.171387Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:48:07.171387Z digest=sha256:95b9864937ff43ddd22f67c3d2ce78a5741e12e8db5cac21410e8cbe81ab6ce7

Observation 7a8c7288-d584-4bb4-95e9-f3b0ba4fb327 · outbound

This paper cites Let's Reinforce Step by Step.

Tiny Reward Models Let's Reinforce Step by Step

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-06T17:48:07.173925Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:48:07.173925Z digest=sha256:a0ac967dad126dabc84c0a96463d884512f55596040c46f9862415f8a853e5ec

Observation af90b18e-5bee-4183-85a5-62a94e9b2624 · outbound

This paper cites Exploring the Limits of Transfer Learning with a Unified Text-to-Text Transformer.

Tiny Reward Models Exploring the Limits of Transfer Learning with a Unified Text-to-Text Transformer

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-06T17:48:07.176314Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:48:07.176314Z digest=sha256:a7d9fd9e18ea7cd16ff3fba912aec2d871b65a42bda7bbd4b68ccd9362ddcf83

Observation c96758dd-57fc-40a8-9c3f-1eaab651fd41 · outbound

This paper cites WARM: On the Benefits of Weight Averaged Reward Models.

Tiny Reward Models WARM: On the Benefits of Weight Averaged Reward Models

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-06T17:48:07.178654Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:48:07.178654Z digest=sha256:20853db715eb10d943d12e814559fcd3568405c818753bbd42b9547c0841cec3

Observation e265842c-7e61-40cc-afc5-6409dd85f99f · outbound

This paper cites BERTs are Generative In-Context Learners.

Tiny Reward Models BERTs are Generative In-Context Learners

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-06T17:48:07.181168Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:48:07.181168Z digest=sha256:d66865d78a06f721df1ab084b8f8f41a2291ed12cf1472c12d8015ca148b29a1

Observation 55c7ffa2-a916-4d5a-9f42-37c62c029a48 · outbound

This paper cites Exploiting Cloze Questions for Few Shot Text Classification and Natural Language Inference.

Tiny Reward Models Exploiting Cloze Questions for Few Shot Text Classification and Natural Language Inference

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-06T17:48:07.183516Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:48:07.183516Z digest=sha256:c57277604d3986a493086a68a6becec6ea6d079e72b8fce17a119d680967fc45

Observation abb3c454-889b-487e-9e31-95e36072bc82 · outbound

This paper cites It's Not Just Size That Matters: Small Language Models Are Also Few-Shot Learners.

Tiny Reward Models It's Not Just Size That Matters: Small Language Models Are Also Few-Shot Learners

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-06T17:48:07.185809Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:48:07.185809Z digest=sha256:84ac45fbf14ecaa9c74f01f3e71d5f6abadb04935d68c01c10a1afeaf43e3089

Observation 2237a1c5-ba21-4367-b8f6-dc294cd7fbbc · outbound

This paper cites The Trickle-down Impact of Reward (In-)consistency on RLHF.

Tiny Reward Models The Trickle-down Impact of Reward (In-)consistency on RLHF

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-06T17:48:07.188338Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:48:07.188338Z digest=sha256:6043caa7a5e6686246c5f8844de825ca5dab682ae84754a84a25fd3c8a9c3dfc

Observation 986bf3e7-c475-4689-8ffb-a0a7508a53d1 · outbound

This paper cites Layer by Layer: Uncovering Hidden Representations in Language Models.

Tiny Reward Models Layer by Layer: Uncovering Hidden Representations in Language Models

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-06T17:48:07.190751Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:48:07.190751Z digest=sha256:e0d5f0a6c62616ee8058414a1289c2d7803037d85ed2e7a182822d72ad428ff0

Observation 60319909-5b82-4069-bfa0-023280b205c8 · outbound

This paper cites D., and Su, W.

Tiny Reward Models D., and Su, W

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-06T17:48:07.193757Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:48:07.193757Z digest=sha256:277d96b5f98aedeb79dafe2f1a1fb9688a2bb49f976bf172d58ee933ef5a4a7a

Observation faba13a6-cc57-4e9c-9c9b-8f8b81993a6d · outbound

This paper cites Large GPT-like Models are Bad Babies: A Closer Look at the Relationship between Linguistic Competence and Psycholinguistic Measures.

Tiny Reward Models Large GPT-like Models are Bad Babies: A Closer Look at the Relationship between Linguistic Competence and Psycholinguistic Measures

Reference 43

Resolution
verified exact
local_arxiv, observed 2026-08-06T17:48:07.285710Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-06T17:48:07.196004Z digest=sha256:619802926311d1d67efea4e4522df961ce0a2554aacb9ab8cc021c3ed9c870aa

Observation d63fe4c4-efae-4003-a284-2e58a6e4d6a0 · outbound

This paper cites How to Fine-Tune BERT for Text Classification?.

Tiny Reward Models How to Fine-Tune BERT for Text Classification?

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-06T17:48:07.198467Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:48:07.198467Z digest=sha256:541e04572faeb708f67a5d569507cb163cd1995232abe9b11512495e1747a610

Observation 5be4873b-40e3-45a8-896b-91a648f1633d · outbound

This paper cites HelpSteer2: Open-source dataset for training top-performing reward models.

Tiny Reward Models HelpSteer2: Open-source dataset for training top-performing reward models

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-06T17:48:07.201291Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:48:07.201291Z digest=sha256:32fce970d649e7e8c173c3d8a72c2816c4e312e48fd7c2cddfc33e9dffa11a8b

Observation 79311ea0-1930-4be7-87c8-9a9fc00b6d25 · outbound

This paper cites Smarter, Better, Faster, Longer: A Modern Bidirectional Encoder for Fast, Memory Efficient, and Long Context Finetuning and Inference.

Tiny Reward Models Smarter, Better, Faster, Longer: A Modern Bidirectional Encoder for Fast, Memory Efficient, and Long Context Finetuning and Inference

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-06T17:48:07.203840Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:48:07.203840Z digest=sha256:ae89f781f6ed2bfb6fcb3c7924c3692561f0719d018b53ce129d31c53016ec3e

Observation de32fa4b-0030-450f-a9c6-4897c78b760c · outbound

This paper cites Finetuned Language Models Are Zero-Shot Learners.

Tiny Reward Models Finetuned Language Models Are Zero-Shot Learners

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-06T17:48:07.206091Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:48:07.206091Z digest=sha256:3cfc6886fd2a453d5c1db0168d38f75c7c6d51152b4fe10a0b3890added3a80c

Observation a71856fb-f940-4e8f-8098-d0b4794b6b94 · outbound

This paper cites MetaMetrics: Calibrating Metrics For Generation Tasks Using Human Preferences.

Tiny Reward Models MetaMetrics: Calibrating Metrics For Generation Tasks Using Human Preferences

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-06T17:48:07.208459Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:48:07.208459Z digest=sha256:6f6de9b2c8fd82c0eeda8d948ad4c7f84b212da7d9af57d1ef9ab519597f602b

Observation b5ece352-57bb-4918-bb9e-91cd6961c748 · outbound

This paper cites How transferable are features in deep neural networks?.

Tiny Reward Models How transferable are features in deep neural networks?

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-06T17:48:07.210826Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:48:07.210826Z digest=sha256:9d0694ccfa054372c3fbfabfd9c855b6c875205e38cded054d410995eaa52805

Pith citing papers

No inbound Pith citation observations are available.