Pith. sign in

Paper Citation Record · LEDGER

Tiny Reward Models

As of 7 August 2026, this Paper Citation Record lists 49 of 49 outbound references and 0 inbound Pith citation observations for arXiv:2507.09973.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.09973 v1

Coverage vector

measured 49 of 49 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T17:48:07.210826Z

measured 49 of 49 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

49 of 49 outbound references displayed

  • verified exact4
  • verified fuzzy2
  • unresolved43
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation bb79d5da-4c7a-40e6-b4d7-0ed61c769ad0 · outbound

This paper cites write newline.

Tiny Reward Models write newline

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-06T17:48:07.089065Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:48:07.089065Z digest=sha256:3dbd3d98d1928b272fcca9f9fcaf558af5cfb7d22c225098cf00bf525c31e922

Observation 5a83696c-c906-427c-90cf-c33df3bad179 · outbound

This paper cites Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback.

Tiny Reward Models Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-06T17:48:07.092692Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:48:07.092692Z digest=sha256:679568f0cd25d85ad9b639578274598fc0b882f01000b9e5fa0eccac815a2cea

Observation 9fc9e5de-1c04-4e9a-8616-6d0672d9b7d1 · outbound

This paper cites PPL-MCTS: Constrained Textual Generation Through Discriminator-Guided MCTS Decoding.

Tiny Reward Models PPL-MCTS: Constrained Textual Generation Through Discriminator-Guided MCTS Decoding

Reference 3

Resolution
verified exact
local_arxiv, observed 2026-08-06T17:48:07.695107Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T17:48:07.095711Z digest=sha256:be3025bd268ae2404e7e018f21d91a1e6875a479bcc23ecd80978ed0b774c84a

Observation 86b77f26-9ca0-4d57-8d65-8bcf86e0ee1e · outbound

This paper cites Rm-r1: Reward modeling as reasoning, 2025.

Tiny Reward Models Rm-r1: Reward modeling as reasoning, 2025

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-06T17:48:07.098573Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:48:07.098573Z digest=sha256:82a8a4cc837bc5de528b8d8e92b954c312355872a170f597a2a6c22cc6454046

Observation 8b4b6dfb-eff9-41c9-a1a2-99a9b1cd50b8 · outbound

This paper cites B., Martic, M., Legg, S., and Amodei, D.

Tiny Reward Models B., Martic, M., Legg, S., and Amodei, D

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:48:07.718376Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T17:48:07.101195Z digest=sha256:b3271dd28deae662b592f85b770fef5cdf41a39ebde5bfe934d3caef368248dd

Observation 5027851b-45cc-41f4-9282-970a71d218c4 · outbound

This paper cites Scaling Instruction-Finetuned Language Models.

Tiny Reward Models Scaling Instruction-Finetuned Language Models

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-06T17:48:07.103616Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:48:07.103616Z digest=sha256:aca159b0f00471a4916de6fb83dfe8f3e760544ed2a3c546bdcc014063d66d73

Observation f86633c8-6b4d-4265-b4bc-f7dbf11bcbe6 · outbound

This paper cites It's All in The [MASK]: Simple Instruction-Tuning Enables BERT-like Masked Language Models As Generative Classifiers.

Tiny Reward Models It's All in The [MASK]: Simple Instruction-Tuning Enables BERT-like Masked Language Models As Generative Classifiers

Reference 7

Resolution
verified exact
local_arxiv, observed 2026-08-06T17:48:07.607922Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T17:48:07.106359Z digest=sha256:b5aeff67ab481e9a060649f3598189b85a5e4be6505dca9722b7b1f5650c3315

Observation c0138e18-c753-41df-a667-1b813b9a507e · outbound

This paper cites BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding.

Tiny Reward Models BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-06T17:48:07.109354Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:48:07.109354Z digest=sha256:5095ae392e238909bcf3d7e118417af499f1816494724e03aab0d13b687a315e

Observation 7ffd1b75-1817-4897-87f9-9b0a4fe3efaa · outbound

This paper cites TinyStories: How Small Can Language Models Be and Still Speak Coherent English?.

Tiny Reward Models TinyStories: How Small Can Language Models Be and Still Speak Coherent English?

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-06T17:48:07.111658Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:48:07.111658Z digest=sha256:3d6b35ee413f435ec0306318b51983bbdd327c51b73be52c094839ceb3b5b856

Observation f98dd840-a992-4e95-b9f3-932fd9983cac · outbound

This paper cites L., and Vigliocco, G.

Tiny Reward Models L., and Vigliocco, G

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:48:07.710863Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T17:48:07.114627Z digest=sha256:f8ebe9152c43f6f9a8da11db84692edf162c304bbd78c647e9810088c1f0506b

Observation 17de84cb-7b24-49a6-ae43-762183081e98 · outbound

This paper cites Scaling Laws for Reward Model Overoptimization.

Tiny Reward Models Scaling Laws for Reward Model Overoptimization

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-06T17:48:07.117097Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:48:07.117097Z digest=sha256:9ec76d70d36665db7b34772ffd9daf5b1a64d8c847d808951646594ade623624

Observation 07f91a81-d6f9-4dca-9acd-b5e61a593cd5 · outbound

This paper cites Making pre-trained language models better few-shot learners.

Tiny Reward Models Making pre-trained language models better few-shot learners

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-06T17:48:07.119680Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:48:07.119680Z digest=sha256:35253b0b0a5f21d153176b8f72cc14fd38c6a6197e3ccad4564bf1253a94afa1

Observation 02458464-c5a9-4475-8514-9a64609c1b63 · outbound

This paper cites Towards Data-Efficient Language Models: A Child-Inspired Approach to Language Learning.

Tiny Reward Models Towards Data-Efficient Language Models: A Child-Inspired Approach to Language Learning

Reference 13

Resolution
verified exact
local_arxiv, observed 2026-08-06T17:48:07.575213Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T17:48:07.122038Z digest=sha256:effdc00bd1efa03f7d58f57a9055e963dadf6d95fb52c3e7374ab1abb92c6141

Observation 80c8c4af-44f5-4f4f-8a0c-f6eb2c026317 · outbound

This paper cites The Llama 3 Herd of Models.

Tiny Reward Models The Llama 3 Herd of Models

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-06T17:48:07.124633Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:48:07.124633Z digest=sha256:7806cf637672fdc61e05755b5b634b4260c313649ca4311b77fb00b4b03909fb

Observation c11d3822-6db0-494b-a4f7-02e261dfed68 · outbound

This paper cites Textbooks Are All You Need.

Tiny Reward Models Textbooks Are All You Need

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-06T17:48:07.127115Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:48:07.127115Z digest=sha256:7810c18c99400386010bb0f10f2ac824ff85a91c8734e1a34cfd794b4eac6803

Observation a15d9dde-7f19-4637-a4f5-ecb0c0c58b8e · outbound

This paper cites Does RLHF Scale? Exploring the Impacts From Data, Model, and Method.

Tiny Reward Models Does RLHF Scale? Exploring the Impacts From Data, Model, and Method

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-06T17:48:07.129543Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:48:07.129543Z digest=sha256:7ab7c4e657801bfeb1352f7adae6e9a99c7e564c95d8b79f893d84ed50c9b8c2

Observation 841364ec-0573-4c03-a26a-f6ec90b140b8 · outbound

This paper cites Universal Language Model Fine-tuning for Text Classification.

Tiny Reward Models Universal Language Model Fine-tuning for Text Classification

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-06T17:48:07.131920Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:48:07.131920Z digest=sha256:4e3973b871315ef5fe3f21d1f4159dfb610f61009bafdb506327bddcc8ff2c22

Observation db601eee-593a-4c61-a1d8-3d38dfbc423c · outbound

This paper cites LoRA: Low-Rank Adaptation of Large Language Models.

Tiny Reward Models LoRA: Low-Rank Adaptation of Large Language Models

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-06T17:48:07.134619Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:48:07.134619Z digest=sha256:f35473c685bf9bfb31e1e0842e703cd0c097b4d1d118d4985c9dc3dbb0f9c2ea

Observation 42fe3188-5383-423d-b63f-299496c0b3a7 · outbound

This paper cites PKU-SafeRLHF: Towards Multi-Level Safety Alignment for LLMs with Human Preference.

Tiny Reward Models PKU-SafeRLHF: Towards Multi-Level Safety Alignment for LLMs with Human Preference

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-06T17:48:07.137146Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:48:07.137146Z digest=sha256:c45f728d693dd0787942ec417f731e03e1715fde6e021051d8a4dece112d42fa

Observation 53dde41c-fc6c-4f35-97a7-5e515f38a3fa · outbound

This paper cites OpenAssistant Conversations -- Democratizing Large Language Model Alignment.

Tiny Reward Models OpenAssistant Conversations -- Democratizing Large Language Model Alignment

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-06T17:48:07.139475Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:48:07.139475Z digest=sha256:c195f0095613339c80576bd95018ca3b63b8ed28f041e5634ad43778d80ef3d2

Observation 459455e7-a116-4b41-a7fc-2cb1cfac5cc5 · outbound

This paper cites RewardBench: Evaluating Reward Models for Language Modeling.

Tiny Reward Models RewardBench: Evaluating Reward Models for Language Modeling

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-06T17:48:07.141889Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:48:07.141889Z digest=sha256:b14ea57f9752167eb64549f48a51121593f16d0ef36eee28d94cad4c13beb75e

Observation 8156c2eb-62bd-4796-9ebe-ca1460e7ab41 · outbound

This paper cites What Would Elsa Do? Freezing Layers During Transformer Fine-Tuning.

Tiny Reward Models What Would Elsa Do? Freezing Layers During Transformer Fine-Tuning

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-06T17:48:07.144511Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:48:07.144511Z digest=sha256:7a8f70912f536992bf205d0efb2d92d7899b9986a052a0cfe4461e09b78075ad

Observation 8e9201f9-ecad-4092-8e74-1c7c14d616c4 · outbound

This paper cites BART: Denoising Sequence-to-Sequence Pre-training for Natural Language Generation, Translation, and Comprehension.

Tiny Reward Models BART: Denoising Sequence-to-Sequence Pre-training for Natural Language Generation, Translation, and Comprehension

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-06T17:48:07.146902Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:48:07.146902Z digest=sha256:0302daf5c90177b17c6e7820d897b364d8de4b49d2ba43ae79d92a22f923ae82

Observation 5a54112a-3525-44c0-9a96-08ef27a9f9dc · outbound

This paper cites Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks.

Tiny Reward Models Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-06T17:48:07.149469Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:48:07.149469Z digest=sha256:05d2587a612a5db777ca66168c5128dbb3613ee1541d616eea5c800e2814766b

Observation 34cc45f9-ca75-409f-9df3-578986314c80 · outbound

This paper cites Let's Verify Step by Step.

Tiny Reward Models Let's Verify Step by Step

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-06T17:48:07.152003Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:48:07.152003Z digest=sha256:4644883bfbf2931f81189ef130d8e653ecef580df163cdec6f87f177777c0ff7

Observation e99e856f-6e33-4b1c-817e-67625d4bcaca · outbound

This paper cites Skywork-Reward: Bag of Tricks for Reward Modeling in LLMs.

Tiny Reward Models Skywork-Reward: Bag of Tricks for Reward Modeling in LLMs

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-06T17:48:07.154579Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:48:07.154579Z digest=sha256:b8edcfe31bd8eaa426908e45fef3c890a7b108fe2236be622b28388d39a612ac

Observation 8c8ae2a7-6cf2-4a81-9f5d-2416827e6c61 · outbound

This paper cites DoRA: Weight-Decomposed Low-Rank Adaptation.

Tiny Reward Models DoRA: Weight-Decomposed Low-Rank Adaptation

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-06T17:48:07.157051Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:48:07.157051Z digest=sha256:95ac0fdba5a53a076b1be47aa63e6ef1bf16b1488b13e917904d682608d87234

Observation e8578aff-60bd-4b87-a2f6-fbdf9c4f6b1c · outbound

This paper cites Routing to the Expert: Efficient Reward-guided Ensemble of Large Language Models.

Tiny Reward Models Routing to the Expert: Efficient Reward-guided Ensemble of Large Language Models

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-06T17:48:07.159727Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:48:07.159727Z digest=sha256:aa6f0a43284d74258c090610884c2252898a4dc1201571c8961b88f835da5dd5

Observation 7cb9e172-efc3-4fa9-932a-ea5822a66474 · outbound

This paper cites Improve Mathematical Reasoning in Language Models by Automated Process Supervision.

Tiny Reward Models Improve Mathematical Reasoning in Language Models by Automated Process Supervision

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-06T17:48:07.162023Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:48:07.162023Z digest=sha256:ee84508314557414922b66da528262782de9585db0638b8639934ed4a4da3412

Observation f02cbb16-bc9d-4043-8623-29554abc7501 · outbound

This paper cites WebGPT: Browser-assisted question-answering with human feedback.

Tiny Reward Models WebGPT: Browser-assisted question-answering with human feedback

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-06T17:48:07.164185Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:48:07.164185Z digest=sha256:214b1b0a222b0cb17245508ea91a58930ca621ce096381ae3b4c76cac616831d

Observation 3279575f-3339-45b6-9782-4f22b9fa39ce · outbound

This paper cites Passage Re-ranking with BERT.

Tiny Reward Models Passage Re-ranking with BERT

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-06T17:48:07.166590Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:48:07.166590Z digest=sha256:d09b177ea47569a799eab7e8af5f3fa1633c11f0b72a05ce06d8bd33862aef04

Observation 57c0428d-05be-40a2-bc31-aae00aa506bc · outbound

This paper cites Nemotron-4 340B Technical Report.

Tiny Reward Models Nemotron-4 340B Technical Report

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-06T17:48:07.168990Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:48:07.168990Z digest=sha256:2bca1fdaefba5af947bdc21e75f8679d9aeb18b00527fdc75fd1d5b4ed46e163

Observation 6a490881-e7e0-4992-9191-c5b2b4fd2089 · outbound

This paper cites Training language models to follow instructions with human feedback.

Tiny Reward Models Training language models to follow instructions with human feedback

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-06T17:48:07.171387Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:48:07.171387Z digest=sha256:1d186f71379387e56005e81f45720085e7d50e9d49811976c595fe1af2f8051d

Observation 7a8c7288-d584-4bb4-95e9-f3b0ba4fb327 · outbound

This paper cites Let's Reinforce Step by Step.

Tiny Reward Models Let's Reinforce Step by Step

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-06T17:48:07.173925Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:48:07.173925Z digest=sha256:dc5ad396ec38c84a5052c6737c14a6ee8e3ef36b92ca77d9294d1e4b779ba395

Observation af90b18e-5bee-4183-85a5-62a94e9b2624 · outbound

This paper cites Exploring the Limits of Transfer Learning with a Unified Text-to-Text Transformer.

Tiny Reward Models Exploring the Limits of Transfer Learning with a Unified Text-to-Text Transformer

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-06T17:48:07.176314Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:48:07.176314Z digest=sha256:522d678ab54f2dc35b3ad55979ba64312403159ce5335969ae291f5df7352e3c

Observation c96758dd-57fc-40a8-9c3f-1eaab651fd41 · outbound

This paper cites WARM: On the Benefits of Weight Averaged Reward Models.

Tiny Reward Models WARM: On the Benefits of Weight Averaged Reward Models

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-06T17:48:07.178654Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:48:07.178654Z digest=sha256:f1bf831258ba9258c23f56e15ad849b99cd9d86ab4840826d37c326a5e763909

Observation e265842c-7e61-40cc-afc5-6409dd85f99f · outbound

This paper cites BERTs are Generative In-Context Learners.

Tiny Reward Models BERTs are Generative In-Context Learners

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-06T17:48:07.181168Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:48:07.181168Z digest=sha256:9e2a9649509cc57e9ad81d3594247b2c3e5b4368a2ab894f18920c52dd196b57

Observation 55c7ffa2-a916-4d5a-9f42-37c62c029a48 · outbound

This paper cites Exploiting Cloze Questions for Few Shot Text Classification and Natural Language Inference.

Tiny Reward Models Exploiting Cloze Questions for Few Shot Text Classification and Natural Language Inference

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-06T17:48:07.183516Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:48:07.183516Z digest=sha256:ca86507bae518a4ea5bef0883e6a036795f0cb62b95f13401c3fa7d60b3efb6c

Observation abb3c454-889b-487e-9e31-95e36072bc82 · outbound

This paper cites It's Not Just Size That Matters: Small Language Models Are Also Few-Shot Learners.

Tiny Reward Models It's Not Just Size That Matters: Small Language Models Are Also Few-Shot Learners

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-06T17:48:07.185809Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:48:07.185809Z digest=sha256:bdb43c5038127a98b1ba0d25b1e71649073c277bec7eac24570999ba38c22c35

Observation 2237a1c5-ba21-4367-b8f6-dc294cd7fbbc · outbound

This paper cites The Trickle-down Impact of Reward (In-)consistency on RLHF.

Tiny Reward Models The Trickle-down Impact of Reward (In-)consistency on RLHF

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-06T17:48:07.188338Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:48:07.188338Z digest=sha256:7797a61e6215c2252bf75af592d0769f3ccaf0c2551c4039e66756726745a5f8

Observation 986bf3e7-c475-4689-8ffb-a0a7508a53d1 · outbound

This paper cites Layer by Layer: Uncovering Hidden Representations in Language Models.

Tiny Reward Models Layer by Layer: Uncovering Hidden Representations in Language Models

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-06T17:48:07.190751Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:48:07.190751Z digest=sha256:93d60e65feaf148d731c7e02e54c30fb663c90b749b67a1a7db5a94d46d28c26

Observation 60319909-5b82-4069-bfa0-023280b205c8 · outbound

This paper cites D., and Su, W.

Tiny Reward Models D., and Su, W

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-06T17:48:07.193757Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:48:07.193757Z digest=sha256:d3ec0937eaf0c1a26edfbdab84ecd6d25f44792d3ab8daaf2f72c2e770efea7b

Observation faba13a6-cc57-4e9c-9c9b-8f8b81993a6d · outbound

This paper cites Large GPT-like Models are Bad Babies: A Closer Look at the Relationship between Linguistic Competence and Psycholinguistic Measures.

Tiny Reward Models Large GPT-like Models are Bad Babies: A Closer Look at the Relationship between Linguistic Competence and Psycholinguistic Measures

Reference 43

Resolution
verified exact
local_arxiv, observed 2026-08-06T17:48:07.285710Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T17:48:07.196004Z digest=sha256:4574f4666335fd9c72926700cd3bc4414aaac5223aa2884d7049564c4c40d1de

Observation d63fe4c4-efae-4003-a284-2e58a6e4d6a0 · outbound

This paper cites How to Fine-Tune BERT for Text Classification?.

Tiny Reward Models How to Fine-Tune BERT for Text Classification?

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-06T17:48:07.198467Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:48:07.198467Z digest=sha256:25dc8eeeb8f316b19660eb900d0c88cc08a728c7597643e73d43d39d226db884

Observation 5be4873b-40e3-45a8-896b-91a648f1633d · outbound

This paper cites HelpSteer2: Open-source dataset for training top-performing reward models.

Tiny Reward Models HelpSteer2: Open-source dataset for training top-performing reward models

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-06T17:48:07.201291Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:48:07.201291Z digest=sha256:5adc1ab7a16fcfd9595961dd76a86cfa9240c8b6064471630aeab2bc4dfadf88

Observation 79311ea0-1930-4be7-87c8-9a9fc00b6d25 · outbound

This paper cites Smarter, Better, Faster, Longer: A Modern Bidirectional Encoder for Fast, Memory Efficient, and Long Context Finetuning and Inference.

Tiny Reward Models Smarter, Better, Faster, Longer: A Modern Bidirectional Encoder for Fast, Memory Efficient, and Long Context Finetuning and Inference

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-06T17:48:07.203840Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:48:07.203840Z digest=sha256:086e679232bb922daf486744e9bfec184c90c986880e6b227d6b47d43d41848a

Observation de32fa4b-0030-450f-a9c6-4897c78b760c · outbound

This paper cites Finetuned Language Models Are Zero-Shot Learners.

Tiny Reward Models Finetuned Language Models Are Zero-Shot Learners

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-06T17:48:07.206091Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:48:07.206091Z digest=sha256:f7b8cbbddd430a8e91ecd3351d53060dae16db912275606c24ad422d9290b6b3

Observation a71856fb-f940-4e8f-8098-d0b4794b6b94 · outbound

This paper cites MetaMetrics: Calibrating Metrics For Generation Tasks Using Human Preferences.

Tiny Reward Models MetaMetrics: Calibrating Metrics For Generation Tasks Using Human Preferences

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-06T17:48:07.208459Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:48:07.208459Z digest=sha256:d54f0ca0269c38952bbd8a639dcae83b1aa24ca157fad428bf4a3826af3cf26b

Observation b5ece352-57bb-4918-bb9e-91cd6961c748 · outbound

This paper cites How transferable are features in deep neural networks?.

Tiny Reward Models How transferable are features in deep neural networks?

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-06T17:48:07.210826Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:48:07.210826Z digest=sha256:95407d730128a39ef3968da9b70928829dfe33f0cc8dac103e10deda3d1cea11

Pith citing papers

No inbound Pith citation observations are available.