Pith. sign in

Paper Citation Record · LEDGER

Self-Rewarding Language Models

As of 6 August 2026, this Paper Citation Record lists 100 of 104 outbound references and 87 inbound Pith citation observations for arXiv:2401.10020.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2401.10020 v3

Coverage vector

measured 100 of 104 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-05-13T12:01:42.290502Z

measured 187 of 187 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-05T06:32:48.257954+00:00

measured 87 of 87 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-06T00:42:59.909123Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-05T02:28:24.338817Z

Reference resolution

100 of 104 outbound references displayed

  • verified exact23
  • verified fuzzy49
  • unresolved6
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch22

External citation measurements

9
pith, observed 2026-08-05T02:28:24.338817Z

Outbound references

Observation 31994b54-2ec2-418c-9940-0e5dbf79ce86 · outbound

This paper cites Advances in Neural Information Processing Systems , volume=.

Self-Rewarding Language Models Advances in Neural Information Processing Systems , volume=

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T12:01:42.726708Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-13T12:01:42.290502Z digest=sha256:ccfd7f972ebc66761259e854cc8562876014132bac6484a7864b20c2aaac1921

Observation 725116c1-126f-4659-a96b-355d62b9efca · outbound

This paper cites Think you have solved question answering?.

Self-Rewarding Language Models Think you have solved question answering?

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T12:01:42.595308Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-13T12:01:42.290502Z digest=sha256:9304160d95ed00a73b3a671126707592e2ce75eaceea9ad967e035ab2787522a

Observation d5b18464-4550-4bde-aee1-8b86426bb5c2 · outbound

This paper cites 2019 , journal =.

Self-Rewarding Language Models 2019 , journal =

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T12:01:42.598333Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-13T12:01:42.290502Z digest=sha256:bb8b4d23f74b5bb993fbf0caa9c2f19ab5ab38f615fa334d078e9b55fa85eeac

Observation 0af92e01-73e4-4365-a060-411220df14de · outbound

This paper cites EMNLP , year=.

Self-Rewarding Language Models EMNLP , year=

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T12:01:42.601211Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-13T12:01:42.290502Z digest=sha256:bbbd4d07733b86eb98809068c6974f4dd81e0caa950ce415416f93851a3bc20e

Observation b1454ac6-0334-4701-9ff2-01acf8007344 · outbound

This paper cites 9th International Conference on Learning Representations.

Self-Rewarding Language Models 9th International Conference on Learning Representations

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T12:01:42.604145Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-13T12:01:42.290502Z digest=sha256:4ba0a9bbab29dc2935d380cabd540c3245df055b19f1ea4155d409ecb1f8eab3

Observation f1b61372-b3e4-48a4-87f8-14087334fbc9 · outbound

This paper cites Thirty-Fourth AAAI Conference on Artificial Intelligence , year =.

Self-Rewarding Language Models Thirty-Fourth AAAI Conference on Artificial Intelligence , year =

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T12:01:42.607340Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-13T12:01:42.290502Z digest=sha256:0c0e394555844a95112012f801c2c56a0dbf8db258803a253036591646ed78b2

Observation adfac2cf-8595-43c9-940a-ae239c7daced · outbound

This paper cites ROUGE : A Package for Automatic Evaluation of Summaries.

Self-Rewarding Language Models ROUGE : A Package for Automatic Evaluation of Summaries

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T12:01:42.610330Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-13T12:01:42.290502Z digest=sha256:7f77c74a831bf6b0695b0a0804a0aaaf9d7956b7349017eec7e1f6d1b0a47b70

Observation 2933b852-04f0-4fac-9f69-a952b05ad815 · outbound

This paper cites 2023 , howpublished =.

Self-Rewarding Language Models 2023 , howpublished =

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T12:01:42.614142Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-13T12:01:42.290502Z digest=sha256:c8d22f5c15c1f47d4be778ad35743c90d247516ef7673e69c27c610711d09c53

Observation 64f1b9d9-6415-40b2-9a78-b7c4d56d4a17 · outbound

This paper cites Improving Neural Machine Translation Models with Monolingual Data.

Self-Rewarding Language Models Improving Neural Machine Translation Models with Monolingual Data

Reference 12

Resolution
metadata mismatch
arxiv_id, observed 2026-05-13T12:01:42.374858Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-13T12:01:42.290502Z digest=sha256:26b2aa3563e1241b25fab9d7629f5850cfbd29e77ee3e7e0eab706eca547cdcb

Observation c52f7301-bc29-450e-9902-5dda04c09567 · outbound

This paper cites Tagged Back-Translation.

Self-Rewarding Language Models Tagged Back-Translation

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-07-04T23:42:53.552734Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-13T12:01:42.290502Z digest=sha256:20331816a60f8dd77154ad4dd8b1529f23421fc1cb6dd357729a348ebcd11d6f

Observation 4fa5bd23-995b-4df4-b2e0-cf9d79204e56 · outbound

This paper cites QLoRA: Efficient Finetuning of Quantized LLMs.

Self-Rewarding Language Models QLoRA: Efficient Finetuning of Quantized LLMs

Reference 16

Resolution
metadata mismatch
local_arxiv, observed 2026-05-13T12:01:42.461178Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-13T12:01:42.290502Z digest=sha256:0cfc978d71546e473445fa8a2000deda22b6fa761f267f98f455e3fd75e9558a

Observation 40ff08f2-550a-491f-98dd-61d8f18cd0ab · outbound

This paper cites LongForm: Effective Instruction Tuning with Reverse Instructions.

Self-Rewarding Language Models LongForm: Effective Instruction Tuning with Reverse Instructions

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-05-13T12:01:42.546745Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-13T12:01:42.290502Z digest=sha256:c01c4c24b51e7ad2887e708dfdaac72f8ce29d6f6817ca7fa022b19ee1717962

Observation 7af784e0-5d8f-47ae-b031-6af3d7ad3cec · outbound

This paper cites Advances in Neural Information Processing Systems , volume=.

Self-Rewarding Language Models Advances in Neural Information Processing Systems , volume=

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T12:01:42.617078Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-13T12:01:42.290502Z digest=sha256:57c02a442d54867439084f656915cf7070e5bb50058680ab60d5eaf7cc9541c5

Observation 9e057a06-22ec-426c-aa0f-14e7df7081ed · outbound

This paper cites LIMA: Less Is More for Alignment.

Self-Rewarding Language Models LIMA: Less Is More for Alignment

Reference 19

Resolution
metadata mismatch
arxiv_id, observed 2026-05-17T11:34:13.088019Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-13T12:01:42.290502Z digest=sha256:a862a15b124410a41c7cbfe552037a3e15113f10110ac6ca6ce92acd04a8cad7

Observation a7fa6976-dfbf-4bc7-93c7-59c428293d37 · outbound

This paper cites The False Promise of Imitating Proprietary LLMs.

Self-Rewarding Language Models The False Promise of Imitating Proprietary LLMs

Reference 20

Resolution
metadata mismatch
arxiv_id, observed 2026-05-18T06:54:31.902059Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-13T12:01:42.290502Z digest=sha256:80fe7c22af2f876b9c7b26b510d3c9f3af5a013dc8e39bb5fe92c2b56a39eac0

Observation 98fcc26b-8ee0-40ed-a112-783ac403c827 · outbound

This paper cites WizardLM: Empowering large pre-trained language models to follow complex instructions.

Self-Rewarding Language Models WizardLM: Empowering large pre-trained language models to follow complex instructions

Reference 21

Resolution
metadata mismatch
local_arxiv, observed 2026-05-13T12:01:42.451682Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-13T12:01:42.290502Z digest=sha256:770e83d64c7364f1535ec6d9393ff6d317cf921bdbc87cb1b122368fc74a82cb

Observation be8aa6ee-88d2-4e99-b978-321456987a9c · outbound

This paper cites and Stoica, Ion and Xing, Eric P.

Self-Rewarding Language Models and Stoica, Ion and Xing, Eric P

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T12:01:42.620279Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-13T12:01:42.290502Z digest=sha256:c0c869061406c2bb4e0782b69fa1dd1c054781f9f173113d4465f4fe84032dcd

Observation d4616484-7a1e-4372-9165-d2878530372c · outbound

This paper cites Hashimoto , title =.

Self-Rewarding Language Models Hashimoto , title =

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T12:01:42.623273Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-13T12:01:42.290502Z digest=sha256:2a0ad6dfb07e05eceb8e79b534e447e28c177ff76adeb3aec15c6f7637517c65

Observation 19e8d7e5-02d2-4c8f-91b9-8ca1703ed691 · outbound

This paper cites Instruction Tuning with GPT-4.

Self-Rewarding Language Models Instruction Tuning with GPT-4

Reference 24

Resolution
metadata mismatch
arxiv_id, observed 2026-05-14T17:04:18.193148Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-13T12:01:42.290502Z digest=sha256:e8b4de9a39cb5151fa20f16c842171d29deba9300e971c2e7ae240ae174655fb

Observation b9509c37-137a-4f8e-b156-24fccf42a9d2 · outbound

This paper cites OPT-IML: Scaling Language Model Instruction Meta Learning through the Lens of Generalization.

Self-Rewarding Language Models OPT-IML: Scaling Language Model Instruction Meta Learning through the Lens of Generalization

Reference 25

Resolution
metadata mismatch
arxiv_id, observed 2026-05-17T06:08:25.849600Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-13T12:01:42.290502Z digest=sha256:ffef27effea9f3e1f8138a04e7a86df025707839b3776ec901b85a9517c5119d

Observation eafb547f-1599-40e1-9392-7deb52fd7233 · outbound

This paper cites arXiv e-prints , pages=.

Self-Rewarding Language Models arXiv e-prints , pages=

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T12:01:42.625735Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-13T12:01:42.290502Z digest=sha256:830cf83442843e441ba732faed54e9ea6a122c98f2738dd49bef8ca96e860641

Observation b9bed052-c2c1-4124-a23e-78a9478662e7 · outbound

This paper cites an unresolved cited work.

Self-Rewarding Language Models Unresolved cited work

Reference 27

Resolution
unresolved
raw_fallback, observed 2026-05-13T12:01:42.628125Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-13T12:01:42.290502Z digest=sha256:3a6f4d3d954fa4f785a2ede46fe4e485986a1064fe48398346e715009c084da6

Observation 763f42ac-7f9f-4424-95d4-44f2f663cf0f · outbound

This paper cites Finetuned Language Models Are Zero-Shot Learners.

Self-Rewarding Language Models Finetuned Language Models Are Zero-Shot Learners

Reference 28

Resolution
metadata mismatch
local_arxiv, observed 2026-05-13T12:01:42.405453Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-13T12:01:42.290502Z digest=sha256:6bb38d9c3aa00dc4ae83b2cf002f674f5e78be9a7616f89e61222c9b79d4093a

Observation 6cd73ad8-b993-4883-87fd-704103e19a2b · outbound

This paper cites Hashimoto , title =.

Self-Rewarding Language Models Hashimoto , title =

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T12:01:42.631506Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-13T12:01:42.290502Z digest=sha256:a02d68d2789cac3321dfcc577d624680323014ba549bc21bfba25e1e374cc70f

Observation 647d196f-cc9c-42cd-969a-cc80e46d4829 · outbound

This paper cites Proceedings of the 45th International ACM SIGIR Conference on Research and Development in Information Retrieval , pages=.

Self-Rewarding Language Models Proceedings of the 45th International ACM SIGIR Conference on Research and Development in Information Retrieval , pages=

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T12:01:42.634714Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-13T12:01:42.290502Z digest=sha256:4ad0c6475875c56aff2a809b9b6e80df5e5495fa48ec1412cd69b96d9e1bdbd6

Observation 7ab2ef3d-d2a8-4202-8b2a-70a8251ae45d · outbound

This paper cites an unresolved cited work.

Self-Rewarding Language Models Unresolved cited work

Reference 31

Resolution
unresolved
raw_fallback, observed 2026-05-13T12:01:42.637591Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-13T12:01:42.290502Z digest=sha256:21bd6f380d6142bf9ae77016b6ad89facd7131f4a12224ce56df1e256df9de39

Observation a544a119-f541-422e-8442-33020c80b8f8 · outbound

This paper cites 2023 , eprint=.

Self-Rewarding Language Models 2023 , eprint=

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T12:01:42.640509Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-13T12:01:42.290502Z digest=sha256:8a89ef324271921a5c0d20fbdefe94482753588f65f2aceabc92c7f02ad4b382

Observation 06370a74-d99d-4f57-af55-7ad9b48d23a0 · outbound

This paper cites 2024 , url=.

Self-Rewarding Language Models 2024 , url=

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T12:01:42.643491Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-13T12:01:42.290502Z digest=sha256:f7c6c2c6f04abb7210029cdf2398c1cf294924e6ee546a36be45a2dad85044db

Observation 6b530596-6414-4a8c-af4a-04358233053c · outbound

This paper cites Multitask Prompted Training Enables Zero-Shot Task Generalization.

Self-Rewarding Language Models Multitask Prompted Training Enables Zero-Shot Task Generalization

Reference 34

Resolution
metadata mismatch
arxiv_id, observed 2026-05-14T17:59:43.395267Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-13T12:01:42.290502Z digest=sha256:9e2b201f5cc678219216cb924069eda309d4c5a2e90ecea831698bfd7822c0b6

Observation 0149cd98-8d07-4f7b-9395-33996d42f4a1 · outbound

This paper cites Cross-Task Generalization via Natural Language Crowdsourcing Instructions.

Self-Rewarding Language Models Cross-Task Generalization via Natural Language Crowdsourcing Instructions

Reference 35

Resolution
metadata mismatch
arxiv_id, observed 2026-05-18T01:57:29.846007Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-13T12:01:42.290502Z digest=sha256:76e4f7eba58376815f2e86064b1340d234fecc03a130b833a0c44c6eb16946df

Observation c2dd0cec-a04b-46da-af65-812fed6f2564 · outbound

This paper cites Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP Tasks.

Self-Rewarding Language Models Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP Tasks

Reference 36

Resolution
metadata mismatch
arxiv_id, observed 2026-05-13T12:01:42.495070Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-13T12:01:42.290502Z digest=sha256:e8afd7c115e93ef2dd906a16855426079435bd0c1c28a97a719260b56456d2b0

Observation ee4707a8-990e-4426-b4aa-8027e2b0051c · outbound

This paper cites Constitutional.

Self-Rewarding Language Models Constitutional

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T12:01:42.646453Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-13T12:01:42.290502Z digest=sha256:b4df06776ab1d95f8c44be81b86c32acc02b1c3d261782355e9a4cc4867057cd

Observation cd30226b-d5df-49ec-b9c9-d52ecdb4ddd9 · outbound

This paper cites Self-critiquing models for assisting human evaluators.

Self-Rewarding Language Models Self-critiquing models for assisting human evaluators

Reference 39

Resolution
metadata mismatch
arxiv_id, observed 2026-05-16T20:25:41.981797Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-13T12:01:42.290502Z digest=sha256:9e318c6757f88773463ac0e045179035213d121f68a25188523c754abeeda4f6

Observation 944ba9d4-8dff-42ff-bbf9-6744e12ce992 · outbound

This paper cites Self-Refine: Iterative Refinement with Self-Feedback.

Self-Rewarding Language Models Self-Refine: Iterative Refinement with Self-Feedback

Reference 40

Resolution
metadata mismatch
local_arxiv, observed 2026-05-13T12:01:42.522298Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-13T12:01:42.290502Z digest=sha256:5a5804c0024879b1cfe462ea1c8d97fe98d8722a4ee326ea0bbb7fdff0436a65

Observation 5d7793c5-945d-4384-81b3-b545bf8c11df · outbound

This paper cites an unresolved cited work.

Self-Rewarding Language Models Unresolved cited work

Reference 41

Resolution
unresolved
raw_fallback, observed 2026-05-13T12:01:42.649362Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-13T12:01:42.290502Z digest=sha256:b727b89c3ee1eb752f0a8e05d3f5f9718ef6e6fed86d528e84bce2f1b8dd6292

Observation 2e3b20d5-15ee-49e4-b670-2681968984bb · outbound

This paper cites 2023 , month =.

Self-Rewarding Language Models 2023 , month =

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T12:01:42.651546Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-13T12:01:42.290502Z digest=sha256:ece631b7a062b023e85ae6331192058e5cb4c3d6971750c2a27dfbe8315703ac

Observation eb8e6915-5f59-40cf-a93f-5257cf003227 · outbound

This paper cites Enhancing Chat Language Models by Scaling High-quality Instructional Conversations.

Self-Rewarding Language Models Enhancing Chat Language Models by Scaling High-quality Instructional Conversations

Reference 43

Resolution
metadata mismatch
arxiv_id, observed 2026-05-15T17:25:08.436185Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-13T12:01:42.290502Z digest=sha256:c1fb9bc70aed3d9c7618ecb6e372547d5336c2215ba49d5a12c6467528cabad5

Observation 3b635495-6310-4e9d-8e82-9a7e81e62345 · outbound

This paper cites The Curious Case of Neural Text Degeneration.

Self-Rewarding Language Models The Curious Case of Neural Text Degeneration

Reference 45

Resolution
metadata mismatch
local_arxiv, observed 2026-05-13T12:01:42.559457Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-13T12:01:42.290502Z digest=sha256:7e95c6dd075d14ebc9d4486d6e76b5441e4aa082140b509e0ece0f4ce45719d8

Observation 55a05645-df9d-4464-a153-b3b80b675902 · outbound

This paper cites arXiv e-prints , pages=.

Self-Rewarding Language Models arXiv e-prints , pages=

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T12:01:42.654175Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-13T12:01:42.290502Z digest=sha256:c50a5503863d5a4e6f38337bc35a7e4efab81b724ab0c5b14aeeffb2e31f23a6

Observation a93c8510-cb55-44e1-8637-8e8266c155c7 · outbound

This paper cites CrowS-Pairs: A Challenge Dataset for Measuring Social Biases in Masked Language Models.

Self-Rewarding Language Models CrowS-Pairs: A Challenge Dataset for Measuring Social Biases in Masked Language Models

Reference 47

Resolution
verified exact
arxiv_id, observed 2026-05-13T12:01:42.380838Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-13T12:01:42.290502Z digest=sha256:b0cb2602734b2e059a9949a06a1c1f1ee133aa9071a8b48e29b117a57ea820a3

Observation 72a95905-f221-497e-9e5c-f88af0db26c3 · outbound

This paper cites How Far Can Camels Go? Exploring the State of Instruction Tuning on Open Resources.

Self-Rewarding Language Models How Far Can Camels Go? Exploring the State of Instruction Tuning on Open Resources

Reference 48

Resolution
verified exact
arxiv_id, observed 2026-05-13T12:01:42.385415Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-13T12:01:42.290502Z digest=sha256:6850ff7a79b1619e7085b979de39d01a81085c705c81b582c8aa0cef3b68df98

Observation eb703f43-de70-429b-a021-c5c08bb8d600 · outbound

This paper cites Proceedings of the AAAI conference on artificial intelligence , volume=.

Self-Rewarding Language Models Proceedings of the AAAI conference on artificial intelligence , volume=

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T12:01:42.656831Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-13T12:01:42.290502Z digest=sha256:e8812cf30e9cbb6256bb208c10f5d48fc5785de5fb5569f2638e202a1416bcb7

Observation 318c89fe-b485-45d8-ba2d-b0110e84726e · outbound

This paper cites Can a Suit of Armor Conduct Electricity? A New Dataset for Open Book Question Answering.

Self-Rewarding Language Models Can a Suit of Armor Conduct Electricity? A New Dataset for Open Book Question Answering

Reference 52

Resolution
metadata mismatch
local_arxiv, observed 2026-05-13T12:01:42.420722Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-13T12:01:42.290502Z digest=sha256:991542af462c84af06ce5e37ae391ea30f124cef07bf5ca7593e63e0572e928b

Observation 6312acd9-a68e-4fbf-9686-f9e530a27c64 · outbound

This paper cites Measuring Massive Multitask Language Understanding.

Self-Rewarding Language Models Measuring Massive Multitask Language Understanding

Reference 53

Resolution
metadata mismatch
local_arxiv, observed 2026-05-13T12:01:42.433928Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-13T12:01:42.290502Z digest=sha256:b6c26eb3bf550911cb25386e1976e388a4ca69cf22ba973a549f6cdb962c029e

Observation 0076f5d9-fd96-43dc-a896-db3d2a8ee47c · outbound

This paper cites 2023 , url =.

Self-Rewarding Language Models 2023 , url =

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T12:01:42.659359Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-13T12:01:42.290502Z digest=sha256:21df570ac3f283d3869d240c899038d3aa61d633455ff352d9e2d65be99957e5

Observation 38d8faae-22d3-4009-b78d-500213db6a87 · outbound

This paper cites Thirty-seventh Conference on Neural Information Processing Systems , year=.

Self-Rewarding Language Models Thirty-seventh Conference on Neural Information Processing Systems , year=

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T12:01:42.661947Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-13T12:01:42.290502Z digest=sha256:12dde9473e6fc9ad03bb559e70457c85e0266fbd9ce8b3f6de53467b7c4b4578

Observation 979de74b-3b4e-4f83-b644-3401c22f9297 · outbound

This paper cites The Twelfth International Conference on Learning Representations , year=.

Self-Rewarding Language Models The Twelfth International Conference on Learning Representations , year=

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T12:01:42.664456Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-13T12:01:42.290502Z digest=sha256:5c1a5ca0b5b80f481015b5b6f611947147dbbe86f026dfaba2b42af38ea3fc38

Observation 267b1716-1260-4025-aa8d-b650af90d68e · outbound

This paper cites Gonzalez and Ion Stoica , booktitle=.

Self-Rewarding Language Models Gonzalez and Ion Stoica , booktitle=

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T12:01:42.667327Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-13T12:01:42.290502Z digest=sha256:e1ff46beab4476b5260c052412ac17cfc481f530bb82b7a2c729a875c73f00f7

Observation 81db4109-a291-40aa-8c62-fd54d523033e · outbound

This paper cites an unresolved cited work.

Self-Rewarding Language Models Unresolved cited work

Reference 60

Resolution
unresolved
raw_fallback, observed 2026-05-13T12:01:42.670297Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-13T12:01:42.290502Z digest=sha256:ed131804dbaabdd75b05d83a55f88bb7a83bd5e62de4c272772ce1acb7d73ac4

Observation 113fe07b-8d74-443a-825f-c8d1106ad10a · outbound

This paper cites Visualizing data using.

Self-Rewarding Language Models Visualizing data using

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T12:01:42.673232Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-13T12:01:42.290502Z digest=sha256:da24fbc7958310f26701ae8398708c6af440eade1b85240d150b52f2dbd95915

Observation f3117206-b70b-4fb1-b066-26a5f495d2e6 · outbound

This paper cites an unresolved cited work.

Self-Rewarding Language Models Unresolved cited work

Reference 63

Resolution
unresolved
raw_fallback, observed 2026-05-13T12:01:42.676247Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-13T12:01:42.290502Z digest=sha256:4443c4b7ace02f62d7caba76c0a87a60d39d50927c23960a25dfa2bb68a96d1b

Observation 1a4f39e1-53b7-42e8-8ba2-8dbe710e112c · outbound

This paper cites Advances in Neural Information Processing Systems , volume=.

Self-Rewarding Language Models Advances in Neural Information Processing Systems , volume=

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T12:01:42.679643Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-13T12:01:42.290502Z digest=sha256:eb6ee3ce4e7ab8b5f99fea71a82fbf531ea0d1a89e8ac9b033f667700d19c783

Observation 61554148-4f04-411a-b2b1-9c9c7c779555 · outbound

This paper cites an unresolved cited work.

Self-Rewarding Language Models Unresolved cited work

Reference 72

Resolution
unresolved
raw_fallback, observed 2026-05-13T12:01:42.682826Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-13T12:01:42.290502Z digest=sha256:45420bf0b05e851e287d0a6b36d949a79da6b4e64effbc32e49fe73a4a722778

Observation 67b346b2-bd1e-48e5-b27f-5e654ed2b8aa · outbound

This paper cites 2023 , url=.

Self-Rewarding Language Models 2023 , url=

Reference 73

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T12:01:42.686899Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-13T12:01:42.290502Z digest=sha256:8d2b46d7cd82aabaca81b324a1bab7377f18237831866e18093a56e7f788d9c2

Observation af17cc94-be69-4df7-9715-dd52c6095034 · outbound

This paper cites Thirty-seventh Conference on Neural Information Processing Systems Datasets and Benchmarks Track , year=.

Self-Rewarding Language Models Thirty-seventh Conference on Neural Information Processing Systems Datasets and Benchmarks Track , year=

Reference 75

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T12:01:42.690354Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-13T12:01:42.290502Z digest=sha256:c7a1c48e0e079646835280e9f50bd1fd8514c09e8502981f9c297d98c6f3ac97

Observation 60f6000e-afc4-4544-ae44-1ac1b3202b15 · outbound

This paper cites LLaMA: Open and Efficient Foundation Language Models.

Self-Rewarding Language Models LLaMA: Open and Efficient Foundation Language Models

Reference 78

Resolution
metadata mismatch
local_arxiv, observed 2026-05-13T12:01:42.532980Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-13T12:01:42.290502Z digest=sha256:2ce04f62ad4723e3ae7a9619bfc98176da4791125cd536efb609f541ee2bbf13

Observation 4dbcd9d2-7203-4380-825d-af05804473e1 · outbound

This paper cites Proceedings of the 25th International Conference on Machine Learning , pages=.

Self-Rewarding Language Models Proceedings of the 25th International Conference on Machine Learning , pages=

Reference 80

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T12:01:42.703356Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-13T12:01:42.290502Z digest=sha256:0741e66555cfe0814671aabfb5b3601709c0aaab0fa02ef0553bcd98b908f7b9

Observation 6b0fc452-10b3-4156-97f4-06fe4158b913 · outbound

This paper cites Machine learning , volume=.

Self-Rewarding Language Models Machine learning , volume=

Reference 81

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T12:01:42.708423Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-13T12:01:42.290502Z digest=sha256:6eac1d7faca73d9e89a2cc08110b2c1a24c1e57c1ff0740efe021d0b51193f0c

Observation 1db4f041-1d64-4170-9258-3463d54c8fee · outbound

This paper cites OpenAI blog , volume=.

Self-Rewarding Language Models OpenAI blog , volume=

Reference 82

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T12:01:42.711957Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-13T12:01:42.290502Z digest=sha256:97c7b91f98d54a204747b8cf5b2905ab076c3368666a9dd9cb4bc7eb3623b0e4

Observation 2b56ee4e-a1e1-4214-8842-259d35c94304 · outbound

This paper cites GPT-4 Technical Report.

Self-Rewarding Language Models GPT-4 Technical Report

Reference 83

Resolution
verified exact
local_arxiv, observed 2026-05-13T12:01:42.555447Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-13T12:01:42.290502Z digest=sha256:d1dd8b86203b4230ca0536d8573a12eab5101096561e9820bcf705c7ea5c5038

Observation 7e47feb6-f96a-4fa4-acf1-15dd074e394f · outbound

This paper cites The CRINGE loss: Learning what language not to model.

Self-Rewarding Language Models The CRINGE loss: Learning what language not to model

Reference 84

Resolution
verified exact
doi, observed 2026-05-13T12:01:42.342416Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-13T12:01:42.290502Z digest=sha256:4684eb2d63df1978db38f897026caeb32b6a175cd5e2431c31e5f88ab19e9bdd

Observation 37dafee1-c794-4784-8c85-d9ee68b506fc · outbound

This paper cites Claude 2.

Self-Rewarding Language Models Claude 2

Reference 85

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T12:01:42.714935Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-13T12:01:42.290502Z digest=sha256:2ccbe2b3cd292862dde64afb2c6c1c1e81c60be71acc1be8578b448a091adb01

Observation 279a7133-f1d4-4e09-8ff2-22512915be9f · outbound

This paper cites Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback.

Self-Rewarding Language Models Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback

Reference 86

Resolution
verified exact
local_arxiv, observed 2026-05-13T12:01:42.563363Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-13T12:01:42.290502Z digest=sha256:65f5fc5dd479961683616ea5513870413aa3104fb92e26e84f6a262423f11abb

Observation 3fce83ed-2ffd-4bba-b7a3-4605e7a517f3 · outbound

This paper cites Constitutional AI: Harmlessness from AI Feedback.

Self-Rewarding Language Models Constitutional AI: Harmlessness from AI Feedback

Reference 87

Resolution
verified exact
local_arxiv, observed 2026-05-13T12:01:42.369226Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-13T12:01:42.290502Z digest=sha256:20f524ba9ffc57b55dc77085d6052e32fbfc19032b23e9f83a5c4888ba766eb0

Observation 3f0f706b-e2db-4877-b4d4-7115fe0dd19e · outbound

This paper cites Benchmarking foundation models with language-model-as-an-examiner.

Self-Rewarding Language Models Benchmarking foundation models with language-model-as-an-examiner

Reference 88

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T12:01:42.722495Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-13T12:01:42.290502Z digest=sha256:6dd1c1c85ced730724f40524decdbed45169cddc3189d8174bfbd3491875463b

Observation 4c230694-83eb-4bec-a20b-dfb20451651b · outbound

This paper cites Piqa: Reasoning about physical commonsense in natural language.

Self-Rewarding Language Models Piqa: Reasoning about physical commonsense in natural language

Reference 89

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T12:01:42.592439Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-13T12:01:42.290502Z digest=sha256:762840c442e0d54875a3933a5732ccdd89a8c7f0aa1f0ae53dd2ee5302fb9dcc

Observation 9c103a83-b4f7-4d84-8aea-f09699b7c539 · outbound

This paper cites AlpaGasus : Training a better alpaca with fewer data.

Self-Rewarding Language Models AlpaGasus : Training a better alpaca with fewer data

Reference 90

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T12:01:42.730778Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-13T12:01:42.290502Z digest=sha256:681fee306189da3024c9eca70f108900eadc621cd9503d1017479e0156a9483a

Observation 7d47b342-7b40-4b9c-ade3-65f777a95b5c · outbound

This paper cites Self-Play Fine-Tuning Converts Weak Language Models to Strong Language Models.

Self-Rewarding Language Models Self-Play Fine-Tuning Converts Weak Language Models to Strong Language Models

Reference 91

Resolution
verified exact
arxiv_id, observed 2026-05-14T23:00:22.024625Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-13T12:01:42.290502Z digest=sha256:4f42fc5f7010cfcb1b7bfdf3db90db187c732c29949071c3e542aae38cf0d4cf

Observation 226a7d65-ffa5-4f6e-8a4a-42f636bdb605 · outbound

This paper cites Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge.

Self-Rewarding Language Models Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge

Reference 92

Resolution
verified exact
local_arxiv, observed 2026-05-13T12:01:42.395007Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-13T12:01:42.290502Z digest=sha256:75eed040c06c5d97c62eafffae0567ca275f72eb691a2c2cf3d9f79337daacc5

Observation e46a2f2d-c1f5-45d5-816e-90e4bc1f6f57 · outbound

This paper cites Training Verifiers to Solve Math Word Problems.

Self-Rewarding Language Models Training Verifiers to Solve Math Word Problems

Reference 93

Resolution
verified exact
local_arxiv, observed 2026-05-13T12:01:42.400273Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-13T12:01:42.290502Z digest=sha256:6ab3a10276b8493876da6d17fd7aa9ca5310410672c4340a58b5c8cbcc83b9ed

Observation c595846b-de6b-4f54-897a-edc878437d05 · outbound

This paper cites A unified architecture for natural language processing: Deep neural networks with multitask learning.

Self-Rewarding Language Models A unified architecture for natural language processing: Deep neural networks with multitask learning

Reference 94

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T12:01:42.734598Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-13T12:01:42.290502Z digest=sha256:95a466df0c9fa5abbfd8433adeaf05ee9e2cb95bd9c21866f840f9779b9d6814

Observation 650f1700-58a3-4bc1-8196-94d98ab8091e · outbound

This paper cites AlpacaFarm: A Simulation Framework for Methods that Learn from Human Feedback.

Self-Rewarding Language Models AlpacaFarm: A Simulation Framework for Methods that Learn from Human Feedback

Reference 95

Resolution
metadata mismatch
arxiv_id, observed 2026-05-13T12:01:42.411055Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-13T12:01:42.290502Z digest=sha256:f41ea97d88b730c80a1cf94f41a0eab2194a850fc894734b87204f46d8be41fd

Observation 40c30f43-a60e-4da2-9bfb-2b749953fefb · outbound

This paper cites The Devil Is in the Errors: Leveraging Large Language Models for Fine-grained Machine Translation Evaluation.

Self-Rewarding Language Models The Devil Is in the Errors: Leveraging Large Language Models for Fine-grained Machine Translation Evaluation

Reference 96

Resolution
verified exact
doi, observed 2026-05-13T12:01:42.359868Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-13T12:01:42.290502Z digest=sha256:0678a0a1d33439d4e31440b2f449d0cde545156248543e0bbbfc345ba13cfbc4

Observation fac3085a-c512-4da3-bb82-1aeb948c1407 · outbound

This paper cites Reinforced Self-Training (ReST) for Language Modeling.

Self-Rewarding Language Models Reinforced Self-Training (ReST) for Language Modeling

Reference 97

Resolution
verified exact
local_arxiv, observed 2026-05-13T12:01:42.415914Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-13T12:01:42.290502Z digest=sha256:8907b607e822774a02baae3fa2de5751e0c02528f3b6dbc6905f6d7e571f2cc5

Observation f0c6dc93-5c2c-47b6-a2d7-f58b65dc8000 · outbound

This paper cites Measuring massive multitask language understanding.

Self-Rewarding Language Models Measuring massive multitask language understanding

Reference 98

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T12:01:42.738256Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-13T12:01:42.290502Z digest=sha256:54356d196a9c82113bbca409491f5020366083538d427c72e3554f3c04530292

Observation 255488b6-6486-47df-8b8e-fe4ecff2c815 · outbound

This paper cites Unnatural Instructions: Tuning Language Models with (Almost) No Human Labor.

Self-Rewarding Language Models Unnatural Instructions: Tuning Language Models with (Almost) No Human Labor

Reference 99

Resolution
verified exact
doi, observed 2026-05-13T12:01:42.364245Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-13T12:01:42.290502Z digest=sha256:dbb325f9ecbfda7568b827b42b81a36f6b6800552a0b89f13f65566ce8bc435d

Observation 027108b2-b368-4f88-889c-94b34120d6c1 · outbound

This paper cites Prometheus: Inducing Fine-grained Evaluation Capability in Language Models.

Self-Rewarding Language Models Prometheus: Inducing Fine-grained Evaluation Capability in Language Models

Reference 100

Resolution
verified exact
arxiv_id, observed 2026-05-13T12:01:42.426442Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-13T12:01:42.290502Z digest=sha256:b8e711fc4af1b9980119c7697b156dba6e470b701ebdb71b7f9ade3aa39ad66a

Observation 5e117b42-fa35-4fe3-976c-c8c37299e9cc · outbound

This paper cites OpenAssistant Conversations -- Democratizing Large Language Model Alignment.

Self-Rewarding Language Models OpenAssistant Conversations -- Democratizing Large Language Model Alignment

Reference 101

Resolution
metadata mismatch
arxiv_id, observed 2026-05-13T12:01:42.430273Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-13T12:01:42.290502Z digest=sha256:63fd3b156ec329a4c0a40e7da8e703466cdffbd5e491affb523d20ba6e76cca3

Observation 938c560b-6c88-41b8-ab7e-b24082894ade · outbound

This paper cites Toutanova, Llion Jones, Ming-Wei Chang, Andrew Dai, Jakob Uszkoreit, Quoc Le, and Slav Petrov.

Self-Rewarding Language Models Toutanova, Llion Jones, Ming-Wei Chang, Andrew Dai, Jakob Uszkoreit, Quoc Le, and Slav Petrov

Reference 102

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T12:01:42.742548Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-13T12:01:42.290502Z digest=sha256:f80b2d04f7f9d251e1d67190689d363c76f2cc0e664739a7069f11dd52941db3

Observation 1703f623-039c-4602-998e-9b96a625b743 · outbound

This paper cites RLAIF vs. RLHF: Scaling Reinforcement Learning from Human Feedback with AI Feedback.

Self-Rewarding Language Models RLAIF vs. RLHF: Scaling Reinforcement Learning from Human Feedback with AI Feedback

Reference 103

Resolution
verified exact
arxiv_id, observed 2026-05-15T21:32:28.362135Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-13T12:01:42.290502Z digest=sha256:99bb83d7f01055b831e86d83672bf2a13e2926d064c7e6d0c0f614ca4e2dd25c

Observation 76b75438-d0c9-4048-abf9-dcd6f057f7a3 · outbound

This paper cites Self-alignment with instruction backtranslation.

Self-Rewarding Language Models Self-alignment with instruction backtranslation

Reference 104

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T12:01:42.746109Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-13T12:01:42.290502Z digest=sha256:387fa8846fbf2a4869f4b4ccca4c82f69534688189041aae24e2ef24aea32dea

Observation 23e4108e-751c-4857-9eed-cc753150129c · outbound

This paper cites Hashimoto.

Self-Rewarding Language Models Hashimoto

Reference 105

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T12:01:42.767942Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-13T12:01:42.290502Z digest=sha256:5c439e3a1b7ca912a31bc566768d42f0cc28de7ccee1ce10e54d141cbcf92650

Observation 1b866eb6-2d81-42e0-be46-b9554cd7985c · outbound

This paper cites ROUGE : A package for automatic evaluation of summaries.

Self-Rewarding Language Models ROUGE : A package for automatic evaluation of summaries

Reference 106

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T12:01:42.773998Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-13T12:01:42.290502Z digest=sha256:8e801fd74ff272ece0f3d96ae990338df6e138826a9ebb68c99cfdf5b68731e3

Observation 5142f443-50db-45a3-a6c3-f960cc0b8358 · outbound

This paper cites Can a suit of armor conduct electricity? a new dataset for open book question answering.

Self-Rewarding Language Models Can a suit of armor conduct electricity? a new dataset for open book question answering

Reference 107

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T12:01:42.777821Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-13T12:01:42.290502Z digest=sha256:99913fd9f94d6d859cefad724fab5604ea0e3ba16874c7e0634f434581f1bfd2

Observation 5f8a03f5-922c-48a2-b2af-b1c2a79ac6ab · outbound

This paper cites Training language models to follow instructions with human feedback.

Self-Rewarding Language Models Training language models to follow instructions with human feedback

Reference 108

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T12:01:42.566640Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-13T12:01:42.290502Z digest=sha256:2793f5df3280002d035e51dcb06e08d3e0951e2e0b9b0c532c59bec934626db4

Observation 6fc8b0e2-3af5-4882-b55f-9c25c10421cd · outbound

This paper cites Automatically Correcting Large Language Models: Surveying the landscape of diverse self-correction strategies.

Self-Rewarding Language Models Automatically Correcting Large Language Models: Surveying the landscape of diverse self-correction strategies

Reference 109

Resolution
verified exact
arxiv_id, observed 2026-05-13T12:01:42.468060Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-13T12:01:42.290502Z digest=sha256:b3ac58ef5e3a094c789cbf4fec80462cd9040390ed6a65dec4d84df5758f949d

Observation 597a14cc-3aa2-40bc-9927-a695fd3ebd26 · outbound

This paper cites Language models are unsupervised multitask learners.

Self-Rewarding Language Models Language models are unsupervised multitask learners

Reference 110

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T12:01:42.570093Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-13T12:01:42.290502Z digest=sha256:52875df457d2e57a9e84f1bb9614a53de65980ccfcd9bd335d830b62410efe02

Observation 0bfc952c-61d4-4ffe-885e-17984c4e86b0 · outbound

This paper cites Direct preference optimization: Your language model is secretly a reward model.

Self-Rewarding Language Models Direct preference optimization: Your language model is secretly a reward model

Reference 111

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T12:01:42.573340Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-13T12:01:42.290502Z digest=sha256:e4f9ab44c16a316035da0fbb50b7d0cccf64336673869cd9da85b3fe9134c8d1

Observation 13dbc281-ed12-4238-b341-4b8a4c0eb765 · outbound

This paper cites Branch-Solve-Merge Improves Large Language Model Evaluation and Generation.

Self-Rewarding Language Models Branch-Solve-Merge Improves Large Language Model Evaluation and Generation

Reference 112

Resolution
verified exact
arxiv_id, observed 2026-05-13T12:01:42.482991Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-13T12:01:42.290502Z digest=sha256:a1e5b96428adc01d7d2a8e9b23c766047fbda0a3867278b6897fd1b9af96b308

Observation 088bd612-4aa5-4c4b-b45a-d8bc5861ae84 · outbound

This paper cites SocialIQA: Commonsense Reasoning about Social Interactions.

Self-Rewarding Language Models SocialIQA: Commonsense Reasoning about Social Interactions

Reference 113

Resolution
verified exact
arxiv_id, observed 2026-05-13T12:22:26.192007Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-13T12:01:42.290502Z digest=sha256:51b53dae53924eb9d386627bcdef6167a8d0510c524b92f0958cf79e5972b3ae

Observation 02e2bcbc-c53a-4f8f-b30f-f535d39c9082 · outbound

This paper cites Proximal Policy Optimization Algorithms.

Self-Rewarding Language Models Proximal Policy Optimization Algorithms

Reference 114

Resolution
verified exact
local_arxiv, observed 2026-05-13T12:01:42.490790Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-13T12:01:42.290502Z digest=sha256:9f6e6a5b1f727d2ed17657dd324f17e85b9ad2b4f67431ce3d37f2f88370d4cb

Observation 9a7d8f85-8216-48fd-9e80-24998d0c5188 · outbound

This paper cites Learning to summarize with human feedback.

Self-Rewarding Language Models Learning to summarize with human feedback

Reference 115

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T12:01:42.576691Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-13T12:01:42.290502Z digest=sha256:bacaad8948d8869eac009bdb7c45d0806f5a7645a80318478a7ae388d56df041

Observation e7a2d0cb-808e-483c-8e33-de3d0441bf7b · outbound

This paper cites Hashimoto.

Self-Rewarding Language Models Hashimoto

Reference 116

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T12:01:42.579741Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-13T12:01:42.290502Z digest=sha256:ed320d19171f538c64d3de537df002c1c6402d3fcec8b94a4564cfb501e34ccb

Observation f70294a5-f8e1-41c3-bdce-ad929aefd96a · outbound

This paper cites Llama 2: Open Foundation and Fine-Tuned Chat Models.

Self-Rewarding Language Models Llama 2: Open Foundation and Fine-Tuned Chat Models

Reference 117

Resolution
verified exact
local_arxiv, observed 2026-05-13T12:01:42.503076Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-13T12:01:42.290502Z digest=sha256:6a961cb8704bdf5d9c93c6273ebf4853fee0edefd1a0b5d2c7883bae9f5f83ec

Observation dc17169d-6b90-404d-9479-58691c9bb72d · outbound

This paper cites Visualizing data using t-SNE.

Self-Rewarding Language Models Visualizing data using t-SNE

Reference 118

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T12:01:42.582928Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-13T12:01:42.290502Z digest=sha256:7ea415c21ce659d863001d4b9b41970f3e3d82835db6f3bb4cd8925b2edb8e68

Observation 542183c1-5c4d-405c-ae3d-19d037588e8f · outbound

This paper cites Smith, Daniel Khashabi, and Hannaneh Hajishirzi.

Self-Rewarding Language Models Smith, Daniel Khashabi, and Hannaneh Hajishirzi

Reference 119

Resolution
verified exact
doi, observed 2026-05-13T12:01:42.351649Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-13T12:01:42.290502Z digest=sha256:f996adaaa00fd966ca16d7be145082aa715112267558836313bc3d09b18a269c

Observation be79fbfa-d555-494a-8c4a-1fae5e1094ea · outbound

This paper cites Iterative Preference Learning from Human Feedback: Bridging Theory and Practice for RLHF under KL-Constraint.

Self-Rewarding Language Models Iterative Preference Learning from Human Feedback: Bridging Theory and Practice for RLHF under KL-Constraint

Reference 120

Resolution
verified exact
arxiv_id, observed 2026-05-13T12:01:42.513253Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-13T12:01:42.290502Z digest=sha256:8c8c335046658edebefd14a714cc6dfb4cae4bda7eb00bc531eb52532c596d51

Observation bdd04f55-6a75-4367-b275-0d6bb227348b · outbound

This paper cites Some things are more CRINGE than others: Iterative Preference Optimization with the Pairwise Cringe Loss.

Self-Rewarding Language Models Some things are more CRINGE than others: Iterative Preference Optimization with the Pairwise Cringe Loss

Reference 121

Resolution
metadata mismatch
arxiv_id, observed 2026-05-13T12:01:42.518264Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-13T12:01:42.290502Z digest=sha256:537fea74fb23a66852f31633b44956ce4b9fff4369645d463f1d23c17c1e76bc

Observation 87ecf850-2d3c-4cb1-b334-e4ba3f24daf1 · outbound

This paper cites RRHF : Rank responses to align language models with human feedback.

Self-Rewarding Language Models RRHF : Rank responses to align language models with human feedback

Reference 122

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T12:01:42.586588Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-13T12:01:42.290502Z digest=sha256:11052f242df5d3d208e93789106f44a2c2e2ffd13c21edaf6a94f2cbf611a690

Observation 9d7f4113-d9d8-4e29-a26f-e2e1170e3243 · outbound

This paper cites URL https:// doi.org/10.18653/v1/p19-1472.

Self-Rewarding Language Models URL https:// doi.org/10.18653/v1/p19-1472

Reference 123

Resolution
metadata mismatch
doi, observed 2026-05-13T12:01:42.347297Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-13T12:01:42.290502Z digest=sha256:5bb041819d21c4abce0d05596dacfb1ec062730f06afe96113c3591a714ac669

Pith citing papers

Observation 0ddbc15f-8636-45db-8b9e-b53b88bcc368 · inbound

Self-Play Fine-Tuning Converts Weak Language Models to Strong Language Models cites this paper.

Self-Play Fine-Tuning Converts Weak Language Models to Strong Language Models Self-Rewarding Language Models

Reference 3

Resolution
metadata mismatch
local_arxiv, observed 2026-05-14T23:00:21.250918Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-14T23:00:20.720030Z digest=sha256:b7cde01a1401924454d777c12ad0ce6355d21915479d104d8fd2b2bfd470e733

Observation b498cc2c-6bf4-4fc5-ad15-a9f91c662656 · inbound

KTO: Model Alignment as Prospect Theoretic Optimization cites this paper.

KTO: Model Alignment as Prospect Theoretic Optimization Self-Rewarding Language Models

Reference 23

Resolution
metadata mismatch
arxiv_id, observed 2026-05-13T12:01:42.779186Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-12T12:17:53.478052Z digest=sha256:8e0f123612f3cfab16ae668efa02241188896b298d34f57c89fe95dadc90ead8

Observation 93501320-7e32-4ca5-885f-c7d18ed1bd81 · inbound

LlamaFactory: Unified Efficient Fine-Tuning of 100+ Language Models cites this paper.

LlamaFactory: Unified Efficient Fine-Tuning of 100+ Language Models Self-Rewarding Language Models

Reference 10

Resolution
metadata mismatch
arxiv_id, observed 2026-05-13T12:01:42.779186Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-11T12:40:02.129674Z digest=sha256:74814384eba4773ace90cb3416dc8c7916476964b80082103b4b8aa90c8fa90b

Observation a1388664-2fa5-4511-bca0-c965b7fa8165 · inbound

TextGrad: Automatic "Differentiation" via Text cites this paper.

TextGrad: Automatic "Differentiation" via Text Self-Rewarding Language Models

Reference 32

Resolution
metadata mismatch
arxiv_id, observed 2026-05-13T12:01:42.779186Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-13T11:27:58.098484Z digest=sha256:3091445832d6c5bd4e4ced2af2835fd3e1d0bd22d93aeffe1393cc769e611605

Observation 759d9ff3-ec21-4452-9a23-6c63f6773f81 · inbound

The FineWeb Datasets: Decanting the Web for the Finest Text Data at Scale cites this paper.

The FineWeb Datasets: Decanting the Web for the Finest Text Data at Scale Self-Rewarding Language Models

Reference 64

Resolution
verified exact
arxiv_id, observed 2026-05-13T12:01:42.779186Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-13T04:36:45.363131Z digest=sha256:2a73b863cc912b48128b4b1aefdbcfd67befd14aba34baf3896243acfe1772e9

Observation d7d6c5fd-dc09-4622-953b-957c1f47e449 · inbound

UNA: A Unified Supervised Framework for Efficient LLM Alignment Across Feedback Types cites this paper.

UNA: A Unified Supervised Framework for Efficient LLM Alignment Across Feedback Types Self-Rewarding Language Models

Reference 23

Resolution
metadata mismatch
local_arxiv, observed 2026-05-23T21:23:27.377474Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-23T21:22:36.970101Z digest=sha256:b7729d328d52fc539955859fcf13faa5256a6bceff6b8cee9ea5f6931d6b194d

Observation 43567dc9-29ee-4850-a561-374a6966c29b · inbound

A Survey on LLM-as-a-Judge cites this paper.

A Survey on LLM-as-a-Judge Self-Rewarding Language Models

Reference 199

Resolution
verified exact
local_arxiv, observed 2026-05-23T17:35:43.863879Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-23T17:33:13.394338Z digest=sha256:303e1051db1cb1a6d9dea1832ae20e617fba401542c90efd051c44d3cb919d9b

Observation 985a03ac-cbb1-4116-aff7-6d47da53886b · inbound

LLMs-as-Judges: A Comprehensive Survey on LLM-based Evaluation Methods cites this paper.

LLMs-as-Judges: A Comprehensive Survey on LLM-based Evaluation Methods Self-Rewarding Language Models

Reference 285

Resolution
verified exact
arxiv_id, observed 2026-05-13T12:01:42.779186Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-11T23:08:34.312466Z digest=sha256:ed268e393ca5433afb897eba3e47acac3d35554089a67d360cb17889049ba68a

Observation 31843f9f-dd1b-4618-a10b-18ab76d89998 · inbound

Reinforcement Learning from Human Feedback cites this paper.

Reinforcement Learning from Human Feedback Self-Rewarding Language Models

Reference 285

Resolution
verified exact
local_arxiv, observed 2026-05-22T19:32:01.174402Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-22T19:27:40.991325Z digest=sha256:195c53c0145cf9b810fd51c56523eeabe0d48adb7c6651be494cd2c234117c4d

Observation 632c9860-33a5-4e8f-b341-5a12f7ac8a02 · inbound

A Survey of Self-Evolving Agents: What, When, How, and Where to Evolve on the Path to Artificial Super Intelligence cites this paper.

A Survey of Self-Evolving Agents: What, When, How, and Where to Evolve on the Path to Artificial Super Intelligence Self-Rewarding Language Models

Reference 195

Resolution
metadata mismatch
local_arxiv, observed 2026-05-14T22:23:15.031489Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-14T22:23:14.621091Z digest=sha256:d84c3511b912e0fb9f9f9534e6f2307c105727dfac169951772ccc682c5453a2

Observation 596f078d-45fb-46bf-a7ad-bcc89012defb · inbound

Temporal Self-Rewarding Language Models: Decoupling Chosen-Rejected via Past-Future cites this paper.

Temporal Self-Rewarding Language Models: Decoupling Chosen-Rejected via Past-Future Self-Rewarding Language Models

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-05T23:06:49.712423Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T23:06:49.712423Z digest=sha256:a0abf2a0215aac285a966a3130630a511070391cca9178f02a52cbf8d54a81a1

Observation b58b8565-ac9d-418f-baf7-a7c69adfc8d5 · inbound

Self-Rewarding Vision-Language Model via Reasoning Decomposition cites this paper.

Self-Rewarding Vision-Language Model via Reasoning Decomposition Self-Rewarding Language Models

Reference 25

Resolution
verified exact
local_arxiv, observed 2026-05-18T21:06:50.919220Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-18T21:03:31.606674Z digest=sha256:8961d52aa30a0ccd2249141801273038eb7c3636e7060a89861db8dcaf676c1b

Observation 6202109f-f96a-4bf0-a89b-444b0fda9ba0 · inbound

Improving Alignment in LVLMs with Debiased Self-Judgment cites this paper.

Improving Alignment in LVLMs with Debiased Self-Judgment Self-Rewarding Language Models

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-05T15:00:53.736544Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T15:00:53.736544Z digest=sha256:ed5dfddfa1062951bb1313f81e38fe78b9c167435f3ceba03e74fdbf6e1f07a2

Observation 8f6d1485-5ea9-4e63-8b1f-def683ab6f0a · inbound

Igniting Creative Writing in Small Language Models: LLM-as-a-Judge versus Multi-Agent Refined Rewards cites this paper.

Igniting Creative Writing in Small Language Models: LLM-as-a-Judge versus Multi-Agent Refined Rewards Self-Rewarding Language Models

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-05T14:22:39.209835Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T14:22:39.209835Z digest=sha256:e3a23b096ec81f7de5f24641349db4f3263e0fb44c8eeb6316d3bde7bf8589f1

Observation 46ff7d88-df70-4e82-ba8b-5a1480e5df3a · inbound

Improving Large Vision and Language Models by Learning from a Panel of Peers cites this paper.

Improving Large Vision and Language Models by Learning from a Panel of Peers Self-Rewarding Language Models

Reference 74

Resolution
unresolved
no resolver link, observed 2026-08-05T12:27:27.690602Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:27:27.690602Z digest=sha256:be49af01c978815b0aa5c2bd56868936a6401d1760ef8626c9e4e8e6704f587a

Observation 457351ee-529c-446f-86c1-9bac5eb10b8b · inbound

Beyond Correctness: Harmonizing Process and Outcome Rewards through RL Training cites this paper.

Beyond Correctness: Harmonizing Process and Outcome Rewards through RL Training Self-Rewarding Language Models

Reference 37

Resolution
verified exact
local_arxiv, observed 2026-05-21T22:40:43.185228Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-21T22:38:57.833414Z digest=sha256:3a964e8485d873b5ec3fdeb9a20e03cea30de852f3bbb4d5d724b582273bde7a

Observation 5f0889fe-3ffc-4605-8f3c-129dde1d2f9e · inbound

Enhancing Speech Large Language Models through Reinforced Behavior Alignment cites this paper.

Enhancing Speech Large Language Models through Reinforced Behavior Alignment Self-Rewarding Language Models

Reference 55

Resolution
metadata mismatch
local_arxiv, observed 2026-05-21T22:24:23.537377Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-21T22:23:52.392075Z digest=sha256:7ca567fce35c9062a7ba62fb9ad116319bfdb23477fa50c576d62cf51c516895

Observation 9072b3cc-72b5-4a2d-8b39-6bfc30a446b7 · inbound

On the Shelf Life of Fine-Tuned LLM-Judges: Future-Proofing, Backward-Compatibility, and Question Generalization cites this paper.

On the Shelf Life of Fine-Tuned LLM-Judges: Future-Proofing, Backward-Compatibility, and Question Generalization Self-Rewarding Language Models

Reference 51

Resolution
verified exact
local_arxiv, observed 2026-05-18T12:56:24.592720Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-18T12:53:45.767341Z digest=sha256:2e37e953329ca80b7571397804bd4aaad7e65e4dbdb42c2b690ca7fb5d109d3c

Observation 0c669d5e-9341-4ae3-9449-1814176dc4a3 · inbound

Breaking the Self-Confirming Loop: Diagnosing and Mitigating Systemic Reward Bias in Self-Rewarding RL cites this paper.

Breaking the Self-Confirming Loop: Diagnosing and Mitigating Systemic Reward Bias in Self-Rewarding RL Self-Rewarding Language Models

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-04T10:44:31.957359Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T10:44:31.957359Z digest=sha256:b8dc458c5b12a42db54ee54b29e37c5dd60bc92c7ceaec6bb12fe9843295c584

Observation 6b66f49f-04bf-4706-aedb-d3cab0e19629 · inbound

Coupled Variational Reinforcement Learning for Language Model General Reasoning cites this paper.

Coupled Variational Reinforcement Learning for Language Model General Reasoning Self-Rewarding Language Models

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-03T16:43:57.799392Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T16:43:57.799392Z digest=sha256:0253b4e7596cf7cca46812b4fe5aee1e224fde03d92c9d58b0a95aa742f256bd

Observation 23a6939e-eb50-4155-bcc8-347fa703c964 · inbound

CPMobius: Iterative Coach-Player Reasoning for Data-Free Reinforcement Learning cites this paper.

CPMobius: Iterative Coach-Player Reasoning for Data-Free Reinforcement Learning Self-Rewarding Language Models

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-03T05:14:21.845716Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T05:14:21.845716Z digest=sha256:37eacb87a5d9c689293c671c3dec134b6483636dd623c3ce551b5e568671e207

Observation 6f426e85-277b-48ca-a7f0-571f93d4a738 · inbound

A Survey of Reinforcement Learning For Economics cites this paper.

A Survey of Reinforcement Learning For Economics Self-Rewarding Language Models

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-02T18:34:55.457454Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T18:34:55.457454Z digest=sha256:827ac9d9ebddb3c6ef91b9d1d030ee3037c013686e73f431ede3c6f77a3bfdd1

Observation 3533ebd3-39cb-403a-aeed-c2988b2bfc98 · inbound

Toward Epistemic Stability: Engineering Consistent Procedures for Industrial LLM Hallucination Reduction cites this paper.

Toward Epistemic Stability: Engineering Consistent Procedures for Industrial LLM Hallucination Reduction Self-Rewarding Language Models

Reference 3

Resolution
metadata mismatch
local_arxiv, observed 2026-05-15T14:30:03.693172Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-15T14:26:09.893096Z digest=sha256:a93f2d3c721f3eed782c82ad9d8f343a8f9d37b93c0ba6dbb1eb801bec81f0b5

Observation eba0c7ab-2891-47b2-8a99-792297af4b3e · inbound

Visual-ERM: Reward Modeling for Visual Equivalence cites this paper.

Visual-ERM: Reward Modeling for Visual Equivalence Self-Rewarding Language Models

Reference 37

Resolution
verified exact
local_arxiv, observed 2026-05-15T11:19:57.863750Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-15T11:19:42.002790Z digest=sha256:f851ef71dd1ac70fdb155cef73255e86ffb8a65b2bc07ee3b26ba3383d9390f3

Observation 9fc9fe92-33da-447d-bf00-19bbdd65bba2 · inbound

AdaRubric: Task-Adaptive Rubrics for Reliable LLM Agent Evaluation and Reward Learning cites this paper.

AdaRubric: Task-Adaptive Rubrics for Reliable LLM Agent Evaluation and Reward Learning Self-Rewarding Language Models

Reference 11

Resolution
verified exact
local_arxiv, observed 2026-05-15T06:49:49.343051Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-15T06:45:59.036473Z digest=sha256:154aa6aca1ae195a70fe625572ddd92ece22ba05952285b9179cf6f880fb4630

Observation c7b12d90-1405-4881-b3fe-0620883b79b4 · inbound

Can LLMs Learn to Reason Robustly under Noisy Supervision? cites this paper.

Can LLMs Learn to Reason Robustly under Noisy Supervision? Self-Rewarding Language Models

Reference 31

Resolution
verified exact
local_arxiv, observed 2026-05-13T17:08:01.273470Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-13T16:58:42.129870Z digest=sha256:ff07ec733bcd68a15707afb73a8239a6d24d6d10e536f67312ee3396e6bdab56

Observation 0b86fdbb-a4f8-40b7-b35f-7d04808a240c · inbound

Pioneer Agent: Continual Improvement of Small Language Models in Production cites this paper.

Pioneer Agent: Continual Improvement of Small Language Models in Production Self-Rewarding Language Models

Reference 102

Resolution
verified exact
arxiv_id, observed 2026-05-13T12:01:42.779186Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-10T17:48:40.520740Z digest=sha256:416dcc006c7532b83e10543ad0363e37d6aa98988529eeb438c598dbceb386d1

Observation b8982aa8-85e4-4c2b-aaa4-469763691a53 · inbound

Utilizing and Calibrating Hindsight Process Rewards via Reinforcement with Mutual Information Self-Evaluation cites this paper.

Utilizing and Calibrating Hindsight Process Rewards via Reinforcement with Mutual Information Self-Evaluation Self-Rewarding Language Models

Reference 7

Resolution
metadata mismatch
arxiv_id, observed 2026-05-13T12:01:42.779186Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T15:32:01.575073Z digest=sha256:3541906ddae610147b3cc2a59a86efafd5b25689fc2668d0130355a41de76d85

Observation d9b6c656-ea4e-4272-8718-995a3aac259d · inbound

Beyond Verifiable Rewards: Rubric-Based GRM for Reinforced Fine-Tuning SWE Agents cites this paper.

Beyond Verifiable Rewards: Rubric-Based GRM for Reinforced Fine-Tuning SWE Agents Self-Rewarding Language Models

Reference 36

Resolution
verified exact
local_arxiv, observed 2026-05-15T12:25:35.780289Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-15T12:22:13.551709Z digest=sha256:1017768e9f88cf9d8544c135eb4746c3fd7ec08ff81958ccbcc0da09eaaf501d

Observation d66bfeb8-2c50-445c-a393-a8a05754e4df · inbound

PoliLegalLM: A Technical Report on a Large Language Model for Political and Legal Affairs cites this paper.

PoliLegalLM: A Technical Report on a Large Language Model for Political and Legal Affairs Self-Rewarding Language Models

Reference 39

Resolution
metadata mismatch
arxiv_id, observed 2026-05-13T12:01:42.779186Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-10T06:03:01.460352Z digest=sha256:f84b5d68f07e33f0bae1c882a46f078fa39b835ea9fdacc12c201d28fb396cbc

Observation edb53ea2-d14a-4f08-a049-cd13b4ddc6ea · inbound

Neural Garbage Collection: Learning to Forget while Learning to Reason cites this paper.

Neural Garbage Collection: Learning to Forget while Learning to Reason Self-Rewarding Language Models

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-05-13T12:01:42.779186Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T04:57:12.923258Z digest=sha256:0fffa2103f38f29225997fd5a59e55c8481872aa43edb6c841faaf6e17363e34

Observation fdc661dd-a1be-4a94-9a16-75e501a7e4e6 · inbound

IRIS: Interpolative R\'enyi Iterative Self-play for Large Language Model Fine-Tuning cites this paper.

IRIS: Interpolative R\'enyi Iterative Self-play for Large Language Model Fine-Tuning Self-Rewarding Language Models

Reference 89

Resolution
verified exact
arxiv_id, observed 2026-05-13T12:01:42.779186Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T01:15:12.985803Z digest=sha256:46ac0a3d07efa66ad8a4002f97a9911bb0acc1dd8717ee401e5dbae5ef8ed185

Observation 541196b6-1eb7-4750-9d55-e6443b126665 · inbound

Skills-Coach: A Self-Evolving Skill Optimizer via Training-Free GRPO cites this paper.

Skills-Coach: A Self-Evolving Skill Optimizer via Training-Free GRPO Self-Rewarding Language Models

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-05-13T12:01:42.779186Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-07T09:56:14.913173Z digest=sha256:8e4aa8cbade2d2af83a590b0e11670b9daf2fc4e4ee1c02ffc7d008259993937

Observation 26262ddf-924c-41f0-817c-81ffcecffe56 · inbound

Enhanced LLM Reasoning by Optimizing Reward Functions with Search-Driven Reinforcement Learning cites this paper.

Enhanced LLM Reasoning by Optimizing Reward Functions with Search-Driven Reinforcement Learning Self-Rewarding Language Models

Reference 24

Resolution
verified exact
arxiv_id, observed 2026-05-13T12:01:42.779186Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-08T19:14:22.162163Z digest=sha256:048a8c961e51d304d7ee952770f5a54db69f05c2de6e5e982f08eaca6b444e0d

Observation bb5d7e04-394e-4cf1-a2a8-843f79faa19a · inbound

Enhanced LLM Reasoning by Optimizing Reward Functions with Search-Driven Reinforcement Learning cites this paper.

Enhanced LLM Reasoning by Optimizing Reward Functions with Search-Driven Reinforcement Learning Self-Rewarding Language Models

Reference 24

Resolution
verified exact
arxiv_id, observed 2026-05-13T12:01:42.779186Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-11T02:10:46.725354Z digest=sha256:aa2c79570dc5eb383de0d048b8f1400a9cc4e23ed5956dce56c23df56a0423f7

Observation 72394f1c-9f91-4a31-872c-bd37d4cad5da · inbound

ARIS: Autonomous Research via Adversarial Multi-Agent Collaboration cites this paper.

ARIS: Autonomous Research via Adversarial Multi-Agent Collaboration Self-Rewarding Language Models

Reference 19

Resolution
verified exact
arxiv_id, observed 2026-05-13T12:01:42.779186Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-08T17:58:27.418883Z digest=sha256:1b06cf20d716515ce4826e695a17b2162f38bf1ad415c10fe27cfb8edee895e9

Observation 2ade9202-b664-440e-8877-40ef044adc74 · inbound

StraTA: Incentivizing Agentic Reinforcement Learning with Strategic Trajectory Abstraction cites this paper.

StraTA: Incentivizing Agentic Reinforcement Learning with Strategic Trajectory Abstraction Self-Rewarding Language Models

Reference 42

Resolution
metadata mismatch
arxiv_id, observed 2026-05-13T12:01:42.779186Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-08T09:59:48.604813Z digest=sha256:c00656b6ea191132a6b4ce7a3fe0e0259bea3e267e14dd6c90067baed467b957

Observation ee02b45c-13a6-472d-b9b2-2ebc2d966ee3 · inbound

SEIF: Self-Evolving Reinforcement Learning for Instruction Following cites this paper.

SEIF: Self-Evolving Reinforcement Learning for Instruction Following Self-Rewarding Language Models

Reference 41

Resolution
verified exact
arxiv_id, observed 2026-05-13T12:01:42.779186Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-11T01:51:09.514927Z digest=sha256:5b2ae7536ea95da3ed605437f64a29e0289862658752509daf57aaa1f80cb177

Observation 74acb89b-9a59-48b2-bcba-d1196ba93e95 · inbound

MedThink: Enhancing Diagnostic Accuracy in Small Models via Teacher-Guided Reasoning Correction cites this paper.

MedThink: Enhancing Diagnostic Accuracy in Small Models via Teacher-Guided Reasoning Correction Self-Rewarding Language Models

Reference 38

Resolution
verified exact
arxiv_id, observed 2026-05-13T12:01:42.779186Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-12T01:22:15.608348Z digest=sha256:e8b924d86b7e353db2478c3d1cfdccaa85bcdfcc557777e9892f385eb032f206

Observation a8e753a1-485a-4b49-b492-234c6d25770d · inbound

Beyond Static Bias: Adaptive Multi-Fidelity Bandits with Improving Proxies cites this paper.

Beyond Static Bias: Adaptive Multi-Fidelity Bandits with Improving Proxies Self-Rewarding Language Models

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-05-13T12:01:42.779186Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-12T01:11:11.354047Z digest=sha256:bd491f70b7e571b5811e0fbf5c5a2bad768bab5ae4373e80807f1ffe40394481

Observation 1744a6ae-c52c-4356-b868-0588f0b8a435 · inbound

CoDistill-GRPO: A Co-Distillation Recipe for Efficient Group Relative Policy Optimization cites this paper.

CoDistill-GRPO: A Co-Distillation Recipe for Efficient Group Relative Policy Optimization Self-Rewarding Language Models

Reference 35

Resolution
verified exact
arxiv_id, observed 2026-05-13T12:01:42.779186Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-12T00:59:44.364491Z digest=sha256:e9e3dd6d64e6cca0c0fcc6ec2c6501dbc18756abb0d0b9da8686cf826f126a86

Observation 33059efc-3cad-4536-8292-f914a96910dc · inbound

Primal Generation, Dual Judgment: Self-Training from Test-Time Scaling cites this paper.

Primal Generation, Dual Judgment: Self-Training from Test-Time Scaling Self-Rewarding Language Models

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-05-13T12:01:42.779186Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-13T02:25:58.830629Z digest=sha256:38b8a411cad614472ec4d5d269551fda2a8fab9bb50ac7b47d2168d78e8f2e46

Observation 88791b15-faf8-4e8d-ae65-0c2d42a27a52 · inbound

Primal Generation, Dual Judgment: Self-Training from Test-Time Scaling cites this paper.

Primal Generation, Dual Judgment: Self-Training from Test-Time Scaling Self-Rewarding Language Models

Reference 92

Resolution
verified exact
local_arxiv, observed 2026-05-14T20:47:58.105294Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-14T20:46:38.558668Z digest=sha256:31cfa70db4ad4898da9fc76531289263c49b450db9fe548a7ecfa2d66f36853d

Observation bd14f2d7-e531-4521-8c78-c8b52055df3b · inbound

From Generic Correlation to Input-Specific Credit in On-Policy Self Distillation cites this paper.

From Generic Correlation to Input-Specific Credit in On-Policy Self Distillation Self-Rewarding Language Models

Reference 31

Resolution
verified exact
arxiv_id, observed 2026-05-13T12:01:42.779186Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-13T01:24:34.093541Z digest=sha256:6fbf2429c043babcfd90dba1942499e94ab816b135f57773d221e804cde227d7

Observation 754e0a15-3f02-4279-b195-0293cbe81b07 · inbound

Towards Order Fairness: Mitigating LLMs Order Sensitivity through Dual Group Advantage Optimization cites this paper.

Towards Order Fairness: Mitigating LLMs Order Sensitivity through Dual Group Advantage Optimization Self-Rewarding Language Models

Reference 45

Resolution
metadata mismatch
arxiv_id, observed 2026-05-13T12:01:42.779186Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-13T07:32:58.404947Z digest=sha256:ec5b1c85a5a79a0ab4d3a13d515fd766d4dc233579637c2b1280235d234635f0

Observation af8ddb94-9ef2-4529-9439-fac1feeffc13 · inbound

TokenRatio: Principled Token-Level Preference Optimization via Ratio Matching cites this paper.

TokenRatio: Principled Token-Level Preference Optimization via Ratio Matching Self-Rewarding Language Models

Reference 140

Resolution
metadata mismatch
arxiv_id, observed 2026-05-13T12:01:42.779186Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-13T04:55:55.013900Z digest=sha256:f732fc32e1a05d8f136b93a671a519e92ce6fa8c18614264452bdf441be647d7

Observation 2fbb6b58-ab9a-49a9-b585-7807f0343689 · inbound

TokenRatio: Principled Token-Level Preference Optimization via Ratio Matching cites this paper.

TokenRatio: Principled Token-Level Preference Optimization via Ratio Matching Self-Rewarding Language Models

Reference 140

Resolution
metadata mismatch
local_arxiv, observed 2026-05-15T05:45:06.693643Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-15T05:41:10.714594Z digest=sha256:900cb71edb360e360ebeb4c1cad08965f3a3f740a28dbe7a2111a25b4be197b3

Observation 00f64250-d0ca-494e-841a-b32e878676b4 · inbound

FrontierSmith: Synthesizing Open-Ended Coding Problems at Scale cites this paper.

FrontierSmith: Synthesizing Open-Ended Coding Problems at Scale Self-Rewarding Language Models

Reference 41

Resolution
verified exact
local_arxiv, observed 2026-05-15T02:03:28.920751Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-15T02:02:25.597640Z digest=sha256:2edaf08c2380ec1a90c2257322624c2bea0ce7912070b72ac793e6bd69fec0b4

Observation 3b80c03c-943c-4d49-a5b8-e2a91782a24b · inbound

Video-Zero: Self-Evolution Video Understanding cites this paper.

Video-Zero: Self-Evolution Video Understanding Self-Rewarding Language Models

Reference 1

Resolution
verified exact
local_arxiv, observed 2026-06-30T21:35:04.364598Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-30T21:32:16.939563Z digest=sha256:55f040284180310d3843d2be248399c33ddddac0032b3e58b4dd5e47ee67584a

Observation a4fd26ba-f14e-45a1-9ef4-da2ad38dc904 · inbound

Deep Pre-Alignment for VLMs cites this paper.

Deep Pre-Alignment for VLMs Self-Rewarding Language Models

Reference 91

Resolution
verified exact
local_arxiv, observed 2026-05-19T16:27:39.147627Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-19T16:26:41.094936Z digest=sha256:c70e8631009a580ec1ead369dd9b4f51c0c95b7f2b1831bb660d46719be10c65

Observation d4bf3bf9-45f9-4391-a1f3-3593c8db321e · inbound

DuIVRS-2: An LLM-based Interactive Voice Response System for Large-scale POI Attribute Acquisition cites this paper.

DuIVRS-2: An LLM-based Interactive Voice Response System for Large-scale POI Attribute Acquisition Self-Rewarding Language Models

Reference 1

Resolution
metadata mismatch
local_arxiv, observed 2026-05-20T10:33:12.946146Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-20T10:30:48.957568Z digest=sha256:9fc44c0c3e8c1cfb4dfc11b929eec604307516d4cffb2f50a63cdce82b5d5660

Observation 8d10e24a-c840-471a-8ef0-a4b4a6b7cef6 · inbound

SOLAR: A Self-Optimizing Open-Ended Autonomous Agent for Lifelong Learning and Continual Adaptation cites this paper.

SOLAR: A Self-Optimizing Open-Ended Autonomous Agent for Lifelong Learning and Continual Adaptation Self-Rewarding Language Models

Reference 22

Resolution
verified exact
local_arxiv, observed 2026-05-21T11:24:08.705016Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-21T11:21:30.867480Z digest=sha256:10784ddd8ade555bef5fa33fea0d9ee2d5b70e91c59559d1756398c02550b97b

Observation 66035c1a-aac5-4f27-bc4d-4f300070f43d · inbound

Reinforcing Human Behavior Simulation via Verbal Feedback cites this paper.

Reinforcing Human Behavior Simulation via Verbal Feedback Self-Rewarding Language Models

Reference 56

Resolution
verified exact
local_arxiv, observed 2026-05-21T07:24:02.403329Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-21T07:21:48.649289Z digest=sha256:66b31bff8a855bd766494f3bbe599c030fd74c642ca822e69700569295dd77d2

Observation 7d53a406-9368-45d7-9cbe-58e7d17e1046 · inbound

Weak Critics Make Strong Learners: On-Policy Critique Distillation for Scalable Oversight cites this paper.

Weak Critics Make Strong Learners: On-Policy Critique Distillation for Scalable Oversight Self-Rewarding Language Models

Reference 23

Resolution
metadata mismatch
local_arxiv, observed 2026-07-01T19:56:11.018058Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-28T21:55:58.645062Z digest=sha256:5d922a8f470d6f046c3b15e05d7b61603b9068a5e5f028f486595ab726177db2

Observation 4c04b105-014c-4767-84ed-167cb91426d6 · inbound

Ryze: Evidence-Enriched Data Synthesis from Biomedical Papers cites this paper.

Ryze: Evidence-Enriched Data Synthesis from Biomedical Papers Self-Rewarding Language Models

Reference 4

Resolution
metadata mismatch
local_arxiv, observed 2026-06-28T20:42:37.884408Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-06-28T18:21:22.421280Z digest=sha256:715110920b52c12aaa8ed654b0aaf65e32cc07c2a18911b66da2c72591414413

Observation 3f85ff3b-6de0-4164-9839-a75efce56570 · inbound

Deep Research as Rubric for Reinforcement Learning cites this paper.

Deep Research as Rubric for Reinforcement Learning Self-Rewarding Language Models

Reference 13

Resolution
verified exact
local_arxiv, observed 2026-06-28T17:22:25.254132Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-28T17:16:16.243967Z digest=sha256:c708f7869a66c1586186d6960bd7810228f33e36bc6d0cadbeeb32a5155db8df

Observation e3ae54f1-7685-4b9e-8614-c3c8a25126c1 · inbound

Trust Region On-Policy Distillation cites this paper.

Trust Region On-Policy Distillation Self-Rewarding Language Models

Reference 75

Resolution
metadata mismatch
local_arxiv, observed 2026-07-01T20:46:14.371457Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-06-28T17:38:50.313305Z digest=sha256:96bf696508be507dafc3b38a967faa9e3f4ea5bcc086cd96682286bfe1d89b1a

Observation a2729d11-3347-4b24-91bd-84380aa0c271 · inbound

Exploiting Verification-Generation Gap: Test-Time Reinforcement Learning with Confidence-Conditioned Verification cites this paper.

Exploiting Verification-Generation Gap: Test-Time Reinforcement Learning with Confidence-Conditioned Verification Self-Rewarding Language Models

Reference 13

Resolution
verified exact
local_arxiv, observed 2026-07-02T02:26:27.255554Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-28T10:53:00.223228Z digest=sha256:25d37528e68238be87634f3edbca81a3d745b47ccd976877aa774912e9f6ca76

Observation f325a507-3750-4a3c-ab35-355bd749c9d0 · inbound

Policy-Conditioned Counterfactual Credit for Verifiable Reinforcement Learning of Long-Horizon Language Agents cites this paper.

Policy-Conditioned Counterfactual Credit for Verifiable Reinforcement Learning of Long-Horizon Language Agents Self-Rewarding Language Models

Reference 60

Resolution
metadata mismatch
local_arxiv, observed 2026-07-02T06:06:41.596133Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-06-28T07:34:18.370135Z digest=sha256:63d6c03d49ca293d6d227f4a619b04fbc543efe8ca4f4aa48bef7f3ea7bfbc5f

Observation c2ef25eb-40bc-4a71-8fea-c8e652265e88 · inbound

Provenance-Grounded Gating and Adaptive Recovery in Synthetic Post-Training Data Curation cites this paper.

Provenance-Grounded Gating and Adaptive Recovery in Synthetic Post-Training Data Curation Self-Rewarding Language Models

Reference 22

Resolution
verified exact
local_arxiv, observed 2026-07-03T06:07:41.401853Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-27T12:55:11.402557Z digest=sha256:9080af9725f98ade749cfad9a3c0bbb0ed09c760803117eb0e8b765ed94f7677

Observation c8f387f7-9032-4d3a-9bc6-ead6f3f9782d · inbound

Ouroboros-Spatial: Closing the Data-Model Loop for Spatial Reasoning cites this paper.

Ouroboros-Spatial: Closing the Data-Model Loop for Spatial Reasoning Self-Rewarding Language Models

Reference 65

Resolution
verified exact
local_arxiv, observed 2026-07-03T09:17:48.632412Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-27T10:26:51.429436Z digest=sha256:4cc9512e3332abd3fb8cdebff40a51749e91120945499a9a3da84eb6c545e410

Observation c47cb9a6-e728-43a6-bc43-322a6d365ff7 · inbound

Ouroboros-Spatial: Closing the Data-Model Loop for Spatial Reasoning cites this paper.

Ouroboros-Spatial: Closing the Data-Model Loop for Spatial Reasoning Self-Rewarding Language Models

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-02T11:51:37.183739Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T11:51:37.183739Z digest=sha256:f0d853715f707edc84620f978dbe5d5e2013de3b4d0516b7dacefe465dfc6b87

Observation c2378224-a61f-43f3-af19-e24cc36c68e7 · inbound

Embodied-BenchClaw: An Autonomous Multi-Agent System for Embodied Spatial Intelligence Benchmark Construction cites this paper.

Embodied-BenchClaw: An Autonomous Multi-Agent System for Embodied Spatial Intelligence Benchmark Construction Self-Rewarding Language Models

Reference 14

Resolution
verified exact
local_arxiv, observed 2026-07-03T10:48:02.883263Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-27T09:48:49.786021Z digest=sha256:a8413aacf92b38863e3b29daf223a3c26c7201a6193803558c88e6650211fcc3

Observation 6957cad2-ce98-4d45-9395-7b1907aa4b4e · inbound

Self-Improvement Can Self-Regress: The Rise-and-Collapse Failure Mode of LLM Self-Training cites this paper.

Self-Improvement Can Self-Regress: The Rise-and-Collapse Failure Mode of LLM Self-Training Self-Rewarding Language Models

Reference 20

Resolution
verified exact
local_arxiv, observed 2026-07-04T00:29:16.725900Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-26T21:07:28.660119Z digest=sha256:d190824c77758acc86a7d5b6681f14ce63fbed7a145533f824f328b7353c9615

Observation 1dc97d87-059a-4ada-8bdc-4668aa4cba7f · inbound

Modularized Reinforcement Learning on LLMs: From MDP Creation to Exploration and Learning cites this paper.

Modularized Reinforcement Learning on LLMs: From MDP Creation to Exploration and Learning Self-Rewarding Language Models

Reference 252

Resolution
verified exact
local_arxiv, observed 2026-07-04T08:09:40.710745Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-26T12:15:08.304150Z digest=sha256:11a0af7918b9c67d9c5d921f9b4fc038637a04073945ec7660d2e5280fb9b092

Observation f4494c5f-6319-4502-acd6-7ff65dac4369 · inbound

Grounded Scaling: Why Agentic AI Needs Deterministic Environments cites this paper.

Grounded Scaling: Why Agentic AI Needs Deterministic Environments Self-Rewarding Language Models

Reference 11

Resolution
verified exact
local_arxiv, observed 2026-07-04T08:49:42.653791Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-26T10:51:34.720931Z digest=sha256:36f3cfe177bc8d6d144fc2ae6bef83c9d5668194c11de61f8d857f8b3e37cad3

Observation d9270c68-6889-43cc-9d2b-564542049922 · inbound

PRIDE: Privileged Information-enhanced Distillation for Empathetic Dialogue Generation cites this paper.

PRIDE: Privileged Information-enhanced Distillation for Empathetic Dialogue Generation Self-Rewarding Language Models

Reference 24

Resolution
verified exact
local_arxiv, observed 2026-07-04T10:39:45.645866Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-26T08:37:10.149519Z digest=sha256:8ab9f2d72f03e3b8faa2a29cec3bc73b78945d1fb9c35062f6c60fad5d7dd421

Observation e49e096e-fe76-41bf-b4d4-74b231a50bbf · inbound

Vision-driven Preference Synthesis for Mitigating Hallucinations in VLMs cites this paper.

Vision-driven Preference Synthesis for Mitigating Hallucinations in VLMs Self-Rewarding Language Models

Reference 72

Resolution
verified exact
local_arxiv, observed 2026-07-01T15:35:47.629107Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-30T01:22:16.176398Z digest=sha256:a7ac7715038aeb5149ddfa993f80e650fa1cfa0e61d6be96699f9521a34b5006

Observation 98e891cc-9b8b-4972-9a39-c755dc36ee82 · inbound

PASTA: A Paraphrasing And Self-Training Approach for Knowledge Updating in LLMs cites this paper.

PASTA: A Paraphrasing And Self-Training Approach for Knowledge Updating in LLMs Self-Rewarding Language Models

Reference 26

Resolution
verified exact
local_arxiv, observed 2026-06-30T09:44:37.488982Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-30T09:42:40.357949Z digest=sha256:2f967e3c90e7d1e9c9571f56ea6750b4abcd1c00648b26acdf1bd00935cf2866

Observation 3f49bf9b-6d01-49a2-aa20-a5051b328941 · inbound

Active-GRPO: Adaptive Imitation and Self-Improving Reasoning for Molecular Optimization cites this paper.

Active-GRPO: Adaptive Imitation and Self-Improving Reasoning for Molecular Optimization Self-Rewarding Language Models

Reference 56

Resolution
verified exact
local_arxiv, observed 2026-07-02T16:17:08.604101Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-07-02T16:07:49.362916Z digest=sha256:cbb0fabc204dd45b40c87b40bff9d0dcf06915eba33ef664b901e75efd302f82

Observation 733794e4-1aa9-4014-8757-65abe37e24ba · inbound

Multi-Turn On-Policy Distillation with Prefix Replay cites this paper.

Multi-Turn On-Policy Distillation with Prefix Replay Self-Rewarding Language Models

Reference 210

Resolution
unresolved
no resolver link, observed 2026-07-11T13:53:36.775836Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-11T13:53:36.775836Z digest=sha256:1d78a9e25abf0ff097a87cb0f14370f2b706db12014672791161d2688519f3b4

Observation d57da815-ead9-4632-b797-8533f778b58e · inbound

Multi-Turn On-Policy Distillation with Prefix Replay cites this paper.

Multi-Turn On-Policy Distillation with Prefix Replay Self-Rewarding Language Models

Reference 211

Resolution
unresolved
no resolver link, observed 2026-08-02T08:40:56.757215Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T08:40:56.757215Z digest=sha256:0eb475cbb97fb6e5283763ff2249d0d341ef5fb35d9eaaec89dc9a1261dcfc69

Observation 615534d4-4def-4d21-ad04-1e6b9a311a42 · inbound

Search Beyond What Can Be Taught: Evolving the Knowledge Boundary in Agentic Visual Generation cites this paper.

Search Beyond What Can Be Taught: Evolving the Knowledge Boundary in Agentic Visual Generation Self-Rewarding Language Models

Reference 53

Resolution
verified exact
local_arxiv, observed 2026-07-07T13:23:47.420272Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-07-07T13:18:50.793715Z digest=sha256:49600213f1b308861d26a72c378b618435935fc3a989b855d44d3efd55f41d56

Observation c51813a2-1fc2-4974-8364-f77c13fce7b9 · inbound

Search Beyond What Can Be Taught: Evolving the Knowledge Boundary in Agentic Visual Generation cites this paper.

Search Beyond What Can Be Taught: Evolving the Knowledge Boundary in Agentic Visual Generation Self-Rewarding Language Models

Reference 56

Resolution
unresolved
no resolver link, observed 2026-07-11T07:09:39.972776Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T07:09:39.972776Z digest=sha256:64618c4d1cd9aa685bae33720b41c0b7720100d81c077b4cc315264b3d380dc7

Observation 3ec5e70a-4643-450b-835c-7f868db92e59 · inbound

Search Beyond What Can Be Taught: Evolving the Knowledge Boundary in Agentic Visual Generation cites this paper.

Search Beyond What Can Be Taught: Evolving the Knowledge Boundary in Agentic Visual Generation Self-Rewarding Language Models

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-02T08:32:10.968750Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T08:32:10.968750Z digest=sha256:d89fc869159991ade7563b2cf696cd6264d1986ecccf1f129b892a5802635ede

Observation 37b4c542-cd58-4b90-b3f2-5b6767c91cf2 · inbound

Weak-to-Strong Generalization via Direct On-Policy Distillation cites this paper.

Weak-to-Strong Generalization via Direct On-Policy Distillation Self-Rewarding Language Models

Reference 97

Resolution
verified exact
local_arxiv, observed 2026-07-07T12:33:45.183403Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-07-07T12:31:42.224094Z digest=sha256:212337064a484ab96e271911a1aa798d2c951119568d3a7365c97581207d69b0

Observation 36f27840-916c-40e0-a746-c4ccd5e21f1b · inbound

Weak-to-Strong Generalization via Direct On-Policy Distillation cites this paper.

Weak-to-Strong Generalization via Direct On-Policy Distillation Self-Rewarding Language Models

Reference 93

Resolution
unresolved
no resolver link, observed 2026-07-11T07:01:56.628017Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T07:01:56.628017Z digest=sha256:2eb2b0ac3a8ab9971f03688d1452b082a49c96b398aef3c5f5392873cd3beace

Observation 3f3cf442-54df-48bf-9996-6e8ce059e79d · inbound

More Convincing, Not More Correct: Self-Play Reward Hacking of Reference-Free LLM Judges cites this paper.

More Convincing, Not More Correct: Self-Play Reward Hacking of Reference-Free LLM Judges Self-Rewarding Language Models

Reference 13

Resolution
metadata mismatch
local_arxiv, observed 2026-07-08T21:25:38.628686Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-07-08T21:20:07.569587Z digest=sha256:f563a6f84edaf29df17aaf07798a295fcfcca06cbd3db794f2fa181dc046f43c

Observation 47f64fd4-1aa6-4d0e-89f7-1d4fef9652ac · inbound

Recursive Self-Improvement in AI: From Bounded Self-Refinement to Autonomous Research Loops cites this paper.

Recursive Self-Improvement in AI: From Bounded Self-Refinement to Autonomous Research Loops Self-Rewarding Language Models

Reference 8

Resolution
verified exact
local_arxiv, observed 2026-07-09T03:45:55.566585Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-07-09T03:36:57.168246Z digest=sha256:8c0654193f2c984257f64f0d2803c62fc7e3f1738953df84168bf23f695731d0

Observation d69d0e54-baf0-44b8-8252-193322e536a8 · inbound

Long-Horizon-Terminal-Bench: Testing the Limits of Agents on Long-Horizon Terminal Tasks with Dense Reward-Based Grading cites this paper.

Long-Horizon-Terminal-Bench: Testing the Limits of Agents on Long-Horizon Terminal Tasks with Dense Reward-Based Grading Self-Rewarding Language Models

Reference 50

Resolution
unresolved
no resolver link, observed 2026-07-13T05:28:45.311405Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T05:28:45.311405Z digest=sha256:b0b621a2ae850ed6d3e464348d519891398495b6c463dad87cc214b207c6f08d

Observation 52d4ddb4-5b6b-4cca-a130-6451205cf9a5 · inbound

Long-Horizon-Terminal-Bench: Testing the Limits of Agents on Long-Horizon Terminal Tasks with Dense Reward-Based Grading cites this paper.

Long-Horizon-Terminal-Bench: Testing the Limits of Agents on Long-Horizon Terminal Tasks with Dense Reward-Based Grading Self-Rewarding Language Models

Reference 51

Resolution
unresolved
no resolver link, observed 2026-07-14T15:24:58.589243Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T15:24:58.589243Z digest=sha256:bb0d395e9816c9b0fb9ecfe90403e016cb075a72bdec8f7eaa0eeeb718df3d0b

Observation e7c1d6fa-76c3-40ae-a3d0-e44b3ee16705 · inbound

Agentic-DPO: From Imitation to Agentic Policy Optimization on Expert Trajectories cites this paper.

Agentic-DPO: From Imitation to Agentic Policy Optimization on Expert Trajectories Self-Rewarding Language Models

Reference 49

Resolution
unresolved
no resolver link, observed 2026-07-14T10:33:54.851493Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T10:33:54.851493Z digest=sha256:704091f3640834e1fd4355eaa84fcfcb53039ee277f8f0d52026057184012718

Observation be6e5875-72f5-48ee-b4e1-3f8dd21fcfb8 · inbound

DeepBias: Adaptive In-depth Probing of Social Biases in LVLMs cites this paper.

DeepBias: Adaptive In-depth Probing of Social Biases in LVLMs Self-Rewarding Language Models

Reference 44

Resolution
unresolved
no resolver link, observed 2026-07-14T06:09:46.933171Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T06:09:46.933171Z digest=sha256:6fe12286604f97bce0e12cc6e6effa07ec9dec64495c2fd0b231124fd3e42e74

Observation 026459f0-1812-4435-92d0-5f8ab495ed82 · inbound

RRPO: Reference-Relative Policy Optimization with Stratified Conditional Rollouts cites this paper.

RRPO: Reference-Relative Policy Optimization with Stratified Conditional Rollouts Self-Rewarding Language Models

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-01T15:21:20.171156Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T15:21:20.171156Z digest=sha256:64321cc6986b674d61df0bdd85b3a60b9dff5fab7695c640e4b33e6ad016f04e

Observation 08c10460-733c-472e-98c8-325aedeb5194 · inbound

FormulaSPIN: Self-Play Fine-Tuning for Natural Language to Spreadsheet Formula Generation cites this paper.

FormulaSPIN: Self-Play Fine-Tuning for Natural Language to Spreadsheet Formula Generation Self-Rewarding Language Models

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-02T13:32:14.589133Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T13:32:14.589133Z digest=sha256:2890c8e929c7d29bb6285ddc4cc8749b8cf44d7f174f036b7c0db3e30dc38f59

Observation 547954f7-9c66-4dd1-af5b-2907541648ac · inbound

Meta-Learned Reward Shaping for Reinforcement Learning from Human Feedback cites this paper.

Meta-Learned Reward Shaping for Reinforcement Learning from Human Feedback Self-Rewarding Language Models

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-01T03:14:13.190121Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T03:14:13.190121Z digest=sha256:65ddcb8cb4173160aacba2f1a786f5696c251bf7572d44799f77190c8a88e0a4

Observation 1b7c10b3-f874-4fe8-8bb0-b59dfb1efac2 · inbound

Judging Is Not Enumerating: Silent Omissions in LLM-Authored Acceptable Sets cites this paper.

Judging Is Not Enumerating: Silent Omissions in LLM-Authored Acceptable Sets Self-Rewarding Language Models

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-06T00:42:59.909123Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T00:42:59.909123Z digest=sha256:c35abb02649fe2b091c2e51c8177a30b3ea8fb3cc9fce73dc3e4e3d0c3be0b3e