Pith. sign in

Paper Citation Record · LEDGER

Mutual-Taught for Co-adapting Policy and Reward Models

As of 19 August 2026, this Paper Citation Record lists 40 of 40 outbound references and 0 inbound Pith citation observations for arXiv:2506.06292.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.06292 v2

Coverage vector

measured 40 of 40 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-15T20:53:39.945503Z

measured 40 of 40 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-19T06:32:44.657259+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

40 of 40 outbound references displayed

  • verified exact0
  • verified fuzzy1
  • unresolved39
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 0e259bdb-1853-4f76-84aa-ec5ff05cb212 · outbound

This paper cites an unresolved cited work.

Mutual-Taught for Co-adapting Policy and Reward Models Unresolved cited work

Reference 1

Resolution
unresolved
raw_fallback, observed 2026-08-15T20:53:41.252199Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-15T20:53:39.302069Z digest=sha256:72d6c611b6dadb4f778d3c7ea91af312bc66bce4d5172fbf886ae06410eeb0d6

Observation 59f42dc6-7d9c-4be1-adc3-97899f90b612 · outbound

This paper cites an unresolved cited work.

Mutual-Taught for Co-adapting Policy and Reward Models Unresolved cited work

Reference 2

Resolution
unresolved
raw_fallback, observed 2026-08-15T20:53:41.220394Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-15T20:53:39.329015Z digest=sha256:93dfa2b86e52734eb8b055c13092f6fc5b61bfc33648d5a6edf1c12031ea4d6b

Observation f3f9575d-4960-4351-974b-6d1dd7ddc3d3 · outbound

This paper cites an unresolved cited work.

Mutual-Taught for Co-adapting Policy and Reward Models Unresolved cited work

Reference 3

Resolution
unresolved
raw_fallback, observed 2026-08-15T20:53:41.161019Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-15T20:53:39.334587Z digest=sha256:8faa19662030d166fecc313ed305431933e2e7a81c7c80e7eb599c3144bd7008

Observation 978bf97d-65b1-438b-94a4-ab59511d346a · outbound

This paper cites an unresolved cited work.

Mutual-Taught for Co-adapting Policy and Reward Models Unresolved cited work

Reference 4

Resolution
unresolved
raw_fallback, observed 2026-08-15T20:53:41.148198Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-15T20:53:39.339235Z digest=sha256:30752f9bb5b078413fbcc4bfe93a29e8c2fb29f7aec19d9ccaeb62a2f04d6b69

Observation a1381f10-04b2-4af2-ae1a-e2b5632dedf6 · outbound

This paper cites an unresolved cited work.

Mutual-Taught for Co-adapting Policy and Reward Models Unresolved cited work

Reference 5

Resolution
unresolved
raw_fallback, observed 2026-08-15T20:53:41.132737Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-15T20:53:39.343889Z digest=sha256:34d0720a7a203c285a0ef540c6407b1d52e606c06f5e57980649bb6342ce9caf

Observation 8a67c46b-6cbb-4da8-906c-726a865b6ef7 · outbound

This paper cites Training Verifiers to Solve Math Word Problems.

Mutual-Taught for Co-adapting Policy and Reward Models Training Verifiers to Solve Math Word Problems

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-15T20:53:39.400888Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:53:39.400888Z digest=sha256:73baf9a6259deb01069866f5131e4e0b993f0df586aa42046e4b30361539b481

Observation 4a44d61a-c70d-4648-b021-b38faf6ebac7 · outbound

This paper cites an unresolved cited work.

Mutual-Taught for Co-adapting Policy and Reward Models Unresolved cited work

Reference 7

Resolution
unresolved
raw_fallback, observed 2026-08-15T20:53:41.101716Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-15T20:53:39.457730Z digest=sha256:6de15bf9923c1fc627518822839ecdf333944dd30a0aef19436e03e942677a7e

Observation 28b5ae49-fd87-46da-a06b-728da4083026 · outbound

This paper cites an unresolved cited work.

Mutual-Taught for Co-adapting Policy and Reward Models Unresolved cited work

Reference 8

Resolution
unresolved
raw_fallback, observed 2026-08-15T20:53:40.967376Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-15T20:53:39.463308Z digest=sha256:f10f365a6738ab4edc094a91dc57bf12ef7aadb01f50d871b993aa5204fd2b6d

Observation b71dc55c-33ca-4308-a853-269496bca3d3 · outbound

This paper cites The Llama 3 Herd of Models.

Mutual-Taught for Co-adapting Policy and Reward Models The Llama 3 Herd of Models

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-15T20:53:39.467713Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:53:39.467713Z digest=sha256:51486b76101f47820d921551d8e2643da90049d285a1c8d9f1cc5d0f392cec99

Observation d1e7c4d8-2c34-497f-b0bb-7a57312fd4e4 · outbound

This paper cites an unresolved cited work.

Mutual-Taught for Co-adapting Policy and Reward Models Unresolved cited work

Reference 10

Resolution
unresolved
raw_fallback, observed 2026-08-15T20:53:40.919735Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-15T20:53:39.510979Z digest=sha256:b4706a705a52b428bfd5d31d562d10e460c862375e83e41a69bb9b2c497bcc94

Observation b7c67607-e814-42c0-a8e0-af4837770113 · outbound

This paper cites an unresolved cited work.

Mutual-Taught for Co-adapting Policy and Reward Models Unresolved cited work

Reference 11

Resolution
unresolved
raw_fallback, observed 2026-08-15T20:53:40.905940Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-15T20:53:39.550477Z digest=sha256:82dfb1aa81d3c66c33061b48300d3a46ef7cf28b0f1264b5e5d0eecfc77a12fc

Observation d5103b4b-4463-4aac-b4bd-6b3cdf90c96a · outbound

This paper cites an unresolved cited work.

Mutual-Taught for Co-adapting Policy and Reward Models Unresolved cited work

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-15T20:53:39.556303Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:53:39.556303Z digest=sha256:39eb861b4a4110c77c2a20016e0b158052c0614e477dfe766d50a36f3798fbfd

Observation e33448f0-f078-42fc-90be-c542b2e24317 · outbound

This paper cites Unsupervised Prompt Learning for Vision-Language Models.

Mutual-Taught for Co-adapting Policy and Reward Models Unsupervised Prompt Learning for Vision-Language Models

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-15T20:53:39.560608Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:53:39.560608Z digest=sha256:1fab6eee40f96047af022c6e9c46fa398812c09b01d9d8063aa2fe7b5200d2c6

Observation d8b4b98e-2343-4dd3-9aa1-d6aa3620b9c0 · outbound

This paper cites Mistral 7B.

Mutual-Taught for Co-adapting Policy and Reward Models Mistral 7B

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-15T20:53:39.565940Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:53:39.565940Z digest=sha256:9043e0b944ec9f01847b178458144ac74f5f8afc9ed8a0a567d224bd2e5d5e66

Observation 8d8c6a49-c328-42d9-9f32-12d33108a241 · outbound

This paper cites o pf, Yannic Kilcher, Dimitri von R \.

Mutual-Taught for Co-adapting Policy and Reward Models o pf, Yannic Kilcher, Dimitri von R \

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:53:40.861856Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-15T20:53:39.586857Z digest=sha256:b12d578c0689bc2093ffed7842ee84a3825a57cf52aa481a178325c15b7b2acf

Observation fd313a3d-62f6-4ce5-91a2-395096329e06 · outbound

This paper cites RewardBench: Evaluating Reward Models for Language Modeling.

Mutual-Taught for Co-adapting Policy and Reward Models RewardBench: Evaluating Reward Models for Language Modeling

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-15T20:53:39.618730Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:53:39.618730Z digest=sha256:7830751cb9deb6179386c4fc4e6005c1d7bfe399a835848289e75c35258c3d4a

Observation 20ca02fc-c8ee-4bf9-8bcc-1864ceaa277b · outbound

This paper cites From Crowdsourced Data to High-Quality Benchmarks: Arena-Hard and BenchBuilder Pipeline.

Mutual-Taught for Co-adapting Policy and Reward Models From Crowdsourced Data to High-Quality Benchmarks: Arena-Hard and BenchBuilder Pipeline

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-15T20:53:39.623147Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:53:39.623147Z digest=sha256:ea0068bbb027260843bf1d935085655299b19f237f7a01386250dd854ea44116

Observation fe9c420a-51b0-48a4-98f5-93beab23e59e · outbound

This paper cites Hashimoto.

Mutual-Taught for Co-adapting Policy and Reward Models Hashimoto

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-15T20:53:39.628011Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:53:39.628011Z digest=sha256:7ef048bd88fc16675430c6b2b09eff1f0729d548d9a53991de476c52a5869206

Observation 46282fff-87fd-41ac-8404-771df9e80b75 · outbound

This paper cites an unresolved cited work.

Mutual-Taught for Co-adapting Policy and Reward Models Unresolved cited work

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-15T20:53:39.632014Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:53:39.632014Z digest=sha256:b10d4913db4263acf4f43c6e474107a643580f5c8238dc187078b272c3fea01d

Observation 277eeee6-42be-4398-a7f1-b5a2eb80f71b · outbound

This paper cites an unresolved cited work.

Mutual-Taught for Co-adapting Policy and Reward Models Unresolved cited work

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-15T20:53:39.670251Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:53:39.670251Z digest=sha256:26a387b22e790ac01649eea7334d7ad51346c74ff30ea597a55760a581f53acc

Observation d01e6eb2-b3b3-4777-bf7b-e3e8ad51cf25 · outbound

This paper cites an unresolved cited work.

Mutual-Taught for Co-adapting Policy and Reward Models Unresolved cited work

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-15T20:53:39.684813Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:53:39.684813Z digest=sha256:ab987feadfbdc3cb0a0562b124a771bdad06706064e7239c79842e59715fe429

Observation 50a0298e-512c-4696-87af-53bfcd238469 · outbound

This paper cites West-of-N: Synthetic Preferences for Self-Improving Reward Models.

Mutual-Taught for Co-adapting Policy and Reward Models West-of-N: Synthetic Preferences for Self-Improving Reward Models

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-15T20:53:39.688797Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:53:39.688797Z digest=sha256:079646965927794c03880212fea9e234f8de63d491889c01223ebac71b40d52f

Observation b4c6a41b-b39b-428f-b89d-50d428430fdd · outbound

This paper cites an unresolved cited work.

Mutual-Taught for Co-adapting Policy and Reward Models Unresolved cited work

Reference 23

Resolution
unresolved
raw_fallback, observed 2026-08-15T20:53:40.765235Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-15T20:53:39.693626Z digest=sha256:1c7be94bda4fb85ad8f1e2b05c1566326dceea0768b85fa240d6093bf4b2ed3b

Observation ca2ed841-9dfc-4063-acfe-bf54bb93437f · outbound

This paper cites Direct Nash Optimization: Teaching Language Models to Self-Improve with General Preferences.

Mutual-Taught for Co-adapting Policy and Reward Models Direct Nash Optimization: Teaching Language Models to Self-Improve with General Preferences

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-15T20:53:39.697397Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:53:39.697397Z digest=sha256:57dd85080dabe1b8097582c22ddc6e9de0c9a75d60852b74c84a196ddc475858

Observation da4f1290-f09a-4fd8-bc74-bc7fb255912a · outbound

This paper cites Proximal Policy Optimization Algorithms.

Mutual-Taught for Co-adapting Policy and Reward Models Proximal Policy Optimization Algorithms

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-15T20:53:39.745017Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:53:39.745017Z digest=sha256:a588a924486a206740882feca282c78ea300a4c3ff526fed587d4adae916217c

Observation 950088f6-7ef1-41a3-b947-585f9315510d · outbound

This paper cites an unresolved cited work.

Mutual-Taught for Co-adapting Policy and Reward Models Unresolved cited work

Reference 26

Resolution
unresolved
raw_fallback, observed 2026-08-15T20:53:40.750998Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-15T20:53:39.758841Z digest=sha256:1cfbc7a378211604fb0dbe6128a9186732bde51186525a25d1820fe7351f2336

Observation a06adbd8-8925-4aa7-8c30-88739ff09d19 · outbound

This paper cites Llama 2: Open Foundation and Fine-Tuned Chat Models.

Mutual-Taught for Co-adapting Policy and Reward Models Llama 2: Open Foundation and Fine-Tuned Chat Models

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-15T20:53:39.764028Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:53:39.764028Z digest=sha256:7c1cd6d11888adf4b2c26f5c55d7438ca42539ca937a464466ab8d0ee3a0325a

Observation 5becf1b0-5333-47e0-a261-560c85222f6e · outbound

This paper cites Meta-Rewarding Language Models: Self-Improving Alignment with LLM-as-a-Meta-Judge.

Mutual-Taught for Co-adapting Policy and Reward Models Meta-Rewarding Language Models: Self-Improving Alignment with LLM-as-a-Meta-Judge

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-15T20:53:39.769219Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:53:39.769219Z digest=sha256:f402ab8e027e05323e9e630d6c95d90978e5ed9f532ff8994435657fc88559b0

Observation 54dcd746-5496-46ad-907d-1bc06e1664b6 · outbound

This paper cites an unresolved cited work.

Mutual-Taught for Co-adapting Policy and Reward Models Unresolved cited work

Reference 29

Resolution
unresolved
raw_fallback, observed 2026-08-15T20:53:40.721762Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-15T20:53:39.773480Z digest=sha256:f56f1907adc370a584930635813c4ff19599b069938947eda24cfe59b543560d

Observation c370cb21-0787-42ff-85a4-5c6ca7a02d41 · outbound

This paper cites an unresolved cited work.

Mutual-Taught for Co-adapting Policy and Reward Models Unresolved cited work

Reference 30

Resolution
unresolved
raw_fallback, observed 2026-08-15T20:53:40.649631Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-15T20:53:39.804455Z digest=sha256:2e67deda1380c875281fbdc8070113262ae87ccb4899939ae8044d95880ae3d7

Observation 56ba907a-9370-4dd8-a591-d43550b8194e · outbound

This paper cites an unresolved cited work.

Mutual-Taught for Co-adapting Policy and Reward Models Unresolved cited work

Reference 31

Resolution
unresolved
raw_fallback, observed 2026-08-15T20:53:40.635594Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-15T20:53:39.821265Z digest=sha256:6da8c9f558ea616627c9f7418b9904987c23cdbd330f532c55a8b76228783085

Observation 65e4bd1c-ee79-473c-bd2b-5549e6116acf · outbound

This paper cites Some things are more CRINGE than others: Iterative Preference Optimization with the Pairwise Cringe Loss.

Mutual-Taught for Co-adapting Policy and Reward Models Some things are more CRINGE than others: Iterative Preference Optimization with the Pairwise Cringe Loss

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-15T20:53:39.825287Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:53:39.825287Z digest=sha256:da9669d541b0e9647349f4ed770a5ca514c472b3129b27b932ab3ce72438178a

Observation 570ee809-c4ba-4efc-ad84-de2f1b6863cd · outbound

This paper cites an unresolved cited work.

Mutual-Taught for Co-adapting Policy and Reward Models Unresolved cited work

Reference 33

Resolution
unresolved
raw_fallback, observed 2026-08-15T20:53:40.552162Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-15T20:53:39.830032Z digest=sha256:4e41cae12542bc8eae21ab81f3e2883366813535622983a733e6b1b2a56d5bd1

Observation 6c4a7e11-e3a4-46cf-8633-c45d1a077b12 · outbound

This paper cites an unresolved cited work.

Mutual-Taught for Co-adapting Policy and Reward Models Unresolved cited work

Reference 34

Resolution
unresolved
raw_fallback, observed 2026-08-15T20:53:40.479115Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-15T20:53:39.833947Z digest=sha256:66f7b1d7ce3cc785f3d3573a9e35f1cad5f21f66f5b72d00819ad7bea724a769

Observation fdaae8eb-a5b5-4938-ab43-48aa74c9bcf8 · outbound

This paper cites an unresolved cited work.

Mutual-Taught for Co-adapting Policy and Reward Models Unresolved cited work

Reference 35

Resolution
unresolved
raw_fallback, observed 2026-08-15T20:53:40.464775Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-15T20:53:39.869192Z digest=sha256:23c825ec3b09f5b5923dc6a1b087b69a9cb263fc5d162a23dcd3735422c89250

Observation 6247927f-d8b1-4abb-a09b-43bc9f775eb8 · outbound

This paper cites Self-Exploring Language Models: Active Preference Elicitation for Online Alignment.

Mutual-Taught for Co-adapting Policy and Reward Models Self-Exploring Language Models: Active Preference Elicitation for Online Alignment

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-15T20:53:39.902673Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:53:39.902673Z digest=sha256:12f2e4527891d97c607868e134aaf4ae93ea1c9037ff9210e16fdb2835bca9aa

Observation 1935e351-7c0b-4431-828e-425bf210485d · outbound

This paper cites an unresolved cited work.

Mutual-Taught for Co-adapting Policy and Reward Models Unresolved cited work

Reference 37

Resolution
unresolved
raw_fallback, observed 2026-08-15T20:53:40.307258Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-15T20:53:39.906790Z digest=sha256:84da2a4e3fbe85e67988ce771e7683c36d9c7df25fb625803b6b878fbf2303a2

Observation 20027d6c-b0f4-4dca-96a2-aa95bbc23a91 · outbound

This paper cites an unresolved cited work.

Mutual-Taught for Co-adapting Policy and Reward Models Unresolved cited work

Reference 38

Resolution
unresolved
raw_fallback, observed 2026-08-15T20:53:40.227908Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-15T20:53:39.910135Z digest=sha256:f19b0db19835b945eceb14e7be6e6d9144d8dd17efebe59530e79c4844b9f388

Observation 05910a02-702f-4b74-91d3-591b3c2acc1d · outbound

This paper cites online" 'onlinestring :=.

Mutual-Taught for Co-adapting Policy and Reward Models online" 'onlinestring :=

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-15T20:53:39.914465Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:53:39.914465Z digest=sha256:241f4f45c4de6a08bca7b05653ff36cdae01cc155d90c792e336fd523c028075

Observation 96265a0b-a2e8-4105-8ca1-1e62070ec375 · outbound

This paper cites write newline.

Mutual-Taught for Co-adapting Policy and Reward Models write newline

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-15T20:53:39.945503Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:53:39.945503Z digest=sha256:b5458a4259ff2fd67c5ecd6080d8fe4ff3b47f96dd193f5c0c8def19bbea0328

Pith citing papers

No inbound Pith citation observations are available.