Pith. sign in

Paper Citation Record · LEDGER

Reward Models Are Secretly Value Functions: Temporally Coherent Reward Modeling

As of 8 August 2026, this Paper Citation Record lists 18 of 18 outbound references and 1 inbound Pith citation observation for arXiv:2604.22981.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2604.22981 v1

Coverage vector

measured 18 of 18 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-05-08T12:05:00.013591Z

measured 19 of 19 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T14:30:31.955883Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-07T18:16:16.850658Z

Reference resolution

18 of 18 outbound references displayed

  • verified exact0
  • verified fuzzy13
  • unresolved4
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation be5091b0-a62b-49bc-ac7d-599812a5804e · outbound

This paper cites If r(x, y0..k) = E[r(x, y)|x, y0..k], then the average value ofr(x, y0..k) −r (x, y)should be close to 0.

Reward Models Are Secretly Value Functions: Temporally Coherent Reward Modeling If r(x, y0..k) = E[r(x, y)|x, y0..k], then the average value ofr(x, y0..k) −r (x, y)should be close to 0

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T13:42:24.557246Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-08T12:05:00.013591Z digest=sha256:9d315fee938edcd7d8a72590f80202f8d7260389aedbd0b7b5d6d7a21ed3271d

Observation 5e4a2f64-1e80-49fd-901a-8fdb925d2287 · outbound

This paper cites Sincey0..k is more informative thany0..k−1, we should expect to see the mean squared prediction error(r(x, y0..k)−r(x, y)) 2 decrease askincreases.

Reward Models Are Secretly Value Functions: Temporally Coherent Reward Modeling Sincey0..k is more informative thany0..k−1, we should expect to see the mean squared prediction error(r(x, y0..k)−r(x, y)) 2 decrease askincreases

Reference 2

Resolution
malformed identifier
raw_fallback, observed 2026-05-26T13:42:24.593496Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-08T12:05:00.013591Z digest=sha256:4b8939f9f317e67c061f5368606cdbd53d14a521954913cdd9f868a1a9a1f6e6

Observation a9bce3cd-d509-47eb-93f0-8cdd50bd9241 · outbound

This paper cites What is the capital of Italy?.

Reward Models Are Secretly Value Functions: Temporally Coherent Reward Modeling What is the capital of Italy?

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T13:42:24.572252Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-08T12:05:00.013591Z digest=sha256:2e3de66496fbf2d6be06cc87e2ab3f0f523663f199d22911b5779ffde1f978e1

Observation cd944571-b660-4f56-9e35-3781291e1025 · outbound

This paper cites It is Rome.

Reward Models Are Secretly Value Functions: Temporally Coherent Reward Modeling It is Rome

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T13:42:24.554097Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-08T12:05:00.013591Z digest=sha256:d2412372b8dd86b44829d032f52a8ceac291c488fa796f2d1ed2897991447cc2

Observation a7ecf277-a44b-4135-a292-35a059b43718 · outbound

This paper cites Rome, also known as The Eternal City.

Reward Models Are Secretly Value Functions: Temporally Coherent Reward Modeling Rome, also known as The Eternal City

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T13:42:24.551205Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-08T12:05:00.013591Z digest=sha256:fc2ddff42e16300d82b8613e80488e4506ac8262f400f5ae34e1ede4c885f04d

Observation 7a241ce9-a992-4319-9a90-e3e02e632abd · outbound

This paper cites Rome, GA.

Reward Models Are Secretly Value Functions: Temporally Coherent Reward Modeling Rome, GA

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T13:42:24.563365Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-08T12:05:00.013591Z digest=sha256:efecc0583e9b5f00a022b785f5c9b4785e065cb20c291ac0bd00ec7c7ffbfe87

Observation 5aaf6c01-d19c-41fd-8727-2f390cf29941 · outbound

This paper cites Rome, also known as The City of Love.

Reward Models Are Secretly Value Functions: Temporally Coherent Reward Modeling Rome, also known as The City of Love

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T13:42:24.596565Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-08T12:05:00.013591Z digest=sha256:0f77b0b27b7381a8ccd43ffe8b2718a45f6bc32344732d36d639e1f7a6290d2f

Observation 67466b84-f137-4545-a1ba-ee6e10a9c17f · outbound

This paper cites \n\n" separator) minus at the end of the previous step was 18 2 1 0 1 2 3 Reward Model Output.

Reward Models Are Secretly Value Functions: Temporally Coherent Reward Modeling \n\n" separator) minus at the end of the previous step was 18 2 1 0 1 2 3 Reward Model Output

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T13:42:24.587513Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-08T12:05:00.013591Z digest=sha256:42be1f89e782a4345823bf3cfb288f983b5601dcf0e9e1879a94c600ffe09bce

Observation 8046382c-4f46-4d8c-95f2-8400448eebb5 · outbound

This paper cites Similar to the previous method, but sigmoid transformation was applied to the reward scores before subtraction.

Reward Models Are Secretly Value Functions: Temporally Coherent Reward Modeling Similar to the previous method, but sigmoid transformation was applied to the reward scores before subtraction

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T13:42:24.566790Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-08T12:05:00.013591Z digest=sha256:9fd63e9a295251e8d02733e1c5f2ecfd5810cfb3320b986b0e184a603d3e7991

Observation 7c574f0e-b9e3-45fa-8bdb-9036d509dcef · outbound

This paper cites Normalized Difference.

Reward Models Are Secretly Value Functions: Temporally Coherent Reward Modeling Normalized Difference

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T13:42:24.581651Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-08T12:05:00.013591Z digest=sha256:574cd822096c3c8ac382f781d3e7af475272e4e02d5f4d7ab21278a2080af6bb

Observation 623d9013-7318-490e-8bff-f6d0e0811832 · outbound

This paper cites This was used to evaluate the checkpoints every 10 training steps to understand training progress.

Reward Models Are Secretly Value Functions: Temporally Coherent Reward Modeling This was used to evaluate the checkpoints every 10 training steps to understand training progress

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T13:42:24.544749Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-08T12:05:00.013591Z digest=sha256:7b21d8408d696cf0d403eb0a3b51b058edcdfc6d5b034e279f399fdb37522cb3

Observation d91248a3-0142-4d9c-874f-adc6005bf530 · outbound

This paper cites an unresolved cited work.

Reward Models Are Secretly Value Functions: Temporally Coherent Reward Modeling Unresolved cited work

Reference 12

Resolution
unresolved
raw_fallback, observed 2026-05-26T13:42:24.590360Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-08T12:05:00.013591Z digest=sha256:616087aa13d25894cbb83ae25b40ab77b8d19fc138148fd83a61ff5d9fbb74c4

Observation 395e19b0-b562-4292-9e50-9b4647b677f3 · outbound

This paper cites The prompts for evaluation are also sourced fromDolci-Instruct-RL, but we only used the 10%validation subset with no overlap with the PPO training subset.

Reward Models Are Secretly Value Functions: Temporally Coherent Reward Modeling The prompts for evaluation are also sourced fromDolci-Instruct-RL, but we only used the 10%validation subset with no overlap with the PPO training subset

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T13:42:24.560193Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-08T12:05:00.013591Z digest=sha256:edba01d053fdf7b3d24d8215c060487028425efca5b11021d923aa58e7392214

Observation 2d8b95e0-d29e-4f26-b95f-e4fd9a0ddac9 · outbound

This paper cites warm start.

Reward Models Are Secretly Value Functions: Temporally Coherent Reward Modeling warm start

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T13:42:24.548230Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-08T12:05:00.013591Z digest=sha256:a67da45c152428a0bb614e97a2c1ab584442045954d0d217b458838b99ac4cf7

Observation 69b1cbeb-0e57-4997-9d1d-c9a0b6609106 · outbound

This paper cites an unresolved cited work.

Reward Models Are Secretly Value Functions: Temporally Coherent Reward Modeling Unresolved cited work

Reference 15

Resolution
unresolved
raw_fallback, observed 2026-05-26T13:42:24.584526Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-08T12:05:00.013591Z digest=sha256:831ae8945a65a43f657153f339ce1a13e77629c043f6e73d02c5cf822d3e74f6

Observation ad260881-b957-49cf-80ae-47b819cdeab2 · outbound

This paper cites an unresolved cited work.

Reward Models Are Secretly Value Functions: Temporally Coherent Reward Modeling Unresolved cited work

Reference 16

Resolution
unresolved
raw_fallback, observed 2026-05-26T13:42:24.578416Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-08T12:05:00.013591Z digest=sha256:ed370a38fee7bd275dbaf4c8c2f3c38ba9aa05f0258287dc490d01bb5def0c1a

Observation c56e4abe-6ae8-45ee-a89f-9ec30623b440 · outbound

This paper cites an unresolved cited work.

Reward Models Are Secretly Value Functions: Temporally Coherent Reward Modeling Unresolved cited work

Reference 17

Resolution
unresolved
raw_fallback, observed 2026-05-26T13:42:24.575416Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-08T12:05:00.013591Z digest=sha256:c6223a07e18bb6d50ffb1bf98e0bf34156abbf5e7535a958c24036dd4ea66dc3

Observation 807ae89b-8af9-40a1-82cb-664db535cdaf · outbound

This paper cites A" (Response A is better).

Reward Models Are Secretly Value Functions: Temporally Coherent Reward Modeling A" (Response A is better)

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T13:42:24.569442Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-08T12:05:00.013591Z digest=sha256:5f2db6a717622c8e98f804edf83a0bd6c952adf4cc07e611fbd975f96bde114f

Pith citing papers

Observation 164682d0-2cef-4822-b803-1943d511c074 · inbound

Sample-Adaptive Latent Rewards for Uncertainty-Guided Diffusion Post-Training cites this paper.

Sample-Adaptive Latent Rewards for Uncertainty-Guided Diffusion Post-Training Reward Models Are Secretly Value Functions: Temporally Coherent Reward Modeling

Reference 22

Resolution
verified exact
local_arxiv, observed 2026-08-07T14:30:35.757355Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T14:30:31.955883Z digest=sha256:f353601462419dc2357db9c995bee8b3b77103e392495ee3f84824d716cf2ac9