Pith. sign in

Paper Citation Record · LEDGER

Learning Latent Reasoning Traces for Scalar Reward Models End-to-End

As of 21 August 2026, this Paper Citation Record lists 47 of 47 outbound references and 0 inbound Pith citation observations for arXiv:2607.29185.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2607.29185 v1

Coverage vector

measured 47 of 47 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-03T12:09:09.970266Z

measured 47 of 47 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-20T06:33:59.587034+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

47 of 47 outbound references displayed

  • verified exact2
  • verified fuzzy0
  • unresolved44
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 0dc2264c-ba80-4fbc-a6a6-4e29844905e9 · outbound

This paper cites Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback.

Learning Latent Reasoning Traces for Scalar Reward Models End-to-End Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-03T12:09:03.872221Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T12:09:03.872221Z digest=sha256:09ee413a5ff73e94e4cb3a1d724c392e91c78ca0cfd3112efc809dfe3353f91e

Observation 41d2d2aa-425d-4996-86ab-e048a4a8edbe · outbound

This paper cites Fine-Tuning Language Models from Human Preferences.

Learning Latent Reasoning Traces for Scalar Reward Models End-to-End Fine-Tuning Language Models from Human Preferences

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-03T12:09:03.997789Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T12:09:03.997789Z digest=sha256:febfbc2a9268c59e1c2315cdad0c995e357dfd72c6f9686d839cd4015321466d

Observation 431449d6-8799-4519-83f6-9f9298c0e0a0 · outbound

This paper cites Fine-Grained Human Feedback Gives Better Rewards for Language Model Training , url =.

Learning Latent Reasoning Traces for Scalar Reward Models End-to-End Fine-Grained Human Feedback Gives Better Rewards for Language Model Training , url =

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-03T12:09:04.149160Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T12:09:04.149160Z digest=sha256:b190dd94e880b7613dbef547d02413253189d6e28db99d0b0b680d8d8698c67f

Observation 98425e2c-411b-4603-ad37-b60a5b6deac3 · outbound

This paper cites Proceedings of the 40th International Conference on Machine Learning , pages =.

Learning Latent Reasoning Traces for Scalar Reward Models End-to-End Proceedings of the 40th International Conference on Machine Learning , pages =

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-03T12:09:04.218493Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T12:09:04.218493Z digest=sha256:8946d90ec988a7b04253c632c4dde48e483a8de3833c5a76556d589a21434c87

Observation 253f7795-9355-42b0-a36b-e9ecf7a3f2e7 · outbound

This paper cites Think-RM: Enabling Long-Horizon Reasoning in Generative Reward Models , url =.

Learning Latent Reasoning Traces for Scalar Reward Models End-to-End Think-RM: Enabling Long-Horizon Reasoning in Generative Reward Models , url =

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-03T12:09:04.366566Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T12:09:04.366566Z digest=sha256:cc7176d2e953a3b5d6bf9e688d4d72bf3bca50965cce68743a8b27a93a625dfb

Observation 999280f8-9d1a-41e4-b281-f26268dbd9c0 · outbound

This paper cites 2026 , url=.

Learning Latent Reasoning Traces for Scalar Reward Models End-to-End 2026 , url=

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-03T12:09:04.482182Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T12:09:04.482182Z digest=sha256:44dc1f95ec5e1f1a738b60bc437453423561710d9149b18f6264f209c9978c3b

Observation 5c0d7c3c-cbfa-4876-8a0a-40a36dd1f089 · outbound

This paper cites J1: Incentivizing Thinking in.

Learning Latent Reasoning Traces for Scalar Reward Models End-to-End J1: Incentivizing Thinking in

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-03T12:09:04.597246Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T12:09:04.597246Z digest=sha256:a99998ada0d8d1bb5768f0ed1b093c9038784d956a8efa140b4de798db4cc6b7

Observation fc97e75a-547a-4643-a6e4-d3e1d554138b · outbound

This paper cites Learning Structured Output Representation using Deep Conditional Generative Models , url =.

Learning Latent Reasoning Traces for Scalar Reward Models End-to-End Learning Structured Output Representation using Deep Conditional Generative Models , url =

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-03T12:09:04.676252Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T12:09:04.676252Z digest=sha256:bce12abea15798ba1c17de58b053ee56f7e7f8524d191b3cf77f241260b2eea0

Observation bbfaef32-3e3b-49d4-95dc-0a9b9cb743b8 · outbound

This paper cites Training language models to follow instructions with human feedback , url =.

Learning Latent Reasoning Traces for Scalar Reward Models End-to-End Training language models to follow instructions with human feedback , url =

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-03T12:09:04.762707Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T12:09:04.762707Z digest=sha256:ba6af3ea896d873b6eb59052c979cc692b363063a9b32ce82028b5ae6b34b54c

Observation bb2aeee2-e7ee-4f98-b88c-ac1539f2f3e4 · outbound

This paper cites InfoRM: Mitigating Reward Hacking in RLHF via Information-Theoretic Reward Modeling , url =.

Learning Latent Reasoning Traces for Scalar Reward Models End-to-End InfoRM: Mitigating Reward Hacking in RLHF via Information-Theoretic Reward Modeling , url =

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-03T12:09:04.944216Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T12:09:04.944216Z digest=sha256:052c9ff09942774cf1451d48f8f32b1d1e2515e20e98301b4d093538bce3a8c7

Observation 2c49e1ae-d586-40e7-958e-1338a92d86f5 · outbound

This paper cites Improving Reward Models with Synthetic Critiques.

Learning Latent Reasoning Traces for Scalar Reward Models End-to-End Improving Reward Models with Synthetic Critiques

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-03T12:09:05.175902Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T12:09:05.175902Z digest=sha256:8b121313f2ce20a07e1da0e76988116f9d79259c4dac950063329c1d07ce2efc

Observation 80d06d8e-4780-42aa-895d-3a4d513f0c70 · outbound

This paper cites Self-Generated Critiques Boost Reward Modeling for Language Models.

Learning Latent Reasoning Traces for Scalar Reward Models End-to-End Self-Generated Critiques Boost Reward Modeling for Language Models

Reference 12

Resolution
malformed identifier
no resolver link, observed 2026-08-03T12:09:05.271074Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T12:09:05.271074Z digest=sha256:01bad587ae8b29d37a39d81389a0bff5a8c88bbde0709065c4f6cb62b2de686d

Observation 485eac96-9550-4488-868a-acade5e44020 · outbound

This paper cites Critique-out-Loud Reward Models.

Learning Latent Reasoning Traces for Scalar Reward Models End-to-End Critique-out-Loud Reward Models

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-03T12:09:05.397076Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T12:09:05.397076Z digest=sha256:5dc28a760daf8843f9220377f1308299489151f6077bce4bd615f568e7bc18b8

Observation 98922cab-3159-4e19-8590-cf66749076a7 · outbound

This paper cites Generative Judge for Evaluating Alignment , url =.

Learning Latent Reasoning Traces for Scalar Reward Models End-to-End Generative Judge for Evaluating Alignment , url =

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-03T12:09:05.493941Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T12:09:05.493941Z digest=sha256:0df903b68b513c0e6073a05bb0d84369d73bfc8c1f879fbfe9c1bcbfcfaa41af

Observation 4f9561c5-399d-492f-be3c-31982f207d63 · outbound

This paper cites Deep Reinforcement Learning from Human Preferences , url =.

Learning Latent Reasoning Traces for Scalar Reward Models End-to-End Deep Reinforcement Learning from Human Preferences , url =

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-03T12:09:05.653095Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T12:09:05.653095Z digest=sha256:d13d279ef04032669089033c71bf64aa0bf5fc6ca5369dcc9cd700030833ea37

Observation c9c45d56-12f8-4fa6-9f17-2e48de8854ed · outbound

This paper cites The Twelfth International Conference on Learning Representations , year=.

Learning Latent Reasoning Traces for Scalar Reward Models End-to-End The Twelfth International Conference on Learning Representations , year=

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-03T12:09:05.784919Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T12:09:05.784919Z digest=sha256:dd3a6f850cf4eeff7492b78b6dd396e5745075fe9f5ab22547c39194c69e3a26

Observation c87ef72f-a19a-4448-b63a-45012dbb70cc · outbound

This paper cites Terry , journal =.

Learning Latent Reasoning Traces for Scalar Reward Models End-to-End Terry , journal =

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-03T12:09:05.923413Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T12:09:05.923413Z digest=sha256:ff2128514c41b5319e2d8506549450910b5db2bba5173c17595f1fdee25d3fbf

Observation 99d8a67f-44b0-43b2-941a-ae5ef0deb5b8 · outbound

This paper cites 1959 , publisher=.

Learning Latent Reasoning Traces for Scalar Reward Models End-to-End 1959 , publisher=

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-03T12:09:06.009105Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T12:09:06.009105Z digest=sha256:28ea1a59c81a8d98ab52f158d224daa57aaad559823a76df95199071192c4bbf

Observation b7c64419-c3b0-4a5b-a4b3-48c316a6f2af · outbound

This paper cites Rationalizing Neural Predictions.

Learning Latent Reasoning Traces for Scalar Reward Models End-to-End Rationalizing Neural Predictions

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-03T12:09:06.107054Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T12:09:06.107054Z digest=sha256:d73f7d059c912ee969867e78e2362821faa92a11ef26e3379989d2411ccd1f03

Observation 73f6cb0e-d506-4901-a314-8e71d4b2359c · outbound

This paper cites STaR: Bootstrapping Reasoning With Reasoning , url =.

Learning Latent Reasoning Traces for Scalar Reward Models End-to-End STaR: Bootstrapping Reasoning With Reasoning , url =

Reference 20

Resolution
verified exact
doi, observed 2026-08-03T12:13:46.150492Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-03T12:09:06.296072Z digest=sha256:189956e7599eabf4030d687038ec6662b05061e7e82f8404d231db3824ae231b

Observation be55845e-65b9-49d6-be88-4b0e8f15c7b4 · outbound

This paper cites Saurous, Rif , booktitle =.

Learning Latent Reasoning Traces for Scalar Reward Models End-to-End Saurous, Rif , booktitle =

Reference 21

Resolution
verified exact
doi, observed 2026-08-03T12:13:45.934005Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-03T12:09:06.404409Z digest=sha256:0e540e0523e8785cfb8f6e24c2214ac16222766ef4ec979a1dc785f3c9879d10

Observation 58a2f00e-2cec-4848-a7b3-052e9fb2700e · outbound

This paper cites Beyond Verifiable Rewards: Scaling Reinforcement Learning in Language Models to Unverifiable Data , url =.

Learning Latent Reasoning Traces for Scalar Reward Models End-to-End Beyond Verifiable Rewards: Scaling Reinforcement Learning in Language Models to Unverifiable Data , url =

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-03T12:09:06.520255Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T12:09:06.520255Z digest=sha256:7f19cc63bffb07fac12765a63ffd0250cf333935476d349d149a6a9496cbcacd

Observation 8c4cbae2-b25b-4815-ac02-52a2e1872a01 · outbound

This paper cites 2026 , eprint=.

Learning Latent Reasoning Traces for Scalar Reward Models End-to-End 2026 , eprint=

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-03T12:09:06.747386Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T12:09:06.747386Z digest=sha256:5fbda229d13c0e287547c2ed86b4ad662cf84fa23f662f3800acc6e3f5cb37f2

Observation 886a19f1-4caa-44b8-b525-559dbd02ae8e · outbound

This paper cites Policy Gradient Methods for Reinforcement Learning with Function Approximation , url =.

Learning Latent Reasoning Traces for Scalar Reward Models End-to-End Policy Gradient Methods for Reinforcement Learning with Function Approximation , url =

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-03T12:09:06.931488Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T12:09:06.931488Z digest=sha256:9a1b3c417e9627a062445e408c023147c6c4f6da8203f88ceb3aa27a3022a9a7

Observation a255d2be-95fb-4824-8bac-d838c86dd918 · outbound

This paper cites 2024 , editor =.

Learning Latent Reasoning Traces for Scalar Reward Models End-to-End 2024 , editor =

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-03T12:09:07.125029Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T12:09:07.125029Z digest=sha256:36bf5a14b72dea865b501946c3aa2e6b23a62091efd6be1f9fb41ba959b613d2

Observation 7d56f8fe-1f77-469b-b8b3-0e315a405ce6 · outbound

This paper cites AIMO-2 Winning Solution: Building State-of-the-Art Mathematical Reasoning Models with OpenMathReasoning dataset.

Learning Latent Reasoning Traces for Scalar Reward Models End-to-End AIMO-2 Winning Solution: Building State-of-the-Art Mathematical Reasoning Models with OpenMathReasoning dataset

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-03T12:09:07.287104Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T12:09:07.287104Z digest=sha256:acb90a73ec3495df502582392044111f816483256413ab399a678756e9fb1b4c

Observation 7c3a15a3-b696-4f7f-bf85-4372992b453a · outbound

This paper cites The Thirty-ninth Annual Conference on Neural Information Processing Systems Datasets and Benchmarks Track , year=.

Learning Latent Reasoning Traces for Scalar Reward Models End-to-End The Thirty-ninth Annual Conference on Neural Information Processing Systems Datasets and Benchmarks Track , year=

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-03T12:09:07.449364Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T12:09:07.449364Z digest=sha256:2a284002b5678a3e895a79bc5f8af5fbf1afaf8287af8287811943e6a5ec1a95

Observation 174e1aaa-f415-415a-bc21-aaa61d76f3f3 · outbound

This paper cites WildGuard: Open One-stop Moderation Tools for Safety Risks, Jailbreaks, and Refusals of LLMs , url =.

Learning Latent Reasoning Traces for Scalar Reward Models End-to-End WildGuard: Open One-stop Moderation Tools for Safety Risks, Jailbreaks, and Refusals of LLMs , url =

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-03T12:09:07.583260Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T12:09:07.583260Z digest=sha256:68e9e93ea9d1d1591ff4ae191e11e56813ce22ab1e2b63116a294ea7af3eb185

Observation b64833c2-6236-48f6-aca4-01d4997b3798 · outbound

This paper cites O ffset B ias: Leveraging Debiased Data for Tuning Evaluators.

Learning Latent Reasoning Traces for Scalar Reward Models End-to-End O ffset B ias: Leveraging Debiased Data for Tuning Evaluators

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-03T12:09:07.793737Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T12:09:07.793737Z digest=sha256:401940e9eaa2e2ab3149f659443de2d19226bfbc3400f413f49ec765b8bd39e1

Observation 293f4311-98c8-4572-b693-47dc9ffd176a · outbound

This paper cites Forty-second International Conference on Machine Learning , year=.

Learning Latent Reasoning Traces for Scalar Reward Models End-to-End Forty-second International Conference on Machine Learning , year=

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-03T12:09:07.896547Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T12:09:07.896547Z digest=sha256:faafff113d4273f67ee80da87dbd2d5d3d1f169a512e2f63a643074bf045a526

Observation 90969518-4203-4f4d-984b-a76806aaa611 · outbound

This paper cites and Lee, Jason.

Learning Latent Reasoning Traces for Scalar Reward Models End-to-End and Lee, Jason

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-03T12:09:07.999068Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T12:09:07.999068Z digest=sha256:f6142f1dfcb100002575a23697e76d43e9330690a79d0aed2e2eeafd3971215b

Observation 4302bb35-3fe1-4775-a68d-d3db24b1de73 · outbound

This paper cites 2025 , eprint=.

Learning Latent Reasoning Traces for Scalar Reward Models End-to-End 2025 , eprint=

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-03T12:09:08.110251Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T12:09:08.110251Z digest=sha256:f192d183dbe0f6bd4968605b10eee02577e3fa84bfafffca8551a8373969986b

Observation c18b9ce1-29b8-457f-af93-eddc64eb3203 · outbound

This paper cites and Hajishirzi, Hannaneh.

Learning Latent Reasoning Traces for Scalar Reward Models End-to-End and Hajishirzi, Hannaneh

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-03T12:09:08.182976Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T12:09:08.182976Z digest=sha256:514fb64dd2718e7cf250faa8bfb9e8070b9e41398b6a2a277e6153ec2ccfff64

Observation b667365d-1ffb-4911-b7b3-517dfc6962aa · outbound

This paper cites The Fourteenth International Conference on Learning Representations , year=.

Learning Latent Reasoning Traces for Scalar Reward Models End-to-End The Fourteenth International Conference on Learning Representations , year=

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-03T12:09:08.277734Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T12:09:08.277734Z digest=sha256:874a2ec1e5a2573bbdfa723580e4550d9069db23e7aef3af0ec3b5c1b0c1fd36

Observation fed329e1-8fd5-4f2f-ac92-9af408c9a6ff · outbound

This paper cites 2025 , url=.

Learning Latent Reasoning Traces for Scalar Reward Models End-to-End 2025 , url=

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-03T12:09:08.375921Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T12:09:08.375921Z digest=sha256:61db5b6a2082e2e0cb9f59367aae731b82790cfae5c4a926e7c4793900af6f46

Observation 79867b3e-eb47-433e-b966-e7bbc593a997 · outbound

This paper cites Gonzalez and Ion Stoica , booktitle=.

Learning Latent Reasoning Traces for Scalar Reward Models End-to-End Gonzalez and Ion Stoica , booktitle=

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-03T12:09:08.484852Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T12:09:08.484852Z digest=sha256:e1f47b3ae823bd02e72a1bbc3a4398397ae4a3236d7afa32355cf5d87ef1aef9

Observation 8b8fa87e-aa81-4eac-a9fe-eaf6ac4e9f53 · outbound

This paper cites JudgeBench: A Benchmark for Evaluating.

Learning Latent Reasoning Traces for Scalar Reward Models End-to-End JudgeBench: A Benchmark for Evaluating

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-03T12:09:08.643164Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T12:09:08.643164Z digest=sha256:b84126dfea2f0ae06b2e639a5b1c95dfa5f0c0cfad979faeafa6c7c9121a193f

Observation 218e1456-c11e-4cdd-8322-a84a0339dcef · outbound

This paper cites Proceedings of the Twentieth European Conference on Computer Systems , pages=.

Learning Latent Reasoning Traces for Scalar Reward Models End-to-End Proceedings of the Twentieth European Conference on Computer Systems , pages=

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-03T12:09:08.759591Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T12:09:08.759591Z digest=sha256:6f075e6080f2bde7879a3971aa4ca62b54313f3a508655a1340f51d1aa10f3ea

Observation fc2aeba1-47cd-4469-acff-8e1c2b21b2e4 · outbound

This paper cites Proceedings of the ACM SIGOPS 29th Symposium on Operating Systems Principles , year=.

Learning Latent Reasoning Traces for Scalar Reward Models End-to-End Proceedings of the ACM SIGOPS 29th Symposium on Operating Systems Principles , year=

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-03T12:09:08.853125Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T12:09:08.853125Z digest=sha256:52af5e8fd2a49c2891e30700a0fb773d616d8299f17d99aa43172d285f806928

Observation 0076fc1e-1ad7-4427-aa0b-1a04f24cd55c · outbound

This paper cites Proximal Policy Optimization Algorithms.

Learning Latent Reasoning Traces for Scalar Reward Models End-to-End Proximal Policy Optimization Algorithms

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-03T12:09:08.946082Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T12:09:08.946082Z digest=sha256:9058f94d63ba343266ab3015473bda2361147145c92690d669b872aefab32db5

Observation 87186d91-407e-45f9-9956-6060da303d7c · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

Learning Latent Reasoning Traces for Scalar Reward Models End-to-End DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-03T12:09:09.065374Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T12:09:09.065374Z digest=sha256:c48135a35f5440501f08612163a07e4beee882bb0aff3ee0eb173f4415c8f3b7

Observation fa4bf8ca-2535-4ee2-8908-0f9ab5e3a769 · outbound

This paper cites Second Conference on Language Modeling , year=.

Learning Latent Reasoning Traces for Scalar Reward Models End-to-End Second Conference on Language Modeling , year=

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-03T12:09:09.185899Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T12:09:09.185899Z digest=sha256:93c8c0217062704f199c3049928fddc7f1d541a6f7f2c8ca23ce57e003c42387

Observation 1c6d41ef-8c97-46da-925a-a457b1746d47 · outbound

This paper cites 2024 , eprint=.

Learning Latent Reasoning Traces for Scalar Reward Models End-to-End 2024 , eprint=

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-03T12:09:09.348666Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T12:09:09.348666Z digest=sha256:30ce22171fbfe274862caf45363dd135c9a72af103e5e890d25426d65712ffca

Observation 78c659ae-7380-4c7b-bb48-9f6a876a6875 · outbound

This paper cites First Conference on Language Modeling , year=.

Learning Latent Reasoning Traces for Scalar Reward Models End-to-End First Conference on Language Modeling , year=

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-03T12:09:09.492302Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T12:09:09.492302Z digest=sha256:7d05bf372b2058beb09dff84d65f82ea1ac64bb8e8a658f55e6b2bd94848cf29

Observation 8ccac31c-3c99-483c-a628-34cdebbdedb8 · outbound

This paper cites 2026 , eprint=.

Learning Latent Reasoning Traces for Scalar Reward Models End-to-End 2026 , eprint=

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-03T12:09:09.677129Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T12:09:09.677129Z digest=sha256:28937e090dd39f2429a57299d74741cffc0654125475fef00168790ec1431454

Observation cef30655-655a-4144-9f6b-2b1535ee6cc5 · outbound

This paper cites and Sreedhar, Makesh Narsimhan and Kuchaiev, Oleksii , booktitle =.

Learning Latent Reasoning Traces for Scalar Reward Models End-to-End and Sreedhar, Makesh Narsimhan and Kuchaiev, Oleksii , booktitle =

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-03T12:09:09.790291Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T12:09:09.790291Z digest=sha256:74963cb91eb711783be8b072f21ac6be1a14cc7c66d33b3855d0638679545d2f

Observation f2de6b5f-bd72-4163-810d-d55b1e16c122 · outbound

This paper cites 2025 , eprint=.

Learning Latent Reasoning Traces for Scalar Reward Models End-to-End 2025 , eprint=

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-03T12:09:09.970266Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T12:09:09.970266Z digest=sha256:d9fdd8cdeccad4d4a5340663f1c5d5cdfab7cf81a7fe696fb3adb66ddc99a5bb

Pith citing papers

No inbound Pith citation observations are available.