Pith. sign in

Paper Citation Record · LEDGER

Intra-Trajectory Consistency for Reward Modeling

As of 8 August 2026, this Paper Citation Record lists 52 of 52 outbound references and 0 inbound Pith citation observations for arXiv:2506.09096.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.09096 v3

Coverage vector

measured 52 of 52 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T05:09:38.960214Z

measured 52 of 52 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

52 of 52 outbound references displayed

  • verified exact0
  • verified fuzzy46
  • unresolved6
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 4132a0e4-49f0-476a-91db-a8c37c4f63d7 · outbound

This paper cites Training language models to follow instructions with human feedback.

Intra-Trajectory Consistency for Reward Modeling Training language models to follow instructions with human feedback

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:09:45.831652Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T05:09:33.802950Z digest=sha256:911699caa60e0aeb9540dcbaf56ba295029d845fff68d37ec4fc4e0c040d299a

Observation bec86bc4-369b-4086-b50c-f16b257b9303 · outbound

This paper cites Safe RLHF : Safe reinforcement learning from human feedback.

Intra-Trajectory Consistency for Reward Modeling Safe RLHF : Safe reinforcement learning from human feedback

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:09:45.778388Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T05:09:33.901900Z digest=sha256:62969570f9b8a882099933fa2f86d21f1a12fe06da7ac4434ff3d0247bf31b78

Observation 2fcd8772-3748-4012-983b-cf3ccc1b3c56 · outbound

This paper cites Model alignment as prospect theoretic optimization.

Intra-Trajectory Consistency for Reward Modeling Model alignment as prospect theoretic optimization

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:09:45.715947Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T05:09:33.972354Z digest=sha256:56cf6347f79d16186550c26e5a785fc14772dc045030f6d6a318537fe861f1ff

Observation 99d4cb26-b296-413f-a0f2-d56a0e88cb02 · outbound

This paper cites A Survey of Direct Preference Optimization.

Intra-Trajectory Consistency for Reward Modeling A Survey of Direct Preference Optimization

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T05:09:34.053645Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:09:34.053645Z digest=sha256:5bed71ed01e8479ab41d517f27be16484cc23b4a9d064bf9e1027d6d04c188a8

Observation 6fde3e4a-7d27-4dd2-8899-cc5c84a54d8f · outbound

This paper cites Generative verifiers: Reward modeling as next-token prediction.

Intra-Trajectory Consistency for Reward Modeling Generative verifiers: Reward modeling as next-token prediction

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:09:45.658480Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T05:09:34.125960Z digest=sha256:7205ebaeefdf96e0f20fc34f80e8d919314bf64e2dbc04d0ae19159ab9bf788d

Observation b9fe2fbb-cf39-4bd0-ab3a-daf565f93a65 · outbound

This paper cites Rewarding progress: Scaling automated process verifiers for LLM reasoning.

Intra-Trajectory Consistency for Reward Modeling Rewarding progress: Scaling automated process verifiers for LLM reasoning

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:09:45.586236Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T05:09:34.198454Z digest=sha256:fb402e5034c50c6b96035c89523c1e10db4875d4207af240312424a9fabfc90e

Observation 6661a49b-0f37-47d9-b026-caa023ae1740 · outbound

This paper cites Scaling laws for reward model overoptimization.

Intra-Trajectory Consistency for Reward Modeling Scaling laws for reward model overoptimization

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:09:45.504214Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T05:09:34.320396Z digest=sha256:861f324aea0f1f672d3ee961d586c46c5b8108ec1c6f91765f798248ba6e562f

Observation 726cb910-5925-4f6c-b4a9-48241819474f · outbound

This paper cites Regularizing hidden states enables learning generalizable reward model for LLM s.

Intra-Trajectory Consistency for Reward Modeling Regularizing hidden states enables learning generalizable reward model for LLM s

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:09:45.412624Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T05:09:34.412805Z digest=sha256:7366ce01ba9648993575aa6e86e9e1448dfa2908d8632e2e578a3604fa2a66c1

Observation 53f1914c-cd13-4cac-aaaa-4e94bc21976d · outbound

This paper cites Reward model ensembles help mitigate overoptimization.

Intra-Trajectory Consistency for Reward Modeling Reward model ensembles help mitigate overoptimization

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:09:45.303562Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T05:09:34.526511Z digest=sha256:1212b5c54139607f6764b376c31635966db4f331c3f2eb220073c0c9ce9a4c59

Observation edff0ebb-c86c-41a1-a74a-25c15ce9b1c3 · outbound

This paper cites Warm: On the benefits of weight averaged reward models.

Intra-Trajectory Consistency for Reward Modeling Warm: On the benefits of weight averaged reward models

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:09:45.175665Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T05:09:34.624566Z digest=sha256:1998580b8d23e8d98400ae0a8a14c54bc11500e61113be54bad971ef25f69a14

Observation 32e99435-33a6-4fd5-a7d6-30f7ff031b90 · outbound

This paper cites The trickle-down impact of reward inconsistency on RLHF.

Intra-Trajectory Consistency for Reward Modeling The trickle-down impact of reward inconsistency on RLHF

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:09:45.062844Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T05:09:34.718080Z digest=sha256:c42325838ba089117d0b13ae73762d7e26a9899bf5553480dcd267bec1902210

Observation 51ecde41-b6f3-43e0-b887-94d27783c4df · outbound

This paper cites Rrm: Robust reward model training mitigates reward hacking.

Intra-Trajectory Consistency for Reward Modeling Rrm: Robust reward model training mitigates reward hacking

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:09:44.952080Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T05:09:34.818224Z digest=sha256:cff8e0ef8c540e0071ebe56ebc01a9cc4acc2e45898386893280d82ed218a2f1

Observation bbc1ae30-c7e8-4a0c-860c-cac1e1c9f90a · outbound

This paper cites Length-controlled alpacaeval: A simple debiasing of automatic evaluators.

Intra-Trajectory Consistency for Reward Modeling Length-controlled alpacaeval: A simple debiasing of automatic evaluators

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:09:44.827940Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T05:09:34.911868Z digest=sha256:fe382dc66e9774af7fe80aa82430728f629f74366da9c7e3763de7be9d7f8484

Observation 2b9af422-9ecf-4211-89ba-30c3ec0a3010 · outbound

This paper cites Odin: Disentangled reward mitigates hacking in rlhf.

Intra-Trajectory Consistency for Reward Modeling Odin: Disentangled reward mitigates hacking in rlhf

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:09:44.701760Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T05:09:35.020744Z digest=sha256:406069d0f3969c66afe9574c6b67a89b1d3013dda5ce3637f81d0aa2a48e5276

Observation fdc4bee4-9ae1-446c-aa16-55969f355381 · outbound

This paper cites Improving discriminative capability of reward models in rlhf using contrastive learning.

Intra-Trajectory Consistency for Reward Modeling Improving discriminative capability of reward models in rlhf using contrastive learning

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:09:44.590224Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T05:09:35.128852Z digest=sha256:8f7d5ff6eea7c3662cb3a48c5d3df6ba7d7c300c4e9cbcc0c3373ec4082144e5

Observation 49f91bbc-1be1-4b00-9db4-52a9d972270c · outbound

This paper cites Rethinking reward modeling in preference-based large language model alignment.

Intra-Trajectory Consistency for Reward Modeling Rethinking reward modeling in preference-based large language model alignment

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:09:44.454000Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T05:09:35.197601Z digest=sha256:dc4d46d2f1a175a96bf4f1cc40fe7c5f7bee7630e6409a01889777ec2d7f4fac

Observation 9049e081-f6f4-4c1f-a5fa-e908588fa405 · outbound

This paper cites Let's verify step by step.

Intra-Trajectory Consistency for Reward Modeling Let's verify step by step

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T05:09:35.330672Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:09:35.330672Z digest=sha256:45e5d547a66958525a41765672b5ebfe0ef04c34c4a78f2a9e2875cee0d7e7a8

Observation 5e927d83-6394-4273-ac20-844e64b50103 · outbound

This paper cites Math-shepherd: Verify and reinforce llms step-by-step without human annotations.

Intra-Trajectory Consistency for Reward Modeling Math-shepherd: Verify and reinforce llms step-by-step without human annotations

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:09:44.340960Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T05:09:35.442708Z digest=sha256:f5d6c2d546f8baabd915353cba1c6aa3e898cab7d77db5321fde79aa6b715274

Observation 7a7ca000-670d-417c-8a6a-7a8db72ec35f · outbound

This paper cites Rest-mcts*: Llm self-training via process reward guided tree search.

Intra-Trajectory Consistency for Reward Modeling Rest-mcts*: Llm self-training via process reward guided tree search

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:09:44.200619Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T05:09:35.554133Z digest=sha256:8377f97f7e2ac207cb20b154f6319a737cf29cf2649eea5a1f164e6fb41cbbb1

Observation 0e4c4813-eb08-4f65-bbb2-cbf80ae1a160 · outbound

This paper cites Gemma: Open models based on gemini research and technology.

Intra-Trajectory Consistency for Reward Modeling Gemma: Open models based on gemini research and technology

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:09:44.040192Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T05:09:35.617671Z digest=sha256:c46aa2d55965351b587d9c6b9ab499a849165affb6cb446ee76130e87c9f8f11

Observation 0b1e1919-0da9-432f-98da-33781cc30cd4 · outbound

This paper cites The llama 3 herd of models.

Intra-Trajectory Consistency for Reward Modeling The llama 3 herd of models

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:09:43.927236Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T05:09:35.758591Z digest=sha256:3bd7aeedcaecf3e41f835edf16a4e39e576820f59a5cb1a70b214030aaad6584

Observation 379d687f-f425-49c7-bf25-29e493da0680 · outbound

This paper cites Training a helpful and harmless assistant with reinforcement learning from human feedback.

Intra-Trajectory Consistency for Reward Modeling Training a helpful and harmless assistant with reinforcement learning from human feedback

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:09:43.851006Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T05:09:35.867529Z digest=sha256:1150efaf4561615ec881292772fbe730e5b06ece44e60435ca97b03c464f6e83

Observation e15458fe-de22-4138-89a3-13ee36986d9f · outbound

This paper cites an unresolved cited work.

Intra-Trajectory Consistency for Reward Modeling Unresolved cited work

Reference 23

Resolution
unresolved
raw_fallback, observed 2026-08-07T05:09:43.590225Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T05:09:35.924570Z digest=sha256:5687586e71e7e877b14b82cadcda0413399aa5c503996afc454494ea4147249f

Observation 2c87082c-53e2-4bcf-8941-90a39fabb102 · outbound

This paper cites Solving math word problems with process-and outcome-based feedback.

Intra-Trajectory Consistency for Reward Modeling Solving math word problems with process-and outcome-based feedback

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:09:43.247491Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T05:09:36.016024Z digest=sha256:6b9cda9a7447829eb0eadb5a01bd8fad446c50dae3b8de49fbf7765f945dcb8c

Observation b6be155d-90e6-4dfb-8466-2f519790f09e · outbound

This paper cites Token-level direct preference optimization.

Intra-Trajectory Consistency for Reward Modeling Token-level direct preference optimization

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:09:42.898366Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T05:09:36.115444Z digest=sha256:238714a4b58e2ac99f7aa194dd8c25d9afc20d5e0c4ff2afe10e2275009fa1e8

Observation 835f56c4-ee2f-4d5e-b96b-004c8aae0fb5 · outbound

This paper cites Process reinforcement through implicit rewards.

Intra-Trajectory Consistency for Reward Modeling Process reinforcement through implicit rewards

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:09:42.614422Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T05:09:36.195715Z digest=sha256:d85a15a0502d300ce3ceb792359bfeaa6d9bbdb28b32f2990cae2c853ff4912b

Observation 521f59ea-afed-41d1-be65-117b7ccc2e9c · outbound

This paper cites Gen ARM : Reward guided generation with autoregressive reward model for test-time alignment.

Intra-Trajectory Consistency for Reward Modeling Gen ARM : Reward guided generation with autoregressive reward model for test-time alignment

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:09:42.503125Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T05:09:36.318045Z digest=sha256:25d20b85a028774d7dd3251322072875797cefa17c160b02df46ad4546c7b60b

Observation 1be149c5-8847-42b9-a1aa-f1115b8585de · outbound

This paper cites Fine-grained human feedback gives better rewards for language model training.

Intra-Trajectory Consistency for Reward Modeling Fine-grained human feedback gives better rewards for language model training

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:09:42.368824Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T05:09:36.407602Z digest=sha256:e6bded730c3a3afe183c86175b56420796c7dbcd383cc256afc24610608653d9

Observation 196c5c63-3519-4ae4-95de-a01b34f6c74a · outbound

This paper cites Safety alignment should be made more than just a few tokens deep.

Intra-Trajectory Consistency for Reward Modeling Safety alignment should be made more than just a few tokens deep

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T05:09:36.488861Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:09:36.488861Z digest=sha256:0c6b9ede347ab0fe93d50d8b5028cf3a1931a67467c422bd84e0c37079364cf5

Observation 6237dda3-1d2e-4d98-b747-9e20ee8a911f · outbound

This paper cites Process reward model with q-value rankings.

Intra-Trajectory Consistency for Reward Modeling Process reward model with q-value rankings

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:09:42.234540Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T05:09:36.596630Z digest=sha256:24f6fc892ce3e7548e98844c8b730badb242de5cdb712839042c1d16f7e90cdf

Observation ce204ed8-5904-45fb-9405-ced160567cf9 · outbound

This paper cites Coda: Contrast-enhanced and diversity-promoting data augmentation for natural language understanding.

Intra-Trajectory Consistency for Reward Modeling Coda: Contrast-enhanced and diversity-promoting data augmentation for natural language understanding

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:09:42.119546Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T05:09:36.685789Z digest=sha256:f67ff907f8eb126169b68a9154b5157a3fceeb80a27f54d765362560629e209b

Observation 98bdd41b-ec82-485a-943d-8e407496663e · outbound

This paper cites A survey on data augmentation for text classification.

Intra-Trajectory Consistency for Reward Modeling A survey on data augmentation for text classification

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:09:41.971159Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T05:09:36.777800Z digest=sha256:1c63a38c5ba52b69530e80c5de3ffa4c7fd4a7ac18a5b2d340450dc878907d7d

Observation 44606918-c76c-4a6c-bd26-703a45d2b9f7 · outbound

This paper cites Cubuk, Alex Kurakin, Han Zhang, and Colin Raffel.

Intra-Trajectory Consistency for Reward Modeling Cubuk, Alex Kurakin, Han Zhang, and Colin Raffel

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:09:41.850902Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T05:09:36.873930Z digest=sha256:d7f3d88cb8cf8ce1484d2ac5055f87d476aace47f4a4535abb03227587b67e74

Observation 8e096c90-a31e-446b-8466-1a2f3dbd2c6d · outbound

This paper cites Flexmatch: Boosting semi-supervised learning with curriculum pseudo labeling.

Intra-Trajectory Consistency for Reward Modeling Flexmatch: Boosting semi-supervised learning with curriculum pseudo labeling

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:09:41.732648Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T05:09:36.958032Z digest=sha256:2d0c3076a5000c32f2e1e3386db02672718a7f33aa071e6e5e5cedce9183dcde

Observation 62cc8540-6c22-42c8-bf54-bd30eaa62a12 · outbound

This paper cites Qwen2 technical report.

Intra-Trajectory Consistency for Reward Modeling Qwen2 technical report

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:09:41.612274Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T05:09:37.074903Z digest=sha256:3008e583b714b7355deeeaf23337df6267377c6d710a0b42124bd159dc99eea8

Observation 5b5989df-c9f7-487c-abac-0ae2dad00254 · outbound

This paper cites Measuring mathematical problem solving with the math dataset.

Intra-Trajectory Consistency for Reward Modeling Measuring mathematical problem solving with the math dataset

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:09:41.503487Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T05:09:37.183892Z digest=sha256:a1b881fc3e7248031b3e6e2170e7fa9ed8e46fbb414d13438d33aba879589417

Observation 73391db7-3f72-47da-8f0c-1a3638a177ca · outbound

This paper cites Free process rewards without process labels.

Intra-Trajectory Consistency for Reward Modeling Free process rewards without process labels

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:09:41.385834Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T05:09:37.281339Z digest=sha256:67615905abc2a2fbfe6687f2392510044534514f07d519baecbac2d7b6d42ce3

Observation 3bc09aff-6ed7-4505-8345-735472625b6b · outbound

This paper cites Mistral 7B.

Intra-Trajectory Consistency for Reward Modeling Mistral 7B

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-07T05:09:37.400183Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:09:37.400183Z digest=sha256:6b81e28e29cf77cfc487c6669b76df3171f8e56fd56984bf08d355aeffce5266

Observation f2697b11-fb0a-4340-8c89-b747571cf43d · outbound

This paper cites Direct preference optimization: Your language model is secretly a reward model.

Intra-Trajectory Consistency for Reward Modeling Direct preference optimization: Your language model is secretly a reward model

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:09:41.255820Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T05:09:37.508705Z digest=sha256:566d347316f3560189768060642d37fc84f8b11e0bf3ab96ce2a96caf0137c3c

Observation 10235350-5bec-4bbc-a2e7-56b114fc2575 · outbound

This paper cites Rlhf workflow: From reward modeling to online rlhf.

Intra-Trajectory Consistency for Reward Modeling Rlhf workflow: From reward modeling to online rlhf

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:09:41.128081Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T05:09:37.639203Z digest=sha256:aeb7429ade77cd858bb5ef582f62eb2abc4b03f8ca6277fd801768f55133a1b2

Observation 6e5a6512-f8d7-4337-b74d-7e6e175ebe34 · outbound

This paper cites Rewardbench: Evaluating reward models for language modeling.

Intra-Trajectory Consistency for Reward Modeling Rewardbench: Evaluating reward models for language modeling

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:09:40.978309Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T05:09:37.747142Z digest=sha256:1426c72fc8fd1733c84a8f79acdcc034b954a9f79d746f1b9025d65ed5a9ed0d

Observation e05e6164-2d74-45cd-a9a0-df3fa4aefba7 · outbound

This paper cites Llama 2: Open foundation and fine-tuned chat models.

Intra-Trajectory Consistency for Reward Modeling Llama 2: Open foundation and fine-tuned chat models

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:09:40.799148Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T05:09:37.856360Z digest=sha256:fdb02a113839ff958730658dafa9d54f19280c88eaae0847d17b7cc2e240648b

Observation 43d67609-7c95-4db5-b8a2-0c528715a0b9 · outbound

This paper cites Secrets of rlhf in large language models part ii: Reward modeling.

Intra-Trajectory Consistency for Reward Modeling Secrets of rlhf in large language models part ii: Reward modeling

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:09:40.614967Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T05:09:38.001579Z digest=sha256:2d25f2eb562ca35076d5c258379c54b437fdc97964277a22f5fa2b27b2384d1d

Observation f81486a0-a0ec-4cec-83a8-fc12d2ec4624 · outbound

This paper cites Reward model ensembles help mitigate overoptimization.

Intra-Trajectory Consistency for Reward Modeling Reward model ensembles help mitigate overoptimization

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:09:40.426846Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T05:09:38.111936Z digest=sha256:37d26b28f9d23b376f39faeccb04b166b8ba4bb6dd4f16d139390cc71f644711

Observation 3c2aebab-b8cb-4e42-b102-a5ccbcf6015e · outbound

This paper cites Alpacafarm: A simulation framework for methods that learn from human feedback.

Intra-Trajectory Consistency for Reward Modeling Alpacafarm: A simulation framework for methods that learn from human feedback

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:09:40.259515Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T05:09:38.239973Z digest=sha256:a2716489c84fda56bc0dd33daf98f8d9cc55a2e8b9bbb9ded4e8c9e8adfd0af6

Observation 30b3becf-60e8-4b2f-b91a-08d6e6de7e84 · outbound

This paper cites Loose lips sink ships: Mitigating length bias in reinforcement learning from human feedback.

Intra-Trajectory Consistency for Reward Modeling Loose lips sink ships: Mitigating length bias in reinforcement learning from human feedback

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:09:40.082421Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T05:09:38.357506Z digest=sha256:3ec42244166118a16c6798fa65fceb370791df80eb5b1c3821aa844ed64047a8

Observation d8e979ef-ed7f-4cfb-ac10-6836fa9f4ea1 · outbound

This paper cites Pfohl, Deepak Ramachandran, Peter Shaw, and Jonathan Berant.

Intra-Trajectory Consistency for Reward Modeling Pfohl, Deepak Ramachandran, Peter Shaw, and Jonathan Berant

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:09:39.901062Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T05:09:38.466365Z digest=sha256:099325ac5129ebc8bd06dd53d4b152de92f58a24ae92ec542fa7f33c6178d3d9

Observation 7d9606f7-8284-4b33-8833-7550696c0906 · outbound

This paper cites Proximal policy optimization algorithms.

Intra-Trajectory Consistency for Reward Modeling Proximal policy optimization algorithms

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:09:39.751532Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T05:09:38.545848Z digest=sha256:7771d07f35d9154b03aa47f7743be39e14a0d5b236631aa02ee9040515c69eb9

Observation 49fbd15f-3dca-4368-b712-494501cc534c · outbound

This paper cites Understanding the learning dynamics of alignment with human feedback.

Intra-Trajectory Consistency for Reward Modeling Understanding the learning dynamics of alignment with human feedback

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:09:39.555383Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T05:09:38.671903Z digest=sha256:b641fa2ebd1165bff49b6429dd7998a48c346058ff14c7360b2166c4cfb2f71b

Observation f245d885-99b7-4bee-be2f-03ae603fcd91 · outbound

This paper cites Improve mathematical reasoning in language models by automated process supervision.

Intra-Trajectory Consistency for Reward Modeling Improve mathematical reasoning in language models by automated process supervision

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:09:39.388219Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T05:09:38.760953Z digest=sha256:7faa3584e4b8cca749372475a1668af53b9e0972a3dfd47d832ef2522813e5f2

Observation dab70296-bd7d-493d-b28f-2bcddad4ce2f · outbound

This paper cites Llamafactory: Unified efficient fine-tuning of 100+ language models.

Intra-Trajectory Consistency for Reward Modeling Llamafactory: Unified efficient fine-tuning of 100+ language models

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:09:39.187419Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T05:09:38.854123Z digest=sha256:cdb4e399ee5ca89af1cced4db9ca6ac464cd108084386813a92e90c86d01d4a3

Observation 1327d8a5-e129-4bc2-a7e8-75d29996d2b4 · outbound

This paper cites Lora: Low-rank adaptation of large language models.

Intra-Trajectory Consistency for Reward Modeling Lora: Low-rank adaptation of large language models

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-07T05:09:38.960214Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:09:38.960214Z digest=sha256:6247b9c68573ba0545acad559a88a7d78f6d7fed2bd098fb8bd45e0b183ca8d0

Pith citing papers

No inbound Pith citation observations are available.