Pith. sign in

Paper Citation Record · LEDGER

ConsistRM: Improving Generative Reward Models via Consistency-Aware Self-Training

As of 4 August 2026, this Paper Citation Record lists 14 of 14 outbound references and 0 inbound Pith citation observations for arXiv:2604.07484.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2604.07484 v2

Coverage vector

measured 14 of 14 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-05-10T18:12:55.726473Z

measured 14 of 14 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-04T06:34:03.388597+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

14 of 14 outbound references displayed

  • verified exact2
  • verified fuzzy2
  • unresolved6
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch4

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation a2c5f952-0625-401f-b6ee-d5d35e16f2ef · outbound

This paper cites Deep Think with Confidence.

ConsistRM: Improving Generative Reward Models via Consistency-Aware Self-Training Deep Think with Confidence

Reference 1

Resolution
verified exact
arxiv_id, observed 2026-05-16T11:30:21.518780Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T18:12:55.726473Z digest=sha256:fbe4c158dbe48c027e42321125c52ca604fe67873a0decbbfb91116b07835f92

Observation 06d2ddd8-1325-45e9-89bb-dc941196c020 · outbound

This paper cites Correlated Proxies: A New Definition and Improved Mitigation for Reward Hacking.

ConsistRM: Improving Generative Reward Models via Consistency-Aware Self-Training Correlated Proxies: A New Definition and Improved Mitigation for Reward Hacking

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-05-11T05:20:57.524394Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T18:12:55.726473Z digest=sha256:9986a106a9b3dfbd47e6efbd4d6b9f37e9e84c1080677f42e806f60440093778

Observation 07d7b8a6-e0a2-4e6d-b4f6-98930e1cad61 · outbound

This paper cites Generative Reward Models.

ConsistRM: Improving Generative Reward Models via Consistency-Aware Self-Training Generative Reward Models

Reference 3

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T05:20:57.515709Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T18:12:55.726473Z digest=sha256:92c28725ecf26d7bd48259e3485456a7a6de8e5be7fcc36f2033798372f2ecd9

Observation 5f890e66-3a40-4204-a50f-f731830dcea3 · outbound

This paper cites Learning to Plan & Reason for Evaluation with Thinking-LLM-as-a-Judge.

ConsistRM: Improving Generative Reward Models via Consistency-Aware Self-Training Learning to Plan & Reason for Evaluation with Thinking-LLM-as-a-Judge

Reference 4

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T05:20:57.479472Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T18:12:55.726473Z digest=sha256:b695e7d550506a2a0cabe5ecae9422685c3043e0b0e20d526923e92a7088e45a

Observation 370d0ac6-6f41-4d36-bca1-4fd51f43450a · outbound

This paper cites Qwen3 Technical Report.

ConsistRM: Improving Generative Reward Models via Consistency-Aware Self-Training Qwen3 Technical Report

Reference 5

Resolution
metadata mismatch
local_arxiv, observed 2026-05-11T05:20:57.502619Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T18:12:55.726473Z digest=sha256:e8e76dc0ad4fa73a702954537a33779b9d6fd58fd05fab0dde7b0f2b2ed99caa

Observation 5ecaf77b-8d39-4576-b78f-eca6fc09c686 · outbound

This paper cites Consistent Paths Lead to Truth: Self-Rewarding Reinforcement Learning for LLM Reasoning.

ConsistRM: Improving Generative Reward Models via Consistency-Aware Self-Training Consistent Paths Lead to Truth: Self-Rewarding Reinforcement Learning for LLM Reasoning

Reference 6

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T05:20:57.469347Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T18:12:55.726473Z digest=sha256:e5574b4bc7a78c3ca49baef6989c9322df212842501a74fad7e5cc194423b6f9

Observation c3c920da-d8a9-44c7-9d48-5b26fcb69339 · outbound

This paper cites an unresolved cited work.

ConsistRM: Improving Generative Reward Models via Consistency-Aware Self-Training Unresolved cited work

Reference 7

Resolution
unresolved
raw_fallback, observed 2026-05-17T06:31:36.906704Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T18:12:55.726473Z digest=sha256:3be0d4be5db5bc618be0efb81d27fb47e8f71995389e4f57eaf8342dbc080b71

Observation 54dd87a1-56bb-4821-a9cf-c777e9f2d5da · outbound

This paper cites an unresolved cited work.

ConsistRM: Improving Generative Reward Models via Consistency-Aware Self-Training Unresolved cited work

Reference 8

Resolution
unresolved
raw_fallback, observed 2026-05-17T06:31:36.887357Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T18:12:55.726473Z digest=sha256:bbf83dbffecf02e39849b6883ebe01040caf933893a54b576e6b4c7e00b73b8b

Observation d4baa38a-00cf-4d09-ac61-44f62aa32209 · outbound

This paper cites an unresolved cited work.

ConsistRM: Improving Generative Reward Models via Consistency-Aware Self-Training Unresolved cited work

Reference 9

Resolution
unresolved
raw_fallback, observed 2026-05-17T06:31:36.889538Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T18:12:55.726473Z digest=sha256:3ff5dcb77810951e6338c21dfd2568af7a0fb198a7d752bd055a4b20b695bca7

Observation 5ff1ec89-606a-46b0-a8fb-680d7a218b7c · outbound

This paper cites </Criterion> <Analysis> Response 1: Response 1 provides a broad overview of NHI systems and correctly explains the concepts of both top-down and bottom-up approaches in healthcare.

ConsistRM: Improving Generative Reward Models via Consistency-Aware Self-Training </Criterion> <Analysis> Response 1: Response 1 provides a broad overview of NHI systems and correctly explains the concepts of both top-down and bottom-up approaches in healthcare

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T06:31:36.897887Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T18:12:55.726473Z digest=sha256:26b14d9099f383f5f27ca2ad69ed0644ce476747d5e7c182fef906d54df9ecf1

Observation e0d1a515-5b34-46f1-b4ab-427baea06963 · outbound

This paper cites an unresolved cited work.

ConsistRM: Improving Generative Reward Models via Consistency-Aware Self-Training Unresolved cited work

Reference 11

Resolution
unresolved
raw_fallback, observed 2026-05-17T06:31:36.903427Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T18:12:55.726473Z digest=sha256:13aa1d9b5acd70d4874a028cbba823ee97f662daf41e5877d6ff1c09f3c317c1

Observation 0b054f66-50e0-4232-974a-5aacc5adf354 · outbound

This paper cites an unresolved cited work.

ConsistRM: Improving Generative Reward Models via Consistency-Aware Self-Training Unresolved cited work

Reference 12

Resolution
unresolved
raw_fallback, observed 2026-05-17T06:31:36.891858Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T18:12:55.726473Z digest=sha256:a7a915ad22e4b7964e6f2170e314e6505d4bc1be19d21551e291a4851cac2997

Observation 8687569f-0162-4e09-a581-2367bd9fcc4a · outbound

This paper cites an unresolved cited work.

ConsistRM: Improving Generative Reward Models via Consistency-Aware Self-Training Unresolved cited work

Reference 13

Resolution
unresolved
raw_fallback, observed 2026-05-17T06:31:36.894820Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T18:12:55.726473Z digest=sha256:024c9c50ef02bafde95ac491fc568cae261c1e0462b598d6916a2eca6477d8f8

Observation 1a0e2a63-45e1-4361-b2bd-dd5b142ed282 · outbound

This paper cites </Criterion> <Analysis> Response 1: Response 1 provides a comprehensive overview of the NHI system and explains both top-down and bottom-up approaches in the context of healthcare.

ConsistRM: Improving Generative Reward Models via Consistency-Aware Self-Training </Criterion> <Analysis> Response 1: Response 1 provides a comprehensive overview of the NHI system and explains both top-down and bottom-up approaches in the context of healthcare

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T06:31:36.884191Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T18:12:55.726473Z digest=sha256:be54b683eb9fd45c4e931f55c2116bb3b87a28a4034cdf56f027a7181994a6a4

Pith citing papers

No inbound Pith citation observations are available.