Pith. sign in

Paper Citation Record · LEDGER

Meta-Rewarding Language Models: Self-Improving Alignment with LLM-as-a-Meta-Judge

As of 5 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 12 inbound Pith citation observations for arXiv:2407.19594.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2407.19594 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 12 of 12 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-05T06:32:48.257954+00:00

measured 12 of 12 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-05T23:06:49.698152Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

2
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation ca078fb9-faad-404f-8986-e3a52e6b7170 · inbound

LLMs-as-Judges: A Comprehensive Survey on LLM-based Evaluation Methods cites this paper.

LLMs-as-Judges: A Comprehensive Survey on LLM-based Evaluation Methods Meta-Rewarding Language Models: Self-Improving Alignment with LLM-as-a-Meta-Judge

Reference 254

Resolution
verified exact
arxiv_id, observed 2026-05-11T23:08:35.709019Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-11T23:08:34.312466Z digest=sha256:7f9974a3d753e054e1d47fc8693dbc55b92be76771bbe301da3ae5596995e35e

Observation 52402e35-8588-418a-9cd3-f34ca807c882 · inbound

LIMO: Less is More for Reasoning cites this paper.

LIMO: Less is More for Reasoning Meta-Rewarding Language Models: Self-Improving Alignment with LLM-as-a-Meta-Judge

Reference 36

Resolution
verified exact
arxiv_id, observed 2026-05-17T02:11:37.211488Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-17T02:11:36.932541Z digest=sha256:98fa8e69566127563bd08cf1e8cc9d8e5b09d31c919ce7442d2199eb51f26f80

Observation 71b2bf67-f471-482f-ba00-5427082fa2ee · inbound

Temporal Self-Rewarding Language Models: Decoupling Chosen-Rejected via Past-Future cites this paper.

Temporal Self-Rewarding Language Models: Decoupling Chosen-Rejected via Past-Future Meta-Rewarding Language Models: Self-Improving Alignment with LLM-as-a-Meta-Judge

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-05T23:06:49.698152Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T23:06:49.698152Z digest=sha256:0649e8958bf75274f4a0aa81c9eed3bc53fe3229e661e68a11611813bebf24b2

Observation c4ab6dcc-ad00-4bb4-8283-79b3df2fef12 · inbound

Overconfidence in LLM-as-a-Judge: Diagnosis and Confidence-Driven Solution cites this paper.

Overconfidence in LLM-as-a-Judge: Diagnosis and Confidence-Driven Solution Meta-Rewarding Language Models: Self-Improving Alignment with LLM-as-a-Meta-Judge

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-05T22:54:25.666027Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T22:54:25.666027Z digest=sha256:d0a07c10e764513cfa260b578e40d8c35c52d17512af371b91205e0d1c51a23b

Observation 388d9b6e-b63c-45f9-90f4-387e4da831d4 · inbound

Scoring, Reasoning, and Selecting the Best! Ensembling Large Language Models via a Peer-Review Process cites this paper.

Scoring, Reasoning, and Selecting the Best! Ensembling Large Language Models via a Peer-Review Process Meta-Rewarding Language Models: Self-Improving Alignment with LLM-as-a-Meta-Judge

Reference 52

Resolution
metadata mismatch
arxiv_id, observed 2026-05-16T19:58:22.687864Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-16T19:57:03.999154Z digest=sha256:bdc95f3d94aaeade35157cba54cf930e4bb6c8c795056d6ca842b105a3165694

Observation 03b422e5-c77a-4319-b3bd-13d999609f2d · inbound

Can LLMs Learn to Reason Robustly under Noisy Supervision? cites this paper.

Can LLMs Learn to Reason Robustly under Noisy Supervision? Meta-Rewarding Language Models: Self-Improving Alignment with LLM-as-a-Meta-Judge

Reference 26

Resolution
metadata mismatch
arxiv_id, observed 2026-05-13T17:08:01.288616Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-13T16:58:42.129870Z digest=sha256:3b058661ad2411d0c24931963ceb878048ff21172c4aa26018bef31b219df64b

Observation fcf5096a-b107-46f8-ab63-0b466d348d59 · inbound

Rethinking Math Reasoning Evaluation: A Robust LLM-as-a-Judge Framework Beyond Symbolic Rigidity cites this paper.

Rethinking Math Reasoning Evaluation: A Robust LLM-as-a-Judge Framework Beyond Symbolic Rigidity Meta-Rewarding Language Models: Self-Improving Alignment with LLM-as-a-Meta-Judge

Reference 30

Resolution
verified exact
arxiv_id, observed 2026-05-11T19:31:09.407696Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-08T11:45:54.181019Z digest=sha256:0ee2d27140dc0e62090c507277c8c627969db405d4ccc28c6f215ff9e50eb3d5

Observation 8da6dbac-c112-4210-853b-3ca4df8d9d87 · inbound

Weak Critics Make Strong Learners: On-Policy Critique Distillation for Scalable Oversight cites this paper.

Weak Critics Make Strong Learners: On-Policy Critique Distillation for Scalable Oversight Meta-Rewarding Language Models: Self-Improving Alignment with LLM-as-a-Meta-Judge

Reference 21

Resolution
metadata mismatch
arxiv_id, observed 2026-07-01T19:56:10.972571Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-28T21:55:58.645062Z digest=sha256:1f0bcf63547308ed3262ad6586e865db828c85ff66ca15791d717abb7a981e00

Observation 3efc4b0a-cf2d-4b52-9170-587accd10d67 · inbound

Trust Region On-Policy Distillation cites this paper.

Trust Region On-Policy Distillation Meta-Rewarding Language Models: Self-Improving Alignment with LLM-as-a-Meta-Judge

Reference 76

Resolution
metadata mismatch
arxiv_id, observed 2026-07-01T20:56:13.295277Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-06-28T17:38:50.313305Z digest=sha256:05e8aaea65049d7c00cd80a35d138dcdcf1d78d633e6139700e03523bbc8bf85

Observation 2f98913a-87b1-4de3-8e53-0b135e21af01 · inbound

The Periodic Table of LLM Reasoning: A Structured Survey of Reasoning Paradigms, Methods, and Failure Modes cites this paper.

The Periodic Table of LLM Reasoning: A Structured Survey of Reasoning Paradigms, Methods, and Failure Modes Meta-Rewarding Language Models: Self-Improving Alignment with LLM-as-a-Meta-Judge

Reference 264

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T05:57:41.551756Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-06-27T12:59:51.091008Z digest=sha256:2ab2fa67d0455add0f73fe76aadd67c9d4e3e568abf97ade16c661277e554da6

Observation 103da097-1211-4c59-961f-d31d121ebf1d · inbound

SingGuard: A Policy-Adaptive Multimodal LLM Guardrail with Dynamic Reasoning cites this paper.

SingGuard: A Policy-Adaptive Multimodal LLM Guardrail with Dynamic Reasoning Meta-Rewarding Language Models: Self-Improving Alignment with LLM-as-a-Meta-Judge

Reference 218

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T09:59:44.789963Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-06-26T09:19:50.623741Z digest=sha256:1a00da279dca6605a5790154209fac3c33ef6b11cfefc56045ccd808dc838dc4

Observation 99a6daa6-1d24-418e-b7ff-d2b1e72ebef6 · inbound

SingGuard: A Policy-Adaptive Multimodal LLM Guardrail with Dynamic Reasoning cites this paper.

SingGuard: A Policy-Adaptive Multimodal LLM Guardrail with Dynamic Reasoning Meta-Rewarding Language Models: Self-Improving Alignment with LLM-as-a-Meta-Judge

Reference 217

Resolution
metadata mismatch
arxiv_id, observed 2026-07-01T18:55:59.731218Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-06-29T01:18:19.195007Z digest=sha256:a757c914319404049f75dc9688db1f7c5d6af8e3a81c3dd92fce8d370bb44f53