Pith. sign in

Paper Citation Record · LEDGER

CritiqueLLM: Towards an Informative Critique Generation Model for Evaluation of Large Language Model Generation

As of 9 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 10 inbound Pith citation observations for arXiv:2311.18702.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2311.18702 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 10 of 10 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 10 of 10 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T14:34:55.013457Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T09:29:43.607406Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 0af6dbad-e5ec-4a80-b768-b723924ff16e · inbound

InternLM2 Technical Report cites this paper.

InternLM2 Technical Report CritiqueLLM: Towards an Informative Critique Generation Model for Evaluation of Large Language Model Generation

Reference 176

Resolution
verified exact
arxiv_id, observed 2026-05-15T11:44:38.326618Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-15T11:44:38.066501Z digest=sha256:e43584f87850892f9d8a328a300e8355d3d8698706a801d7fd7eb74ecc1dfd74

Observation dbb066f1-98e7-43fe-bf8a-e72f3c266edb · inbound

Do NOT Think That Much for 2+3=? On the Overthinking of o1-Like LLMs cites this paper.

Do NOT Think That Much for 2+3=? On the Overthinking of o1-Like LLMs CritiqueLLM: Towards an Informative Critique Generation Model for Evaluation of Large Language Model Generation

Reference 117

Resolution
verified exact
arxiv_id, observed 2026-05-13T15:51:29.214794Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-13T15:51:29.022336Z digest=sha256:0bb6bed86124dc8fb889e4898071763e1f67ece5d4f81eac485b83faa7732eee

Observation 0611415f-2911-4494-a03d-deda01d92a83 · inbound

Towards an AI co-scientist cites this paper.

Towards an AI co-scientist CritiqueLLM: Towards an Informative Critique Generation Model for Evaluation of Large Language Model Generation

Reference 30

Resolution
verified exact
arxiv_id, observed 2026-05-11T13:02:44.505754Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-11T13:02:43.571234Z digest=sha256:75f81f52dc088db199ca15be537e9de5c5456e45d4543aee1d0de1608c9c0b8d

Observation cc82f5a6-d743-41f8-be7d-16ad5d48d062 · inbound

Generative RLHF-V: Learning Principles from Multi-modal Human Preference cites this paper.

Generative RLHF-V: Learning Principles from Multi-modal Human Preference CritiqueLLM: Towards an Informative Critique Generation Model for Evaluation of Large Language Model Generation

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-07T14:34:55.013457Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:34:55.013457Z digest=sha256:f0453c7b8117ff98f8b1ce732906c5113d6bd486f992c0cd30625f1ce329fb1b

Observation 2ef0ae42-c98e-4034-b278-6014bd173cdd · inbound

A Survey of Automatic Evaluation Methods on Text, Visual and Speech Generations cites this paper.

A Survey of Automatic Evaluation Methods on Text, Visual and Speech Generations CritiqueLLM: Towards an Informative Critique Generation Model for Evaluation of Large Language Model Generation

Reference 139

Resolution
unresolved
no resolver link, observed 2026-08-07T10:17:46.702243Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:17:46.702243Z digest=sha256:543fc1d4de0e9e90667b30ef51868e7a566767b9ad6fc12fce67361ce863e48c

Observation b938a2bd-bebd-42e9-a32d-77a7003ab2ec · inbound

R4ec: A Reasoning, Reflection, and Refinement Framework for Recommendation Systems cites this paper.

R4ec: A Reasoning, Reflection, and Refinement Framework for Recommendation Systems CritiqueLLM: Towards an Informative Critique Generation Model for Evaluation of Large Language Model Generation

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-06T14:58:24.553833Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:58:24.553833Z digest=sha256:96665569e56892f9f196fd2b17446a4569fc6e127a3fea923a1a0c19c113cfb3

Observation 26313a46-29a8-4b33-9102-c6bf9d104811 · inbound

HumanAgencyBench: Scalable Evaluation of Human Agency Support in AI Assistants cites this paper.

HumanAgencyBench: Scalable Evaluation of Human Agency Support in AI Assistants CritiqueLLM: Towards an Informative Critique Generation Model for Evaluation of Large Language Model Generation

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-04T20:35:54.456352Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:35:54.456352Z digest=sha256:63f42401683aab2de574f15064bfcc1cab2ea35b8dcb52f59dfb3353cc1ec381

Observation 5fa5a145-6d26-41d6-a271-bcf7b38baa39 · inbound

Judging the Judges: A Systematic Evaluation of Bias Mitigation Strategies in LLM-as-a-Judge Pipelines cites this paper.

Judging the Judges: A Systematic Evaluation of Bias Mitigation Strategies in LLM-as-a-Judge Pipelines CritiqueLLM: Towards an Informative Critique Generation Model for Evaluation of Large Language Model Generation

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-05-11T20:41:14.167128Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-08T08:14:18.535385Z digest=sha256:60f8bcc70b55b7c4363d93c85de94256b74f1c8df5163b2dd07206a253c668bf

Observation 0c3cac2a-3b65-46a3-a080-cd4c29c1389c · inbound

BLUEX v2: Benchmarking LLMs on Open-Ended Questions from Brazilian University Entrance Exams cites this paper.

BLUEX v2: Benchmarking LLMs on Open-Ended Questions from Brazilian University Entrance Exams CritiqueLLM: Towards an Informative Critique Generation Model for Evaluation of Large Language Model Generation

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-07-04T09:29:43.608779Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-26T09:59:24.703388Z digest=sha256:68e9197b60c567e67026652cae8ed522511a6bafc82923faf26cedc1d4e795d2

Observation c50c3c1a-dc09-4438-b367-781a115730b7 · inbound

BLUEX v2: Benchmarking LLMs on Open-Ended Questions from Brazilian University Entrance Exams cites this paper.

BLUEX v2: Benchmarking LLMs on Open-Ended Questions from Brazilian University Entrance Exams CritiqueLLM: Towards an Informative Critique Generation Model for Evaluation of Large Language Model Generation

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-07-01T07:05:28.220496Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-07-01T07:04:21.555971Z digest=sha256:f7566936d6df5f8a2e8f904f1b94b8471c1c2bf039729197eb1a0be4496c3e74