Pith. sign in

REVIEW 8 cited by

TICKing All the Boxes: Generated Checklists Improve LLM Evaluation and Generation

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2410.03608 v1 pith:54SYME3V submitted 2024-10-04 cs.AI cs.CLcs.HCcs.LG

classification cs.AIcs.CLcs.HCcs.LG
keywords checklistsevaluationinstructionhumanllmsstickabsolutebest-of-n
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
abstract

Given the widespread adoption and usage of Large Language Models (LLMs), it is crucial to have flexible and interpretable evaluations of their instruction-following ability. Preference judgments between model outputs have become the de facto evaluation standard, despite distilling complex, multi-faceted preferences into a single ranking. Furthermore, as human annotation is slow and costly, LLMs are increasingly used to make these judgments, at the expense of reliability and interpretability. In this work, we propose TICK (Targeted Instruct-evaluation with ChecKlists), a fully automated, interpretable evaluation protocol that structures evaluations with LLM-generated, instruction-specific checklists. We first show that, given an instruction, LLMs can reliably produce high-quality, tailored evaluation checklists that decompose the instruction into a series of YES/NO questions. Each question asks whether a candidate response meets a specific requirement of the instruction. We demonstrate that using TICK leads to a significant increase (46.4% $\to$ 52.2%) in the frequency of exact agreements between LLM judgements and human preferences, as compared to having an LLM directly score an output. We then show that STICK (Self-TICK) can be used to improve generation quality across multiple benchmarks via self-refinement and Best-of-N selection. STICK self-refinement on LiveBench reasoning tasks leads to an absolute gain of $+$7.8%, whilst Best-of-N selection with STICK attains $+$6.3% absolute improvement on the real-world instruction dataset, WildBench. In light of this, structured, multi-faceted self-improvement is shown to be a promising way to further advance LLM capabilities. Finally, by providing LLM-generated checklists to human evaluators tasked with directly scoring LLM responses to WildBench instructions, we notably increase inter-annotator agreement (0.194 $\to$ 0.256).

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 8 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Training AI Scientists to Replicate Research

    cs.LG 2026-08 conditional novelty 6.0 of 10

    A 27B-parameter post-trained agent, Faraday, outperforms frontier coding agents at replicating held-out research figures by directing a larger coding model as a tool.

  2. Graph-Structured Rubrics: Compiling Rubrics into Typed Evaluation Graphs for LLM Judges

    cs.AI 2026-08 conditional novelty 6.0 of 10

    Compiling rubrics into typed evaluation graphs before seeing responses improves LLM judge agreement on four pointwise and two pairwise benchmarks over Prometheus-style and checklist baselines.

  3. CalibratedRubric: Task-Adaptive Rubric Banks for Open-Ended LLM Evaluation

    cs.CL 2026-07 conditional novelty 6.0 of 10

    A task-adaptive rubric-selection method using Bayesian measurability and IRT-based greedy assembly that compresses rubric banks while improving agreement and preserving rank fidelity.

  4. Rubrics on Trial: Evolving Rubrics from a Single Query via Synthetic Pairwise Evidence

    cs.CL 2026-07 conditional novelty 6.0 of 10

    A query-only, annotation-free method that evolves rubrics via synthetic pairwise answer comparisons achieves best average preference-judgment accuracy on seven evaluation sets.

  5. Different Teachers, Different Capabilities: Sub-1B On-Device Distillation for Structured Text Enrichment

    cs.AI 2026-07 conditional novelty 6.0 of 10

    Distilling an 8B reasoning teacher into a 0.6B student recovers most summary quality at ~50× speed, but teacher type—not scale alone—determines which capabilities transfer.

  6. When Life Gives You Samples: The Benefits of Scaling up Inference Compute for Multilingual LLMs

    cs.CL 2025-06 conditional novelty 6.0 of 10

    Hedged sampling, checklist-based one-pass selection (CHOPS), and cross-lingual MBR (X-MBR) improve multilingual LLM output quality when scaling from one to five samples.

  7. MinosEval: Distinguishing Factoid and Non-Factoid for Tailored Open-Ended QA Evaluation with LLMs

    cs.CL 2025-06 conditional novelty 6.0 of 10

    MinosEval improves open-ended QA evaluation by sorting questions into factoid and non-factoid and applying tailored scoring, outperforming baselines on four datasets.

  8. EvalAgent: Discovering Implicit Evaluation Criteria from the Web

    cs.CL 2025-04 conditional novelty 6.0 of 10

    EvalAgent automatically mines expert web advice to generate implicit, actionable evaluation criteria for AI writing, and combining these with LLM-generated criteria increases recall of human-valued criteria compared t...

Pith tools