REVIEW 9 cited by
ROSCOE: A Suite of Metrics for Scoring Step-by-Step Reasoning
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Large language models show improved downstream task performance when prompted to generate step-by-step reasoning to justify their final answers. These reasoning steps greatly improve model interpretability and verification, but objectively studying their correctness (independent of the final answer) is difficult without reliable methods for automatic evaluation. We simply do not know how often the stated reasoning steps actually support the final end task predictions. In this work, we present ROSCOE, a suite of interpretable, unsupervised automatic scores that improve and extend previous text generation evaluation metrics. To evaluate ROSCOE against baseline metrics, we design a typology of reasoning errors and collect synthetic and human evaluation scores on commonly used reasoning datasets. In contrast with existing metrics, ROSCOE can measure semantic consistency, logicality, informativeness, fluency, and factuality - among other traits - by leveraging properties of step-by-step rationales. We empirically verify the strength of our metrics on five human annotated and six programmatically perturbed diagnostics datasets - covering a diverse set of tasks that require reasoning skills and show that ROSCOE can consistently outperform baseline metrics.
Forward citations
Cited by 9 Pith papers
-
More Debate, Same Evidence: Structural Limits of Homogeneous Multi-Agent Groundedness
A three-agent homogeneous debate panel does not reliably improve groundedness verification: gains and losses are task-dependent, and its main mechanism is threshold recalibration rather than evidence acquisition.
-
Step-Tagging: Toward controlling the generation of Language Reasoning Models through step monitoring
Monitoring the semantic type of reasoning steps with a lightweight BERT classifier can drive interpretable early stopping that cuts generation tokens by 20-50% at modest accuracy cost.
-
Rethinking Human Preference Evaluation of LLM Rationales
A fine-grained attribute-based evaluation of LLM rationales can explain human preferences and reveal model trade-offs that binary comparisons obscure.
-
VIS-Shepherd: Constructing Critic for LLM-based Data Visualization Generation
A 7-billion-parameter multimodal model fine-tuned on 2,500 expert critiques of data visualizations matches or beats much larger models at identifying visualization defects.
-
Knowledge or Reasoning? A Close Look at How LLMs Think Across Domains
LLM reasoning can be scored separately for knowledge and step-by-step information gain, and doing so shows SFT and RL affect these two capacities differently across medicine and math.
-
ARB: A Comprehensive Arabic Multimodal Reasoning Benchmark
ARB provides 1,356 Arabic multimodal questions with 5,119 human-reviewed reasoning steps and shows leading models score much higher on reasoning fluency than on correct answers.
-
Diagnosing Pathological Chain-of-Thought in Reasoning Models
Three log-probability-difference metrics — Necessity, Paraphrasability, Substantivity — are proposed and tested on deliberately fine-tuned 'model organisms' to detect post-hoc, encoded, and internalized chain-of-thoug...
-
Chain-of-Code Collapse: Reasoning Failures in LLMs via Adversarial Prompting in Code Generation
Prompt rewrites of LeetCode problems cause large accuracy swings in nine LLMs, but invalid negation test cases and inconsistent tables make the headline numbers unreliable.
-
Chain-of-Thought for Autonomous Driving: A Comprehensive Survey and Future Prospects
A survey that classifies chain-of-thought methods for autonomous driving into modular, logical, and reflective pipelines, and proposes three evolutionary stages from direct prompting to reinforcement learning.
Discussion (0). Continue with ORCID to comment.