Pith. sign in

REVIEW 1 cited by

STREET: A Multi-Task Structured Reasoning and Explanation Benchmark

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2302.06729 v1 pith:CA2TI6JP submitted 2023-02-13 cs.CL cs.AI

classification cs.CLcs.AI
keywords reasoninglanguagemodelsstructuredanswerbenchmarkexplanationexplanations
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

We introduce STREET, a unified multi-task and multi-domain natural language reasoning and explanation benchmark. Unlike most existing question-answering (QA) datasets, we expect models to not only answer questions, but also produce step-by-step structured explanations describing how premises in the question are used to produce intermediate conclusions that can prove the correctness of a certain answer. We perform extensive evaluation with popular language models such as few-shot prompting GPT-3 and fine-tuned T5. We find that these models still lag behind human performance when producing such structured reasoning steps. We believe this work will provide a way for the community to better train and test systems on multi-step reasoning and explanations in natural language.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Evaluating LLM Reasoning in the Operations Research Domain with ORQA

    cs.CL 2024-12 conditional novelty 6.0 of 10

    ORQA is a new 1,513-question multiple-choice benchmark showing that open-source LLMs score up to 77% on operations research modeling questions, well below a 93% expert baseline.

Pith tools