Pith. sign in

REVIEW 2 cited by

First Tragedy, then Parse: History Repeats Itself in the New Era of Large Language Models

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2311.05020 v2 pith:6VPWAHAH submitted 2023-11-08 cs.CL

classification cs.CL
keywords firstlargellmsmodelsresearchersstillidentifylanguage
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
abstract

Many NLP researchers are experiencing an existential crisis triggered by the astonishing success of ChatGPT and other systems based on large language models (LLMs). After such a disruptive change to our understanding of the field, what is left to do? Taking a historical lens, we look for guidance from the first era of LLMs, which began in 2005 with large $n$-gram models for machine translation (MT). We identify durable lessons from the first era, and more importantly, we identify evergreen problems where NLP researchers can continue to make meaningful contributions in areas where LLMs are ascendant. We argue that disparities in scale are transient and researchers can work to reduce them; that data, rather than hardware, is still a bottleneck for many applications; that meaningful realistic evaluation is still an open problem; and that there is still room for speculative approaches.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. LLM-Rubric: A Multidimensional, Calibrated Approach to Automated Evaluation of Natural Language Texts

    cs.CL 2024-12 accept novelty 6.0 of 10

    A personalized calibration network that combines LLM answers to multiple rubric questions predicted human judges' overall satisfaction scores on dialogues about twice as accurately as the uncalibrated LLM.

  2. MEQA: A Meta-Evaluation Framework for Question & Answer LLM Benchmarks

    cs.CL 2025-04 conditional novelty 5.0 of 10

    MEQA scores eight cybersecurity QA benchmarks against a 44-sub-criteria rubric, finding strengths in reproducibility and comparability and weaknesses in prompt robustness and reliability.

Pith tools