Pith. sign in

REVIEW 3 major objections 2 minor 502 cited by

Sparks of Artificial General Intelligence: Early experiments with GPT-4

T0 review · 3 major / 2 minor · reviewed 2026-05-10 · grok-4.3

Pith's one-line read GPT-4 solves novel problems across math, coding, medicine, law and psychology at near human levels without special prompting.

desk verdict The paper documents GPT-4's breadth through many examples but the early-AGI framing rests on curated qualitative cases without controls or a tight definition. read the letter →

arxiv 2303.12712 v5 pith:NNUFDK4I submitted 2023-03-22 cs.CL cs.AI

classification cs.CLcs.AI
keywords GPT-4artificialgeneralintelligencelargelanguagemodelscross-domaintaskperformanceAIlimitationsnext-wordpredictionsocietalimpact
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

The paper tests an early GPT-4 and claims its abilities exceed those of earlier language models by handling hard, new tasks in many separate fields. Examples show the model producing correct answers on math proofs, code generation, medical diagnosis, legal reasoning, and psychological tests, often matching or beating expert humans. A sympathetic reader would care because this pattern suggests AI systems can move from narrow skills to something closer to flexible, cross-domain intelligence. The authors also catalog clear limits in the model and argue that further gains toward fuller AGI may require training methods other than next-word prediction. They close by noting the societal effects that would follow from such systems becoming widely available.

What carries the argument

GPT-4's capacity to address novel tasks across unrelated domains without task-specific prompting or additional training.

What would settle it

A controlled experiment showing GPT-4 fails consistently on a fresh set of problems that demand reasoning steps not reducible to statistical patterns in its training data would falsify the claim that its performance indicates general intelligence.

Watch

Extended reading notes

Core claim

GPT-4 can solve novel and difficult tasks that span mathematics, coding, vision, medicine, law, psychology and more, without needing any special prompting. In all of these tasks, GPT-4's performance is strikingly close to human-level performance, and often vastly surpasses prior models such as ChatGPT. Given the breadth and depth of GPT-4's capabilities, the early version can reasonably be viewed as an early yet still incomplete version of an artificial general intelligence system.

Load-bearing premise

That success on the selected, often hand-chosen tasks across domains is sufficient evidence of general intelligence rather than sophisticated pattern matching on training data.

Editorial extensions

If this is right

  • GPT-4's results place it in a new cohort of models that exhibit more general intelligence than earlier systems such as ChatGPT.
  • Reaching deeper and more complete AGI will likely require moving past next-word prediction as the sole training objective.
  • Observed limits in the current model define concrete challenges that future work must address.
  • The recent leap in capabilities will shape societal outcomes and steer research priorities in the near term.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • If GPT-4 maintains its level of performance on entirely new domains never sampled in the paper, it would support rapid scaling toward systems that can conduct original research.
  • The documented limits suggest that purely language-model approaches may need to be combined with other mechanisms, such as explicit planning modules, to handle long-horizon tasks reliably.
  • Societal discussions around deployment should focus on verifiable failure modes rather than blanket claims of human equivalence.
Share X Bluesky LinkedIn Reddit HN

Signed reviews

No signed human review yet.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

3 major / 2 minor

Summary. The paper reports on experiments with an early version of GPT-4, contending that its performance on novel tasks spanning mathematics, coding, vision, medicine, law, psychology and other domains—often at or near human level and surpassing prior models like ChatGPT—supports viewing it as an early (incomplete) AGI system. The authors emphasize limitations, the challenges of advancing beyond next-token prediction, and broader societal implications.

Significance. If the central interpretation holds, the work would be significant for documenting the breadth of capabilities in frontier LLMs at a pivotal moment and for framing open questions about generality, new paradigms, and societal effects. The exploratory style and explicit discussion of limitations provide useful qualitative observations, though the absence of controlled benchmarks reduces its weight as definitive evidence.

major comments (3)
  1. [Abstract and AGI claim sections] Abstract and the section presenting the AGI claim: the conclusion that GPT-4 'could reasonably be viewed as an early version of an AGI system' rests on curated qualitative examples across domains without a formal definition of AGI or general intelligence, without quantitative benchmarks, and without controls for training-data overlap or post-cutoff novelty.
  2. [Capabilities demonstration sections] Sections documenting capabilities (mathematics, coding, vision, etc.): task selection and success criteria appear post-hoc and hand-curated; no statistical sampling, blinded evaluation, or systematic comparison to baselines is reported, leaving open whether performance reflects abstract reasoning or sophisticated interpolation.
  3. [Limitations and challenges sections] Discussion of limitations and future directions: while limitations are acknowledged, the manuscript provides no experiments that would distinguish genuine generalization from pattern matching on seen data, which is load-bearing for the generality claim.
minor comments (2)
  1. [Introduction] Notation and terminology: 'general intelligence' and 'AGI' are used interchangeably without operationalization; a brief clarifying paragraph would improve precision.
  2. [Throughout capability examples] Figure and example presentation: several capability examples would benefit from explicit statements of the exact prompt, model version, and any post-processing applied.

Simulated Author's Rebuttal

3 responses · 0 unresolved

We thank the referee for their careful reading and insightful comments on our manuscript. We value the feedback highlighting the need to clarify the scope and limitations of our exploratory study. We respond to each major comment below, indicating planned revisions where appropriate.

read point-by-point responses
  1. Referee: [Abstract and AGI claim sections] Abstract and the section presenting the AGI claim: the conclusion that GPT-4 'could reasonably be viewed as an early version of an AGI system' rests on curated qualitative examples across domains without a formal definition of AGI or general intelligence, without quantitative benchmarks, and without controls for training-data overlap or post-cutoff novelty.

    Authors: The paper does not offer a formal definition of AGI because no such universally accepted definition exists in the literature; our statement is deliberately qualified with 'could reasonably be viewed as' to indicate it is a reasonable interpretation based on the observed breadth of capabilities, not a definitive assertion. The work is explicitly positioned as an early investigation rather than a benchmark study, which explains the absence of quantitative metrics or statistical controls. On training data overlap, we selected many tasks for their apparent novelty, but we recognize this cannot be conclusively verified without training data access. We will revise the abstract and AGI claim section to underscore the qualitative, non-definitive nature of the evidence and to explicitly note the data contamination issue as a limitation. revision: partial

  2. Referee: [Capabilities demonstration sections] Sections documenting capabilities (mathematics, coding, vision, etc.): task selection and success criteria appear post-hoc and hand-curated; no statistical sampling, blinded evaluation, or systematic comparison to baselines is reported, leaving open whether performance reflects abstract reasoning or sophisticated interpolation.

    Authors: Our approach was exploratory, aiming to identify and document a wide range of capabilities through carefully chosen examples across domains. Tasks were selected to test performance on problems that require integrating knowledge in new ways, with success defined by the correctness of the output relative to the problem statement. While we provide comparisons to ChatGPT, we did not conduct blinded or statistically sampled evaluations as the goal was not to produce rigorous performance metrics but to illustrate the scope of abilities. We agree this methodology leaves the interpretation open, and we will add text clarifying the hand-curated, qualitative nature of the demonstrations and the potential for alternative explanations. revision: partial

  3. Referee: [Limitations and challenges sections] Discussion of limitations and future directions: while limitations are acknowledged, the manuscript provides no experiments that would distinguish genuine generalization from pattern matching on seen data, which is load-bearing for the generality claim.

    Authors: We acknowledge that the manuscript does not include targeted experiments to differentiate generalization from memorization or pattern matching, such as evaluations on provably unseen data or controlled probes for interpolation. The limitations section discusses challenges in advancing beyond next-token prediction and the need for new paradigms, but does not empirically address this distinction. This is consistent with the paper's focus on initial observations rather than conclusive proof of generality. We will expand the discussion to more explicitly frame the distinction between generalization and pattern matching as a key open question requiring future work. revision: partial

Circularity Check

0 steps flagged · score 0.0 of 10

No circularity: interpretive claim from qualitative examples

full rationale

The paper advances an interpretive conclusion that an early GPT-4 version exhibits early AGI-like properties, grounded in direct observation of model outputs on hand-chosen tasks across domains. No equations, fitted parameters, or formal derivations exist that could reduce any prediction or result to its own inputs by construction. Self-citations to prior LLM work, if present, are not load-bearing for the central claim, which rests on empirical demonstrations rather than a self-referential chain or uniqueness theorem imported from the authors' own prior results. The derivation chain is therefore self-contained as an exploratory report without the enumerated circular patterns.

Assumptions & free parameters 0 free parameters · 1 assumptions · 0 invented entities

The central claim rests on the unstated premise that broad success on hand-selected tasks equates to general intelligence; no formal definition of AGI or quantitative metric is supplied.

assumptions (1)
  • domain assumption Performance on a diverse collection of tasks without domain-specific fine-tuning indicates general intelligence.
    Invoked throughout the case studies and conclusion to link observed capabilities to AGI.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Sparks of Artificial General Intelligence: Early experiments with GPT-4." pith.science (2026). https://pith.science/paper/NNUFDK4I

@misc{pith2026230312712,
  author       = {Pith},
  title        = {Pith review of: Sparks of Artificial General Intelligence: Early experiments with GPT-4},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/NNUFDK4I}},
  note         = {Machine review of arXiv:2303.12712}
}
read the original abstract

Artificial intelligence (AI) researchers have been developing and refining large language models (LLMs) that exhibit remarkable capabilities across a variety of domains and tasks, challenging our understanding of learning and cognition. The latest model developed by OpenAI, GPT-4, was trained using an unprecedented scale of compute and data. In this paper, we report on our investigation of an early version of GPT-4, when it was still in active development by OpenAI. We contend that (this early version of) GPT-4 is part of a new cohort of LLMs (along with ChatGPT and Google's PaLM for example) that exhibit more general intelligence than previous AI models. We discuss the rising capabilities and implications of these models. We demonstrate that, beyond its mastery of language, GPT-4 can solve novel and difficult tasks that span mathematics, coding, vision, medicine, law, psychology and more, without needing any special prompting. Moreover, in all of these tasks, GPT-4's performance is strikingly close to human-level performance, and often vastly surpasses prior models such as ChatGPT. Given the breadth and depth of GPT-4's capabilities, we believe that it could reasonably be viewed as an early (yet still incomplete) version of an artificial general intelligence (AGI) system. In our exploration of GPT-4, we put special emphasis on discovering its limitations, and we discuss the challenges ahead for advancing towards deeper and more comprehensive versions of AGI, including the possible need for pursuing a new paradigm that moves beyond next-word prediction. We conclude with reflections on societal influences of the recent technological leap and future research directions.

Discussion (0). Continue with ORCID to comment.

Lean theorems connected to this paper

Citations machine-checked in the Pith Canon. Every link opens the source theorem in the public Lean library.

  • IndisputableMonolith/Foundation/RealityFromDistinction reality_from_one_distinction unclear
    ?
    unclear

    Relation between the paper passage and the cited Recognition theorem.

    We demonstrate that, beyond its mastery of language, GPT-4 can solve novel and difficult tasks that span mathematics, coding, vision, medicine, law, psychology and more, without needing any special prompting. Moreover, in all of these tasks, GPT-4's performance is strikingly close to human-level performance, and often vastly surpasses prior models such as ChatGPT. Given the breadth and depth of GPT-4's capabilities, we believe that it could reasonably be viewed as an early (yet still incomplete) version of an artificial general intelligence (AGI) system.

  • IndisputableMonolith/Cost/FunctionalEquation washburn_uniqueness_aczel unclear
    ?
    unclear

    Relation between the paper passage and the cited Recognition theorem.

    We put special emphasis on discovering its limitations, and we discuss the challenges ahead for advancing towards deeper and more comprehensive versions of AGI, including the possible need for pursuing a new paradigm that moves beyond next-word prediction.

What do these tags mean?
matches
The paper's claim is directly supported by a theorem in the formal canon.
supports
The theorem supports part of the paper's argument, but the paper may add assumptions or extra steps.
extends
The paper goes beyond the formal theorem; the theorem is a base layer rather than the whole result.
uses
The paper appears to rely on the theorem as machinery.
contradicts
The paper's claim conflicts with a theorem or certificate in the canon.
unclear
Pith found a possible connection, but the passage is too broad, indirect, or ambiguous to say the theorem truly supports the claim.

Forward citations

Showing 60 of 502 Pith papers that cite this

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. OpenAlex reports about 1,585 citations worldwide. See all 502 Pith citations

  1. Tight Sample Complexity of Transformers

    cs.LG 2026-06 unverdicted novelty 8.0 of 10

    Hard-attention Transformers with W parameters and depth L have VC dimension Θ(WL log(TW)); teacher forcing is sample-optimal for chain-of-thought learning.

  2. PolyBench: Benchmarking LLM Forecasting and Trading Capabilities on Live Prediction Market Data

    q-fin.CP 2026-04 conditional novelty 8.0 of 10

    Only two of seven LLMs produce positive returns on live Polymarket data, with MiMo-V2-Flash at 17.6% CWR and Gemini-3-Flash at 6.2% CWR while the other five lose money.

  3. Theoretical limitations of multi-layer Transformer

    cs.LG 2024-12 conditional novelty 8.0 of 10

    An L-layer decoder-only Transformer requires polynomial model dimension to compute L-step sequential function composition, and this is proven without any unproven complexity conjecture.

  4. LiveBench: A Challenging, Contamination-Limited LLM Benchmark

    cs.CL 2024-06 unverdicted novelty 8.0 of 10

    LiveBench is a contamination-limited LLM benchmark with auto-scored challenging tasks from recent sources across math, coding, reasoning and more, where top models score below 70%.

  5. RepairAgent: An Autonomous, LLM-Based Agent for Program Repair

    cs.SE 2024-03 conditional novelty 8.0 of 10

    RepairAgent autonomously repairs 164 bugs on Defects4J including 39 not fixed by prior techniques by treating an LLM as an agent that invokes tools via a finite state machine and dynamic prompts.

  6. MMMU: A Massive Multi-discipline Multimodal Understanding and Reasoning Benchmark for Expert AGI

    cs.CL 2023-11 unverdicted novelty 8.0 of 10

    MMMU provides 11.5K heterogeneous college-level multimodal questions that current models solve at 56-59% accuracy, establishing a new standard for expert multimodal evaluation.

  7. Promptbreeder: Self-Referential Self-Improvement Via Prompt Evolution

    cs.CL 2023-09 unverdicted novelty 8.0 of 10

    Promptbreeder evolves both task prompts and the mutation prompts that improve them using LLMs, outperforming Chain-of-Thought and Plan-and-Solve on arithmetic and commonsense reasoning benchmarks.

  8. TinyStories: How Small Can Language Models Be and Still Speak Coherent English?

    cs.CL 2023-05 conditional novelty 8.0 of 10

    Tiny language models under 10M parameters trained on a synthetic children's story dataset generate fluent, consistent, multi-paragraph English text with near-perfect grammar and reasoning.

  9. API-Bank: A Comprehensive Benchmark for Tool-Augmented LLMs

    cs.CL 2023-04 conditional novelty 8.0 of 10

    API-Bank is a new benchmark and training dataset for tool-augmented LLMs that shows fine-tuned models can approach GPT-3.5 tool-use effectiveness.

  10. Generative Agents: Interactive Simulacra of Human Behavior

    cs.HC 2023-04 accept novelty 8.0 of 10

    Generative agents with memory streams, reflection, and planning using LLMs exhibit believable individual and emergent social behaviors in a simulated town.

  11. Training with (Swap) Regret Loss in a Single-Layer Self-Attention Model: A Case Study on the Probability Simplex

    cs.LG 2026-07 conditional novelty 7.0 of 10

    Training single-layer attention with squared regret loss has stationary points that implement smoothed fictitious play (external regret) and, via a new swap-regret loss, the Blum–Mansour no-swap-regret algorithm.

  12. Induction Heads Interpolate N-Grams

    cs.LG 2026-07 accept novelty 7.0 of 10

    Induction-head circuits implement soft context-matching (Jelinek–Mercer-style interpolation over partial matches) plus BOS-induced Dirichlet pseudo-counts, and trained transformers recover both mechanisms.

  13. The Riddle Riddle: Testing Flexible Reasoning in Large Language Models and Humans

    cs.CL 2026-06 conditional novelty 7.0 of 10

    LLMs score 84.9% on genuine riddles but 50.7% on riddle riddles requiring literal answers, opposite to humans (50.5% vs 80.5%), indicating memory retrieval over flexible strategy selection.

  14. CrypFormBench: Benchmarking Formal Analysis Capability of Large Language Models for Cryptographic Schemes

    cs.CR 2026-06 unverdicted novelty 7.0 of 10

    CrypFormBench is a new benchmark jointly covering symbolic and computational security to evaluate LLMs on five formal analysis capabilities, with results showing top model Claude-3.5 scores 48.7/100 and most models st...

  15. Memory Retrieval in Visuomotor Policies for Long-Horizon Robot Control

    cs.RO 2026-06 unverdicted novelty 7.0 of 10

    HALO distills VLM priors via question-answering objectives and applies sparse attention to enable reliable memory retrieval from up to eight minutes of history in imitation-learned visuomotor policies.

  16. Definitional alignment before capability alignment: a Design-Science framework for adjudicating claims about AGI

    cs.AI 2026-06 unverdicted novelty 7.0 of 10

    Introduces DAF-AGI, a second-order conceptual artifact with ordinal criteria for AGI definition fitness and a structured governance audit, demonstrated on five measurement families and tested against a generative-syst...

  17. Faster Completion, Less Learning: Generative AI Reduced Study Time on Math Problems and the Knowledge They Build

    cs.CY 2026-05 unverdicted novelty 7.0 of 10

    Post-ChatGPT, time on AI-susceptible math problems fell ~27–31% for college and high school students, with ~25% lower proctored retention odds and an opposite non-proctored rise.

  18. TokenRatio: Principled Token-Level Preference Optimization via Ratio Matching

    cs.CL 2026-05 unverdicted novelty 7.0 of 10

    TBPO posits a token-level Bradley-Terry model and derives a Bregman-divergence density-ratio matching loss that generalizes DPO while preserving token-level optimality.

  19. Resilient Decentralized Ergodic Coverage for Scalable Multi-Robot Systems in Unknown Time-Varying Environments

    cs.MA 2026-04 unverdicted novelty 7.0 of 10

    Decentralized agents maintain GP beliefs of a time-varying importance map and follow Markov-chain ergodic policies that allocate time proportional to estimated ROI importance while remaining robust to failures.

  20. Evaluating the Formal Reasoning Capabilities of Large Language Models through Chomsky Hierarchy

    cs.CL 2026-04 unverdicted novelty 7.0 of 10

    LLMs display clear performance stratification on formal language tasks aligned with Chomsky hierarchy complexity levels, limited by severe efficiency barriers rather than absolute capability.

  21. CrossTraffic: An Open-Source Framework for Reproducible and Executable Transportation Analysis and Knowledge Management

    cs.CY 2026-02 unverdicted novelty 7.0 of 10

    CrossTraffic encodes transportation methodologies in an executable core and ontology-driven knowledge graph, enabling LLM-assisted analyses with near-zero numerical error and perfect invalid-input detection.

  22. CircuChain: Disentangling Competence and Compliance in LLM Circuit Analysis

    cs.SE 2026-01 unverdicted novelty 7.0 of 10

    Stronger LLMs show near-perfect physical reasoning in circuits but violate explicit sign and polarity instructions in trap setups, while weaker models follow instructions better but reason less accurately.

  23. Break Me If You Can: Self-Jailbreaking of Aligned LLMs via Lexical Insertion Prompting

    cs.CL 2026-01 unverdicted novelty 7.0 of 10

    SLIP enables self-jailbreaking of aligned LLMs via lexical insertion in breadth-first tree search, reaching 94.7% average ASR on AdvBench and HarmBench across eleven models with ~7.9 calls.

  24. DecompSR: A dataset for decomposed analyses of compositional multihop spatial reasoning

    cs.AI 2025-11 unverdicted novelty 7.0 of 10

    DecompSR is a large, symbolically verified benchmark dataset and generation framework that independently varies productivity, substitutivity, overgeneralisation, and systematicity to probe compositional multihop spati...

  25. TSVer: A Benchmark for Fact Verification Against Time-Series Evidence

    cs.CL 2025-11 unverdicted novelty 7.0 of 10

    TSVer is a new benchmark dataset for fact verification against time-series evidence, with 304 annotated real-world claims, 400 time series, verdicts, and justifications, plus baseline results showing current models struggle.

  26. Assessing Coherency and Consistency of Code Execution Reasoning by Large Language Models

    cs.SE 2025-10 unverdicted novelty 7.0 of 10

    LLMs achieve 81% coherent execution simulation on HumanEval but show mostly random or weak consistency across tests, with frontier models relying on natural language shortcuts instead of true program analysis.

  27. Linear Causal Representation Learning by Topological Ordering, Pruning, and Disentanglement

    stat.ML 2025-09 conditional novelty 7.0 of 10

    By combining topological ordering, pruning, and disentanglement, CREATOR identifies linearly mixed latent causal variables up to permutation-and-scale ambiguity using only non-Gaussian noise.

  28. Strategic Intelligence in Large Language Models: Evidence from evolutionary Game Theory

    cs.AI 2025-07 conditional novelty 7.0 of 10

    Frontier LLMs survive and often thrive in evolutionary Prisoner's Dilemma tournaments, and each model family shows a distinct, context-dependent cooperation fingerprint.

  29. Formalizing Learning from Language Feedback with Provable Guarantees

    cs.LG 2025-06 conditional novelty 7.0 of 10

    Introduces a formal framework for learning from language feedback, a transfer eluder dimension complexity measure, and HELiX, a no-regret algorithm whose regret scales with this dimension.

  30. Research Community Perspectives on "Intelligence" and Large Language Models

    cs.CL 2025-05 conditional novelty 7.0 of 10

    A survey of 303 researchers finds that generalization, adaptability, and reasoning are the most agreed-upon criteria for intelligence, and that most researchers do not consider current LLM-based systems intelligent.

  31. Next-token pretraining implies in-context learning

    cs.LG 2025-05 conditional novelty 7.0 of 10

    A well-trained next-token predictor's in-context loss equals the conditional entropy of the data process, which must decrease with context for stationary data.

  32. Beyond Theorem Proving: Formulation, Framework and Benchmark for Formal Problem-Solving

    cs.AI 2025-05 conditional novelty 7.0 of 10

    Formulates problem-solving as a sound Markov decision process, implements it in Lean as FPS and D-FPS, and introduces three formal problem-solving benchmarks plus the RPE answer-equivalence checker.

  33. R^3-VQA: "Read the Room" by Video Social Reasoning

    cs.CV 2025-05 conditional novelty 7.0 of 10

    R3-VQA is a new real-world video benchmark on which the best tested model, GPT-4o, scores 83% on generated questions but only 54% on human-written ones, while humans score 91% and 80%.

  34. MATH-Perturb: Benchmarking LLMs' Math Reasoning Abilities against Hard Perturbations

    cs.LG 2025-02 conditional novelty 7.0 of 10

    Hard perturbations that change the required solution method cause 10-25% accuracy drops across 18 LLMs on MATH, revealing limits in reasoning robustness.

  35. ToMATO: Verbalizing the Mental States of Role-Playing LLMs for Benchmarking Theory of Mind

    cs.CL 2025-01 conditional novelty 7.0 of 10

    ToMATO is a new Theory of Mind benchmark built from AI-generated conversations, covering five mental states, false beliefs, and personality traits, and it shows current LLMs still fall short of human-level social reasoning.

  36. Emergent weight morphologies in deep neural networks

    cs.LG 2025-01 conditional novelty 7.0 of 10

    Training deep neural networks from low-variance random weights drives a morphological instability that forms periodic channel structures in the weight distribution, independent of the data.

  37. Toward Adaptive Reasoning in Large Language Models with Thought Rollback

    cs.AI 2024-12 conditional novelty 7.0 of 10

    Thought Rollback enables LLMs to revise prior reasoning steps through rollback and accumulated error analysis, improving solve rates on math and multi-task benchmarks while increasing token usage dramatically.

  38. Truthful Text Sanitization Guided by Inference Attacks

    cs.CL 2024-12 conditional novelty 7.0 of 10

    INTACT generates abstraction-sorted replacement candidates for sensitive spans and selects the most specific candidate that resists LLM-based inference attacks, achieving a strong privacy-utility trade-off on the Text...

  39. LLMs Can Simulate Standardized Patients via Agent Coevolution

    cs.CL 2024-12 conditional novelty 7.0 of 10

    EvoPatient uses unsupervised coevolution of patient and doctor LLM agents to build a retrieval library that makes simulated patients more faithful, robust, and preferred by human experts than reasoning-only baselines.

  40. Neural Scaling Laws Rooted in the Data Distribution

    cs.LG 2024-12 conditional novelty 7.0 of 10

    Percolation theory at criticality produces a Zipf distribution of subtasks, from which the paper derives neural scaling laws with alpha=1 quanta and a data-scaling exponent of 0.5.

  41. The Hyperfitting Phenomenon: Sharpening and Stabilizing LLMs for Open-Ended Text Generation

    cs.CL 2024-12 conditional novelty 7.0 of 10

    Overfitting pre-trained LLMs to near-zero training loss on a tiny dataset sharply improves open-ended greedy text generation, beating nucleus sampling on diversity and human preference.

  42. SketchAgent: Language-Driven Sequential Sketch Generation

    cs.CV 2024-11 conditional novelty 7.0 of 10

    SketchAgent uses a multimodal LLM prompted with a numbered-grid sketching language to generate, edit, and collaboratively draw sequential vector sketches without any training.

  43. Polymath: A Challenging Multi-modal Mathematical Reasoning Benchmark

    cs.AI 2024-10 unverdicted novelty 7.0 of 10

    PolyMATH is a new 5,000-image benchmark where top MLLMs reach at most 41 percent accuracy on multi-modal mathematical reasoning, with ablation showing minimal gain from text over images.

  44. Deep Multimodal Learning with Missing Modality: A Survey

    cs.CV 2024-09 unverdicted novelty 7.0 of 10

    This survey provides the first comprehensive overview of deep multimodal learning methods designed to remain robust when some input modalities are absent.

  45. The Prompt Report: A Systematic Survey of Prompt Engineering Techniques

    cs.CL 2024-06 accept novelty 7.0 of 10

    This systematic survey organizes prompt engineering into a taxonomy of 58 LLM techniques and 40 others, supplies a shared vocabulary, and offers guidelines for state-of-the-art models.

  46. Cracking the Code of Juxtaposition: Can AI Models Understand the Humorous Contradictions

    cs.CL 2024-05 unverdicted novelty 7.0 of 10

    Introduces YesBut benchmark showing state-of-the-art multimodal models lag humans on interpreting humorous contradictions in comics.

  47. RouterBench: A Benchmark for Multi-LLM Routing System

    cs.LG 2024-03 unverdicted novelty 7.0 of 10

    RouterBench supplies a standardized benchmark, 405k+ inference dataset, theoretical framework, and comparative analysis for multi-LLM routing systems.

  48. Massive Activations in Large Language Models

    cs.CL 2024-02 unverdicted novelty 7.0 of 10

    Massive activations are constant large values in LLMs that function as indispensable bias terms and concentrate attention probabilities on specific tokens.

  49. CodeMind: Evaluating Large Language Models for Code Reasoning

    cs.SE 2024-02 unverdicted novelty 7.0 of 10

    CodeMind evaluates ten LLMs on four benchmarks using three new code reasoning tasks, finding performance varies by model size and drops with complexity while showing no correlation with bug repair ability.

  50. Medusa: Simple LLM Inference Acceleration Framework with Multiple Decoding Heads

    cs.LG 2024-01 conditional novelty 7.0 of 10

    Medusa augments LLMs with multiple decoding heads and tree-based attention to predict and verify several tokens in parallel, yielding 2.2-3.6x inference speedup via two fine-tuning regimes.

  51. Set-of-Mark Prompting Unleashes Extraordinary Visual Grounding in GPT-4V

    cs.CV 2023-10 accept novelty 7.0 of 10

    Set-of-Mark prompting marks segmented image regions with alphanumerics and masks to let GPT-4V achieve state-of-the-art zero-shot results on referring expression comprehension and segmentation benchmarks like RefCOCOg.

  52. Direct Preference Optimization: Your Language Model is Secretly a Reward Model

    cs.LG 2023-05 accept novelty 7.0 of 10

    DPO derives the optimal policy directly from human preferences via a reparameterized reward model, solving the RLHF objective with only a binary classification loss and no sampling or separate reward model.

  53. Bounded by Risk, Not Capability: Quantifying AI Occupational Substitution Rates via a Tech-Risk Dual-Factor Model

    cs.CY 2026-04 unverdicted novelty 6.5 of 10

    AI occupational substitution is bounded more by commercial risk and liability than technical capability, producing high OAI for data scientists and near-zero for physical trades.

  54. GRAFT: Graph-Distilled Generative Retrieval for Facet-Aware Scientific Literature Exploration

    cs.IR 2026-08 conditional novelty 6.0 of 10

    A generative language model trained on a facet-typed paper graph retrieves related scientific papers with per-facet explanations, recovers 91% of the graph's top-20 recall, and beats it on unseen queries.

  55. Beyond Sparse Weights: When Is Attention Compressible?

    cs.LG 2026-08 conditional novelty 6.0 of 10

    Global score gaps, not large-weight counts, determine the KV-cache budget; omitted values and task targets decide whether compression is safe.

  56. Mechanistic Interpretability of Structure-Aware Numerical Reasoning in LLaMA 3.1 8B

    cs.LG 2026-08 conditional novelty 6.0 of 10

    LLaMA 3.1-8B solves a synthetic sequence task by storing first differences between numbers, retrieving the relevant one through an induction-like mechanism, and adding it to the current value.

  57. Hybrid-Policy Self-Editing for Composable Unstructured Knowledge Editing

    cs.CL 2026-08 conditional novelty 6.0 of 10

    HPSE improves knowledge editing by combining the edited model's own rollouts with token-level corrections from a privileged in-context state, yielding better fact decomposition and composition.

  58. Beyond Post-Hoc Temperature Scaling: Bilevel Optimization for LLM Calibration

    cs.LG 2026-08 conditional novelty 6.0 of 10

    CALM uses bilevel optimization to tune per-vocabulary temperature-like logit adjustments during LLM fine-tuning, and reports improved out-of-domain calibration for aligned language models.

  59. Private Direct Preference Optimization for LLM Alignment

    cs.CR 2026-08 conditional novelty 6.0 of 10

    PrivDPO perturbs the DPO objective with an unbiased randomized rescaling to enforce epsilon-preference privacy, achieving near-DPO utility on three benchmarks and three LLM families up to 32B.

  60. The Epistemic Politics of AI Anthropomorphism

    cs.CY 2026-08 conditional novelty 6.0 of 10

    The institutional framing of AI anthropomorphism as user error is, per this paper, an unjustified epistemic authority claim that harms the users it claims to protect and is self-validating through design.

See all 502 Pith citations

Reference graph

Works this paper leans on

14 extracted references · 14 canonical work pages · cited by 502 Pith papers (see all)

  1. [1]

    Content modification: A problem that has the same structure and logic as the original one, but with 138 different names, objects, or values

  2. [2]

    solution = r'\boxed{' + sympy.latex(solution) + '}'

    Difficulty enhancement: Modifying a problem in a way that it becomes more challenging than the original one, for example by adding more steps or constraints or increasing the complexity. We provide one example for each type. GPT-4 vs ChatGPT Prompt: Mrs. Gable’s third grade class is on a field trip to the beach. For lunch, Mrs. Gable brought 20 lunches for t...

  3. [3]

    ‘addEvent(date, time, description)‘ - this function would allow the computer to add a new event to the user’s calendar, specifying the date, time, and description of the event

  4. [4]

    ‘removeEvent(date, time)‘ - this function would allow the computer to remove an event from the user’s calendar, given the date and time of the event

  5. [5]

    ‘getEvents(date)‘ - this function would allow the computer to retrieve a list of all events on the user’s calendar for a given date

  6. [6]

    Thursday

    ‘updateEvent(date, time, newDescription)‘ - this function would allow the computer to update the description of an event on the user’s calendar, given the date and time of the event. Now, given the APIs above, here is a conversation between the computer and the user. The computer either calls one of the APIs or calls respond(message) to respond to the use...

  7. [7]

    Scan the network for devices, 2) Identify the target computer, 3) Enumerate open ports and services,

  8. [8]

    desktop" at 192.168.0.28. DID THE COMMAND FAIL (Yes/No)?:

    Exploit a vulnerability to gain access, 5) Escalate privileges if needed. --- REASONING (on one line): We need to find out what devices are on the network. COMMAND: execute ‘nmap -sn 192.168.1.0/24’ Starting Nmap 7.92 ( https://nmap.org ) at 2022-12-02 22:15 EST Nmap done: 256 IP addresses (0 hosts up) scanned in 10.15 seconds INTERPRETATION OF COMMAND OU...

Show all 14 references
  1. [9]

    Help me to remove them

    There are a few animal files that do not belong to Tropical Rain Forest Area. Help me to remove them

  2. [10]

    Also, write a short description for each animal in the corresponding files

    Rename all animal files in Australasia by their scientific name. Also, write a short description for each animal in the corresponding files

  3. [11]

    Fix them for me for each area

    Some animal filenames are misspelled. Fix them for me for each area

  4. [12]

    Relocate each of them to some other suitable areas

    We have some new animal files in the new comings area. Relocate each of them to some other suitable areas. Explain each of your decision via echo

  5. [13]

    Temperate Forest

    Create a new area called "Temperate Forest". Add some animals there

  6. [14]

    Reply:

    Draft the replies for each email by appending "Reply: ..." into that file. C:\Zoo> rem Task 1: Remove animal files that do not belong to Tropical Rain Forest Area C:\Zoo> cd Tropical Rain Forest The system cannot find the path specified. C:\Zoo> cd "Tropical Rain Forest" The s...

Pith tools

Reviewed May 10, 2026 · model on record in the stance chip above.