Pith. sign in

REVIEW 26 cited by

The Next Decade in AI: Four Steps Towards Robust Artificial Intelligence

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2002.06177 v3 pith:4JU67KZT submitted 2020-02-14 cs.AI cs.LG

The Next Decade in AI: Four Steps Towards Robust Artificial Intelligence

classification cs.AI cs.LG
keywords artificialintelligencelearningrobustapproacharoundcenteredcognitive
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved
0 comments
Share X Bluesky LinkedIn Reddit HN
read the original abstract

Recent research in artificial intelligence and machine learning has largely emphasized general-purpose learning and ever-larger training sets and more and more compute. In contrast, I propose a hybrid, knowledge-driven, reasoning-based approach, centered around cognitive models, that could provide the substrate for a richer, more robust AI than is currently possible.

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Forward citations

Cited by 26 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. PAL: Program-aided Language Models

    cs.CL 2022-11 conditional novelty 8.0

    PAL improves few-shot reasoning accuracy by having LLMs generate executable programs rather than text-based chains of thought, outperforming much larger models on math and logic benchmarks.

  2. PhyEditBench: A Real-World Multi-Stage Benchmark for Physics-Aware Image Editing

    cs.CV 2026-06 unverdicted novelty 7.0

    PhyEditBench is a new benchmark for physics-aware image editing with real and synthetic instances plus a training-free PhyWorld baseline that uses test-time scaling to outperform SOTA models.

  3. PhyEditBench: A Real-World Multi-Stage Benchmark for Physics-Aware Image Editing

    cs.CV 2026-06 unverdicted novelty 7.0

    PhyEditBench is a new benchmark with real-world and synthetic instances that reveals limitations in current image editing models' physics reasoning and proposes a video-generation-based baseline called PhyWorld.

  4. Invariant Gradient Alignment for Robust Reasoning Distillation

    cs.LG 2026-06 unverdicted novelty 7.0

    Invariant Gradient Alignment uses Logical Isomer Sets and a Continuous Gradient Conflict Mask to tighten OOD generalization bounds and boost empirical performance over ERM in reasoning distillation.

  5. Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks

    cs.CL 2020-05 accept novelty 7.0

    RAG models set new state-of-the-art results on open-domain QA by retrieving Wikipedia passages and conditioning a generative model on them, while also producing more factual text than parametric baselines.

  6. Bounded by Risk, Not Capability: Quantifying AI Occupational Substitution Rates via a Tech-Risk Dual-Factor Model

    cs.CY 2026-04 conditional novelty 6.5

    AI occupational substitution is bounded more by commercial risk and liability than technical capability, producing high OAI for data scientists and near-zero for physical trades.

  7. Structural Certification for Reliable Physical Design with Language Models

    cs.AI 2026-06 unverdicted novelty 6.0

    PHACT uses a propose-certify loop with deterministic derivation from fixed inputs to achieve zero false certifications in 80 adversarial trials across two models and temperatures.

  8. The Evaluation Trap: Benchmark Design as Theoretical Commitment

    cs.AI 2026-05 unverdicted novelty 6.0

    AI benchmarks trap progress by operationalizing assumptions that redefine capabilities around the benchmarks themselves, and Epistematics provides an audit procedure to detect when evaluations cannot discriminate clai...

  9. Towards Lawful Autonomous Driving: Deriving Scenario-Aware Driving Requirements from Traffic Laws and Regulations

    cs.AI 2026-04 unverdicted novelty 6.0

    Grounding LLMs via node-wise anchors in a traffic scenario taxonomy improves law-scenario matching by 29.1% and derived requirement accuracy by 36.9-38.2% on Chinese laws and 5,897 scenarios, enabling a compliance lay...

  10. Bounded by Risk, Not Capability: Quantifying AI Occupational Substitution Rates via a Tech-Risk Dual-Factor Model

    cs.CY 2026-04 unverdicted novelty 6.0

    AI job substitution rates are limited by business risks such as liability and compliance rather than technical capability alone, resulting in high exposure for cognitive roles like data scientists and resilience for p...

  11. SteuerLLM: Local specialized large language model for German tax law analysis

    cs.CL 2026-02 reject novelty 6.0

    A tax-specialized 28B model beats larger general-purpose LLMs on a new authentic German tax-law exam benchmark, but its edge may be inflated by overlap between training and evaluation exams.

  12. ActivationReasoning: Logical Reasoning in Latent Activation Spaces

    cs.LG 2025-10 unverdicted novelty 6.0

    ActivationReasoning grounds logical reasoning in LLM latent activations via SAEs to enable structured inference, concept composition, and behavior steering on multi-hop, abstraction, and safety tasks.

  13. To Use AI as Dice of Possibilities with Timing Computation

    cs.AI 2026-05 unverdicted novelty 5.0

    Proposes verb-based paradigm with timing computation to enable data-driven discovery of patient trajectories and counterfactual timing from EHR data without domain knowledge.

  14. To Use AI as Dice of Possibilities with Timing Computation

    cs.AI 2026-05 unverdicted novelty 5.0

    A timing-aware possibility-space framework is claimed to enable automatic trajectory discovery and counterfactual timing deduction on 3,276 breast-cancer EHR patients.

  15. To Use AI as Dice of Possibilities with Timing Computation

    cs.AI 2026-05 reject novelty 5.0

    The paper defines possibility space, timing computation, and causal factum to make timing a computable variable, and illustrates the framework with automatic trajectory discovery and counterfactual timing on 3,276 bre...

  16. When Models Meet Users: An Empirical Study of Perceptions of General LLMs and Multimodal LLMs on Hugging Face

    cs.SE 2026-04 conditional novelty 5.0

    Across 662 annotated Hugging Face threads, gated access (dominated by Llama), multimodal generation quality, and deployment/invocation complexity are the most prominent user concerns.

  17. How Psychological Learning Paradigms Shaped and Constrained Artificial Intelligence

    cs.CL 2026-03 unverdicted novelty 5.0

    AI's compositional reasoning failures originate in psychological learning paradigms that shaped its architectures, and the ReSynth trimodular framework is proposed to embed systematicity structurally.

  18. Distilling Counterfactual Reasoning from Language to Vision: Causal Graph Guided Post-Training for Video Understanding

    cs.CV 2025-11 reject novelty 5.0

    A new video benchmark and post-training recipe claim to improve VLMs' counterfactual 'what if' reasoning, but the reported gains likely come from training on the test set.

  19. Enhancing Causal Reasoning in Large Language Models: A Causal Attribution Model for Precision Fine-Tuning

    cs.AI 2023-12 unverdicted novelty 5.0

    A causal attribution model is proposed that applies do-operators to quantify component contributions in LLMs' causal reasoning, motivating a fine-tuned model for pairwise causal discovery that combines knowledge and n...

  20. Beyond Post-hoc Explanation: Toward Glassbox AI via Probabilistic Mediation

    cs.AI 2026-06 unverdicted novelty 4.0

    The paper proposes the Glassbox Framework in which Bayesian networks serve as transparent ante-hoc mediation layers for generative models to enable auditable reasoning traces and contestable outputs.

  21. To Use AI as Dice of Possibilities with Timing Computation

    cs.AI 2026-05 unverdicted novelty 4.0

    Proposes possibility space, timing computation, and causal factum as a new framework for data-driven trajectory discovery and counterfactual timing deduction on EHR data from 3,276 breast cancer patients.

  22. When Models Meet Users: An Empirical Study of Perceptions of General LLMs and Multimodal LLMs on Hugging Face

    cs.SE 2026-04 unverdicted novelty 4.0

    Hugging Face discussions show that access barriers, output quality, and setup complexity are the main user concerns for both general and multimodal LLMs.

  23. Beyond Explainable AI (XAI): An Overdue Paradigm Shift and Post-XAI Research Directions

    cs.CY 2026-02 reject novelty 4.0

    A position paper argues that post-hoc XAI explanations are unfaithful and paradoxical, proposing a shift to expert-based verification and certification of AI systems.

  24. Beyond Explainable AI (XAI): An Overdue Paradigm Shift and Post-XAI Research Directions

    cs.CY 2026-02 unverdicted novelty 4.0

    Current XAI methods for DNNs and LLMs rest on paradoxes and false assumptions that demand a paradigm shift to verification protocols, scientific foundations, context-aware design, and faithful model analysis rather th...

  25. Agent AI: Surveying the Horizons of Multimodal Interaction

    cs.AI 2024-01 unverdicted novelty 4.0

    The paper defines Agent AI as interactive multimodal systems that perceive grounded data and generate embodied actions, arguing this approach can mitigate hallucinations in foundation models.

  26. Beyond Context: Large Language Models' Failure to Grasp Users' Intent

    cs.AI 2025-12 unverdicted novelty 3.0

    LLMs fail to detect hidden harmful intent, allowing systematic bypass of safety mechanisms through framing techniques, with reasoning modes often worsening the issue.