Pith. sign in

REVIEW 3 cited by

Combining Induction and Transduction for Abstract Reasoning

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2411.02272 v4 pith:6LMGASGW submitted 2024-11-04 cs.LG cs.AIcs.CL

classification cs.LGcs.AIcs.CL
keywords testneuraltrainingtransductionbetterconceptsdirectlyexamples
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

When learning an input-output mapping from very few examples, is it better to first infer a latent function that explains the examples, or is it better to directly predict new test outputs, e.g. using a neural network? We study this question on ARC by training neural models for induction (inferring latent functions) and transduction (directly predicting the test output for a given test input). We train on synthetically generated variations of Python programs that solve ARC training tasks. We find inductive and transductive models solve different kinds of test problems, despite having the same training problems and sharing the same neural architecture: Inductive program synthesis excels at precise computations, and at composing multiple concepts, while transduction succeeds on fuzzier perceptual concepts. Ensembling them approaches human-level performance on ARC.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Re:Form -- Reducing Human Annotations in Scalable Formal Software Verification with RL in LLMs: A Preliminary Study on Dafny

    cs.CL 2025-07 conditional novelty 6.0 of 10

    A minimal-prior pipeline with automated data curation and verifier-driven RL lets small LLMs generate verifiable Dafny specifications and beat larger proprietary models on a synthetic compositional benchmark.

  2. EasyARC: Evaluating Vision Language Models on True Visual Reasoning

    cs.CV 2025-06 conditional novelty 6.0 of 10

    EasyARC is a new procedurally generated visual reasoning benchmark where state-of-the-art vision-language models score below 20%, despite tasks designed to be easy.

  3. GIFARC: Synthetic Dataset for Leveraging Human-Intuitive Analogies to Elevate AI Reasoning

    cs.AI 2025-05 conditional novelty 6.0 of 10

    The authors build a pipeline that converts GIFs into ARC-style puzzles with analogy labels and executable solutions, and report small in-context experiments suggesting the analogy labels shift an LLM's stated reasoning style.

Pith tools