Learning Compositional Neural Programs with Recursive Tree Search and Planning

Alexandre Laterre; David Kas; Guillaume Ligner; Karim Beguir; Nando de Freitas; Nicolas Perrin; Olivier Sigaud; Scott Reed; Thomas Pierrot

arxiv: 1905.12941 · v2 · pith:UEYFIZHRnew · submitted 2019-05-30 · 💻 cs.AI

Learning Compositional Neural Programs with Recursive Tree Search and Planning

Thomas Pierrot , Guillaume Ligner , Scott Reed , Olivier Sigaud , Nicolas Perrin , Alexandre Laterre , David Kas , Karim Beguir

show 1 more author

Nando de Freitas

This is my paper

classification 💻 cs.AI

keywords alphanpineuralspecificationalphazerocontributesexecutionformlearning

0 comments

read the original abstract

We propose a novel reinforcement learning algorithm, AlphaNPI, that incorporates the strengths of Neural Programmer-Interpreters (NPI) and AlphaZero. NPI contributes structural biases in the form of modularity, hierarchy and recursion, which are helpful to reduce sample complexity, improve generalization and increase interpretability. AlphaZero contributes powerful neural network guided search algorithms, which we augment with recursion. AlphaNPI only assumes a hierarchical program specification with sparse rewards: 1 when the program execution satisfies the specification, and 0 otherwise. Using this specification, AlphaNPI is able to train NPI models effectively with RL for the first time, completely eliminating the need for strong supervision in the form of execution traces. The experiments show that AlphaNPI can sort as well as previous strongly supervised NPI variants. The AlphaNPI agent is also trained on a Tower of Hanoi puzzle with two disks and is shown to generalize to puzzles with an arbitrary number of disk

This paper has not been read by Pith yet.

discussion (0)

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

Training Transformers as a Universal Computer
cs.AI 2026-04 unverdicted novelty 7.0

A transformer trained on random meaningless MicroPy programs generalizes to execute diverse human-written programs, providing empirical evidence it can act as a universal computer.