Pith. sign in

REVIEW 13 cited by

PonderNet: Learning to Ponder

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2107.05407 v2 pith:66CPJOCB submitted 2021-07-12 cs.LG cs.AIcs.CC

classification cs.LGcs.AIcs.CC
keywords pondernetcomputationnetworksneuralproblemamountcomplexcomplexity
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

In standard neural networks the amount of computation used grows with the size of the inputs, but not with the complexity of the problem being learnt. To overcome this limitation we introduce PonderNet, a new algorithm that learns to adapt the amount of computation based on the complexity of the problem at hand. PonderNet learns end-to-end the number of computational steps to achieve an effective compromise between training prediction accuracy, computational cost and generalization. On a complex synthetic problem, PonderNet dramatically improves performance over previous adaptive computation methods and additionally succeeds at extrapolation tests where traditional neural networks fail. Also, our method matched the current state of the art results on a real world question and answering dataset, but using less compute. Finally, PonderNet reached state of the art results on a complex task designed to test the reasoning capabilities of neural networks.1

Discussion (0). Sign in to comment.

Forward citations

Cited by 13 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Adaptive Compute in Latent World Models: When Depth Helps, Hurts, or Doesn't Matter

    cs.LG 2026-07 unverdicted novelty 8.0 of 10

    Predictor depth survives autoregressive composition on 6/9 DMC tasks but inverts on 2/9 because per-step deep supervision trains shallow exits to out-roll the full stack (routability catch-22).

  2. The Equilibrium Is the Initialization: Lazy Identity Collapse in Physics-Structured Deep Equilibrium Reasoning

    cs.LG 2026-07 accept novelty 7.0 of 10

    In a port-Hamiltonian DEQ with learned initialization, the equilibrium equals the start to numerical precision and contributes +0.00 pp accuracy in 18 of 19 runs; a four-test diagnostic exposes the no-op.

  3. Turbo Connection: Reasoning as Information Flow from Higher to Lower Layers

    cs.LG 2026-02 conditional novelty 7.0 of 10

    An architecture that routes hidden states from higher to lower layers between consecutive tokens improves reasoning accuracy and length generalization in fine-tuned LLMs.

  4. Scaling Latent Reasoning via Looped Language Models

    cs.CL 2025-10 unverdicted novelty 7.0 of 10

    Looped language models with latent iterative computation and entropy-regularized depth allocation achieve performance matching up to 12B standard LLMs through superior knowledge manipulation.

  5. LoopMTP: A looped transformer guided by latent multi-token prediction

    cs.CL 2026-08 conditional novelty 6.0 of 10

    Aligning each loop iteration's hidden state with a future token's embedding improves looped transformer accuracy by up to 8.1% relative over a non-looped baseline.

  6. Dynamic Parameterization Is Not Dynamic Inference

    cs.LG 2026-07 conditional novelty 6.0 of 10

    Dynamic parameterization alone does not establish dynamic inference or computational savings, as demonstrated by a frozen-controller audit on FeatureGate and MUDDPythia.

  7. Per-Token Fixed-Point Convergence in Depth-Recurrent Transformers

    cs.AI 2026-07 conditional novelty 6.0 of 10

    In a depth-recurrent transformer, each token converges to a fixed point at its own rate; a parameter-free early-exit rule reads this and matches depth-8 quality at 4.94 average loops, beating a learned router.

  8. DeepLoop: Depth Scaling for Looped Transformers

    cs.LG 2026-07 conditional novelty 6.0 of 10

    Looped Transformers need residual-scaling exponent p=1/2 instead of DeepNorm's 1/4 when shared blocks are revisited with aligned visit-wise gradients.

  9. Adaptive Depth in Looped Transformers: Diagnosing Learned Halting Gates and Trajectory Readouts

    cs.LG 2026-07 conditional novelty 6.0 of 10

    In looped transformers, halting-gate failures come mainly from how gate training reshapes the trajectory; fixed-prior depth supervision plus simple confidence readouts yields better accuracy per unit of compute.

  10. Listen, Think, Transcribe: Continuous Latent Test-Time Scaling for ASR

    cs.SD 2026-07 conditional novelty 6.0 of 10

    Two small modules enable continuous latent test-time refinement on a frozen ASR backbone, cutting error on hard speech under a 500-utterance regime where fine-tuning, LoRA and prompt tuning all regress.

  11. Latency Coding for Efficient and Low-Latency Deep Spiking Neural Networks

    cs.NE 2026-03 conditional novelty 5.0 of 10

    Deep TTFS/latency-coded SNNs can be trained directly with backpropagation, reaching ~93.6% on CIFAR-10 with an average inference latency near one timestep.

  12. Entropy-Guided Loop: Achieving Reasoning through Uncertainty-Aware Generation

    cs.AI 2025-08 conditional novelty 5.0 of 10

    A lightweight entropy-triggered refinement loop improves a small LLM's answer quality to roughly 95% of a reasoning model's, at about one-third the cost.

  13. Change of Thought: Adaptive Test-Time Computation

    cs.LG 2025-07 reject novelty 4.0 of 10

    A transformer layer that iteratively refines its attention matrix to a fixed point is claimed to improve accuracy with no extra parameters, but the benchmark evidence is not reproducible.

Pith tools