Pith. sign in

URL http://ieeexplore.ieee.org/document/ 7900006/

2 Pith papers cite this work, alongside 1,193 external citations. Polarity classification is still indexing.

2 Pith papers citing it
1,193 external citations · external index

citation-role summary

method 1

citation-polarity summary

fields

cs.AI 1 cs.LG 1

years

2026 2

verdicts

UNVERDICTED 2

roles

method 1

polarities

use method 1

representative citing papers

The Shape of Overthinking: Backtracking Bursts in Long Reasoning Traces

cs.AI · 2026-05-27 · unverdicted · novelty 7.0

On 6000 Qwen3-8B AIME traces, late-clustered moderate-to-severe backtracks are more common in incorrect outputs, enabling prefix-causal burst-aware filtering that outperforms fixed-length cutoffs at shallow and intermediate depths.

Sparse Layers are Critical to Scaling Looped Language Models

cs.LG · 2026-05-09 · unverdicted · novelty 5.0 · 2 refs

Looped-MoE models scale better than dense looped or standard transformers because routing changes across loops, and they enable stronger compute-quality trade-offs via early exits at loop boundaries.

citing papers explorer

Showing 2 of 2 citing papers.

  • The Shape of Overthinking: Backtracking Bursts in Long Reasoning Traces cs.AI · 2026-05-27 · unverdicted · none · ref 5

    On 6000 Qwen3-8B AIME traces, late-clustered moderate-to-severe backtracks are more common in incorrect outputs, enabling prefix-causal burst-aware filtering that outperforms fixed-length cutoffs at shallow and intermediate depths.

  • Sparse Layers are Critical to Scaling Looped Language Models cs.LG · 2026-05-09 · unverdicted · none · ref 8 · 2 links

    Looped-MoE models scale better than dense looped or standard transformers because routing changes across loops, and they enable stronger compute-quality trade-offs via early exits at loop boundaries.