Pith. sign in

hub

The swe-bench illusion: When state-of-the-art llms remember instead of reason

20 Pith papers cite this work. Polarity classification is still indexing.

20 Pith papers citing it

hub tools

citation-role summary

other 2 background 1

citation-polarity summary

years

2026 19 2025 1

polarities

unclear 2 background 1

representative citing papers

Memorization Diagnostics for Code LLMs Should be Scale-Aware

cs.SE · 2026-08-13 · conditional · novelty 7.0

Encoder-side synonym fuzzing and decoder-side log-likelihood probes lose discriminative power on large dense code LLMs, while reversible I/O transforms show scaled models preserve algorithmic structure and fail mainly on output serialization.

Reproduction Test Generation for Java SWE Issues

cs.SE · 2026-05-05 · unverdicted · novelty 6.0 · 2 refs

Introduces the first benchmark for Java reproduction test generation from repository issues and adapts a prior Python tool to produce high performance on it.

Diagnosing CFG Interpretation in LLMs

cs.AI · 2026-04-22 · unverdicted · novelty 6.0

LLMs maintain surface syntax for novel CFGs but fail to preserve semantics under recursion and branching, relying on keyword bootstrapping rather than pure symbolic reasoning.

citing papers explorer

Showing 20 of 20 citing papers.