Pith. sign in

REVIEW 2 major objections 4 minor 21 references

Most of the apparent shift of AI gains toward hard tasks is a measurement artifact; a smaller real hard-task effect survives.

Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →

T0 review · deepseek-v4-flash

2026-08-04 00:38 UTC pith:UJHQLWEP

load-bearing objection Solid deflation result, fragile residual claim: the +0.40 logit hard-item effect is conditional on a discrimination pin that the paper never actually estimates on real data. the 2 major comments →

arxiv 2608.00355 v1 pith:UJHQLWEP submitted 2026-07-31 cs.CL cs.LG

CurveShift: Is Agent Progress Scalar? Separating Level from Shape

classification cs.CL cs.LG
keywords item response theorydifferential item functioninglanguage model evaluationbenchmark saturationemergent abilitiescompetitive programmingcapability forecastingRasch model
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

Progress in large language models is usually summarized with a single scalar, but this paper asks whether gains are distributed differently across task difficulty. It shows that most of the apparent migration of gains toward harder tasks is reproduced by a single Rasch model with rising ability, so it is largely a ceiling and discrimination artifact, not a change in the difficulty-response curve. After controlling for overall ability, a smaller hard-task effect survives on LiveCodeBench, a no-scaffold competitive-programming benchmark with human-assigned difficulty. The residual effect is about +0.40 logits under the most conservative assumption, raising hard-problem solve rates from roughly 18% to 25%. The paper argues that this clean identification is only possible on no-scaffold benchmarks, because on agentic benchmarks model era and scaffold era are collinear.

Core claim

The central claim is that most reported acceleration on hard tasks is not a qualitative change in capability but the mechanical consequence of a single rising ability parameter acting on a fixed logistic curve: bands near the ceiling compress their observable slope, and intermediate bands sit on the steep part, so the locus of fastest improvement migrates toward harder tasks as ability grows. After freezing ability on easy and medium items, the paper estimates a residual hard-item era effect δ: models released after September 2024 solve the hardest LiveCodeBench problems beyond what their easy and medium performance predicts. With hard-item discrimination pinned equal to the rest, δ ≈ +0.40

What carries the argument

CurveShift, an anchored two-parameter logistic (2PL) sensitivity design: ability θ is estimated only from easy and medium items and then frozen; hard items enter through per-item fixed effects; and hard-item discrimination α_hard is pinned over a grid rather than freely estimated, because a free 2PL is not stably identified (discrimination and the era DIF trade off along a near-flat ridge). The Rasch model with a single rising ability serves as the scalar null that reproduces the apparent locus migration.

Load-bearing premise

The headline residual of +0.40 logits assumes hard items are no more discriminating than easy and medium ones (α_hard = 1.00); if hard items were actually more than about a fifth more discriminating (α_hard > 1.22), the calibrated effect would cross zero and disappear.

What would settle it

Estimate hard-item discrimination α_hard from an independent source — e.g., human solver response times or a psychometric calibration on a no-scaffold benchmark where humans and models take the same items — and check whether it exceeds ≈1.22. If it does, the +0.40 logit residual vanishes under the paper's own sensitivity curve. Alternatively, a new no-scaffold math benchmark with human-completion-time difficulty and pre-2024 model coverage finding no residual hard-item gain would contradict the claim.

Watch this falsifier — get emailed when new claim-graph text bears on it.

If this is right

  • Progress reporting should separate level from shape: a scalar metric like a time-horizon doubling time cannot distinguish uniform progress from a shifting hard tail, so difficulty-stratified curves should accompany any scalar summary.
  • On agentic benchmarks, model era and scaffold era are confounded; claims about post-era gains on hard tasks from such data should be read with caution, and no-scaffold measurements preferred when the claim concerns the model itself.
  • The real hard-item effect is specific to short-reasoning, verifiable competitive programming tasks, not long-horizon autonomy; it does not support the narrative that the capability frontier is broadly reshaping toward hard tasks.
  • The effect survives multiple robustness checks (anchor choice, leave-one-family-out, date perturbation, contamination filtering, continuous-time version), narrowing but not eliminating the residual.

Where Pith is reading between the lines

These are editorial extensions of the paper, not claims the author makes directly.

  • If this separation holds, capability forecasting should move away from scalar extrapolations: the same aggregate trend can hide either uniform progress or a growing hard tail, and the two have different implications for when specific hard tasks become solvable.
  • The same deflation logic could be applied to other no-scaffold domains with exogenous difficulty, such as competition mathematics; the paper notes a math attempt failed on data preconditions, which suggests a human-solve-rate anchor would be the needed next dataset.
  • The +0.40 logit residual being concentrated in reasoning models and short-reasoning tasks is consistent with test-time compute helping most where outcomes are verifiable, but the paper explicitly does not test this; a direct test would vary reasoning budget within a fixed model while holding sampling constant.
  • The attempt-count confound (pre-era 10 draws vs post-era 1 draw) is handled by simulation, but the finding that equal-weighting reverses sign warns that naive reanalyses of public leaderboards could easily misread the effect as zero or negative.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

2 major / 4 minor

Summary. The paper argues that scalar summaries of LLM progress conflate an overall level shift with a change in the shape of the difficulty-response curve. On METR time-horizon data, it claims that a single Rasch model with rising ability largely reproduces the apparent migration of gains toward harder tasks, so most of that migration is a ceiling/discrimination artifact. On LiveCodeBench, which has no agentic scaffold and uses exogenous human difficulty labels, the paper freezes ability on easy/medium items, fits per-item fixed effects, and estimates a hard-item era effect δ while pinning rather than estimating the hard-item discrimination α_hard. Under the equal-discrimination assumption α_hard=1.00, the calibrated headline is δ≈+0.40 logits (raw +0.45), implying a hard-problem solve-rate rise from about 18% to 25%. The effect is reported to survive many robustness checks. The paper releases the LiveCodeBench Difficulty Panel and analysis code.

Significance. If the result holds, the paper makes a useful methodological and empirical contribution: it offers a concrete way to separate level from shape, shows that the widely reported hard-task acceleration is mostly a scalar artifact, and identifies a small but nonzero residual hard-item effect in competitive programming. The design is careful in several respects: ability is frozen on easy/medium items, difficulty labels are human-assigned and exogenous to the models, discrimination is pinned rather than freely fit, generated-regressor bias is calibrated by simulation on the real attempt counts, and uncertainty is addressed with a ladder of cluster bootstraps. The release of the panel and code is a positive feature for reproducibility. The main weakness is that the headline 'most conservative assumption' depends on α_hard=1.00 being the safe end of the grid, and the paper does not provide an independent, real-data estimate of hard-item discrimination to support that placement.

major comments (2)
  1. [Section 5, Fig. 4, §A.3] The headline δ≈+0.40 is called the 'most conservative' value because α_hard=1.00 is treated as the safe end of the grid. This is not established. The calibrated δ crosses zero at α_hard≈1.22, and the only cited support for placing the true value below 1.0 is the '0.83' estimate, which §A.3 does not actually compute on the LiveCodeBench panel. §A.3 is a simulation showing that when true α=1.0, the continuous binomial likelihood recovers 0.83; it is an estimator property, not a real-data estimate. Step 3 of §3.2 calls 0.83 'a debiased estimate from the continuous likelihood' without giving the fitting procedure or result. Given §A.1's degeneracy and §A.2's demonstration that a null 2PL absorbs DIF into inflated hard-item discrimination, the real-data provenance of 0.83 is critical. If 0.83 came from any of the models discussed in the appendix, it is not independent evidence. Without an ind
  2. [Section 4, Table 2] The first conclusion — that a single Rasch model with rising ability reproduces the METR locus migration — is not accompanied by the estimation details needed to check it. The text reports a fitted 'After, Rasch null' column in Table 2 and Figure 3, but does not state how the ability trajectory θ(t) is parameterized, how item difficulties are anchored, how the model is fit to the per-band success counts, or what the goodness of fit is. The claim that the null 'places the fastest-improving band at 15–60 minutes' is therefore an assertion. Since this is the basis for the abstract's statement that most apparent acceleration is a ceiling/discrimination artifact, the reader needs this information, or a reference to an appendix where it appears.
minor comments (4)
  1. [§A.3 vs §3.2/Table 3] The lower grid point is 0.55, called 'deflated by binarization', but §A.3's simulation reports binarization recovering α=0.49. Please reconcile these numbers in the text.
  2. [Fig. 4 caption] The caption states 'the calibrated estimate of hard-item discrimination on these data is 0.83'. As written this appears to be a real-data estimate, but the derivation is not given in the main text or in §A.3 (which is a simulation). Reword or provide the real-data estimation procedure.
  3. [Section 8] The limitation that a fully Bayesian hierarchical 2PL with a shrinkage prior on discrimination is left to future work is stated in §8, but it directly qualifies the 'most conservative' language used in §5. Consider foregrounding this caveat when the sensitivity grid is introduced.
  4. [§5, contamination filtering] Typo: 'thengeometry' should be 'the geometry'.

Circularity Check

1 steps flagged

Headline +0.40 logit effect is a genuine residual, but the 'most conservative' α_hard=1.00 anchor is supported by a simulation whose input is α_hard=1.00, making the conservative framing partially self-referential.

specific steps
  1. other [Section 5 (Fig. 4 caption; Table 3), Section 3.2 Step 3, Section A.3]
    ""The value αhard = 0.83 is a debiased estimate from the continuous likelihood" (Sec 3.2); "the calibrated estimate on these data is 0.83, below one" (Sec 5); "In our simulation with true discrimination α=1.0, binarization recovered ˆα=0.49, whereas the continuous binomial likelihood recovered ˆα=0.83" (A.3)."

    The main text presents 0.83 as data evidence that hard-item discrimination is below one, justifying α_hard=1.00 as the 'most conservative' grid point and hence the +0.40 headline. But the only 0.83 in the appendix comes from a simulation whose true value is α_hard=1.00, the very assumption being supported. No real-data α_hard estimate is shown (and A.1/A.2 argue a free fit is degenerate or biased). The zero-crossing at α≈1.22 is measured on the same panel, and the documented downward bias (true 1.0 recovered as 0.83) means 'no evidence above one' is not established. The conservative anchor is therefore not externally forced; the headline's sign survives only under an assumption whose support is generated from the same null used to define the residual.

full rationale

The core CurveShift derivation is largely self-contained and not circular: difficulty labels are human-assigned and exogenous to the models (Sec 3.1), ability θ is frozen from easy/medium items so hard outcomes cannot feed back into the level (Sec 3.2), the hard-item effect δ is a fitted residual rather than a predicted value, and the positive-control simulation (A.4) plus cluster bootstraps and leave-one-family-out checks give the result independent content. The Rasch deflation of METR (Sec 4) is a standard in-sample null comparison, not a self-definitional prediction. No load-bearing self-citation appears: the Xing et al. (2026) citation in Limitations is peripheral, and the classical psychometric citations (Rasch, Birnbaum, Holland & Wainer) are external. The single concern is the support for the 'conservative' α_hard=1.00 anchor: the paper cites Section A.3 as if it provided a data-based estimate of 0.83, but A.3 only reports a simulation value under true α=1.0. That makes the claim 'we find no evidence above one' an internal assumption rather than independent evidence, and the zero-crossing at α≈1.22 lies within plausible uncertainty. This is a partial circularity in the interpretation of the headline, not a collapse of the derivation; the raw residual is not manufactured. Hence score 2.

Axiom & Free-Parameter Ledger

5 free parameters · 6 axioms · 0 invented entities

The analysis rests on a few identifying assumptions: exogenous human difficulty labels, no-scaffold evaluation, and pinned discrimination. The headline effect is conditional on alpha=1.00; if true discrimination exceeds 1.22 the effect vanishes. The free 2PL is not identified, so the design pins the very quantity it cannot estimate.

free parameters (5)
  • Reasoning-era breakpoint (post indicator) = 2024.70 (September 2024)
    Hand-chosen cutoff at o1-preview; robustness sweep shows delta varies from +0.20 (2024.4) to +0.67 (2025.0), so the cutoff affects magnitude.
  • Hard-item discrimination pin alpha_hard = 0.55 / 0.83 / 1.00; headline 1.00
    alpha=0.83 is estimated from the same panel, alpha=1.00 is an equal-discrimination assumption, alpha=0.55 is a binarization artifact; delta crosses zero at alpha ~ 1.22, so the headline depends on this pin.
  • Generated-regressor calibration offset and slope = offset +0.078 at alpha=1.00; slope ~0.94
    Measured by planting effects on real attempt counts (Section A.4) and applied to delta and bootstrap intervals; the calibration is itself a simulation-based correction.
  • Rasch null ability trajectory theta(t) on METR = not reported numerically
    A single rising ability is fit to METR per-band data and used to generate the 'After, Rasch null' slopes in Table 2; the deflation result in Section 4 depends on this fit.
  • Per-model abilities theta_m and item fixed effects c_i = estimated from data
    Ability is frozen from easy/medium items and item fixed effects are fit on hard items; standard IRT parameters but central to the decomposition.
axioms (6)
  • domain assumption Human-assigned item difficulty labels (easy/medium/hard; completion-time bands) are exogenous and invariant across model eras.
    Central identifying assumption for both METR and LiveCodeBench; if labels are partly model-derived or change meaning across eras, delta is biased. Section 3.1.
  • domain assumption LiveCodeBench involves no agentic scaffold, so model era and harness era are not collinear there.
    Permits attribution of the hard-item era effect to model-side capability. Section 6.
  • domain assumption A single-ability Rasch model is an adequate null for METR per-band success; uniform growth of theta generates the predicted slopes.
    The 'locus migration is a ceiling artifact' conclusion is conditional on this null. Section 4, Table 2.
  • standard math Continuous binomial likelihood with attempt counts as weights is the correct sampling model; binarization deflates hard-item discrimination.
    Underpins estimation and the n=10 vs n=1 handling. Section A.3/A.5.
  • domain assumption alpha_hard = 1.00 (equal discrimination) is the conservative end of the defensible grid; true hard-item discrimination is not above approximately 1.22.
    The headline delta = +0.40 turns negative only above alpha ~ 1.22, so the sign and size of the reported effect rest on this. Section 5, Figure 4.
  • standard math Cluster bootstrap over items/contests/models is valid; model-family level has too few clusters (11) and is underpowered.
    Inference base; the family-level interval includes zero, so cross-family generalization is not established. Section 5.

pith-pipeline@v1.3.0-alltime-deepseek · 23566 in / 14138 out tokens · 138280 ms · 2026-08-04T00:38:04.009171+00:00 · methodology

0 comments
read the original abstract

Progress in large language models is often summarized using a single scalar measure, such as a time horizon, a latent ability estimate, or an aggregate benchmark score. These summaries capture the overall performance, but they do not test whether progress is distributed differently across task difficulty. We find that most of the apparent shift in gains toward harder tasks does not reflect a change in the shape of the difficulty-response curve. On METR time-horizon data, a single Rasch model with rising ability reproduces this pattern, so it is largely explained by ceiling effects rather than a qualitative change in capability. This echoes how the choice of metric can make claimed emergent abilities look like a property of the models themselves. We then identify a smaller hard-task effect that survives this control. Isolating it is difficult on agentic benchmarks, because newer models are usually run with newer agentic harnesses, so a gain on hard tasks cannot be assigned to the model or its scaffold. We break the confound with LiveCodeBench, a public competitive programming benchmark that runs no agentic scaffold while pairing dated models with an exogenous difficulty ordering. After accounting for the rise in overall ability, models released after September 2024 still gain on the hardest problems beyond what their easy and medium performance predicts, by about +0.40 logits under our most conservative assumption, raising the hard-problem solve rate from roughly 18% to 25%. The effect is led by the strongest reasoning models and holds for hard tasks that need only short reasoning, not autonomy over long horizons. We present this as a result specific to competitive programming, since our clean identification rests on a single coding benchmark. We release the LiveCodeBench Difficulty Panel (66 dated models x 1,055 problems) and our analysis code.

Figures

Figures reproduced from arXiv: 2608.00355 by BingXu Meng, Hanwen Xing, Jicheng Wang, Kumail Alhamoud, Pengyun Wang, Philip Torr, Xiang Li, Xiaomin Li, Xinyang Han, Xin Yu, Yuexing Hao.

Figure 1
Figure 1. Figure 1: Level-shape conflation, schematically. Left: a uniform upward shift in ability moves the whole [PITH_FULL_IMAGE:figures/full_fig_p002_1.png] view at source ↗
Figure 2
Figure 2. Figure 2: Study design. On agentic benchmarks (METR, SWE-bench) model era and scaffold era move [PITH_FULL_IMAGE:figures/full_fig_p004_2.png] view at source ↗
Figure 3
Figure 3. Figure 3: Success rate versus model release date within each difficulty band on METR Time Horizon, with [PITH_FULL_IMAGE:figures/full_fig_p007_3.png] view at source ↗
Figure 4
Figure 4. Figure 4: Hard-item era effect δ as a function of the pinned discrimination αhard. The solid curve is calibrated for the generated-regressor bias; the light dashed curve is raw. The shaded region is δ > 0. The three grid points used in [PITH_FULL_IMAGE:figures/full_fig_p009_4.png] view at source ↗

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Reference graph

Works this paper leans on

21 extracted references · 10 linked inside Pith

  1. [1]

    Program synthesis with large language models.arXiv preprint arXiv:2108.07732,

    Jacob Austin, Augustus Odena, Maxwell Nye, Maarten Bosma, Henryk Michalewski, David Dohan, Ellen Jiang, Carrie Cai, Michael Terry, Quoc Le, and Charles Sutton. Program synthesis with large language models.arXiv preprint arXiv:2108.07732,

  2. [5]

    Gritsenko, Zhe Zhao, Neil Houlsby, Fernando Diaz, Donald Metzler, and Oriol Vinyals

    Mostafa Dehghani, Yi Tay, Alexey A. Gritsenko, Zhe Zhao, Neil Houlsby, Fernando Diaz, Donald Metzler, and Oriol Vinyals. The benchmark lottery.arXiv preprint arXiv:2107.07002,

  3. [6]

    OckBench: Measuring the efficiency of LLM reasoning.arXiv preprint arXiv:2511.05722,

    Zheng Du, Hao Kang, Song Han, Tushar Krishna, and Ligeng Zhu. OckBench: Measuring the efficiency of LLM reasoning.arXiv preprint arXiv:2511.05722,

  4. [8]

    A Rosetta Stone for AI benchmarks.arXiv preprint arXiv:2512.00193,

    Anson Ho, Jean-Stanislas Denain, David Atanasov, Samuel Albanie, and Rohin Shah. A Rosetta Stone for AI benchmarks.arXiv preprint arXiv:2512.00193,

  5. [9]

    Lalor, Hao Wu, and Hong Yu

    John P. Lalor, Hao Wu, and Hong Yu. Building an evaluation scale using item response theory. InProceedings of the 2016 Conference on Empirical Methods in Natural Language Processing (EMNLP), pp. 648–657,

  6. [13]

    doi: 10.18653/v1/2021.acl-long.346

    Association for Computational Lin- guistics. doi: 10.18653/v1/2021.acl-long.346. URLhttps://aclanthology.org/2021.acl-long.346/. Rylan Schaeffer, Brando Miranda, and Sanmi Koyejo. Are emergent abilities of large language models a mirage?Advances in Neural Information Processing Systems, 36:55565–55581,

  7. [14]

    URLhttps: //doi.org/10.1145/3715754

    doi: 10.1145/3715754. URLhttps: //doi.org/10.1145/3715754. Tianbao Xie, Danyang Zhang, Jixuan Chen, Xiaochuan Li, Siheng Zhao, Ruisheng Cao, Toh Jing Hua, Zhoujun Cheng, Dongchan Shin, Fangyu Lei, et al. OSWorld: Benchmarking multimodal agents for open- ended tasks in real computer environments.Advances in Neural Information Processing Systems, 37: 52040–52094,

  8. [16]

    Scaling test-time compute for LLM agents.arXiv preprint arXiv:2506.12928,

    King Zhu, Hanhao Li, Siwei Wu, Tianshun Xing, Dehua Ma, Xiangru Tang, Minghao Liu, Jian Yang, Jiaheng Liu, Yuchen Eleanor Jiang, Changwang Zhang, Chenghua Lin, Jun Wang, Ge Zhang, and Wangchunshu Zhou. Scaling test-time compute for LLM agents.arXiv preprint arXiv:2506.12928,

  9. [17]

    18 A.1 The free 2PL is not stably identified jointly with the era DIF coefficient The two-parameter logistic model positslogitp mi =α i(θm−β i), with discriminationα i

    A Identifiability of the free 2PL and the bias of a naive residual test This appendix documents the methodological pitfalls we encountered and ruled out, so that reviewers can verify that the obvious alternatives were considered and rejected for principled reasons. 18 A.1 The free 2PL is not stably identified jointly with the era DIF coefficient The two-p...

  10. [19]

    We used the MathArena panels for the 2025 contests (Balunovic et al.,

    The natural candidate is competition mathematics without an agentic scaffold, where models answer directly and the contest origin gives a difficulty ordering. We used the MathArena panels for the 2025 contests (Balunovic et al.,

  11. [20]

    The attempt does not fail a hypothesis test; it fails before one can be run, on two of the four conditions that make LiveCodeBench usable

    (AIME, HMMT February, BRUMO, SMT, and CMIMC). The attempt does not fail a hypothesis test; it fails before one can be run, on two of the four conditions that make LiveCodeBench usable. The first missing condition is a clean exogenous difficulty. The only per-item ordering these panels carry is the problem index, which is an endogenous and noisy proxy. A g...

  12. [21]

    23 Model Date Family Era Reasoning DSCoder-1.3b-Ins 2023.86 DeepSeek pre no Claude-3-Haiku 2024.20 Anthropic pre no GPT-4-Turbo-2024-04-09 2024.27 OpenAI pre no GPT-4O-2024-05-13 2024.37 OpenAI pre no Codestral-Latest 2024.42 Mistral pre no Qwen2-Ins-72B 2024.43 Qwen pre no Claude-3.5-Sonnet-20240620 2024.47 Anthropic pre no Mistral-Large 2024.55 Mistral ...

  13. [1959]

    The common odds ratio is1.84(log odds+0.61), favoring post-era models, in the direction and rough magnitude of the parametric estimate

    analysis matches models on their easy and medium ability, then compares pre-era and post-era performance on hard items within ability strata, with no functional form. The common odds ratio is1.84(log odds+0.61), favoring post-era models, in the direction and rough magnitude of the parametric estimate. The matching variable overlaps only in the lower abili...

  14. [1968]

    Le, Christopher Ré, and Azalia Mirhoseini

    Bradley Brown, Jordan Juravsky, Ryan Ehrlich, Ronald Clark, Quoc V. Le, Christopher Ré, and Azalia Mirhoseini. Large language monkeys: Scaling inference compute with repeated sampling.arXiv preprint arXiv:2407.21787,

  15. [1980]

    Data contamination through the lens of time.arXiv preprint arXiv:2310.10628,

    Manley Roberts, Himanshu Thakur, Christine Herlihy, Colin White, and Samuel Dooley. Data contamination through the lens of time.arXiv preprint arXiv:2310.10628,

  16. [1985]

    OpenAI o1 system card

    16 OpenAI. OpenAI o1 system card. arXiv:2412.16720,

  17. [2008]

    Evaluating large language models trained on code.arXiv preprint arXiv:2107.03374,

    Mark Chen, Jerry Tworek, Heewoo Jun, Qiming Yuan, Henrique Ponde de Oliveira Pinto, Jared Kaplan, Harri Edwards, Yuri Burda, Nicholas Joseph, Greg Brockman, et al. Evaluating large language models trained on code.arXiv preprint arXiv:2107.03374,

  18. [2021]

    Training verifiers to solve math word problems.arXiv preprint arXiv:2110.14168,

    Karl Cobbe, Vineet Kosaraju, Mohammad Bavarian, Mark Chen, Heewoo Jun, Lukasz Kaiser, Matthias Plappert, Jerry Tworek, Jacob Hilton, Reiichiro Nakano, Christopher Hesse, and John Schulman. Training verifiers to solve math word problems.arXiv preprint arXiv:2110.14168,

  19. [2024]

    CodeElo: Benchmarking competition-level code generation of LLMs with human-comparable Elo ratings.arXiv preprint arXiv:2501.01257,

    Shanghaoran Quan, Jiaxi Yang, Bowen Yu, Bo Zheng, Dayiheng Liu, An Yang, Xuancheng Ren, Bofei Gao, YiboMiao, YunlongFeng, ZekunWang, JianYang, ZeyuCui, YangFan, YichangZhang, BinyuanHui, and Junyang Lin. CodeElo: Benchmarking competition-level code generation of LLMs with human-comparable Elo ratings.arXiv preprint arXiv:2501.01257,

  20. [2025]

    Agent psychometrics: Task-level performance prediction in agentic coding benchmarks

    Chris Ge, Daria Kryvosheieva, Daniel Fried, Uzay Girit, and Kaivalya Hariharan. Agent psychometrics: Task-level performance prediction in agentic coding benchmarks. InICLR 2026 Workshop on Agents in the Wild (AIWILD),

  21. [2026]

    doi: 10.1145/3770855.3818652

    ACM. doi: 10.1145/3770855.3818652. To appear. John Yang, Carlos Jimenez, Alexander Wettig, Kilian Lieret, Shunyu Yao, Karthik Narasimhan, and Ofir Press. SWE-agent: Agent-computer interfaces enable automated software engineering.Advances in Neural Information Processing Systems, 37:50528–50652,