Pith. sign in

REVIEW 2 major objections 5 minor 28 references

On a controlled software-engineering agent benchmark, how you format a skill does not beat the raw skill; which model runs the task does.

Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →

T0 review · grok-4.5

2026-07-12 05:13 UTC pith:LPXOOB3Y

load-bearing objection Clean factor-separated measurement: on this substrate executor tier moves quality and real dollars; skill rewriting does not beat raw skills or break even. the 2 major comments →

arxiv 2607.03048 v1 pith:LPXOOB3Y submitted 2026-07-03 cs.SE

Compression, structure, and executor capability: a controlled real-cost decomposition of language-model agent skill optimisation

classification cs.SE
keywords large language model agentsagent skillscontrolled experimentcost measurementnon-inferioritysoftware qualityexecutor capabilityskill representation
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

Agent builders often shorten, restructure, recompile, or selectively load reusable instruction packages ("skills") on the belief that a smaller or better-organised skill will cut cost without hurting success. This paper tests that belief with quality and real money measured on the same runs, and with the usual confounded factors separated. Across ten delivery conditions, forty verifier-scored software tasks, and 1,200 rollouts, deterministic shortening stays near the raw skill without proving non-inferiority, structured and scoped versions do not raise pass rate or lower real cost, and structure is no better than plain text once content is held fixed. The only contrast that survives multiplicity control is executor capability: moving to the stronger model raises pass rate by about 27 percentage points at roughly five times the real cost. Under published on-demand prices, no optimised representation reaches a practical break-even, because none lowers mean solve-stage cost below the raw baseline. A sympathetic reader cares because the result reframes skill optimisation as a quality-and-currency measurement problem rather than a default engineering win.

Core claim

Executor capability is the dominant lever, and none of the tested representation strategies improves over the raw skill on either executor tier. After task-clustered inference and multiplicity control, the sole surviving quality contrast is the executor change (+26.7 percentage points pass rate at roughly five times real cost); compiler tier has no robust effect, and under real prices no optimised representation reaches a practical break-even because none reduces mean solve-stage cost below the raw baseline.

What carries the argument

A factor-separated skill-delivery design with a shared semantic ledger: linear and structured views are rendered from identical unit payloads so presentation is isolated from content, quality is gated by a fixed non-inferiority margin, and cost is priced in currency at on-demand rates with compilation amortised separately.

Load-bearing premise

The finding that representation tricks do not beat the raw skill rests on forty tasks and two model tiers from one provider family under one harness; if other tasks, models, or skill formats behave differently, the null representation result need not transfer.

What would settle it

Re-run the same factor-separated design on a larger or different task set and model family; if any representation (shortening, structure, scoped loading, or strong compilation) both preserves or improves pass rate within the non-inferiority margin and lowers real solve-stage cost enough for a finite practical break-even, the central claim fails.

Watch this falsifier — get emailed when new claim-graph text bears on it.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

2 major / 5 minor

Summary. The paper reports a controlled 10-condition decomposition of language-model agent skill delivery on 40 software-engineering tasks (1,200 rollouts), separating no-skill, raw skills, deterministic extractive shortening, content-matched linear vs structured rendering from a shared semantic ledger, scoped loading, and a compiler imes executor tier factorial. Quality is verifier pass rate at task level; cost is solve-stage token cost at on-demand provider prices (with token-volume normalisation and separately amortised compilation). With task-level inference, task-clustered intervals, a pre-fixed contrast family under Holm control, and a −5 pp non-inferiority margin, the only multiplicity-corrected quality effect is executor capability (+26.7 pp at roughly five times real cost). Deterministic shortening is descriptively near the raw baseline without establishing non-inferiority; structured rendering and scoped loading do not improve (and tend to lower) pass rate on the compact executor without lowering cost; structured is indistinguishable from linear at matched content; compiler tier is null; and under real prices no optimised representation has a finite mean break-even against the raw skill.

Significance. If the result holds on the reported substrate, it is a useful corrective for a growing practice literature that treats skill shortening, structural rewriting, stronger-model compilation, and scoped loading as efficiency wins. The contribution is methodological as much as empirical: factor-separated conditions, payload-equality for presentation vs content, sentinel no-skill audit, pre-registered margins and contrast family, real-price primary accounting that exposes the divergence from token-volume views, and open artefacts (Zenodo 10.5281/zenodo.21148499). That package makes the null representation result and the executor-dominance claim falsifiable and reusable, which is rarer than end-to-end skill-system demos in this area.

major comments (2)
  1. [Abstract, §6, §8, Table 1] Abstract, §6, and §8 state that “no representation strategy improves over the raw skill on either executor tier.” Only the structured-rendering family (S3-*-*) is fully crossed with executor tier; S2E-M, S2L-55-M, and S4-55-M are run only on the compact executor (Table 1). Raw still beats structured on the strong tier (S1-55 64.2% vs S3-55-55 55.8%, Table 2), so the structured null is supported on both tiers, but the stronger wording for all representation strategies is not. Please restate the claim to match the design (e.g., no tested representation improves on the compact tier; structured does not improve on either tier) and flag the uncrossed arms as a design limit in §7.
  2. [Abstract, §4.3, Table 3, §7–§8] With n=40 tasks the task-clustered intervals are wide (e.g., S2E-M−S1-M: [−7.5, +9.2] pp, Table 3), so several “null” representation contrasts are correctly read as insufficient evidence within the pre-fixed margins rather than as evidence of exact equality (§7 Conclusion validity). The headline phrasing in the abstract and conclusion (“does not improve,” “no representation strategy improves”) can be read as stronger than that. Align the abstract/conclusion language with the non-inferiority and multiplicity framing already used in §4.3–§5 so that readers do not over-read the nulls.
minor comments (5)
  1. [Fig. 3] Fig. 1 and Fig. 2 are clear; Fig. 3 would benefit from an explicit note in the caption that the y-scale is dominated by the strong-executor real-price bars, so compact-executor differences are hard to read visually (the table already carries the numbers).
  2. [§4.2, Eq. (1)] Eq. (1) is standard; a one-line reminder that a non-positive denominator implies no finite break-even would help readers who skip §4.2.
  3. [Table 4] Table 4 groups conditions by executor tier with an “internal break”; a horizontal rule or bold subhead in the table body would make that grouping easier to scan in print.
  4. [§3.2, §3.4] The model labels “gpt-5.4-mini” / “gpt-5.5” and the shorthand “M” / “55” should be tied once to dated snapshot identifiers (as the harness description implies) so the experiment remains reproducible if product names drift.
  5. [Front matter / tables] Minor copy: “F unding” and “T able” spacing artefacts appear in the compiled text; clean before camera-ready.

Circularity Check

0 steps flagged

No circularity: empirical measurement study whose quality and cost outcomes are external to the interventions under test.

full rationale

The paper is a controlled software-engineering experiment, not a first-principles derivation. Quality is the binary verifier pass on an external benchmark; cost is solve-stage token spend at published provider prices. The ten conditions, shared semantic ledger with payload-equality checks, sentinel no-skill audit, pre-fixed non-inferiority and cost margins, task-clustered intervals, and Holm-controlled contrast family are design choices that separate factors rather than define outcomes in terms of those factors. No parameter is fitted to a subset of the confirmatory rollouts and then re-presented as a prediction; Eq. (1) break-even is an accounting identity applied after measurement, not a circular claim. Self-citations are ordinary related-work pointers and do not load-bear the central result. The derivation chain therefore does not reduce by construction to its inputs; the honest finding is no significant circularity.

Axiom & Free-Parameter Ledger

4 free parameters · 5 axioms · 2 invented entities

The central claim rests on an empirical protocol rather than free-form theory. Load-bearing choices are the quality construct (verifier pass), the cost construct (on-demand USD with separate amortisation), pre-set decision margins, the two model tiers, the shared semantic ledger as content-matching device, and the assumption that this task/harness slice speaks for skill-optimisation practice. No new physical entities are postulated; the ledger and condition labels are methodological constructs.

free parameters (4)
  • non-inferiority margin on pass rate = -5 percentage points
    Fixed at −5 percentage points as the smallest practically meaningful pass-rate difference; gates whether a strategy is treated as quality-preserving (Section 4.3).
  • cost practical-equivalence margin = 10%
    Cost difference within 10% of baseline solve cost treated as practical equivalence; efficiency gains reported only if quality criterion holds (Section 4.3).
  • compact and strong model tier identities = gpt-5.4-mini / gpt-5.5
    Executor/compiler tiers fixed as gpt-5.4-mini (M) and gpt-5.5 (55); all tier contrasts and the ~5–7× real-price gap depend on these choices and their published rates (Section 3.2, 5.5).
  • repetitions per task-condition cell = 3
    Three repetitions per cell fix the per-task pass probability estimator and contribute to interval width (Section 3.4).
axioms (5)
  • domain assumption Executable verifier pass (with timeout as failure) is an adequate primary quality measure for skill-optimisation claims on this benchmark.
    Section 4.1 defines quality as binary verifier pass aggregated per task; partial progress and solution quality beyond the verifier are out of scope (also Threats, Construct validity).
  • domain assumption Standard on-demand provider prices on solve-stage tokens are the primary cost truth; token-volume normalisation is only a robustness view.
    Section 4.2 makes real-price cached-observed cost primary; comparative conclusions depend on the large per-token tier price ratio.
  • standard math The task is the unit of inference; task-clustered intervals and Holm control on the fixed eight-contrast quality family are appropriate.
    Section 4.4 follows standard controlled SE experiment practice (bootstrap over tasks, sign-flip permutation, Holm).
  • domain assumption Shared semantic ledger plus payload-equality check makes the structured-vs-linear contrast a pure presentation contrast.
    Section 3.3; residual risk that model-generated ledgers differ in uncaptured ways is noted in Threats.
  • domain assumption The 40 analysed tasks across eight categories, after excluding GPU/resource-capped and harness-tuning tasks, support confirmatory inference about representation strategies.
    Section 3.1 and Appendix A; external-validity threat acknowledges specificity of this set.
invented entities (2)
  • shared semantic ledger (atomic semantic units + deterministic linear/structured renderers) no independent evidence
    purpose: Hold skill content fixed while varying presentation so structure can be isolated from content.
    Methodological intermediate representation introduced for the S2L/S3 family; not an independent scientific object beyond this protocol.
  • ten skill-delivery condition labels (S0–S4 factorial over representation, compiler, executor) no independent evidence
    purpose: Factor-separate no-skill, raw, extractive shortening, linear/structured renders, scoped loading, and model tiers.
    Experimental design constructs; effects are measured outcomes, not postulated mechanisms.

pith-pipeline@v1.1.0-grok45 · 17971 in / 3525 out tokens · 42067 ms · 2026-07-12T05:13:15.051132+00:00 · methodology

0 comments
read the original abstract

Agent skills, reusable instruction artefacts supplied to a tool-using language model, are increasingly optimised by shortening, structural rewriting, stronger-model compilation, and scoped loading, on the assumption that a smaller or better-organised skill lowers cost while preserving success. That assumption is rarely tested with quality and real monetary cost measured on the same runs and the contributing factors separated. This study reports a controlled decomposition over ten skill-delivery conditions, 40 software-engineering tasks, and three repetitions per cell (1,200 rollouts), separating no-skill execution, raw skills, deterministic shortening, linear and structured rendering from a shared semantic ledger, scoped loading, and the compiler and executor model tiers. Quality is the verifier pass rate at task level; cost is solve-stage token cost at standard provider prices, with a token-volume-normalised view for robustness and compilation cost amortised separately. The task is the unit of inference, intervals are task-clustered, and the contrast family is multiplicity-controlled. Deterministic shortening is close to the raw baseline but does not establish non-inferiority within the preset margin. Structured rendering and scoped loading lower pass rate on the compact executor without lowering cost, and structured rendering is indistinguishable from linear text at matched content. The only contrast surviving correction is executor capability, which raises pass rate by 27 percentage points at roughly five times the real cost, with compiler tier showing no robust effect. Under real prices no optimised representation reaches a practical break-even. The evidence indicates that executor capability is the dominant lever and that no representation strategy improves over the raw skill on either executor tier.

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Reference graph

Works this paper leans on

28 extracted references · 1 canonical work pages

  1. [1]

    Anthropic Engineering Blog (2025)

    Zhang, B., Lazuka, K., Murag, M.: Equipping agents for the real world with Agent Skills. Anthropic Engineering Blog (2025). https://www.anthropic.com/engineering/equipping-agents-for-the-real-world- with-agent-skills

  2. [2]

    arXiv:2604.08224 (2026)

    Zhou, C., Chai, H., Chen, W., Guo, Z., Shan, R., Song, Y., Xu, T., Yang, Y., Yu, A., Zhang, W., Zheng, C., Zhu, J., Zheng, Z., Zhang, Z., Lou, X., Zhang, C., Fu, Z., Wang, J., Liu, W., Lin, J., Zhang, W.: Externalization in LLM agents: a unified review of memory, skills, protocols and harness engineering. arXiv:2604.08224 (2026)

  3. [3]

    arXiv:2602.12430 (2026)

    Xu, R., Yan, Y.: Agent skills for large language models: architecture, acquisition, security, and the path forward. arXiv:2602.12430 (2026)

  4. [4]

    arXiv:2305.16291 (2023)

    Wang, G., Xie, Y., Jiang, Y., Mandlekar, A., Xiao, C., Zhu, Y., Fan, L., Anand- kumar, A.: Voyager: an open-ended embodied agent with large language models. arXiv:2305.16291 (2023)

  5. [5]

    arXiv:2405.15793 (2024)

    Yang, J., Jimenez, C.E., Wettig, A., Lieret, K., Yao, S., Narasimhan, K., Press, O.: SWE-agent: agent-computer interfaces enable automated software engineering. arXiv:2405.15793 (2024)

  6. [6]

    arXiv:2310.06770

    Jimenez, C.E., Yang, J., Wettig, A., Yao, S., Pei, K., Press, O., Narasimhan, K.: SWE-bench: can language models resolve real-world GitHub issues? In: International Conference on Learning Representations (2024). arXiv:2310.06770

  7. [7]

    In: International Conference on Machine Learning (2024)

    Huang, Q., Vora, J., Liang, P., Leskovec, J.: MLAgentBench: evaluating lan- guage agents on machine learning experimentation. In: International Conference on Machine Learning (2024). arXiv:2310.03302

  8. [8]

    In: International Conference on Learning Representations (2023)

    Yao, S., Zhao, J., Yu, D., Du, N., Shafran, I., Narasimhan, K., Cao, Y.: ReAct: synergizing reasoning and acting in language models. In: International Conference on Learning Representations (2023). arXiv:2210.03629

  9. [9]

    In: Advances in Neural Information Processing Systems (2023)

    Schick, T., Dwivedi-Yu, J., Dess` ı, R., Raileanu, R., Lomeli, M., Zettlemoyer, L., Cancedda, N., Scialom, T.: Toolformer: language models can teach themselves to use tools. In: Advances in Neural Information Processing Systems (2023). arXiv:2302.04761

  10. [10]

    In: Advances in Neural Information Processing Systems (2023)

    Shinn, N., Cassano, F., Gopinath, A., Narasimhan, K., Yao, S.: Reflexion: language agents with verbal reinforcement learning. In: Advances in Neural Information Processing Systems (2023). arXiv:2303.11366

  11. [11]

    In: Proceedings of 15 EMNLP (2023)

    Jiang, H., Wu, Q., Lin, C.-Y., Yang, Y., Qiu, L.: LLMLingua: compressing prompts for accelerated inference of large language models. In: Proceedings of 15 EMNLP (2023). arXiv:2310.05736

  12. [12]

    arXiv:2305.05176 (2023)

    Chen, L., Zaharia, M., Zou, J.: FrugalGPT: how to use large language models while reducing cost and improving performance. arXiv:2305.05176 (2023)

  13. [13]

    arXiv:2406.18665 (2024)

    Ong, I., Almahairi, A., Wu, V., Zhang, W.-L., Willmott, D., Stoica, I., Zaharia, M.: RouteLLM: learning to route LLMs with preference data. arXiv:2406.18665 (2024)

  14. [14]

    Transactions of the Association for Computational Linguistics (2024)

    Liu, N.F., Lin, K., Hewitt, J., Paranjape, A., Bevilacqua, M., Petroni, F., Liang, P.: Lost in the middle: how language models use long contexts. Transactions of the Association for Computational Linguistics (2024). arXiv:2307.03172

  15. [15]

    arXiv:2505.06120 (2025)

    Laban, P., Hayashi, H., Zhou, Y., Neville, J.: LLMs get lost in multi-turn conversation. arXiv:2505.06120 (2025)

  16. [16]

    arXiv:2604.12301 (2026)

    Agyemang, J.O., Kponyo, J.J., Amponsah, E., Boakye, G.M.A., Agyekum, K.O.- B.O.: Local-Splitter: a measurement study of seven tactics for reducing cloud LLM token usage on coding-agent workloads. arXiv:2604.12301 (2026)

  17. [17]

    In: Advances in Neural Information Processing Systems (2025)

    Xu, W., Liang, Z., Mei, K., Gao, H., Tan, J., Zhang, Y.: A-MEM: agentic memory for LLM agents. In: Advances in Neural Information Processing Systems (2025). arXiv:2502.12110

  18. [18]

    arXiv:2504.06188 (2026)

    Li, F., Tagkopoulos, P., Tagkopoulos, I.: SkillFlow: scalable and efficient agent skill retrieval system. arXiv:2504.06188 (2026)

  19. [19]

    Journal of General Internal Medicine26(2), 192–196 (2011)

    Walker, E., Nowacki, A.S.: Understanding equivalence and noninferiority testing. Journal of General Internal Medicine26(2), 192–196 (2011)

  20. [20]

    Chapman & Hall, New York (1993)

    Efron, B., Tibshirani, R.J.: An Introduction to the Bootstrap. Chapman & Hall, New York (1993)

  21. [21]

    Scandinavian Journal of Statistics6(2), 65–70 (1979)

    Holm, S.: A simple sequentially rejective multiple test procedure. Scandinavian Journal of Statistics6(2), 65–70 (1979)

  22. [22]

    Springer, Berlin (2012)

    Wohlin, C., Runeson, P., H¨ ost, M., Ohlsson, M.C., Regnell, B., Wessl´ en, A.: Experimentation in Software Engineering. Springer, Berlin (2012)

  23. [23]

    Empirical Software Engineering27(4), 94 (2022)

    Baltes, S., Ralph, P.: Sampling in software engineering research: a critical review and guidelines. Empirical Software Engineering27(4), 94 (2022)

  24. [24]

    Transactions on Machine Learning Research (2024)

    Sumers, T.R., Yao, S., Narasimhan, K., Griffiths, T.L.: Cognitive architec- tures for language agents. Transactions on Machine Learning Research (2024). arXiv:2309.02427

  25. [25]

    IEEE Transactions on 16 Software Engineering50(4), 911–936 (2024)

    Wang, J., Huang, Y., Chen, C., Liu, Z., Wang, S., Wang, Q.: Software testing with large language models: survey, landscape, and vision. IEEE Transactions on 16 Software Engineering50(4), 911–936 (2024)

  26. [26]

    Software Quality Journal34, 8 (2026)

    Haldar, S., Capretz, L.F.: Automated test plan generation using large language models. Software Quality Journal34, 8 (2026). https://doi.org/10.1007/s11219- 026-09744-9

  27. [27]

    He, J., Rungta, M., Koleczek, D., Sekhon, A., Wang, F.X., Hasan, S.: Does prompt formatting have any impact on LLM performance? arXiv:2411.10541 (2024)

  28. [28]

    Tam, Z.R., Wu, C.-K., Tsai, Y.-L., Lin, C.-Y., Lee, H.-Y., Chen, Y.-N.: Let me speak freely? A study on the impact of format restrictions on the performance of large language models. arXiv:2408.02442 (2024) Appendix A Task manifest 17 T able A1Tasks in the analysis set, with category and difficulty Task Category Difficulty 3d-scan-calc industrial and phys...