REVIEW 2 major objections 5 minor 28 references
On a controlled software-engineering agent benchmark, how you format a skill does not beat the raw skill; which model runs the task does.
Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →
T0 review · grok-4.5
2026-07-12 05:13 UTC pith:LPXOOB3Y
load-bearing objection Clean factor-separated measurement: on this substrate executor tier moves quality and real dollars; skill rewriting does not beat raw skills or break even. the 2 major comments →
Compression, structure, and executor capability: a controlled real-cost decomposition of language-model agent skill optimisation
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
Core claim
Executor capability is the dominant lever, and none of the tested representation strategies improves over the raw skill on either executor tier. After task-clustered inference and multiplicity control, the sole surviving quality contrast is the executor change (+26.7 percentage points pass rate at roughly five times real cost); compiler tier has no robust effect, and under real prices no optimised representation reaches a practical break-even because none reduces mean solve-stage cost below the raw baseline.
What carries the argument
A factor-separated skill-delivery design with a shared semantic ledger: linear and structured views are rendered from identical unit payloads so presentation is isolated from content, quality is gated by a fixed non-inferiority margin, and cost is priced in currency at on-demand rates with compilation amortised separately.
Load-bearing premise
The finding that representation tricks do not beat the raw skill rests on forty tasks and two model tiers from one provider family under one harness; if other tasks, models, or skill formats behave differently, the null representation result need not transfer.
What would settle it
Re-run the same factor-separated design on a larger or different task set and model family; if any representation (shortening, structure, scoped loading, or strong compilation) both preserves or improves pass rate within the non-inferiority margin and lowers real solve-stage cost enough for a finite practical break-even, the central claim fails.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper reports a controlled 10-condition decomposition of language-model agent skill delivery on 40 software-engineering tasks (1,200 rollouts), separating no-skill, raw skills, deterministic extractive shortening, content-matched linear vs structured rendering from a shared semantic ledger, scoped loading, and a compiler imes executor tier factorial. Quality is verifier pass rate at task level; cost is solve-stage token cost at on-demand provider prices (with token-volume normalisation and separately amortised compilation). With task-level inference, task-clustered intervals, a pre-fixed contrast family under Holm control, and a −5 pp non-inferiority margin, the only multiplicity-corrected quality effect is executor capability (+26.7 pp at roughly five times real cost). Deterministic shortening is descriptively near the raw baseline without establishing non-inferiority; structured rendering and scoped loading do not improve (and tend to lower) pass rate on the compact executor without lowering cost; structured is indistinguishable from linear at matched content; compiler tier is null; and under real prices no optimised representation has a finite mean break-even against the raw skill.
Significance. If the result holds on the reported substrate, it is a useful corrective for a growing practice literature that treats skill shortening, structural rewriting, stronger-model compilation, and scoped loading as efficiency wins. The contribution is methodological as much as empirical: factor-separated conditions, payload-equality for presentation vs content, sentinel no-skill audit, pre-registered margins and contrast family, real-price primary accounting that exposes the divergence from token-volume views, and open artefacts (Zenodo 10.5281/zenodo.21148499). That package makes the null representation result and the executor-dominance claim falsifiable and reusable, which is rarer than end-to-end skill-system demos in this area.
major comments (2)
- [Abstract, §6, §8, Table 1] Abstract, §6, and §8 state that “no representation strategy improves over the raw skill on either executor tier.” Only the structured-rendering family (S3-*-*) is fully crossed with executor tier; S2E-M, S2L-55-M, and S4-55-M are run only on the compact executor (Table 1). Raw still beats structured on the strong tier (S1-55 64.2% vs S3-55-55 55.8%, Table 2), so the structured null is supported on both tiers, but the stronger wording for all representation strategies is not. Please restate the claim to match the design (e.g., no tested representation improves on the compact tier; structured does not improve on either tier) and flag the uncrossed arms as a design limit in §7.
- [Abstract, §4.3, Table 3, §7–§8] With n=40 tasks the task-clustered intervals are wide (e.g., S2E-M−S1-M: [−7.5, +9.2] pp, Table 3), so several “null” representation contrasts are correctly read as insufficient evidence within the pre-fixed margins rather than as evidence of exact equality (§7 Conclusion validity). The headline phrasing in the abstract and conclusion (“does not improve,” “no representation strategy improves”) can be read as stronger than that. Align the abstract/conclusion language with the non-inferiority and multiplicity framing already used in §4.3–§5 so that readers do not over-read the nulls.
minor comments (5)
- [Fig. 3] Fig. 1 and Fig. 2 are clear; Fig. 3 would benefit from an explicit note in the caption that the y-scale is dominated by the strong-executor real-price bars, so compact-executor differences are hard to read visually (the table already carries the numbers).
- [§4.2, Eq. (1)] Eq. (1) is standard; a one-line reminder that a non-positive denominator implies no finite break-even would help readers who skip §4.2.
- [Table 4] Table 4 groups conditions by executor tier with an “internal break”; a horizontal rule or bold subhead in the table body would make that grouping easier to scan in print.
- [§3.2, §3.4] The model labels “gpt-5.4-mini” / “gpt-5.5” and the shorthand “M” / “55” should be tied once to dated snapshot identifiers (as the harness description implies) so the experiment remains reproducible if product names drift.
- [Front matter / tables] Minor copy: “F unding” and “T able” spacing artefacts appear in the compiled text; clean before camera-ready.
Circularity Check
No circularity: empirical measurement study whose quality and cost outcomes are external to the interventions under test.
full rationale
The paper is a controlled software-engineering experiment, not a first-principles derivation. Quality is the binary verifier pass on an external benchmark; cost is solve-stage token spend at published provider prices. The ten conditions, shared semantic ledger with payload-equality checks, sentinel no-skill audit, pre-fixed non-inferiority and cost margins, task-clustered intervals, and Holm-controlled contrast family are design choices that separate factors rather than define outcomes in terms of those factors. No parameter is fitted to a subset of the confirmatory rollouts and then re-presented as a prediction; Eq. (1) break-even is an accounting identity applied after measurement, not a circular claim. Self-citations are ordinary related-work pointers and do not load-bear the central result. The derivation chain therefore does not reduce by construction to its inputs; the honest finding is no significant circularity.
Axiom & Free-Parameter Ledger
free parameters (4)
- non-inferiority margin on pass rate =
-5 percentage points
- cost practical-equivalence margin =
10%
- compact and strong model tier identities =
gpt-5.4-mini / gpt-5.5
- repetitions per task-condition cell =
3
axioms (5)
- domain assumption Executable verifier pass (with timeout as failure) is an adequate primary quality measure for skill-optimisation claims on this benchmark.
- domain assumption Standard on-demand provider prices on solve-stage tokens are the primary cost truth; token-volume normalisation is only a robustness view.
- standard math The task is the unit of inference; task-clustered intervals and Holm control on the fixed eight-contrast quality family are appropriate.
- domain assumption Shared semantic ledger plus payload-equality check makes the structured-vs-linear contrast a pure presentation contrast.
- domain assumption The 40 analysed tasks across eight categories, after excluding GPU/resource-capped and harness-tuning tasks, support confirmatory inference about representation strategies.
invented entities (2)
-
shared semantic ledger (atomic semantic units + deterministic linear/structured renderers)
no independent evidence
-
ten skill-delivery condition labels (S0–S4 factorial over representation, compiler, executor)
no independent evidence
read the original abstract
Agent skills, reusable instruction artefacts supplied to a tool-using language model, are increasingly optimised by shortening, structural rewriting, stronger-model compilation, and scoped loading, on the assumption that a smaller or better-organised skill lowers cost while preserving success. That assumption is rarely tested with quality and real monetary cost measured on the same runs and the contributing factors separated. This study reports a controlled decomposition over ten skill-delivery conditions, 40 software-engineering tasks, and three repetitions per cell (1,200 rollouts), separating no-skill execution, raw skills, deterministic shortening, linear and structured rendering from a shared semantic ledger, scoped loading, and the compiler and executor model tiers. Quality is the verifier pass rate at task level; cost is solve-stage token cost at standard provider prices, with a token-volume-normalised view for robustness and compilation cost amortised separately. The task is the unit of inference, intervals are task-clustered, and the contrast family is multiplicity-controlled. Deterministic shortening is close to the raw baseline but does not establish non-inferiority within the preset margin. Structured rendering and scoped loading lower pass rate on the compact executor without lowering cost, and structured rendering is indistinguishable from linear text at matched content. The only contrast surviving correction is executor capability, which raises pass rate by 27 percentage points at roughly five times the real cost, with compiler tier showing no robust effect. Under real prices no optimised representation reaches a practical break-even. The evidence indicates that executor capability is the dominant lever and that no representation strategy improves over the raw skill on either executor tier.
Reference graph
Works this paper leans on
-
[1]
Anthropic Engineering Blog (2025)
Zhang, B., Lazuka, K., Murag, M.: Equipping agents for the real world with Agent Skills. Anthropic Engineering Blog (2025). https://www.anthropic.com/engineering/equipping-agents-for-the-real-world- with-agent-skills
2025
-
[2]
Zhou, C., Chai, H., Chen, W., Guo, Z., Shan, R., Song, Y., Xu, T., Yang, Y., Yu, A., Zhang, W., Zheng, C., Zhu, J., Zheng, Z., Zhang, Z., Lou, X., Zhang, C., Fu, Z., Wang, J., Liu, W., Lin, J., Zhang, W.: Externalization in LLM agents: a unified review of memory, skills, protocols and harness engineering. arXiv:2604.08224 (2026)
Pith/arXiv arXiv 2026
-
[3]
Xu, R., Yan, Y.: Agent skills for large language models: architecture, acquisition, security, and the path forward. arXiv:2602.12430 (2026)
Pith/arXiv arXiv 2026
-
[4]
Wang, G., Xie, Y., Jiang, Y., Mandlekar, A., Xiao, C., Zhu, Y., Fan, L., Anand- kumar, A.: Voyager: an open-ended embodied agent with large language models. arXiv:2305.16291 (2023)
Pith/arXiv arXiv 2023
-
[5]
Yang, J., Jimenez, C.E., Wettig, A., Lieret, K., Yao, S., Narasimhan, K., Press, O.: SWE-agent: agent-computer interfaces enable automated software engineering. arXiv:2405.15793 (2024)
Pith/arXiv arXiv 2024
-
[6]
Jimenez, C.E., Yang, J., Wettig, A., Yao, S., Pei, K., Press, O., Narasimhan, K.: SWE-bench: can language models resolve real-world GitHub issues? In: International Conference on Learning Representations (2024). arXiv:2310.06770
Pith/arXiv arXiv 2024
-
[7]
In: International Conference on Machine Learning (2024)
Huang, Q., Vora, J., Liang, P., Leskovec, J.: MLAgentBench: evaluating lan- guage agents on machine learning experimentation. In: International Conference on Machine Learning (2024). arXiv:2310.03302
Pith/arXiv arXiv 2024
-
[8]
In: International Conference on Learning Representations (2023)
Yao, S., Zhao, J., Yu, D., Du, N., Shafran, I., Narasimhan, K., Cao, Y.: ReAct: synergizing reasoning and acting in language models. In: International Conference on Learning Representations (2023). arXiv:2210.03629
Pith/arXiv arXiv 2023
-
[9]
In: Advances in Neural Information Processing Systems (2023)
Schick, T., Dwivedi-Yu, J., Dess` ı, R., Raileanu, R., Lomeli, M., Zettlemoyer, L., Cancedda, N., Scialom, T.: Toolformer: language models can teach themselves to use tools. In: Advances in Neural Information Processing Systems (2023). arXiv:2302.04761
Pith/arXiv arXiv 2023
-
[10]
In: Advances in Neural Information Processing Systems (2023)
Shinn, N., Cassano, F., Gopinath, A., Narasimhan, K., Yao, S.: Reflexion: language agents with verbal reinforcement learning. In: Advances in Neural Information Processing Systems (2023). arXiv:2303.11366
Pith/arXiv arXiv 2023
-
[11]
In: Proceedings of 15 EMNLP (2023)
Jiang, H., Wu, Q., Lin, C.-Y., Yang, Y., Qiu, L.: LLMLingua: compressing prompts for accelerated inference of large language models. In: Proceedings of 15 EMNLP (2023). arXiv:2310.05736
Pith/arXiv arXiv 2023
-
[12]
Chen, L., Zaharia, M., Zou, J.: FrugalGPT: how to use large language models while reducing cost and improving performance. arXiv:2305.05176 (2023)
Pith/arXiv arXiv 2023
-
[13]
Ong, I., Almahairi, A., Wu, V., Zhang, W.-L., Willmott, D., Stoica, I., Zaharia, M.: RouteLLM: learning to route LLMs with preference data. arXiv:2406.18665 (2024)
Pith/arXiv arXiv 2024
-
[14]
Transactions of the Association for Computational Linguistics (2024)
Liu, N.F., Lin, K., Hewitt, J., Paranjape, A., Bevilacqua, M., Petroni, F., Liang, P.: Lost in the middle: how language models use long contexts. Transactions of the Association for Computational Linguistics (2024). arXiv:2307.03172
Pith/arXiv arXiv 2024
-
[15]
Laban, P., Hayashi, H., Zhou, Y., Neville, J.: LLMs get lost in multi-turn conversation. arXiv:2505.06120 (2025)
Pith/arXiv arXiv 2025
-
[16]
Agyemang, J.O., Kponyo, J.J., Amponsah, E., Boakye, G.M.A., Agyekum, K.O.- B.O.: Local-Splitter: a measurement study of seven tactics for reducing cloud LLM token usage on coding-agent workloads. arXiv:2604.12301 (2026)
Pith/arXiv arXiv 2026
-
[17]
In: Advances in Neural Information Processing Systems (2025)
Xu, W., Liang, Z., Mei, K., Gao, H., Tan, J., Zhang, Y.: A-MEM: agentic memory for LLM agents. In: Advances in Neural Information Processing Systems (2025). arXiv:2502.12110
Pith/arXiv arXiv 2025
-
[18]
Li, F., Tagkopoulos, P., Tagkopoulos, I.: SkillFlow: scalable and efficient agent skill retrieval system. arXiv:2504.06188 (2026)
arXiv 2026
-
[19]
Journal of General Internal Medicine26(2), 192–196 (2011)
Walker, E., Nowacki, A.S.: Understanding equivalence and noninferiority testing. Journal of General Internal Medicine26(2), 192–196 (2011)
2011
-
[20]
Chapman & Hall, New York (1993)
Efron, B., Tibshirani, R.J.: An Introduction to the Bootstrap. Chapman & Hall, New York (1993)
1993
-
[21]
Scandinavian Journal of Statistics6(2), 65–70 (1979)
Holm, S.: A simple sequentially rejective multiple test procedure. Scandinavian Journal of Statistics6(2), 65–70 (1979)
1979
-
[22]
Springer, Berlin (2012)
Wohlin, C., Runeson, P., H¨ ost, M., Ohlsson, M.C., Regnell, B., Wessl´ en, A.: Experimentation in Software Engineering. Springer, Berlin (2012)
2012
-
[23]
Empirical Software Engineering27(4), 94 (2022)
Baltes, S., Ralph, P.: Sampling in software engineering research: a critical review and guidelines. Empirical Software Engineering27(4), 94 (2022)
2022
-
[24]
Transactions on Machine Learning Research (2024)
Sumers, T.R., Yao, S., Narasimhan, K., Griffiths, T.L.: Cognitive architec- tures for language agents. Transactions on Machine Learning Research (2024). arXiv:2309.02427
Pith/arXiv arXiv 2024
-
[25]
IEEE Transactions on 16 Software Engineering50(4), 911–936 (2024)
Wang, J., Huang, Y., Chen, C., Liu, Z., Wang, S., Wang, Q.: Software testing with large language models: survey, landscape, and vision. IEEE Transactions on 16 Software Engineering50(4), 911–936 (2024)
2024
-
[26]
Software Quality Journal34, 8 (2026)
Haldar, S., Capretz, L.F.: Automated test plan generation using large language models. Software Quality Journal34, 8 (2026). https://doi.org/10.1007/s11219- 026-09744-9
-
[27]
He, J., Rungta, M., Koleczek, D., Sekhon, A., Wang, F.X., Hasan, S.: Does prompt formatting have any impact on LLM performance? arXiv:2411.10541 (2024)
Pith/arXiv arXiv 2024
-
[28]
Tam, Z.R., Wu, C.-K., Tsai, Y.-L., Lin, C.-Y., Lee, H.-Y., Chen, Y.-N.: Let me speak freely? A study on the impact of format restrictions on the performance of large language models. arXiv:2408.02442 (2024) Appendix A Task manifest 17 T able A1Tasks in the analysis set, with category and difficulty Task Category Difficulty 3d-scan-calc industrial and phys...
Pith/arXiv arXiv 2024
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.