Pith. sign in

REVIEW 3 major objections 5 minor 40 references

Agent Skills often claim one capability and implement another; progressive contrastive learning can catch that mismatch before a skill is loaded or run.

Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →

T0 review · grok-4.5

2026-07-14 10:59 UTC pith:IRZEFDU3

load-bearing objection Solid pre-execution screening paper for Agent Skills: real corpus, clean staged ablations, and large misaligned-class gains; the main open risk is how the Challenge Set was seeded. the 3 major comments →

arxiv 2607.10534 v1 pith:IRZEFDU3 submitted 2026-07-12 cs.AI cs.CRcs.LG

Cross-Layer Misalignment Detection in Agent Skills: A Progressive Loading-Aware Contrastive Learning Approach

classification cs.AI cs.CRcs.LG
keywords Agent Skillsagentic AItrustworthinesscross-layer misalignmentcontrastive learningcontinued pretrainingpre-execution screeninghierarchical artifacts
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

Open skill marketplaces let users and agents choose third-party Agent Skills from short metadata, even though the real behavior lives in deeper instruction and resource layers that load later. The paper names this gap cross-layer misalignment: surface claims that are unsupported, exaggerated, or actively hide unsafe behavior. It argues that ordinary language models, even security-tuned ones, mostly treat plausible packages as aligned and do not systematically check claim-versus-evidence. Using a large normalized corpus of open-source skills and a human-verified challenge set of real packages, the authors show that skill-format adaptation alone is weak, while a progressive contrastive method that trains on aligned packages versus metadata-swapped and corrupted-claim variants turns cross-layer consistency into a usable pre-execution signal. A sympathetic reader should care because trust decisions about skills are made before full inspection or execution, and a reliable screen would protect marketplaces, users, and operators without waiting for runtime failure.

Core claim

Unadapted general and cybersecurity LLMs barely detect misaligned Agent Skills (Macro-F1 near 0.45, misaligned-class F1 below 0.15), two-stage continued pretraining only modestly helps format familiarity, and Progressive Loading-Aware Hierarchical Contrastive Learning on top of that raises Macro-F1 to about 0.87–0.89 and misaligned-class F1 to roughly 0.79–0.82 on a held-out human-verified challenge set of naturally occurring skills.

What carries the argument

Progressive Loading-Aware Hierarchical Contrastive Learning (PL-HCL): after two-stage continued pretraining on short metadata–instruction views then full package views, the model scores same-skill layers higher than metadata-swapped and corrupted-claim negatives so it learns whether user-facing claims are supported by instructions and resources.

Load-bearing premise

The method assumes that synthetic training pairs—swapped metadata and corrupted capability claims—teach a claim-versus-evidence signal that still works on real misaligned skills found in the wild.

What would settle it

Hold out a fresh set of naturally occurring misaligned skills whose failure modes were not used to design the synthetic swap or corruption recipes; if CPT+PL-HCL then fails to raise misaligned-class F1 well above CPT-only and unadapted baselines under the same fixed decision rule, the central transfer claim fails.

Watch this falsifier — get emailed when new claim-graph text bears on it.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

3 major / 5 minor

Summary. The paper formalizes cross-layer misalignment in Agent Skills—mismatch between user-facing metadata and deeper instruction/resource evidence—as a pre-execution, artifact-level detection problem. It constructs a large open-source corpus (≈265k normalized packages) and a human-verified Challenge Set (1,444 skills; 1,150 aligned / 294 misaligned), then proposes PL-HCL: two-stage continued pretraining on short then full skill views, followed by hierarchical contrastive learning that pulls same-skill layers together and pushes metadata-swapped and corrupted-claim negatives apart (Eqs. 1–3, 7–8). Staged experiments E0–E2 show unadapted 8B LLMs yield Macro-F1 ≈0.45–0.52 with F1_m <0.15; CPT alone yields modest format adaptation; CPT+PL-HCL raises Macro-F1 to 0.87–0.89 and F1_m to ≈0.79–0.82 on the Challenge Set across Llama-3.1-8B and Foundation-Sec-8B, with qualitative claim–body audits illustrating the learned signal.

Significance. If the reported gains transfer beyond the Challenge Set construction pipeline, this is a timely and practically useful contribution to trustworthy agentic AI: it targets a real supply-side risk in skill marketplaces, separates misalignment from pure malice, and supplies a deployable pre-execution screen rather than post-hoc task success. Strengths include the large normalized corpus, the staged E0–E2 isolation of CPT vs. contrastive learning, dual-annotator agreement (97.2% on a 144-sample re-code), intrinsic NTP and contrastive diagnostics, qualitative audits that separate malicious and benign-misaligned cases, an honest limitations section, and released dataset/code links. The progressive-loading framing also generalizes to other layered artifacts (extensions, MCP servers, package registries).

major comments (3)
  1. [§4.3–4.4, Supplement B, Tables 1–3] Supplement B states that Challenge Set gold labels begin with MASB scanner severity-ranked candidate flags, after which only flagged skills receive sandbox inspection and human coding. The main text (§4.3–4.4, §5) repeatedly calls the 294 misaligned packages “naturally occurring, in-the-wild” and asserts that training negatives are synthetic while evaluation is not. The manuscript does not report (i) the fraction of the corpus that was scanner-flagged, (ii) how many unflagged packages entered the Challenge Set, or (iii) F1_m / Macro-F1 stratified by whether a skill would have been MASB-flagged. Because the central claim is transfer from synthetic swap/corruption negatives (Eqs. 2–3) to real marketplace misalignment, this selection mechanism is load-bearing: if misaligned examples are disproportionately surface-detectable by a security scanner, the jump from F1_m <0.15 to ≈0.79–0.82 may n
  2. [§4.3–4.4, Eqs. (6)–(8), Tables 1–3] PL-HCL is trained with a multi-view contrastive objective (Eqs. 6–8) over aligned, swap, and corruption views, yet extrinsic evaluation (§4.4, Tables 1–3) is zero-/few-shot prompting that asks the model to predict aligned vs. misaligned. The manuscript never states whether inference uses the learned cross-layer score a_θ,ϕ(S), a classification head on projected embeddings, or only the CPT+contrastively-adapted decoder under natural-language prompts. Without this mapping, it is unclear what signal is actually being measured in E2 and whether the large F1_m gains require the contrastive geometry at test time or only representation adaptation. Please specify the exact inference procedure used for all reported Challenge Set numbers and, if prompting is the sole path, justify why the contrastive score is not used as a primary detector.
  3. [§5, Tables 1–3] E0–E2 compare base, CPT-only, and CPT+PL-HCL LLM checkpoints, but there is no non-LLM or lightweight baseline that directly implements the claimed claim-versus-evidence idea (e.g., cosine similarity between independently encoded M and I/R layers, a simple NLI-style entailment check of description against SKILL.md, or keyword/permission heuristics). Given that PL-HCL’s design hypothesis is precisely cross-layer consistency, such baselines are needed to show that the full hierarchical contrastive pipeline is necessary for the reported Macro-F1 range rather than that any skill-aware embedding comparison would suffice. An ablation removing swap vs. corruption negatives (or short-view vs. full-view PL-HCL) would similarly strengthen the causal attribution in Table 3.
minor comments (5)
  1. [Abstract, Table 1, §5.1] Table 1 reports the “strongest prompting setting” per backbone; although Tables 2–3 give full k∈{0,2,5} results, the abstract and main-findings prose still lead with the peak numbers. Consider leading with a fixed-k (e.g., k=2) or mean-over-k summary to avoid any appearance of test-set configuration selection.
  2. [§5.4, Figure 2] In §5.4 / Figure 2 the base model is named “Foundation-Sec-8B-Reasoning” while tables use “Foundation-Sec-8B”; please unify backbone names and state whether reasoning-mode decoding was used only for the qualitative audit.
  3. [Eq. (7), Table 7, Supplement E] Layer weights α_MI, α_MR, α_IR are fixed at (1/3,1/3,1/3) and pair-kind weights at 1.0:1.5:1.0 (Table 7) with no sensitivity analysis. A short appendix sweep would help readers assess robustness of free parameters listed in the training setup.
  4. [Supplement A, Table 4, §4.2] Corpus construction (Supplement A) reports Stage-2 coverage of only 34.4% of normalized packages due to the 10,240-token budget and empty resource layers. Briefly discuss in the main text whether full-view PL-HCL gains are driven by the resource-rich subset and how empty-R packages are scored at inference.
  5. [Abstract, Supplement F] Minor wording: “cross-layer misalignment” is hyphenated inconsistently in the abstract line break; “Foundation-Sec-8B” vs “FS-8B-R” abbreviations in Supplement F should be defined once in the main text.

Circularity Check

0 steps flagged

No circularity: synthetic contrastive training is held out from a human-labeled natural Challenge Set; reported gains are not forced by construction.

full rationale

The paper's load-bearing chain is: (i) formalize cross-layer misalignment as pre-execution inconsistency between metadata M and deeper layers I/R; (ii) adapt LLMs via two-stage CPT on unlabeled skill packages; (iii) train PL-HCL by contrasting same-package positives P_i=(M_i,I_i,R_i) against synthetic swap negatives N^A_i=(M_j,I_i,R_i) and corruption negatives N^B_i=(M̃_i,I_i,R_i) (Eqs. 1–3, 7–8); (iv) evaluate on a held-out Challenge Set of 1,444 naturally occurring SkillsMP packages with human gold labels, excluded from CPT and PL-HCL. None of these steps reduces a claimed prediction to its inputs by definition. Training negatives are synthetic and never use Challenge Set labels; extrinsic metrics (Macro-F1, F1_m) are computed against independent human ALIGNED/MISALIGNED labels, not against the contrastive score itself. CPT is ordinary causal LM adaptation, not a fit of the evaluation target. There is no uniqueness theorem, ansatz smuggled via self-citation, or renaming of a known result as a first-principles derivation. The MASB-scanner seeding of Challenge Set candidates (Supplement B) is a possible selection/generalization concern, not circularity: it does not make the reported F1 jump equal the training objective by construction. Self-citation of Zhang et al. (2026) on soft prompting is peripheral related-work only. Score 0.

Axiom & Free-Parameter Ledger

6 free parameters · 5 axioms · 3 invented entities

The central claim rests on empirical ML assumptions and design choices rather than a closed-form derivation. Load-bearing free parameters are training hyperparameters of CPT/PL-HCL. Domain axioms concern the layered progressive-loading structure of skills and the validity of synthetic negatives as proxies for real misalignment. Invented entities are the problem framing and the PL-HCL pipeline itself; they are operationalized by data and code rather than postulated unobserved physics-like objects.

free parameters (6)
  • contrastive temperature τ
    Fixed at 0.07 in PL-HCL loss (Eq. 8 / Table 7); scales the softmax over aligned vs negative package scores and affects separation strength.
  • layer similarity weights α_MI, α_MR, α_IR
    Set to (1/3,1/3,1/3) in the cross-layer score (Eq. 7); hand-chosen equal weighting of metadata–instruction, metadata–resource, and instruction–resource similarities.
  • type-B corruption rates (stage schedule)
    0.5 / 0.3 / 0.2 across stages (Table 7); controls how often corrupted-claim negatives are sampled and thus the training mixture.
  • CPT and PL-HCL peak learning rates
    2e-5 (CPT) and 1e-5 (PL-HCL) with AdamW cosine schedule; standard but free choices that affect adaptation quality.
  • token context budgets (4,096 / 10,240)
    Stage-1 metadata+instruction and Stage-2 full-package caps determine which packages enter training and how much resource evidence is seen.
  • pair-kind weights 1.0 : 1.5 : 1.0
    Relative weighting of positive / swap / corruption pair kinds during PL-HCL (Table 7); hand-set training emphasis.
axioms (5)
  • domain assumption Agent Skills are hierarchical progressive-loading artifacts S=(M,I,R) where metadata is typically available before instructions and resources.
    Stated throughout §1 and §4.1; structures both the curriculum CPT stages and the contrastive views.
  • ad hoc to paper Metadata-swapped and metadata-corrupted packages are valid negative proxies for cross-layer misalignment.
    Core of PL-HCL negative construction (§4.3, Eqs. 2–3); the transfer claim depends on this proxy being close enough to natural misalignment.
  • domain assumption Pre-execution package artifacts (without runtime traces) are sufficient to assess whether surface claims are supported by deeper evidence.
    Defines the evaluation target in §4.1 and limitations in §6; excludes dynamic remote payloads and environment-specific effects by design.
  • domain assumption Causal language modeling on skill text adapts format/vocabulary but does not by itself teach claim–evidence consistency.
    Motivates adding PL-HCL after CPT (§4.2–4.3); empirically supported by E1 vs E2 but still an architectural assumption about what CLM learns.
  • standard math Standard contrastive multi-view learning (pull positives, push negatives via cosine similarity and temperature softmax) is a valid objective for layer consistency.
    Uses conventional SimCLR/CLIP-style machinery (Eqs. 6–8) without new theoretical guarantees.
invented entities (3)
  • cross-layer misalignment (for Agent Skills) independent evidence
    purpose: Names the pre-execution inconsistency between user-facing metadata claims and instruction/resource evidence, including both malicious and benign overclaim cases.
    Problem definition introduced in §1; operationalized via labels on the Challenge Set rather than an external physical entity.
  • PL-HCL (Progressive Loading-Aware Hierarchical Contrastive Learning) no independent evidence
    purpose: Training framework that stages short then full package views and contrasts aligned packages against swap and corruption negatives.
    Method contribution of the paper (§4.3); existence is the proposed algorithm plus trained checkpoints, not an unobserved mediator.
  • Normalized layered skill representation S=(M,I,R) no independent evidence
    purpose: Canonical package view used for corpus construction, CPT, contrastive learning, and evaluation.
    Modeling choice in §4.1; useful abstraction of marketplace+GitHub packages, not independently measured outside this pipeline.

pith-pipeline@v1.1.0-grok45 · 26966 in / 3957 out tokens · 48571 ms · 2026-07-14T10:59:48.380972+00:00 · methodology

0 comments
read the original abstract

Large language model (LLM) agents are increasingly extended through Agent Skills, reusable artifacts that package natural-language metadata, procedural instructions, and execution-time resources for runtime use. As open-source skill marketplaces expand, users and agents increasingly rely on brief metadata to select third-party skills, making it difficult to detect inconsistencies between a skill's description and its true behavior, a problem we call cross-layer misalignment. To address this issue, we propose Progressive Loading-Aware Hierarchical Contrastive Learning (PL-HCL), an LLM-based framework that detects misalignment by modeling the layered structure of Agent Skills and learning cross-layer consistency. Using a normalized corpus of over 264,000 open-source skills and a human-verified challenge set, PL-HCL improves Macro-F1 from approximately 0.45 for unadapted baselines to 0.87-0.89 across evaluated LLM backbones. This approach offers an effective screening tool for users and operators, as well as design principles for detecting inconsistencies in layered digital artifacts.

Figures

Figures reproduced from arXiv: 2607.10534 by ChengJun Zhang, Jianna Hur, Jingjing Zhang, Sagar Samtani, Yang Gao.

Figure 1
Figure 1. Figure 1: Overview of the PL-HCL framework. Normalized skill packages are represented as metadata [PITH_FULL_IMAGE:figures/full_fig_p005_1.png] view at source ↗
Figure 2
Figure 2. Figure 2: Description–body audit for two Challenge Set skills flipped from [PITH_FULL_IMAGE:figures/full_fig_p008_2.png] view at source ↗

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Reference graph

Works this paper leans on

40 extracted references · 14 linked inside Pith

  1. [1]

    Anthropic. 2025. Equipping Agents for the Real World with Agent Skills. En- gineering at Anthropic. https://www.anthropic.com/engineering/equipping- agents-for-the-real-world-with-agent-skills Accessed: 2026-06-08

  2. [2]

    Iz Beltagy, Kyle Lo, and Arman Cohan. 2019. SciBERT: A Pretrained Language Model for Scientific Text. InProceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing. Association for Computational Linguistics, Hong Kong, China, 3615–3620

  3. [3]

    Yoshua Bengio, Jérôme Louradour, Ronan Collobert, and Jason Weston. 2009. Curriculum Learning. InProceedings of the 26th Annual International Conference on Machine Learning. Association for Computing Machinery, New York, NY, USA, 41–48

  4. [4]

    Ilias Chalkidis, Manos Fergadiotis, Prodromos Malakasiotis, Nikolaos Aletras, and Ion Androutsopoulos. 2020. LEGAL-BERT: The Muppets Straight Out of Law School. InFindings of the Association for Computational Linguistics: EMNLP

  5. [5]

    Association for Computational Linguistics, Online, 2898–2904

  6. [6]

    Ting Chen, Simon Kornblith, Mohammad Norouzi, and Geoffrey Hinton. 2020. A Simple Framework for Contrastive Learning of Visual Representations. In Proceedings of the 37th International Conference on Machine Learning (Proceedings of Machine Learning Research, Vol. 119). PMLR, 1597–1607

  7. [7]

    Edoardo Debenedetti, Jie Zhang, Mislav Balunovic, Luca Beurer-Kellner, Marc Fischer, and Florian Tramèr. 2024. AgentDojo: A Dynamic Environment to Evaluate Attacks and Defenses for LLM Agents.arXiv preprint arXiv:2406.13352 (2024)

  8. [8]

    Kai Greshake, Sahar Abdelnabi, Shailesh Mishra, Christoph Endres, Thorsten Holz, and Mario Fritz. 2023. Not What You’ve Signed Up For: Compromising Real-World LLM-Integrated Applications with Indirect Prompt Injection. In Proceedings of the 16th ACM Workshop on Artificial Intelligence and Security. 79–90

  9. [9]

    Suchin Gururangan, Ana Marasović, Swabha Swayamdipta, Kyle Lo, Iz Beltagy, Doug Downey, and Noah A. Smith. 2020. Don’t Stop Pretraining: Adapt Language Models to Domains and Tasks. InProceedings of the 58th Annual Meeting of the Association for Computational Linguistics. Association for Computational Linguistics, Online, 8342–8360

  10. [10]

    Tingxu Han, Yi Zhang, Wei Song, Chunrong Fang, Zhenyu Chen, Youcheng Sun, and Lijie Hu. 2026. SWE-Skills-Bench: Do Agent Skills Actually Help in Real-World Software Engineering?arXiv preprint arXiv:2603.15401(2026)

  11. [11]

    Chang Jin, An Wang, Zeming Wei, Kai Wang, Biaojie Zeng, Qiaosheng Zhang, Chao Yang, Jingjing Qu, Xia Hu, and Xingcheng Xu. 2026. SkillSafetyBench: Evaluating Agent Safety under Skill-Facing Attack Surfaces.arXiv preprint arXiv:2605.12015(2026)

  12. [12]

    Joyce, Gideon Miller, Phil Roth, Richard Zak, Elliott Zaresky-Williams, Hyrum Anderson, Edward Raff, and James Holt

    Robert J. Joyce, Gideon Miller, Phil Roth, Richard Zak, Elliott Zaresky-Williams, Hyrum Anderson, Edward Raff, and James Holt. 2025. EMBER2024: A Bench- mark Dataset for Holistic Evaluation of Malware Classifiers.arXiv preprint arXiv:2506.05074(2025)

  13. [13]

    Prannay Khosla, Piotr Teterwak, Chen Wang, Aaron Sarna, Yonglong Tian, Phillip Isola, Aaron Maschinot, Ce Liu, and Dilip Krishnan. 2020. Supervised Contrastive Learning. InAdvances in Neural Information Processing Systems, Vol. 33. 18661–18673

  14. [14]

    Pang Wei Koh, Shiori Sagawa, Henrik Marklund, Sang Michael Xie, Marvin Zhang, Akshay Balsubramani, Weihua Hu, Michihiro Yasunaga, Richard Lanas Phillips, Irena Gao, et al. 2021. WILDS: A Benchmark of in-the-Wild Distribution Shifts. InProceedings of the 38th International Conference on Machine Learning. 5637–5664

  15. [15]

    Pawan Kumar, Benjamin Packer, and Daphne Koller

    M. Pawan Kumar, Benjamin Packer, and Daphne Koller. 2010. Self-Paced Learning for Latent Variable Models. InAdvances in Neural Information Processing Systems, Vol. 23. Curran Associates, Inc

  16. [16]

    Jinhyuk Lee, Wonjin Yoon, Sungdong Kim, Donghyeon Kim, Sunkyu Kim, Chan Ho So, and Jaewoo Kang. 2020. BioBERT: A Pre-trained Biomedical Lan- guage Representation Model for Biomedical Text Mining.Bioinformatics36, 4 (2020), 1234–1240

  17. [17]

    Dongyang Li, Taolin Zhang, Nan Hu, Chengyu Wang, and Xiaofeng He. 2022. Hi- CLRE: A Hierarchical Contrastive Learning Framework for Distantly Supervised Relation Extraction.arXiv preprint arXiv:2202.13352(2022)

  18. [18]

    Xiangyi Li, Wenbo Chen, Yimin Liu, Shenghan Zheng, Xiaokun Chen, Yifeng He, Yubo Li, Bingran You, Haotian Shen, Jiankai Sun, Shuyi Wang, Binxu Li, Qunhong Zeng, Di Wang, Xuandong Zhao, Yuanli Wang, Roey Ben Chaim, Zonglin Di, Yipeng Gao, Junwei He, Yizhuo He, Liqiang Jing, Luyang Kong, Xin Lan, Jiachen Li, Songlin Li, Yijiang Li, Yueqian Lin, Xinyi Liu, X...

  19. [19]

    Yi Liu, Zhihao Chen, Yanjun Zhang, Gelei Deng, Yuekang Li, Jianting Ning, and Leo Yu Zhang. 2026. Malicious Agent Skills in the Wild: A Large-Scale Security Empirical Study.arXiv preprint arXiv:2602.06547(2026)

  20. [20]

    Yi Liu, Weizhe Wang, Ruitao Feng, Yao Zhang, Guangquan Xu, Gelei Deng, Yuekang Li, and Leo Zhang. 2026. Agent Skills in the Wild: An Empirical Study of Security Vulnerabilities at Scale.arXiv preprint arXiv:2601.10338(2026)

  21. [21]

    Hassan, David Lo, Zhenchang Xing, and Xinyu Wang

    Zhongxin Liu, Xin Xia, Ahmed E. Hassan, David Lo, Zhenchang Xing, and Xinyu Wang. 2018. Neural-Machine-Translation-Based Commit Message Generation: How Far Are We?. InProceedings of the 33rd ACM/IEEE International Conference on Automated Software Engineering. Association for Computing Machinery, New York, NY, USA, 373–384

  22. [22]

    Grégoire Mialon, Clémentine Fourrier, Craig Swift, Thomas Wolf, Yann LeCun, and Thomas Scialom. 2023. GAIA: A Benchmark for General AI Assistants.arXiv preprint arXiv:2311.12983(2023)

  23. [23]

    Mitchell

    Emmanouil Antonios Platanios, Otilia Stretcu, Graham Neubig, Barnabas Poczos, and Tom M. Mitchell. 2019. Competence-based Curriculum Learning for Neural Machine Translation. InProceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Tech- nologies. Association for Computational Lingu...

  24. [24]

    Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, Gretchen Krueger, and Ilya Sutskever. 2021. Learning Transferable Visual Models From Natural Language Supervision. InProceedings of the 38th International Conference on Machine Learning. PMLR, 8748–8763

  25. [25]

    Timo Schick, Jane Dwivedi-Yu, Roberto Dessì, Roberta Raileanu, Maria Lomeli, Luke Zettlemoyer, Nicola Cancedda, and Thomas Scialom. 2023. Toolformer: Language Models Can Teach Themselves to Use Tools. InAdvances in Neural Information Processing Systems, Vol. 36

  26. [26]

    David Schmotz, Luca Beurer-Kellner, Sahar Abdelnabi, and Maksym An- driushchenko. 2026. Skill-Inject: Measuring Agent Vulnerability to Skill File Attacks.arXiv preprint arXiv:2602.20156(2026)

  27. [27]

    Natalie Shapira, Chris Wendler, Avery Yen, Gabriele Sarti, Koyena Pal, Olivia Floody, Adam Belfki, Alex Loftus, Aditya Ratan Jannali, Nikhil Prakash, Jas- mine Cui, Giordano Rogers, Jannik Brinkmann, Can Rager, Amir Zur, Michael Ripa, Aruna Sankaranarayanan, David Atkinson, Rohit Gandikota, Jaden Fiotto- Kaufman, EunJeong Hwang, Hadas Orgad, P. Sam Sahil,...

  28. [28]

    SkillsMP. 2026. SkillsMP: Agent Skills Marketplace. https://skillsmp.com/. Ac- cessed: 2026-06-09

  29. [29]

    Lin Tan, Ding Yuan, and Yu Zhou. 2012. iComment: Bugs or Bad Comments?. InProceedings of the 21st ACM SIGSOFT International Symposium on Software Testing and Analysis. Association for Computing Machinery, New York, NY, USA, 145–155

  30. [30]

    Yonglong Tian, Dilip Krishnan, and Phillip Isola. 2020. Contrastive Multiview Coding. InComputer Vision – ECCV 2020. Springer, 776–794

  31. [31]

    Zihan Wang, Peiyi Wang, Lianzhe Huang, Xin Sun, and Houfeng Wang. 2022. Incorporating Hierarchy into Text Encoder: A Contrastive Learning Approach for Hierarchical Text Classification.arXiv preprint arXiv:2203.03825(2022)

  32. [32]

    Bingchao Wu, Yangyuxuan Kang, Daoguang Zan, Bei Guan, and Yongji Wang

  33. [33]

    Hierarchical and Contrastive Representation Learning for Knowledge- Aware Recommendation.arXiv preprint arXiv:2304.07506(2023)

  34. [34]

    Shunyu Yao, Jeffrey Zhao, Dian Yu, Nan Du, Izhak Shafran, Karthik Narasimhan, and Yuan Cao. 2023. ReAct: Synergizing Reasoning and Acting in Language Models. InInternational Conference on Learning Representations

  35. [35]

    Qiusi Zhan, Zhixiang Liang, Zifan Ying, and Daniel Kang. 2024. InjecAgent: Benchmarking Indirect Prompt Injections in Tool-Integrated Large Language Model Agents.arXiv preprint arXiv:2403.02691(2024)

  36. [36]

    Chengjun Zhang, Benjamin Martin Ampel, and Sagar Samtani. 2026. A Domain- Adaptive Soft Prompting Framework for Multi-Type Bias Detection in News. In Proceedings of the Hawaii International Conference on System Sciences

  37. [37]

    Shu Zhang, Ran Xu, Caiming Xiong, and Chetan Ramaiah. 2022. Use All the Labels: A Hierarchical Multi-Label Contrastive Learning Framework.arXiv preprint arXiv:2204.13207(2022)

  38. [38]

    Shurui Zhou, Yang Liu, Jing Kai Siow, Xiaoning Du, and Yang Liu. 2017. Auto- matically Detecting API Misuses Using Statistical Learning. InProceedings of the 2017 11th Joint Meeting on Foundations of Software Engineering. Association for Computing Machinery, New York, NY, USA, 843–853. Cross-Layer Misalignment Detection in Agent Skills: A Progressive Load...

  39. [39]

    For both full-pipeline backbones, CPT Sub-stage 2 enhances final held-out skill-text metrics, reducing perplexity and boosting top-1 accuracy

    Next-Token Prediction:Table 8 reports NTP metrics before and after CPT. For both full-pipeline backbones, CPT Sub-stage 2 enhances final held-out skill-text metrics, reducing perplexity and boosting top-1 accuracy

  40. [40]

    PL-HCL Contrastive Objective:Table 9 reports held-out PL-HCL performance by sub-stage and pair kind. The pair-kind decomposition shows that corrupted type-B pairs are more difficult than swapped type-A pairs, particularly for Foundation-Sec-8B-R, while Llama-3.1-8B performs strongly across all pair types. G Supplementary Results of Full Evaluation Table 1...