Duet instrumentation uses LLM-driven code analysis to instrument performance-relevant changes between two app versions, detecting regressions at up to 5x lower severity than standard duet benchmarks in a testbed evaluation.
An evaluation of open-source software microbenchmark suites for continuous performance assessment
6 Pith papers cite this work, alongside 222 external citations. Polarity classification is still indexing.
citation-role summary
citation-polarity summary
roles
background 1polarities
background 1representative citing papers
A controlled benchmark on 2040 problems reveals poor generalization and high interference in model editing for API updates in code LLMs, with many successes being workarounds rather than true migrations.
CodeXGLUE supplies a standardized collection of 10 code-related tasks, 14 datasets, an evaluation platform, and BERT-, GPT-, and encoder-decoder-style baselines.
Quantized LLMs diverge from their base models at the decision level even when accuracy is preserved, with query and key attention projections showing the greatest structural distortion under low-bit compression.
Hidden dependencies and component variants in SBOMs cause inconsistent vulnerability reporting and VEX handling across scanners.
Systematic review of 97 studies on breaking changes in five software ecosystems, producing a four-dimensional taxonomy, reason/impact categories, 43 detection approaches, and 66 mitigation strategies.
citing papers explorer
-
Duet instrumentation: An Agentic Approach to Improving Sensitivity in Cloud Service Benchmarking
Duet instrumentation uses LLM-driven code analysis to instrument performance-relevant changes between two app versions, detecting regressions at up to 5x lower severity than standard duet benchmarks in a testbed evaluation.
-
Understanding Robustness of Model Editing in Code LLMs
A controlled benchmark on 2040 problems reveals poor generalization and high interference in model editing for API updates in code LLMs, with many successes being workarounds rather than true migrations.
-
CodeXGLUE: A Machine Learning Benchmark Dataset for Code Understanding and Generation
CodeXGLUE supplies a standardized collection of 10 code-related tasks, 14 datasets, an evaluation platform, and BERT-, GPT-, and encoder-decoder-style baselines.
-
The Illusion of Equivalency: Statistical Characterization of Quantization Effects in LLMs
Quantized LLMs diverge from their base models at the decision level even when accuracy is preserved, with query and key attention projections showing the greatest structural distortion under low-bit compression.
-
Hidden Dependencies and Component Variants in SBOM-Based Software Composition Analysis
Hidden dependencies and component variants in SBOMs cause inconsistent vulnerability reporting and VEX handling across scanners.
-
Breaking Changes in Software Ecosystems: A Systematic Literature Review
Systematic review of 97 studies on breaking changes in five software ecosystems, producing a four-dimensional taxonomy, reason/impact categories, 43 detection approaches, and 66 mitigation strategies.