Pith. sign in

cs.SE

Software Engineering

Covers design tools, software metrics, testing and debugging, programming environments, etc. Roughly includes material in all of ACM Subject Classes D.2, except that D.2.4 (program verification) should probably have Logics in Computer Science as the primary subject area.

Papers reviewed in the last 7 days lead, then the papers readers actually read. Ranking is not a quality score.

sort pith recommended most recent

Neuro-symbolic method detects threats in stripped industrial binaries

Lifting semantics into knowledge graphs improves CVE coverage and cuts false positives on real ICS hardware from multiple vendors.

· “Securing the Dark Matter: A Semantic-Enhanced Neuro-Symbolic Framework for Supply Chain Analysis of Opaque Industrial Software”

open re-runnable review →
Figure from the paper

Prefill signals from small LLMs locate root failures in agent traces

Two prefill passes on a 0.6B model rank failure sources in long multi-agent logs with no output tokens and seconds of latency.

· “MASPrism: Lightweight Failure Attribution for Multi-Agent Systems Using Prefill-Stage Signals”

open re-runnable review →
Figure from the paper

Nygard's ADR template outperforms MADR in student usability test

Controlled experiment shows higher overall scores for concise documentation over structured detail capture in architectural decision records

· “One Size Fits All? An Empirical Comparison of ADR Templates regarding Comprehension, Usability, and Ease of Adoption”

open re-runnable review →
Figure from the paper

SWE-TRACE lifts agent success rates while cutting tokens and latency

Shortest-path distillation plus rubric rewards used for both training and step-wise pruning let agents resolve more issues with less compute

· “SWE-TRACE: Optimizing Long-Horizon SWE Agents Through Rubric Process Reward Models and Heuristic Test-Time Scaling”

open re-runnable review →
Figure from the paper

Tool merging and retrieval lifts LLM tool accuracy by up to 38%

Reducing overlapping tools and picking only relevant ones for each query improves selection on standard benchmarks.

· “ToolScope: Enhancing LLM Agent Tool Use through Tool Merging and Context-Aware Filtering”

open re-runnable review →

Late Shopify apps outpaced early movers in 88% of categories

Seven-year panel shows platform governance, not entry timing, decides who grows and who silently fails.

· “Scale, Concentration, and Entry Timing in the Shopify App Ecosystem: A Longitudinal Study of Platform Governance and Application Survival”

open re-runnable review →
Figure from the paper

In-car policies compile into checks that gate every draft move

Compiles policy text into rules, simulates the draft's effects, catching hallucinated IDs and missed side-effects.

· “From Natural Language Policies to Executable Obligations: A Verification Harness for Dependable In-Car LLM Agents”

open re-runnable review →
Figure from the paper

Two legacy science apps became reusable cloud workflows

An LLM guided by architecture reviews and dependency matrices kept each app's outputs identical to the originals.

· “An AI-Assisted Migration Framework for Transforming Legacy Scientific Applications into Reusable Cloud-Based Workflows”

open re-runnable review →
Figure from the paper

Metric feedback loop cuts duplicate code in research notebooks

Six static-analysis metrics, fed back over five refinement rounds, drive measurable structural gains and expose quality trade-offs.

· “From Metrics to Improvement: A Lifecycle-Aware LLM Feedback Framework for Research Software Quality”

open re-runnable review →

Fairness hazard analysis passes real-world hiring test

Two case studies show FHA surfaces recurring bias patterns that HR practitioners call realistic and mostly fixable.

· “Fairness Hazard Analysis for Socio-Technical Processes: A Multiple-Case Study in Bias-sensitive Organisational Settings”

open re-runnable review →

Standard-library imports inflate LLM hallucination rates by 9.4 points

Correcting for Python's built-in modules like os and math changes both hallucination numbers and defense rankings.

· “Evaluating Inference-Time Defenses Against Package Hallucination in LLM-Generated Code”

open re-runnable review →
Figure from the paper

Passing tests says little about code quality in LLM output

Across 340 C# solutions from four LLMs, correctness and quality correlate at just 0.075—Pass@k alone misleads.

· “Benchmarking the Titans: A Multi-Dimensional Empirical Evaluation of LLM Code Generation Quality in the .NET Ecosystem”

open re-runnable review →
Figure from the paper

A model export shifted 12% of alerts while mean score barely moved

Paper argues serving-format conversion is a model change until decisions are re-checked at the deployed threshold.

· “KONTOGRAPH: Verified Point-in-Time Feature Consistency and Amortised Explanation for Real-Time Anti-Money Laundering under a 200 ms Decision Budget”

open re-runnable review →

browse all of cs.SE → full archive · search · sub-categories