Pith. sign in

REVIEW 3 cited by

Large Language Models (LLMs) for Source Code Analysis: applications, models and datasets

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2503.17502 v1 pith:IUQBXN6W submitted 2025-03-21 cs.SE cs.AIcs.CL

classification cs.SEcs.AIcs.CL
keywords analysiscodellmsmodelsdatasetssourcewhatapplications
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Large language models (LLMs) and transformer-based architectures are increasingly utilized for source code analysis. As software systems grow in complexity, integrating LLMs into code analysis workflows becomes essential for enhancing efficiency, accuracy, and automation. This paper explores the role of LLMs for different code analysis tasks, focusing on three key aspects: 1) what they can analyze and their applications, 2) what models are used and 3) what datasets are used, and the challenges they face. Regarding the goal of this research, we investigate scholarly articles that explore the use of LLMs for source code analysis to uncover research developments, current trends, and the intellectual structure of this emerging field. Additionally, we summarize limitations and highlight essential tools, datasets, and key challenges, which could be valuable for future work.

Discussion (0). Sign in to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Evaluating Fine-Tuning and Metrics for Neural Decompilation of Dart AOT Binaries

    cs.SE 2026-07 accept novelty 6.0 of 10

    Fine-tuning small LLMs for Dart decompilation yields no functional improvement and surface metrics can diverge from correctness.

  2. Can Small GenAI Language Models Rival Large Language Models in Understanding Application Behavior?

    cs.SE 2025-11 reject novelty 3.0 of 10

    A five-model benchmark on a self-built malware dataset shows the smallest strong model, Phi-4-mini, outperforming 7-8B models, contradicting the paper's own 'larger models win' framing.

  3. Position Paper: Programming Language Techniques for Bridging LLM Code Generation Semantic Gaps

    cs.SE 2025-07 unverdicted novelty 2.0 of 10

    A position paper arguing that PL techniques, especially formal verification and structure-aware representations, should be deeply integrated into LLM code generation.

Pith tools