Pith. sign in

REVIEW 1 cited by

DependEval: Benchmarking LLMs for Repository Dependency Understanding

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2503.06689 v1 pith:6RHRJBZH submitted 2025-03-09 cs.SE cs.CL

classification cs.SEcs.CL
keywords codellmsunderstandingdependencyrepositoriesrepositorybenchmarkdependeval
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

While large language models (LLMs) have shown considerable promise in code generation, real-world software development demands advanced repository-level reasoning. This includes understanding dependencies, project structures, and managing multi-file changes. However, the ability of LLMs to effectively comprehend and handle complex code repositories has yet to be fully explored. To address challenges, we introduce a hierarchical benchmark designed to evaluate repository dependency understanding (DependEval). Benchmark is based on 15,576 repositories collected from real-world websites. It evaluates models on three core tasks: Dependency Recognition, Repository Construction, and Multi-file Editing, across 8 programming languages from actual code repositories. Our evaluation of over 25 LLMs reveals substantial performance gaps and provides valuable insights into repository-level code understanding.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. VistaDPO: Video Hierarchical Spatial-Temporal Direct Preference Optimization for Large Video Models

    cs.CV 2025-04 conditional novelty 6.0 of 10

    A hierarchical DPO training method and dataset reduce hallucination in video LLMs by aligning preferences at video, clip, object, and token levels.

Pith tools