Pith. sign in

Beyond accuracy: Evaluating self-consistency of code large language models with identitychain

5 Pith papers cite this work. Polarity classification is still indexing.

5 Pith papers citing it

fields

cs.CL 3 cs.SE 2

verdicts

UNVERDICTED 5

representative citing papers

CodeMind: Evaluating Large Language Models for Code Reasoning

cs.SE · 2024-02-15 · unverdicted · novelty 7.0

CodeMind evaluates ten LLMs on four benchmarks using three new code reasoning tasks, finding performance varies by model size and drops with complexity while showing no correlation with bug repair ability.

LLMs Corrupt Your Documents When You Delegate

cs.CL · 2026-04-17 · unverdicted · novelty 6.0

LLMs corrupt an average of 25% of document content during long delegated editing workflows across 52 domains, even frontier models, and agentic tools do not mitigate the issue.

citing papers explorer

Showing 5 of 5 citing papers.