Pith. sign in

REVIEW 1 cited by

Quality evaluation of Tabby coding assistant using real source code snippets

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2504.08650 v1 pith:OS5Z3RAD submitted 2025-04-11 cs.SE

Quality evaluation of Tabby coding assistant using real source code snippets

classification cs.SE
keywords codecodinglanguageassistanceassistantqualitytabbymlaccuracy
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved
0 comments
read the original abstract

Large language models have become a popular tool in software development, providing coding assistance. The proper measurement of the accuracy and reliability of the code produced by such tools is a challenge due to natural language prompts. We propose a simple pipeline that uses state-of-the-art implementation of classic and universal genres of algorithms and data structures. We focus on measuring the quality of TabbyML code assistant due to its open licence and the flexibility in the choice of the language model. Our results presented as cyclomatic complexity, Halstead's Bugs \& Effort and four text-based similarity matrices depict the usability of TabbyML in coding assistance tasks.

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. Do These Violent Delights Have Violent Ends? Measuring the Post-Merge Fate of Agentic Code

    cs.SE 2026-07 conditional novelty 6.5

    Post-merge, agentic code needs ~46–51% more corrective/bug-fix maintenance and introduces more security and dependency findings than human code, with higher burden in low-review projects.