PAIR-Bench defines a progressive hinting protocol with failure-region and hint-depth controls to measure LLM code refinement trajectories in detail.
Self-edit: Fault-aware code editor for code generation,
4 Pith papers cite this work, alongside 44 external citations. Polarity classification is still indexing.
representative citing papers
Three code-specific uncertainty axes (lexical, algorithmic, functional) yield an ensemble that raises average AUROC from 0.696 to 0.776 across five code LLMs, with one single-pass signal matching multi-pass baselines at lower cost.
SLoW selects low-frequency word dictionaries to boost LLM translation quality and efficiency across 100 languages from FLORES.
DIP interleaves English word translations into non-English prompts to boost multilingual reasoning on synthetic benchmarks spanning 10-200 languages.
citing papers explorer
-
Benchmarking Code Improvement with Progressive, Adaptive, and Interactive Feedback
PAIR-Bench defines a progressive hinting protocol with failure-region and hint-depth controls to measure LLM code refinement trajectories in detail.
-
Code Is More Than Text: Uncertainty Estimation for Code Generation
Three code-specific uncertainty axes (lexical, algorithmic, functional) yield an ensemble that raises average AUROC from 0.696 to 0.776 across five code LLMs, with one single-pass signal matching multi-pass baselines at lower cost.
-
SLoW: Select Low-frequency Words! Automatic Dictionary Selection for Translation on Large Language Models
SLoW selects low-frequency word dictionaries to boost LLM translation quality and efficiency across 100 languages from FLORES.
-
Dictionary Insertion Prompting for Multilingual Reasoning on Multilingual Large Language Models
DIP interleaves English word translations into non-English prompts to boost multilingual reasoning on synthetic benchmarks spanning 10-200 languages.