A new benchmark, CoCoNUT, shows that LLMs that can generate working Python code often fail to list the exact lines of code that execute, especially for long or advanced programs.
Introducing bloomberggpt, bloomberg’s 50-billion param- eter large language model, purpose-built from scratch for finance,
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
citation-role summary
background 1
citation-polarity summary
fields
cs.LG 1years
2025 1verdicts
CONDITIONAL 1roles
background 1polarities
background 1representative citing papers
citing papers explorer
-
CoCoNUT: Structural Code Understanding does not fall out of a tree
A new benchmark, CoCoNUT, shows that LLMs that can generate working Python code often fail to list the exact lines of code that execute, especially for long or advanced programs.