SyncBench quantifies agent out-of-sync recovery from 21 real repositories and finds state-of-the-art LLM agents succeed in under 34% of tasks and ask for help in under 5% of turns.
Are they the same picture? adapting concept bottleneck models for human-ai collaboration in image retrieval
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.SE 1years
2025 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
SyncMind: Measuring Agent Out-of-Sync Recovery in Collaborative Software Engineering
SyncBench quantifies agent out-of-sync recovery from 21 real repositories and finds state-of-the-art LLM agents succeed in under 34% of tasks and ask for help in under 5% of turns.