CL-Bench is the first expert-validated benchmark for continual learning in frontier LLMs across six real-world domains, showing limited gains and that naive in-context learning outperforms dedicated memory systems.
hub
American Journal of Physics 66 (1): 64–74
6 Pith papers cite this work, alongside 5,927 external citations. Polarity classification is still indexing.
hub tools
representative citing papers
A first concept inventory for dynamic programming was constructed and preliminarily validated with 172 students, though the 'validated' label overstates the current evidence.
ClueNetwork ranks semantic network construction pipelines by multiplying keyphrase quality, edge-weighting interpretability, and community-detection scores into one product objective.
OpenAI's o4-mini solves ~90% of introductory Halliday & Resnick problems, dropping from 96% on text-only to 79% on image-based problems and declining with difficulty.
Introduces a matched four-condition protocol and ONCU metric to diagnose evidence utilization in long-context and RAG models across synthetic and multi-hop QA tasks.
The authors benchmark and release two numerical routes to ratio-distribution PDFs, a 1D double-exponential Mellin convolution and a 2D vectorized Broda-Khan characteristic-function inversion, applied to Hake's ratio and one non-normal case.
citing papers explorer
-
Continual Learning Bench: Evaluating Frontier AI Systems in Real-World Stateful Environments
CL-Bench is the first expert-validated benchmark for continual learning in frontier LLMs across six real-world domains, showing limited gains and that naive in-context learning outperforms dedicated memory systems.
-
Construction and Preliminary Validation of a Dynamic Programming Concept Inventory
A first concept inventory for dynamic programming was constructed and preliminarily validated with 172 students, though the 'validated' label overstates the current evidence.
-
Semantic Networks as Clues: A Theoretical Foundation and Process Optimization for Semantic Network Construction
ClueNetwork ranks semantic network construction pipelines by multiplying keyphrase quality, edge-weighting interpretability, and community-detection scores into one product objective.
-
Assessing AI in Introductory Physics Problem Solving
OpenAI's o4-mini solves ~90% of introductory Halliday & Resnick problems, dropping from 96% on text-only to 79% on image-based problems and declining with difficulty.
-
Diagnosing Evidence Utilization in Long-Context and Retrieval-Augmented Language Models under Matched Evidence Conditions
Introduces a matched four-condition protocol and ONCU metric to diagnose evidence utilization in long-context and RAG models across synthetic and multi-hop QA tasks.
-
Novel computational approaches for ratio distributions with an application to Hake's ratio in effect size measurement
The authors benchmark and release two numerical routes to ratio-distribution PDFs, a 1D double-exponential Mellin convolution and a 2D vectorized Broda-Khan characteristic-function inversion, applied to Hake's ratio and one non-normal case.