REVIEW 16 cited by
Let's Learn Step by Step: Enhancing In-Context Learning Ability with Curriculum Learning
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Demonstration ordering, which is an important strategy for in-context learning (ICL), can significantly affects the performance of large language models (LLMs). However, most of the current approaches of ordering require high computational costs to introduce the priori knowledge. In this paper, inspired by the human learning process, we propose a simple but effective demonstration ordering method for ICL, named the few-shot In-Context Curriculum Learning (ICCL). The ICCL implies gradually increasing the complexity of prompt demonstrations during the inference process. The difficulty can be assessed by human experts or LLMs-driven metrics, such as perplexity. Then we design extensive experiments to discuss the effectiveness of the ICCL at both corpus-level and instance-level. Moreover, we also investigate the formation mechanism of LLM's ICCL capability. Experimental results demonstrate that ICCL, developed during the instruction-tuning stage, is effective for representative open-source LLMs. To facilitate further research and applications by other scholars, we make the code publicly available.
Forward citations
Cited by 16 Pith papers
-
IRGPT: Understanding Real-world Infrared Image with Bi-cross-modal Curriculum on Large-scale Benchmark
A vision-language model trained on a new 260K-pair real infrared-text dataset beats general VLMs on a 9-task infrared Q&A benchmark, but the benchmark is in-distribution.
-
ThinkRetrieve: Retrieval-Augmented Reasoning Traces for Test-Time Scaling
Per-step retrieval of solved exemplars injected into the reasoning trace improves test-time scaling accuracy, with up to 13.4 absolute points gained on AIME 2025.
-
A global log for medical AI
MedLog defines a nine-field, syslog-style event log for clinical AI, intended to support real-world surveillance and auditing; the four-deployment validation claimed in the abstract is absent from the body.
-
VL-Cogito: Progressive Curriculum Reinforcement Learning for Advanced Multimodal Reasoning
VL-Cogito, trained with progressive curriculum RL, online difficulty weighting, and dynamic length rewards, matches or beats prior reasoning MLLMs on ten multimodal benchmarks.
-
Exploring Imbalanced Annotations for Effective In-Context Learning
Class-imbalanced annotation sets degrade in-context learning, and reweighting demonstration scores by class weights plus a validation-fitted conditional bias term (RCB) recovers most of the loss.
-
OptiSeq: Ordering Examples On-The-Fly for In-Context Learning
OptiSeq selects the in-context example ordering whose output gets the highest zero-shot log-likelihood, improving few-shot accuracy by up to 10.5 points in tests on API sequencing and classification.
-
Revisiting In-Context Learning with Long Context Language Models
With long-context language models, in-context example selection methods give no reliable gain over random sampling, while adding synthetic examples improves performance by about five percent.
-
PromptRefine: Enhancing Few-Shot Performance on Low-Resource Indic Languages with Example Selection from Related Example Banks
PromptRefine uses alternating minimization over language-specific retrievers plus diversity-aware DPP fine-tuning to select cross-lingual in-context examples, improving few-shot generation in low-resource Indic languages.
-
Does Few-Shot Learning Help LLM Performance in Code Synthesis?
Few-shot example choice measurably affects LLM code output, and two proposed selectors (a perplexity ranker and a trained MLP ranker) each improve CodeLlama's Pass@1 on HumanEval+ by about five points.
-
Learning to Select Visual In-Context Demonstrations
A Dueling-DQN agent selects visual in-context demonstrations and outperforms kNN retrieval on objective regression benchmarks but not on subjective preference tasks, per the paper's main table.
-
The Few-shot Dilemma: Over-prompting Large Language Models
Across seven LLMs on two requirements datasets, F1 scores rise then fall as more few-shot examples are added, and TF-IDF-selected examples at small counts match or beat larger prompts, including a 1% gain over prior SOTA.
-
DICE: Dynamic In-Context Example Selection in LLM Agents via Efficient Knowledge Transfer
DICE dynamically retrieves the most relevant in-context demonstrations at each agent step, and in this preprint it raises exact-match and success-rate scores on HotpotQA, ALFWorld, and Webshop across ReAct, Reflexion,...
-
StaICC: Standardized Evaluation for Classification Task in In-context Learning
StaICC standardizes in-context classification evaluation with fixed prompts and splits, then measures 29 LMs and 10 inference methods under those fixed settings.
-
RoboTron-Drive: All-in-One Large Multimodal Model for Autonomous Driving
One 8B multimodal model trained jointly on six driving datasets outperforms individual specialists on average and transfers zero-shot to three unseen driving benchmarks.
-
Curriculum Demonstration Selection for In-Context Learning
A demonstration selection method that samples one example per difficulty tier improves few-shot language model performance by small and sometimes inconsistent margins.
-
Bridging the Gap: In-Context Learning for Modeling Human Disagreement
Across four open-source LLMs and three subjective-task datasets, in-context learning with multi-perspective prompts improves aggregated-label predictions in zero-shot, but disaggregated hard and soft label predictions...
Discussion (0). Continue with ORCID to comment.