Pith. sign in

REVIEW 6 cited by

LAMOL: LAnguage MOdeling for Lifelong Language Learning

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 1909.03329 v2 pith:GB2NMYRA submitted 2019-09-07 cs.CL cs.AI

LAMOL: LAnguage MOdeling for Lifelong Language Learning

classification cs.CL cs.AI
keywords lamollanguagemodeltaskslearninglifelongpreviousmodeling
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved
0 comments
Share X Bluesky LinkedIn Reddit HN
read the original abstract

Most research on lifelong learning applies to images or games, but not language. We present LAMOL, a simple yet effective method for lifelong language learning (LLL) based on language modeling. LAMOL replays pseudo-samples of previous tasks while requiring no extra memory or model capacity. Specifically, LAMOL is a language model that simultaneously learns to solve the tasks and generate training samples. When the model is trained for a new task, it generates pseudo-samples of previous tasks for training alongside data for the new task. The results show that LAMOL prevents catastrophic forgetting without any sign of intransigence and can perform five very different language tasks sequentially with only one model. Overall, LAMOL outperforms previous methods by a considerable margin and is only 2-3% worse than multitasking, which is usually considered the LLL upper bound. The source code is available at https://github.com/jojotenya/LAMOL.

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Forward citations

Cited by 6 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. ReCoLoRA: Spectrum-Aware Recursive Consolidation for Continual LLM Fine-Tuning

    cs.LG 2026-07 conditional novelty 6.0

    Spectrum-initialized LoRA with elbow ranks and recursive SVD consolidation of the effective weight beats rank-swept PEFT baselines on three of four 7–8B models in continual GLUE fine-tuning.

  2. Task-Differentiated Atomic Skill Expansion and Routing for Continual Learning Across Highly Heterogeneous Tasks

    cs.LG 2026-06 unverdicted novelty 6.0

    TASER dynamically expands and orthogonality-constrains atomic skills then routes them with task-conditioned gating, outperforming baselines on the new 19-task HeteroCLBench benchmark for heterogeneous continual learning.

  3. Robust Policy Optimization to Prevent Catastrophic Forgetting

    cs.LG 2026-02 unverdicted novelty 6.0

    FRPO applies a max-min robust optimization over KL-bounded policy neighborhoods during RLHF to reduce catastrophic forgetting of safety and accuracy under subsequent SFT or RL fine-tuning.

  4. Attribution-Guided Continual Learning for Large Language Models

    cs.LG 2026-05 unverdicted novelty 5.0

    An attribution-based continual learning framework for LLMs modulates per-parameter gradients using task-specific importance scores to reduce forgetting of prior tasks.

  5. Attribution-Guided Continual Learning for Large Language Models

    cs.LG 2026-05 conditional novelty 5.0

    LRP-derived element-wise parameter importance scores gate gradients so parameters critical to earlier tasks receive smaller updates during continual LLM fine-tuning.

  6. OrthoPhys: Physically Plausible Video Generation with Orthogonal-View Geometry Guidance

    cs.CV 2026-03 unverdicted novelty 5.0

    Generating synchronized four-view orthogonal foreground videos with geometry-enhanced attention, then using them as rigid guidance, improves physical realism in video generation over direct 2D methods.