Pith. sign in

REVIEW 3 cited by

Is Your LLM Outdated? A Deep Look at Temporal Generalization

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2405.08460 v3 pith:RC3WKOKA submitted 2024-05-14 cs.CL cs.AI

classification cs.CLcs.AI
keywords temporalmodelsgeneralizationllmsadaptabilitybiasdeclineevaluation
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

The rapid advancement of Large Language Models (LLMs) has led to the development of benchmarks that consider temporal dynamics, however, there remains a gap in understanding how well these models can generalize across temporal contexts due to the inherent dynamic nature of language and information. This paper introduces the concept of temporal generalization in LLMs, including bias in past and future generalizations. Then we introduce FreshBench, a new evaluation framework that employs fresh text and event prediction for assessing LLMs' temporal adaptability, ensuring the evaluation process free from data leakage and subjective bias. The experiment shows significant temporal biases and a decline in performance over time. Our findings reveal that powerful models, while initially superior, tend to decline more rapidly in future generalization. Additionally, powerful open-source models demonstrate better long-term adaptability compared to their closed-source counterparts. Our code is available at https://github.com/FreedomIntelligence/FreshBench.

Discussion (0). Sign in to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. MiCoTA: Bridging the Learnability Gap with Intermediate CoT and Teacher Assistants

    cs.CL 2025-07 conditional novelty 6.0 of 10

    Training small language models on intermediate-length reasoning chains from a merged mid-sized teacher assistant improves their math reasoning scores over direct distillation from a large teacher.

  2. Unleashing the Potential of Large Language Models: A Blueprint for Real-Time, Enterprise-Ready Deployments

    cs.LG 2026-08 reject novelty 5.0 of 10

    A blueprint architecture integrating ingestion, continual learning, retrieval, and feedback is claimed to improve LLM latency-cost-freshness, without supporting experimental detail.

  3. NeuralDB: Scaling Knowledge Editing in LLMs to 100,000 Facts with Neural KV Database

    cs.CL 2025-07 conditional novelty 4.0 of 10

    NeuralDB edits up to 100,000 facts in an LLM by storing keys and residuals externally and gating retrieval with cosine similarity, preserving general task performance.

Pith tools