Pith. sign in

REVIEW 10 cited by

DyKnow: Dynamically Verifying Time-Sensitive Factual Knowledge in LLMs

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2404.08700 v3 pith:JW3QW4R3 submitted 2024-04-10 cs.CL cs.AI

classification cs.CLcs.AI
keywords knowledgellmstime-sensitivebenchmarkscasesdynamicallyeditingevaluate
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

LLMs acquire knowledge from massive data snapshots collected at different timestamps. Their knowledge is then commonly evaluated using static benchmarks. However, factual knowledge is generally subject to time-sensitive changes, and static benchmarks cannot address those cases. We present an approach to dynamically evaluate the knowledge in LLMs and their time-sensitiveness against Wikidata, a publicly available up-to-date knowledge graph. We evaluate the time-sensitive knowledge in twenty-four private and open-source LLMs, as well as the effectiveness of four editing methods in updating the outdated facts. Our results show that 1) outdatedness is a critical problem across state-of-the-art LLMs; 2) LLMs output inconsistent answers when prompted with slight variations of the question prompt; and 3) the performance of the state-of-the-art knowledge editing algorithms is very limited, as they can not reduce the cases of outdatedness and output inconsistency.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 10 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Beyond Correctness: Rewarding Faithful Reasoning in Retrieval-Augmented Generation

    cs.CL 2025-10 reject novelty 6.0 of 10

    A turn-level faithfulness reward improves a Search-R1-style agent's Information-Think and Think-Answer faithfulness as judged by the same reward model used for training, while task accuracy is roughly unchanged.

  2. An Agile Method for Implementing Retrieval Augmented Generation Tools in Industrial SMEs

    cs.CL 2025-08 conditional novelty 6.0 of 10

    EASI-RAG is a structured agile method for deploying RAG tools in industrial SMEs, validated by one case study where a no-experience team built a working assistant in three weeks.

  3. Training an LLM-as-a-Judge Model: Pipeline, Insights, and Practical Lessons

    cs.CL 2025-02 conditional novelty 6.0 of 10

    A fine-tuned 14B LLM judge, trained with scenario-based prompts and controlled instruction generation, approaches GPT-4's human-agreement performance, and the paper documents why scaling distillation data can fail.

  4. ImageRef-VL: Enabling Contextual Image Referencing in Vision-Language Models

    cs.CV 2025-01 conditional novelty 6.0 of 10

    Fine-tuning InternVL2 on generated interleaved image-text conversations substantially improves contextual image referencing in RAG chatbots, with a new benchmark and metrics.

  5. AntiLeakBench: Preventing Data Contamination by Automatically Constructing Benchmarks with Updated Real-World Knowledge

    cs.CL 2024-12 conditional novelty 6.0 of 10

    AntiLeakBench automatically constructs QA benchmarks from knowledge updated after each model's cutoff, and its experiments suggest that pre-cutoff evaluation overstates LLM ability.

  6. EvoWiki: Evaluating LLMs on Evolving Knowledge

    cs.CL 2024-12 conditional novelty 6.0 of 10

    EvoWiki categorizes facts as stable, evolved, or uncharted and shows that LLMs perform much worse on evolved and uncharted knowledge, with RAG plus continual learning providing the best adaptation.

  7. Benchmarking and Rethinking Knowledge Editing for Large Language Models

    cs.CL 2025-05 conditional novelty 5.0 of 10

    Under autoregressive and sequential editing, parameter-based knowledge editing methods perform poorly, while the retrieval-based SCR baseline consistently outperforms them across datasets and models.

  8. LLM-Assisted Question-Answering on Technical Documents Using Structured Data-Aware Retrieval Augmented Generation

    cs.CL 2025-06 reject novelty 4.0 of 10

    A RAG pipeline with OCR, table and image to text conversion, and a RAFT-tuned reranker reports high QA scores, but its 50-question evaluation overlaps with its training manuals and its baseline comparison uses only 5 ...

  9. Multiple Abstraction Level Retrieve Augment Generation

    cs.CL 2025-01 conditional novelty 4.0 of 10

    MAL-RAG retrieves document, section, paragraph, and multi-sentence chunks together and claims a 25.7% improvement in AI-judged answer correctness on glycoscience questions over single-level RAG.

  10. Challenges in Guardrailing Large Language Models for Science

    cs.AI 2024-11 conditional novelty 3.0 of 10

    A position paper proposing a guardrail framework with four dimensions (trustworthiness, ethics & bias, safety, legal) and implementation strategies for scientific LLM use.

Pith tools