Pith. sign in

REVIEW 8 cited by

Large Language Models Can Learn Temporal Reasoning

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2401.06853 v6 pith:6CJC73BU submitted 2024-01-12 cs.CL

classification cs.CL
keywords reasoningtemporalllmsdatasetgraphlanguagelargemodels
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

While large language models (LLMs) have demonstrated remarkable reasoning capabilities, they are not without their flaws and inaccuracies. Recent studies have introduced various methods to mitigate these limitations. Temporal reasoning (TR), in particular, presents a significant challenge for LLMs due to its reliance on diverse temporal concepts and intricate temporal logic. In this paper, we propose TG-LLM, a novel framework towards language-based TR. Instead of reasoning over the original context, we adopt a latent representation, temporal graph (TG) that enhances the learning of TR. A synthetic dataset (TGQA), which is fully controllable and requires minimal supervision, is constructed for fine-tuning LLMs on this text-to-TG translation task. We confirmed in experiments that the capability of TG translation learned on our dataset can be transferred to other TR tasks and benchmarks. On top of that, we teach LLM to perform deliberate reasoning over the TGs via Chain-of-Thought (CoT) bootstrapping and graph data augmentation. We observed that those strategies, which maintain a balance between usefulness and diversity, bring more reliable CoTs and final results than the vanilla CoT distillation.

Discussion (0). Sign in to comment.

Forward citations

Cited by 8 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. LTLZinc: a Benchmarking Framework for Continual Learning and Neuro-Symbolic Temporal Reasoning

    cs.AI 2025-07 conditional novelty 7.0 of 10

    LTLZinc generates image-based temporal reasoning and continual learning benchmarks from LTLf formulas over MiniZinc constraints, and experiments show existing methods often fail.

  2. ExAnte: A Benchmark for Ex-Ante Inference in Large Language Models

    cs.LG 2025-05 conditional novelty 7.0 of 10

    Models leak future knowledge despite explicit temporal cutoffs, as quantified by the ExAnte benchmark across four tasks.

  3. TempoBench: Evaluating Temporal Causal Reasoning in Large Language Models

    cs.AI 2025-10 conditional novelty 6.0 of 10

    In TempoBench's formally verified temporal-causality tests, frontier LLMs score 65.6% F1 on normal causal attribution but 7.5% on hard instances, while trace simulation stays far higher.

  4. MDBench: A Synthetic Multi-Document Reasoning Benchmark Generated with Knowledge Guidance

    cs.CL 2025-06 conditional novelty 6.0 of 10

    MDBench is a synthetically generated, knowledge-guided benchmark for multi-document QA on which frontier LLMs achieve only about 60% exact match.

  5. ConceptBot: Enhancing Robot's Autonomy through Task Decomposition with Large Language Models and Knowledge Graph

    cs.RO 2025-08 conditional novelty 5.0 of 10

    Using ConceptNet-augmented prompts, ConceptBot reports 87% vs 31% success on implicit tasks and 76% vs 15% on risk-aware tasks over a re-implemented SayCan baseline, with an 80% SafeAgentBench score.

  6. Uncertainty-quantified Rollout Policy Adaptation for Unlabelled Cross-domain Temporal Grounding

    cs.CV 2025-08 conditional novelty 5.0 of 10

    Rollout-averaged pseudo labels with variance-based confidence weighting let a GRPO-trained temporal grounding model adapt to an unlabelled target domain from only 100-200 videos.

  7. T-GRAG: A Dynamic GraphRAG Framework for Resolving Temporal Conflicts and Redundancy in Knowledge Retrieval

    cs.AI 2025-08 conditional novelty 5.0 of 10

    A temporal GraphRAG framework that partitions knowledge graphs by timestamp and retrieves at subgraph, node, and knowledge levels outperforms RAG baselines on a new Audi annual-report QA benchmark.

  8. Pluri-perspectivism in Human-robot Co-creativity with Older Adults

    cs.HC 2025-07 conditional novelty 5.0 of 10

    A five-dimensional pluri-perspectivist model is introduced to guide context-sensitive, co-creative human-robot interaction, grounded in theory and interviews with artists and art teachers.

Pith tools