Pith. sign in

REVIEW 13 cited by

TEOChat: A Large Vision-Language Assistant for Temporal Earth Observation Data

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2410.06234 v2 pith:KJW7PJBL submitted 2024-10-08 cs.CV cs.AIcs.LG

TEOChat: A Large Vision-Language Assistant for Temporal Earth Observation Data

classification cs.CV cs.AIcs.LG
keywords temporalteochattaskschangedataimagesingleearth
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved
0 comments
Share X Bluesky LinkedIn Reddit HN
read the original abstract

Large vision and language assistants have enabled new capabilities for interpreting natural images. These approaches have recently been adapted to earth observation data, but they are only able to handle single image inputs, limiting their use for many real-world tasks. In this work, we develop a new vision and language assistant called TEOChat that can engage in conversations about temporal sequences of earth observation data. To train TEOChat, we curate an instruction-following dataset composed of many single image and temporal tasks including building change and damage assessment, semantic change detection, and temporal scene classification. We show that TEOChat can perform a wide variety of spatial and temporal reasoning tasks, substantially outperforming previous vision and language assistants, and even achieving comparable or better performance than several specialist models trained to perform specific tasks. Furthermore, TEOChat achieves impressive zero-shot performance on a change detection and change question answering dataset, outperforms GPT-4o and Gemini 1.5 Pro on multiple temporal tasks, and exhibits stronger single image capabilities than a comparable single image instruction-following model on scene classification, visual question answering, and captioning. We publicly release our data, model, and code at https://github.com/ermongroup/TEOChat .

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Forward citations

Cited by 13 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. SenseBench: A Benchmark for Remote Sensing Low-Level Visual Perception and Description in Large Vision-Language Models

    cs.CV 2026-05 unverdicted novelty 8.0

    SenseBench is the first physics-based benchmark with 10K+ instances and dual protocols to evaluate VLMs on remote sensing low-level perception and diagnostic description, revealing domain bias and specific failure modes.

  2. A Multi-Agent Feedback System for Detecting and Describing News Events in Satellite Imagery

    cs.CV 2026-04 unverdicted novelty 7.0

    SkyScraper uses iterative multi-agent feedback to geocode news articles and synthesize captions for satellite image sequences, locating 5x more events than traditional methods and producing a new 5,000-sequence dataset.

  3. SkyVLaM: Multimodal Large Language Model for UAV Video Understanding in Remote Sensing

    cs.CV 2026-07 conditional novelty 6.0

    SkyVLaM introduces a temporal basis perceiver and adaptive dense selection to improve language-conditioned video segmentation in UAV scenes, and contributes the SkyVid dataset.

  4. GeoChrono: Benchmarking and Rethinking Long-Term Temporal Understanding in Remote Sensing

    cs.CV 2026-07 conditional novelty 6.0

    ChronoBench decomposes long-term remote sensing understanding into four cognitive levels, and the GeoChrono model, using per-location temporal trajectories, achieves 78.34% accuracy—over 20 points above prior MLLMs—bu...

  5. RemoteShield: Enable Robust Multimodal Large Language Models for Earth Observation

    cs.CV 2026-04 unverdicted novelty 6.0

    RemoteShield improves robustness of Earth observation MLLMs by training on semantic equivalence clusters of clean and perturbed inputs via preference learning to maintain consistent reasoning under noise.

  6. Decoding the Delta: Unifying Remote Sensing Change Detection and Understanding with Multimodal Large Language Models

    cs.CV 2026-04 unverdicted novelty 6.0

    Delta-LLaVA adds Change-Enhanced Attention, Change-SEG with prior embeddings, and Local Causal Attention to MLLMs to overcome temporal blindness, outperforming general models on a new unified benchmark for bi- and tri...

  7. A Multi-Agent Feedback System for Detecting and Describing News Events in Satellite Imagery

    cs.CV 2026-04 conditional novelty 6.0

    An iterative multi-agent feedback pipeline finds ~5× more multi-temporal news events in satellite imagery than traditional geocoding and yields a 5,000-sequence captioning dataset.

  8. RSICCLLM: A Multimodal Large Language Model for Remote Sensing Image Change Captioning

    cs.CV 2026-06 unverdicted novelty 5.0

    RSICCLLM introduces a post-training framework with RSICI dataset, difference-aware supervised fine-tuning, and dual-negative preference optimization that claims to outperform much larger models on remote sensing image...

  9. HiSem: Hierarchical Semantic Disentangling for Remote Sensing Image Change Captioning

    cs.CV 2026-05 unverdicted novelty 5.0

    HiSem adds bidirectional differential attention and a two-level hierarchical routing module with MoE to handle semantic granularity differences in remote sensing change captioning, reporting +7.52% BLEU-4 on WHU-CDC.

  10. ChatENV: An Interactive Vision-Language Model for Sensor-Guided Environmental Monitoring and Scenario Simulation

    cs.CV 2025-08 unverdicted novelty 5.0

    ChatENV fine-tunes Qwen-2.5-VL on a 177k-image dataset of temporal satellite pairs with sensor metadata to support interactive temporal and what-if reasoning for environmental monitoring.

  11. UniReason-Med: A Shared Grounded Reasoning Interface for 2D-to-3D Transfer in Medical VQA

    cs.CV 2026-06 unverdicted novelty 4.0

    UniReason-Med introduces a unified framework for 2D and 3D medical VQA with shared grounded reasoning, trained on a 220K dataset, claiming that joint 2D+3D supervision improves 3D performance over 3D-only training.

  12. Earth Science Foundation Models: From Perception to Reasoning and Discovery

    astro-ph.IM 2026-05 unverdicted novelty 3.0

    The paper delivers a unified review and roadmap of Earth science foundation models, structured by capability depth from perception to agentic reasoning and by application breadth across atmosphere, hydrosphere, lithos...

  13. Earth Science Foundation Models: From Perception to Reasoning and Discovery

    astro-ph.IM 2026-05 unverdicted novelty 2.0

    A review of Earth science foundation models covering capability evolution from perception to discovery, applications across atmosphere/hydrosphere/lithosphere/biosphere/anthroposphere/cryosphere, over 200 datasets, an...