Large language models achieve macro F1 scores above 0.85 on binary nominal-versus-danger classification from CTAF radio transcripts and METAR weather data using a new synthetic dataset with a 12-category hazard taxonomy.
Title resolution pending
3 Pith papers cite this work. Polarity classification is still indexing.
citation-role summary
citation-polarity summary
roles
background 1polarities
background 1representative citing papers
PilotBench reveals that LLMs follow safety instructions well in flight trajectory prediction but deliver lower numerical precision than traditional forecasters, exposing a precision-controllability tradeoff.
A frozen LLM plus a pretrained trajectory encoder, joined by a small adapter, predicts remaining terminal-area time with ~0.92-minute MAE on Incheon 2022 data.
citing papers explorer
-
Towards Automated Air Traffic Safety Assessment Around Non-Towered Airports Using Large Language Models
Large language models achieve macro F1 scores above 0.85 on binary nominal-versus-danger classification from CTAF radio transcripts and METAR weather data using a new synthetic dataset with a 12-category hazard taxonomy.
-
PilotBench: A Benchmark for General Aviation Agents with Safety Constraints
PilotBench reveals that LLMs follow safety instructions well in flight trajectory prediction but deliver lower numerical precision than traditional forecasters, exposing a precision-controllability tradeoff.
-
LLM4Delay: Flight Delay Prediction via Cross-Modality Adaptation of Large Language Models and Aircraft Trajectory Representation
A frozen LLM plus a pretrained trajectory encoder, joined by a small adapter, predicts remaining terminal-area time with ~0.92-minute MAE on Incheon 2022 data.