{"id":"a0df7d64-c753-420a-ba12-d3b9d16f27d4","arxiv_id":"2505.00989","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"A maritime LLM agent that converts controller queries into SQL improves risk-vessel identification on the authors' new VTS-SQL benchmark, and all models tested score worse on terse command-style queries.","lead":"This paper builds an AI assistant that turns plain-language questions from vessel traffic controllers into database queries that spot risky ships in busy shipping lanes. It reports better accuracy than general-purpose chatbots on a new maritime benchmark, though the benchmark and test data were created by the same team.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Headline 77.80% is internally inconsistent with Table IV's Ours row (avg 68.91) and ablation text (74.25) contradicts Table VI (72.45); central performance claim unverifiable.","rationale":"I read the paper in good faith. The proposed domain application is plausible and the authors have built a working demonstration, which is real engineering effort. However, the central numerical claim is internally inconsistent: the 'Ours' row in Table IV averages 68.91 while the text and a detached line claim 77.80, and the ablation text (74.25%) does not match Table VI (72.45%). These are not matters of external validity or benchmark representativeness; they are internal contradictions that make it impossible to determine the actual performance of VTS-LLM from the manuscript. The reader's weakest assumption about benchmark fidelity is also legitimate, but the internal inconsistency is a more immediate blocker: even if the benchmark were perfect, the reported superiority of VTS-LLM is not reliably established because the 77.80% number is not consistently tied to a specific configuration in the tables. A focused check that reconciles the tables and raw logs would settle whether this is a reporting error or a substantive discrepancy. If the numbers reconcile, the central claim may survive; if not, the claimed superiority is unsupported. Given this ambiguity, the paper as written cannot be fairly accepted or rejected, hence UNVERDICTED. I partially agree with the reader: their benchmark-validity assumption is a valid concern, but I identify the internal numerical inconsistency as the most load-bearing issue for the central claim.","tokens_in":10087,"tokens_out":5447,"duration_ms":52767,"concrete_test":"Produce a single consolidated table for the operational-style comparison in which each configuration (backbone model, prompt format, module ablations) is a labeled row with its score; confirm which configuration yields 77.80, and recompute all averages and ablation deltas from raw execution logs. If 77.80 corresponds to Ours with GPT-4o, make that explicit and re-run the comparison so the headline number is either in the table or removed. For the ablation discrepancy, run the NER-ablation condition again and ensure the reported score matches Table VI.","verdict_should_be":"UNVERDICTED","load_bearing_attack":"The paper's central claim that VTS-LLM 'achieves an overall accuracy of 77.80%, significantly outperforming SOTA baselines' is not supported by the paper's own tables. In Table IV, the row labeled 'Ours' lists scores 68.69, 68.76, 66.29, 74.30, 66.53, with an average of 68.91; the number 77.80 appears only in a detached line 'GPT-4o Score 77.80' below the table, with no clear row or configuration label. Table V repeats the same 'Ours' row (avg 68.91) and again lists 'Ours GPT-4o 77.80' separately. If the 77.80 result is obtained with GPT-4o as the backbone, the tables must say so; if the 'Ours' row is the complete system, then the headline number is inconsistent with the reported data. The ablation study compounds the problem: Section IV-D states that removing NER drops accuracy to 74.25%, but Table VI row #2 reports 72.45% (rows #3 and #4 match text). Additionally, Section IV-E text says command-style score is 72.40% while Table VII and the conclusion state 72.60%. These inconsistencies mean a reader cannot determine the actual performance of the proposed system. This is an internal correctness issue, distinct from benchmark representativeness, and it blocks verification of the strongest claim.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes VTS-LLM, an LLM-based agent for vessel traffic services (VTS), framing risk-prone vessel identification as a knowledge-augmented Text-to-SQL task. It introduces a self-constructed benchmark (schema, corpus, and query–SQL pairs in three linguistic styles) and an agent architecture with NER-based relational reasoning, agent-driven domain knowledge injection, a semantic algebra intermediate representation, and a query-rethink mechanism. The authors report that VTS-LLM achieves 77.80% on the operational query style and outperforms general-purpose and SQL-specialized baselines; they also claim the first empirical evidence that linguistic style variation systematically affects Text-to-SQL performance. The paper includes a comparison study, ablations, a sensitivity analysis across linguistic styles, and a custom penalty-based evaluation metric.","tokens_in":10368,"tokens_out":3212,"duration_ms":31388,"significance":"If the empirical claims are correct, this is a useful step toward applying LLM agents in a genuinely underexplored maritime domain. The paper contributes a new benchmark and a domain-adapted architecture, and the finding that terse command-style queries degrade text-to-SQL quality is practically important. The manuscript also contains concrete, falsifiable comparisons against several strong baselines (GPT-4o, DeepSeek, Gemini, Claude, and SQL-focused models). However, the significance is currently conditional on resolving several internal numerical inconsistencies that prevent verification of the central performance claim.","major_comments":[{"comment":"The headline claim that 'our VTS-LLM achieves an overall accuracy of 77.80%' is not supported by the data as presented. In Table IV, the row labeled 'Ours' lists five scores (68.69, 68.76, 66.29, 74.30, 66.53) with an average of 68.91, and the figure 77.80 appears only in a detached line below the table reading 'GPT-4o Score 77.80' with no row or configuration label. Table V repeats the same 'Ours' row and adds 'Ours GPT-4o 77.80', but the paper nowhere states explicitly that the complete VTS-LLM system uses GPT-4o as its backbone. The authors must state which configuration produces 77.80%, whether it is the full system, and why the 'Ours' row in the same tables shows a different average. Without this clarification the central performance claim is unverifiable.","section":"Section IV-C, Tables IV and V"},{"comment":"The ablation text and table contradict each other. The text states that removing the NER module (row #2) drops accuracy to 74.25%, but Table VI reports 72.45% for that row. The text and table agree for rows #3 (74.06%) and #4 (70.22%), which makes the NER discrepancy more likely a typographical error rather than a deliberate difference. Which number is correct must be fixed, since the ablation is the only direct evidence for the contribution of the NER module.","section":"Section IV-D, Table VI"},{"comment":"The command-style score for VTS-LLM is reported inconsistently: Section IV-E text says '72.40%', while Table VII and Section V both report 72.60%. This is another instance where the reader cannot determine the real result. The inconsistency affects the sensitivity analysis, which is one of the paper's advertised contributions.","section":"Section IV-E, Table VII, and Section V"},{"comment":"The claim that the paper provides 'the first empirical evidence that linguistic style variation can introduce significant and systematic challenges in Text-to-SQL modeling' is too strong given the evidence presented. The analysis in Table VII covers only two Claude models plus VTS-LLM, on a self-constructed benchmark, with no statistical significance testing or error bars. At most this is preliminary evidence for the studied models and data; the 'first' and 'systematic' wording should be tempered, and the authors should discuss prior work on paraphrased or adversarial text-to-SQL queries.","section":"Section V and Section IV-F"},{"comment":"The evaluation rests entirely on a self-built benchmark (VTS-SQL) whose queries were designed by the authors and whose gold SQL was not validated by external VTS operators. No dataset or code link is provided in the paper, and no inter-annotator agreement or external validation is reported. Since every reported accuracy number is relative to this benchmark, the paper should include a more detailed description of dataset construction, release plans, and at least one form of external validation or a public release to enable independent checking.","section":"Section II-C and IV-C"}],"minor_comments":[{"comment":"There are typographical errors: 'noval' should be 'novel' in the Introduction; 'DeepSeep-R1' should be 'DeepSeek-R1' in Table IV; 'Sensitive Analysis' should be 'Sensitivity Analysis' in Section IV-E.","section":"Throughout"},{"comment":"The dataset statement says 'Dataset is available at VTS-SQL', but no URL or repository identifier is given. The same applies to the code and prompt details referenced in Section IV-C.","section":"Section II-C"},{"comment":"The evaluation metric definition is unclear for the 'otherwise' branch: if |GT| = 0 and |GP| > 0, the denominator in Bs is undefined. The authors should specify the boundary behavior of the metric.","section":"Section IV-B, Eq. (3)"},{"comment":"The formatting of the 'Ours' rows is confusing: the row is not aligned with the column of model names, and the separate 'GPT-4o Score' line would be better integrated into the table with a clear footnote or caption indicating the backbone.","section":"Tables IV and V"}],"recommendation":"major_revision","confidential_remarks":"The paper has a useful application and a plausible architecture, but the internal numeric inconsistencies (77.80 vs. 68.91, 74.25 vs. 72.45, 72.40 vs. 72.60) are the kind of issue that, if they survived into the published version, would seriously undermine reader trust. I would encourage the editor to require the authors to (i) rerun or clearly relabel the experiments so all tables and text agree, (ii) specify the backbone model for every reported configuration, and (iii) describe the dataset release and validation. The 'first empirical evidence' claim also needs scaling back or better support. These are fixable, so I do not recommend rejection, but the revision should be checked carefully."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"You should know this paper before citing it: it proposes the first VTS-domain LLM agent and a VTS-specific Text-to-SQL benchmark, which is genuinely useful for an underserved area. But the headline number is not supported by the paper's own tables. Table IV's 'Ours' row averages 68.91, not 77.80; the 77.80 appears in a detached line labeled 'GPT-4o Score' with no clear configuration. The ablation text says dropping NER gives 74.25, while Table VI reports 72.45. The command-style score is 72.40 in one place and 72.60 in the conclusion. These are not typos you can wave away; they block verification of the strongest claim.\n\nThe new contribution is real: formalizing risk-prone vessel identification as knowledge-augmented Text-to-SQL, building a small but multi-style query-SQL test set, and showing that concise, operator-style queries are systematically harder for LLMs. That last finding is plausible and worth checking, though it rests on a self-authored benchmark with no external validation. The comparison against several strong baselines (GPT-4o, Claude, DeepSeek, etc.) is a plus, and the architecture—NER-based reasoning, RAG, semantic algebra, query rethink—is a reasonable combination of established pieces.\n\nThe soft spots are significant. The benchmark, metric, and gold SQL were all authored by the same team, so the evaluation is partly self-referential. The custom match score is defined but not justified against standard execution accuracy; no error bars or significance tests are reported. Code and data are promised but not actually linked in the text, so nothing is independently checkable. The 'first empirical evidence' claim is overreach given the small, self-created dataset.\n\nWho is this for? Anyone working on Text-to-SQL in specialized, safety-critical domains, or on linguistic robustness of LLMs, might find the benchmark and the style-sensitivity observation useful after the numbers are sorted out. As written, I would not cite it yet. But I would send it to peer review: the topic deserves referee time, and a good referee could force the authors to reconcile the inconsistencies, release the artifacts, and tone down the claims. With those fixes, this could become a solid contribution to maritime NLP.\n\nMy honest recommendation: engage with it, but treat every reported number as provisional until the authors clarify what exactly the 77.80% corresponds to and make the code and data available.","headline":"The VTS-LLM paper brings a genuinely new domain benchmark and a sensible agent architecture, but its headline 77.80% result conflicts with its own tables, so the central claim is currently unverifiable.","tokens_in":10906,"tokens_out":2109,"would_cite":false,"duration_ms":22617,"reading_group":"maybe","serious_thinker":"no","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"This paper claims that an LLM agent adapted to vessel traffic services can translate operators' natural-language queries into SQL with 77.80% accuracy on a new benchmark, outperforming general and SQL-focused models, and that query style…","keywords":["vessel traffic services","Text-to-SQL","LLM agent","maritime safety","domain adaptation","knowledge-augmented querying","linguistic style variation","retrieval-augmented generation"],"falsifier":"Collect real VTS operator queries from operational logs, have independent VTS officers write the expected SQL answers, and rerun VTS-LLM and the baselines on that set; if the agent's accuracy drops to baseline levels on live queries, the reported benchmark advantage does not transfer to practice.","tokens_in":9877,"feed_emoji":"⚓","tokens_out":7046,"duration_ms":67370,"temperature":0.7,"pith_summary":"The paper tries to establish that natural-language querying of vessel traffic databases is feasible enough for real operations. It reformulates risk-prone vessel identification as a knowledge-augmented Text-to-SQL task—turning operator requests into database queries with extra maritime rules and knowledge—and builds VTS-LLM, an agent with four adaptations: entity-recognition-based reasoning, domain-knowledge injection, a semantic algebra intermediate representation, and a query-rethink step. On its own VTS-SQL benchmark, the agent reports 77.80% accuracy on operational-style queries and beats general-purpose and SQL-focused baselines. The paper also claims to present the first empirical evidence that linguistic style—command, operational, or formal—systematically changes Text-to-SQL performance. If true, the work opens a path toward natural-language decision support for maritime traffic control and style-aware evaluation of text-to-SQL systems.","feed_headline":"LLM agent turns vessel traffic queries into SQL at 77.8%","feed_subtitle":"New maritime benchmark shows the system beats general models, especially on terse command-style queries.","key_machinery":"The load-bearing mechanism is the four-module VTS-LLM agent pipeline. Named-entity recognition performs hierarchical relational reasoning, resolving query entities such as place names and vessel types against the database before SQL generation. An agent-based domain-knowledge injection stage decomposes each query into subtasks and retrieves relevant maritime rules through retrieval-augmented generation. The semantic algebra intermediate representation maps the query onto relational-algebra operators—selection, projection, join—plus spatial predicates, bridging natural language and executable SQL. A query-rethink module then checks the draft SQL for semantic and logical consistency and corrects it. The evaluation uses a penalty-based match score that discounts predictions that over-select rows, reflecting the safety-critical cost of returning extra vessels.","core_discovery":"The central claim is that risk-prone vessel identification in Vessel Traffic Services can be cast as a knowledge-augmented Text-to-SQL task and solved by a domain-adaptive LLM agent more accurately than off-the-shelf models. VTS-LLM combines hierarchical named-entity-recognition-based relational reasoning, agent-based injection of maritime knowledge, a semantic algebra intermediate representation that inserts spatial predicates into the query plan, and a query-rethink mechanism that validates and corrects draft SQL. The paper reports 72.60%, 77.80%, and 89.72% on command, operational, and formal linguistic styles respectively, with command-style queries hardest for all models; VTS-LLM nevertheless holds its largest margin there. The paper further claims this is the first empirical evidence that linguistic style variation introduces significant and systematic challenges in Text-to-SQL modeling, and that concise, fragmented operator phrasing is precisely where general models degrade most.","pith_inferences":["Beyond maritime, the style-sensitivity pattern likely appears in any operational field with compressed procedural language—aviation ground control, emergency dispatch, industrial process control—so future Text-to-SQL benchmarks should deliberately vary utterance style.","A testable extension would compare VTS-LLM against its own modules stripped one at a time on a much larger query set, and separately against a tuned retrieval-augmented baseline, to separate the contribution of retrieval, prompting, and the semantic algebra representation.","The paper's penalty metric suggests a broader principle for safety-critical Text-to-SQL evaluation: penalizing false-positive selections asymmetrically, since returning extra vessels wastes operator attention more than missing formatting nuances.","If the benchmark were rebuilt from live operator logs with gold SQL validated by independent VTS officers, the same agent could plausibly be extended to spoken input and automated radio responses, as the paper lists as future work."],"forward_implications":["VTS operators could ask short, urgent questions in their own working style and receive database-backed answers without writing SQL or navigating rigid interfaces.","Text-to-SQL systems evaluated only on formal, well-formed queries will overstate their usefulness in operational settings where commands are terse and fragmented.","The VTS-SQL dataset gives maritime and safety researchers a shared testbed for knowledge-augmented Text-to-SQL, with three stylistic variants of every query.","The four-module recipe—entity reasoning, knowledge injection, semantic algebra, query rethink—is transferable to other regulated, knowledge-intensive domains."],"supporting_citations":[{"why":"supplies the motivating example of an LLM agent for real-time traffic surveillance through intelligent querying.","marker":"[10]"},{"why":"frames the generalization challenge across specialized schemas and query intents that the agent must address.","marker":"[11]"},{"why":"provides a standard large-scale Text-to-SQL benchmark used as a general-purpose comparison point.","marker":"[14]"},{"why":"provides the Spider benchmark as a reference evaluation set for general Text-to-SQL models.","marker":"[18]"},{"why":"grounds the retrieval-augmented generation approach used for injecting domain knowledge.","marker":"[20]"},{"why":"supports the semantic algebra intermediate representation for unifying hybrid question answering.","marker":"[22]"},{"why":"motivates the query rethink mechanism as an iterative reasoning and correction step.","marker":"[23]"},{"why":"supplies the exact-match and execution-accuracy evaluation ideas from which the penalty-based match score is derived.","marker":"[24]"}],"fun_headline_variants":["VTS-LLM: Maritime SQL agent beats general models","Vessel traffic queries to SQL: LLM agent wins","Domain-adaptive LLM turns maritime queries into SQL","VTS-LLM: Best at terse maritime commands","Maritime LLM agent outperforms on command-style queries"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The result depends on the VTS-SQL benchmark faithfully representing real VTS operator work: the queries were designed with operator input rather than taken from live operations, and the gold SQL was authored by the research team without independent operator validation.","fun_headline_variants_meta":{"raw":{"variants":["VTS-LLM: Maritime SQL agent beats general models","Vessel traffic queries to SQL: LLM agent wins","Domain-adaptive LLM turns maritime queries into SQL","VTS-LLM: Best at terse maritime commands","Maritime LLM agent outperforms on command-style queries"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.00062,"raw_usage":{"total_tokens":2888,"prompt_tokens":974,"completion_tokens":1914,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":590,"completion_tokens_details":{"reasoning_tokens":1832}},"tokens_in":590,"tokens_out":1914,"duration_ms":15035,"temperature":1.0,"reasoning_tokens":1832,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-16T04:30:05.226066+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Collect real VTS operator queries from operational logs, have independent VTS officers write the expected SQL answers, and rerun VTS-LLM and the baselines on that set; if the agent's accuracy drops to baseline levels on live queries, the reported benchmark advantage does not transfer to practice.","supporting_citations":[{"cited_title":"Can llm already serve as a database interface? a big bench for large- scale database grounded text-to-sqls,","cited_arxiv_id":null,"evidence_quote":"provides a standard large-scale Text-to-SQL benchmark used as a general-purpose comparison point."},{"cited_title":"Spider: A large-scale human-labeled dataset for complex and cross- domain semantic parsing and text-to-sql task,","cited_arxiv_id":null,"evidence_quote":"provides the Spider benchmark as a reference evaluation set for general Text-to-SQL models."},{"cited_title":"Improving the domain adaptation of retrieval augmented generation (rag) models for open domain question answering,","cited_arxiv_id":null,"evidence_quote":"grounds the retrieval-augmented generation approach used for injecting domain knowledge."},{"cited_title":"Blendsql: A scalable dialect for unifying hybrid question answering in relational algebra,","cited_arxiv_id":null,"evidence_quote":"supports the semantic algebra intermediate representation for unifying hybrid question answering."},{"cited_title":"Rethinking the bounds of llm reasoning: Are multi-agent discussions the key?,","cited_arxiv_id":null,"evidence_quote":"motivates the query rethink mechanism as an iterative reasoning and correction step."},{"cited_title":"Semantic evaluation for text-to- sql with distilled test suites,","cited_arxiv_id":null,"evidence_quote":"supplies the exact-match and execution-accuracy evaluation ideas from which the penalty-based match score is derived."}],"review_version":1}