REVIEW 3 major objections 2 minor 36 references
Longest arcs exist for some left-invariant three-dimensional contact sub-Lorentzian structures, under sufficient conditions on solvable Lie groups and the universal cover of SL(2, R).
Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →
Sufficient conditions are given for existence of longest arcs in left-invariant three-dimensional contact sub-Lorentzian structures on solvable Lie groups and the universal cover of SL(2,R).
T0 review reviewed 2026-07-15 challenge →
load-bearing objection We only have the abstract for the sub-Lorentzian existence paper; the supplied body is a different AVSR manuscript, so the claims cannot be audited. the 3 major comments →
Existence of the longest arcs for left-invariant three-dimensional contact sub-Lorentzian structures
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
Core claim
For some left-invariant three-dimensional contact sub-Lorentzian structures (with a known classification), longest arcs exist as solutions of the associated unbounded-control concave optimal-control problem; more generally, the authors give sufficient conditions ensuring existence of longest arcs for left-invariant (sub-)Lorentzian structures on solvable Lie groups and on the universal cover of SL(2, R).
What carries the argument
Reduction, via the known classification of left-invariant three-dimensional contact sub-Lorentzian structures, of the existence question to proposed sufficient conditions on solvable Lie groups and on the universal cover of SL(2, R); those conditions make the unbounded-control concave problem admit maximizers.
Load-bearing premise
The known classification of left-invariant three-dimensional contact sub-Lorentzian structures, together with left-invariance and the contact assumption, is enough to reduce existence to the proposed conditions on solvable groups and the universal cover of SL(2, R).
What would settle it
Produce a left-invariant three-dimensional contact sub-Lorentzian structure covered by the classification for which no longest arc exists between some pair of points, or exhibit a solvable Lie group (or the cover of SL(2, R)) where the stated sufficient conditions hold yet a maximizer fails, or where a maximizer fails while the conditions are the only obstruction claimed.
If this is right
- On the listed groups and structures the longest-arc problem is well-posed: maximizers exist, not merely formal critical points.
- Classification-based case analysis can settle existence without first solving every boundary-value problem.
- Existence statements extend from ordinary Lorentzian structures to contact sub-Lorentzian left-invariant structures in dimension three.
- Under the proposed conditions one may search for maximizers on solvable groups and on the cover of SL(2, R) knowing they exist.
Where Pith is reading between the lines
- The same sufficient conditions may carry over to higher-dimensional left-invariant contact sub-Lorentzian structures once usable normal forms exist.
- A solvable group that violates the conditions would give a concrete class of counter-examples where longest arcs need not exist despite left-invariance.
- The role of the contact assumption in securing existence may suggest parallel results for related sub-Riemannian energy or length problems with unbounded controls.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The abstract claims that existence of longest arcs (optimal curves) is established for some left-invariant three-dimensional contact sub-Lorentzian structures, using a known classification, and that sufficient conditions are given for left-invariant (sub-)Lorentzian structures on solvable Lie groups and on the universal cover of SL(2,R). The problem is framed as an optimal-control problem with unbounded controls and a concave cost, for which existence is nontrivial. The supplied full manuscript body, however, is an unrelated paper on context-aware audio-visual speech recognition (VASR / AV-CoT, multimodal LLMs, CER benchmarks). No statements of the sufficient conditions, no reduction from the classification, no control-theoretic arguments, and no proofs of existence appear in the provided text.
Significance. If the mathematical claims in the abstract hold, the work would address a genuine gap: existence of maximizers for sub-Lorentzian length with unbounded controls is not automatic, and a clean treatment for the classified left-invariant 3D contact family plus sufficient conditions on solvable groups and the universal cover of SL(2,R) would be a useful contribution to geometric control and sub-Lorentzian geometry. That significance cannot be assessed from the supplied body, which contains none of the claimed results.
major comments (3)
- The full manuscript text does not match the title, abstract, or arXiv identifier of the paper under review. The body is a complete, self-contained paper on CA VSR / VASR (Audio-Visual Chain-of-Thought, data pipeline, Chinese-LiPS and VASR test sets, Tables 2–3, CER results). No sub-Lorentzian structures, Lie groups, longest arcs, or optimal-control existence arguments appear. The central existence claim is therefore uncheckable from the submission as provided.
- Because the body is the wrong paper, the abstract’s load-bearing reduction—that left-invariance, contact, and the known classification of 3D contact sub-Lorentzian structures reduce existence to the proposed sufficient conditions on solvable groups and the universal cover of SL(2,R)—cannot be verified. Degenerate, non-contact, or non-solvable cases may or may not be covered; the manuscript supplies no lemmas or statements with which to decide.
- No equations, theorems, or proofs related to the claimed existence result are present. A referee report on soundness of the longest-arc existence argument is impossible until the correct mathematical manuscript is supplied.
minor comments (2)
- The abstract alone is internally coherent as a problem statement, but it is not a substitute for the missing theorems and proofs.
- The provided body (VASR) has its own presentation issues (e.g., OCR/garbled math tokens in displayed equations, figure placeholders), but those are irrelevant to the paper that was supposed to be reviewed.
Circularity Check
No circularity can be exhibited: the supplied full text is an unrelated AVSR paper, and the target math abstract shows no self-definitional or fitted-as-prediction reduction.
full rationale
The load-bearing claim of arXiv:2603.07262 is existence of longest arcs for some left-invariant 3D contact sub-Lorentzian structures (via a known classification) plus sufficient conditions on solvable Lie groups and the universal cover of SL(2,R). The CACHEABLE full manuscript is not that paper; it is an unrelated multimodal speech-recognition work (VASR / Context AVSR). Consequently no equations, reductions, self-citations, uniqueness theorems, or control-theoretic arguments from 2603.07262 are available to quote. From the abstract alone there is no self-definitional loop (existence is not defined as the conditions), no fitted parameter renamed as a prediction, and no uniqueness imported from the authors. Per the hard rules, circularity may be claimed only when a specific reduction can be quoted; absence of the correct body is not circularity. Score 0 with empty steps is therefore the only honest outcome.
Axiom & Free-Parameter Ledger
axioms (3)
- domain assumption Classification of left-invariant three-dimensional contact sub-Lorentzian structures is known and can be used as the starting point for existence analysis.
- domain assumption Finding longest arcs is an optimal control problem with unbounded control set and concave cost functional, for which existence is nontrivial.
- standard math Left-invariant (sub-)Lorentzian structures on solvable Lie groups and on the universal cover of SL(2,R) admit analysis via Lie-algebraic data.
Cite this review
Pith. "Pith review of Existence of the longest arcs for left-invariant three-dimensional contact sub-Lorentzian structures." pith.science (2026). https://pith.science/paper/UCZFGWTV
@misc{pith2026260307262,
author = {Pith},
title = {Pith review of: Existence of the longest arcs for left-invariant three-dimensional contact sub-Lorentzian structures},
year = {2026},
howpublished = {\url{https://pith.science/paper/UCZFGWTV}},
note = {Machine review of arXiv:2603.07262}
}
read the original abstract
The problem of finding optimal curves (the longest arcs) for sub-Lorentzian structures is an optimal control problem with an unbounded control set and a concave cost functional. The question of existence of an optimal solution is nontrivial for such problems. We solve here this question for some left-invariant three-dimensional contact sub-Lorentzian structures, whose classification is known. We propose sufficient conditions for the existence of the longest arcs for left-invariant (sub-)Lorentzian structures on solvable Lie groups and on the universal cover of the Lie group SL(2, R).
Reference graph
Works this paper leans on
-
[1]
Introduction Automatic Speech Recognition (ASR) [1, 2, 3] has achieved re- markable progress in recent years, largely driven by scaling up training data and model capacity. However, such audio-only systems still struggle in scenarios requiring context-aware dis- ambiguation, such as recognizing homophones, named entities, and domain-specific terms. This l...
arXiv 2026
-
[2]
Overall Architecture V ASR aims to achieve CA VSR in diverse real-world scenarios through structured multimodal reasoning
Framework 2.1. Overall Architecture V ASR aims to achieve CA VSR in diverse real-world scenarios through structured multimodal reasoning. Equipped with A V- CoT, V ASR reformulates CA VSR task as a structured Percep- tion–Reasoning–Transcription pipeline. As illustrated in Fig- ure 2, given multimodal inputs consisting of video�and au- dio signals�, the f...
-
[3]
Data Pipeline To overcome the scarcity of open-source datasets for CA VSR, we design an automated, scalable data curation pipeline
Dataset 3.1. Data Pipeline To overcome the scarcity of open-source datasets for CA VSR, we design an automated, scalable data curation pipeline. The objective is to identify challenging, visually rich data in lin- guistic ambiguity and to generate high-quality multimodal an- notations. As shown on the right of Figure 2, we split the data process into two ...
-
[4]
w/o A V-CoT
Experiments 4.1. Datasets We utilize open-source chinese datasets, including MEIJU [28] and MER25 [29], ch-sims-v2 [30], and the Chinese-LiPS [16] training set. All open-source data are uniformly processed using our proposed data processing pipeline, totaling approximately 230 hours. We use Chinese-LiPS test set as the open-source test set. For general-sc...
-
[5]
We propose V ASR, which leverages A V- CoT to explicitly model the transcription process to address the single-modality dominance
Conclusion and Limitation In this work, we introduce the CA VSR task, extending A VSR beyond lip-reading to leverage rich visual context for resolving linguistic ambiguities. We propose V ASR, which leverages A V- CoT to explicitly model the transcription process to address the single-modality dominance. In addition, we develop and release a scalable data...
-
[6]
They play no role in methodology, experimenta- tion, interpretation, or the production of scientific results
Generative AI Use Disclosure Generative AI tools are used for writing refinement and data generation. They play no role in methodology, experimenta- tion, interpretation, or the production of scientific results. The authors bear full intellectual responsibility for all content in this manuscript. All authors agree with the submission of this paper
-
[7]
X. Shi, X. Wang, Z. Guo, Y . Wang, P. Zhang, X. Zhang, Z. Guo, H. Hao, Y . Xi, B. Yang, J. Xu, J. Zhou, and J. Lin, “Qwen3-asr technical report,”CoRR, vol. abs/2601.21337, 2026
Pith/arXiv arXiv 2026
-
[8]
K. Xu, F. Xie, X. Tang, and Y . Hu, “Fireredasr: Open- source industrial-grade mandarin speech recognition mod- els from encoder-decoder to LLM integration,”CoRR, vol. abs/2501.14350, 2025
Pith/arXiv arXiv 2025
-
[9]
Paraformer: Fast and accurate parallel transformer for non-autoregressive end-to- end speech recognition,
Z. Gao, S. Zhang, I. McLoughlin, and Z. Yan, “Paraformer: Fast and accurate parallel transformer for non-autoregressive end-to- end speech recognition,” inProc.Interspeech, 2022, pp. 2063– 2067
2022
-
[10]
Omni- avsr: Towards unified multimodal speech recognition with large language models,
U. Cappellazzo, X. Liu, P. Ma, S. Petridis, and M. Pantic, “Omni- avsr: Towards unified multimodal speech recognition with large language models,”CoRR, vol. abs/2511.07253, 2025
arXiv 2025
-
[11]
AD- A VSR: asymmetric dual-stream enhancement for robust audio- visual speech recognition,
J. Xue, X. Liu, X. Wu, X. Yin, D. Huang, and F. Yu, “AD- A VSR: asymmetric dual-stream enhancement for robust audio- visual speech recognition,”CoRR, vol. abs/2508.07608, 2025
Pith/arXiv arXiv 2025
-
[12]
MLCA-A VSR: multi-layer cross attention fusion based audio-visual speech recognition,
H. Wang, P. Guo, P. Zhou, and L. Xie, “MLCA-A VSR: multi-layer cross attention fusion based audio-visual speech recognition,” in Proc.ICASSP, 2024, pp. 8150–8154
2024
-
[13]
Adaptive audio-visual speech recognition via matryoshka-based multimodal llms,
U. Cappellazzo, M. Kim, and S. Petridis, “Adaptive audio-visual speech recognition via matryoshka-based multimodal llms,” CoRR, vol. abs/2503.06362, 2025
Pith/arXiv arXiv 2025
-
[14]
Mms-llama: Ef- ficient llm-based audio-visual speech recognition with minimal multimodal speech tokens,
J. H. Yeo, H. Rha, S. J. Park, and Y . M. Ro, “Mms-llama: Ef- ficient llm-based audio-visual speech recognition with minimal multimodal speech tokens,” inFindings of the Association for Computational Linguistics, ACL 2025, Vienna, Austria, July 27 - August 1, 2025, ser. Findings of ACL, vol. ACL 2025, 2025, pp. 20 724–20 735
2025
-
[15]
J. H. Yeo, M. Kim, C. W. Kim, S. Petridis, and Y . M. Ro, “Zero- avsr: Zero-shot audio-visual speech recognition with llms by learning language-agnostic speech representations,”CoRR, vol. abs/2503.06273, 2025
Pith/arXiv arXiv 2025
-
[16]
Qwen2.5-omni technical report,
J. Xu, Z. Guo, J. He, H. Hu, T. He, S. Bai, K. Chen, J. Wang, Y . Fan, K. Dang, B. Zhang, X. Wang, Y . Chu, and J. Lin, “Qwen2.5-omni technical report,”CoRR, vol. abs/2503.20215, 2025
Pith/arXiv arXiv 2025
-
[17]
J. Xu, Z. Guo, H. Hu, Y . Chu, X. Wang, J. Heet al., “Qwen3-omni technical report,”CoRR, vol. abs/2509.17765, 2025
Pith/arXiv arXiv 2025
-
[18]
Intern-s1: A scientific multimodal foundation model,
L. Bai, Z. Cai, Y . Cao, M. Cao, W. Cao, C. Chenet al., “Intern-s1: A scientific multimodal foundation model,”CoRR, vol. abs/2508.15763, 2025
arXiv 2025
-
[19]
Minicpm-v: A GPT-4V level MLLM on your phone,
Y . Yao, T. Yu, A. Zhang, C. Wang, J. Cui, H. Zhu, T. Cai, H. Li, W. Zhao, Z. He, Q. Chen, H. Zhou, Z. Zou, H. Zhang, S. Hu, Z. Zheng, J. Zhou, J. Cai, X. Han, G. Zeng, D. Li, Z. Liu, and M. Sun, “Minicpm-v: A GPT-4V level MLLM on your phone,” CoRR, vol. abs/2408.01800, 2024
Pith/arXiv arXiv 2024
-
[20]
G. Comanici, E. Bieber, M. Schaekermann, I. Pasupat, N. Sachdevaet al., “Gemini 2.5: Pushing the frontier with ad- vanced reasoning, multimodality, long context, and next genera- tion agentic capabilities,”CoRR, vol. abs/2507.06261, 2025
Pith/arXiv arXiv 2025
-
[21]
Slidespeech: A large scale slide-enriched audio-visual corpus,
H. Wang, F. Yu, X. Shi, Y . Wang, S. Zhang, and M. Li, “Slidespeech: A large scale slide-enriched audio-visual corpus,” inProc.ICASSP, 2024, pp. 11 076–11 080
2024
-
[22]
Chinese- lips: A chinese audio-visual speech recognition dataset with lip- reading and presentation slides,
J. Zhao, Y . Jia, S. Wang, J. Zhou, H. Wang, and Y . Qin, “Chinese- lips: A chinese audio-visual speech recognition dataset with lip- reading and presentation slides,” inProc.ICME, 2025, pp. 1–6
2025
-
[23]
How2: A large-scale dataset for mul- timodal language understanding,
R. Sanabria, O. Caglayan, S. Palaskar, D. Elliott, L. Barrault, L. Specia, and F. Metze, “How2: A large-scale dataset for mul- timodal language understanding,”CoRR, vol. abs/1811.00347, 2018
Pith/arXiv arXiv 2018
-
[24]
A V ATAR: unconstrained audiovisual speech recog- nition,
V . Gabeur, P. H. Seo, A. Nagrani, C. Sun, K. Alahari, and C. Schmid, “A V ATAR: unconstrained audiovisual speech recog- nition,” inProc.Interspeech, 2022, pp. 2818–2822
2022
-
[25]
Slideavsr: A dataset of paper explanation videos for audio-visual speech recog- nition,
H. Wang, S. Kurita, S. Shimizu, and D. Kawahara, “Slideavsr: A dataset of paper explanation videos for audio-visual speech recog- nition,”CoRR, vol. abs/2401.09759, 2024
Pith/arXiv arXiv 2024
-
[26]
Multi-modality speech recognition driven by background visual scenes,
C. Luo, Y . Liu, W. Sun, and Z. Sun, “Multi-modality speech recognition driven by background visual scenes,” inProc.ICASSP, 2024, pp. 10 926–10 930
2024
-
[27]
Avformer: Injecting vision into frozen speech models for zero-shot av-asr,
P. H. Seo, A. Nagrani, and C. Schmid, “Avformer: Injecting vision into frozen speech models for zero-shot av-asr,” inProc.CVPR, 2023, pp. 22 922–22 931
2023
-
[28]
Lip reading sentences in the wild,
J. S. Chung, A. W. Senior, O. Vinyals, and A. Zisserman, “Lip reading sentences in the wild,” inProc.CVPR, 2017, pp. 3444– 3453
2017
-
[29]
A cascade sequence-to-sequence model for chinese mandarin lip reading,
Y . Zhao, R. Xu, and M. Song, “A cascade sequence-to-sequence model for chinese mandarin lip reading,” inProc.MMAsia, 2019, pp. 32:1–32:6
2019
-
[30]
CN-CVS: A mandarin audio- visual dataset for large vocabulary continuous visual to speech synthesis,
C. Chen, D. Wang, and T. F. Zheng, “CN-CVS: A mandarin audio- visual dataset for large vocabulary continuous visual to speech synthesis,” inProc.ICASSP, 2023, pp. 1–5
2023
-
[31]
V oxceleb2: Deep speaker recognition,
J. S. Chung, A. Nagrani, and A. Zisserman, “V oxceleb2: Deep speaker recognition,” inProc.Interspeech, 2018, pp. 1086–1090
2018
-
[32]
Robust speech recognition via large-scale weak su- pervision,
A. Radford, J. W. Kim, T. Xu, G. Brockman, C. McLeavey, and I. Sutskever, “Robust speech recognition via large-scale weak su- pervision,” inProc.ICML, ser. Proceedings of Machine Learning Research, vol. 202, 2023, pp. 28 492–28 518
2023
-
[33]
S. Bai, K. Chen, X. Liu, J. Wang, W. Ge, S. Song, K. Dang, P. Wang, S. Wang, J. Tang, H. Zhong, Y . Zhu, M. Yang, Z. Li, J. Wan, P. Wang, W. Ding, Z. Fu, Y . Xu, J. Ye, X. Zhang, T. Xie, Z. Cheng, H. Zhang, Z. Yang, H. Xu, and J. Lin, “Qwen2.5-vl technical report,”CoRR, vol. abs/2502.13923, 2025
Pith/arXiv arXiv 2025
-
[34]
MEIJU - the 1st multimodal emotion and intent joint understand- ing challenge,
R. Liu, X. Xing, Z. Lian, H. Li, B. W. Schuller, and H. Zuo, “MEIJU - the 1st multimodal emotion and intent joint understand- ing challenge,” inProc.ICASSP, 2025, pp. 1–2
2025
-
[35]
MER 2025: When affective computing meets large language models,
Z. Lian, R. Liu, K. Xu, B. Liu, X. Liu, Y . Zhang, X. Liu, Y . Li, Z. Cheng, H. Zuo, Z. Ma, X. Peng, X. Chen, Y . Li, E. Cam- bria, G. Zhao, B. W. Schuller, and J. Tao, “MER 2025: When affective computing meets large language models,”CoRR, vol. abs/2504.19423, 2025
Pith/arXiv arXiv 2025
-
[36]
Make acous- tic and visual cues matter: CH-SIMS v2.0 dataset and av-mixup consistent module,
Y . Liu, Z. Yuan, H. Mao, Z. Liang, W. Yanget al., “Make acous- tic and visual cues matter: CH-SIMS v2.0 dataset and av-mixup consistent module,” inProc.ICMI, 2022, pp. 247–258
2022
This paper was first reviewed by grok-4.5 on July 15, 2026.
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.