REVIEW 5 major objections 8 minor 52 references
Current multimodal models cannot reliably keep degraded evidence, security rules, and safe UAV actions coupled at the decision point.
Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →
T0 review · grok-4.5
2026-07-30 13:45 UTC pith:RBI5XKEM
load-bearing objection Solid offline UAV decision benchmark with a real multi-model gap; the headline numbers are only as strong as unvalidated action/policy labels. the 5 major comments →
MulRobBench: A Decision-Level Benchmark for Safe and Security-Policy-Compliant Multimodal UAV Agents
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
Core claim
Across 17 uniformly audited multimodal models on MulRobBench’s 3,024 strict samples, current systems remain far from reliable protocol-conditioned UAV decision making: the best semantic protocol-decision score is 0.5141 and the best strict mean scoring-dimension accuracy is 0.1599. Models handle coarse context and targets relatively well but break on modality-trust selection, constraint extraction, collaboration and abstention triggers, and risk-aware action planning. Modality-removal on a matched 20-anchor subset changes 4–15 action selections per model, showing both vision and text shape decisions while exposing unstable combination of those inputs.
What carries the argument
MulRobBench: an offline, protocol-conditioned Vision-Language-Action decision contract that binds real UAV multimodal observations, injected security-policy mission context, and a closed ten-action safety vocabulary, scored along four links (context, evidence arbitration, degradation-aware reasoning, risk-aware action) with 12 dimensions and both semantic scores and strict structural diagnostics.
Load-bearing premise
Treating airport boundaries, sensitive-place rules, privacy limits, and other policies as injected mission context—rather than labels verified in the source imagery—still fairly measures whether a model preserves security-policy compliance when it chooses an action.
What would settle it
Find a model that, on the same 3,024-sample strict split and closed action contract, simultaneously posts high semantic protocol-decision score, high strict mean dimension accuracy, low unsafe-action rate, and near-zero normalization failures, with modality removal no longer flipping many of the 20-anchor actions.
If this is right
- UAV VLA evaluation must report semantic validity, structural parseability, and action safety side by side; a single accuracy score hides decision-chain breaks.
- Improving coarse scene or target recognition will not close the gap; gains must target modality-trust, constraint extraction, and abstention or collaboration triggers.
- Glare, missing data, and operator shorthand are priority stress cases because they systematically decouple evidence quality from allowed actions.
- Offline protocol-conditioned next-action audits become a necessary gate before claims of smart-city UAV cyber-physical safety.
- Human multi-expert pilots on stratified subsets remain far above models, setting a concrete gap for future systems work.
Where Pith is reading between the lines
- Training objectives that reward fluent captions or open-ended plans may actively work against the structured action contract this benchmark requires.
- If injected policies are later grounded in verified maps and entity labels, the same four-link chain could become a live mission-audit layer rather than only an offline test.
- The non-monotonic modality-ablation results suggest some models are already over-relying on text priors; denser paired vision-text counterfactuals would expose that shortcut more sharply.
- Similar decision-contract thinking likely transfers to other rule-bound embodied settings (ground robots in restricted sites, inspection agents) where evidence degradation and forbidden actions coexist.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper introduces MulRobBench, an offline benchmark for protocol-conditioned safe next-action selection by multimodal UAV agents in smart-city settings. Built on UAVScenes-derived observations with injected protocol semantics (restricted zones, privacy/media rules, temporary mission rules), the benchmark comprises 3,024 strict samples organized into 17 primary-attribution taxonomy nodes and 12 scoring dimensions (D1–D12) across four decision links. Seventeen multimodal models are evaluated under a shared decision contract with separated semantic scoring (Eq. 10–11) and strict structural diagnostics (Eq. 17, 20), plus action-level metrics (SafeAcc, UnsafeRate, MAD, PACS). Headline results: best semantic protocol-decision score 0.5141, best strict mean dimension accuracy 0.1599; a 20-anchor modality-removal study changes 4–15/20 actions per model. The authors attribute failures primarily to modality-trust selection, constraint extraction, collaboration/abstention thresholds, and action-rationale consistency rather than coarse context recognition.
Significance. If the results hold, this is a useful and timely contribution. The decision-level endpoint (protocol-conditioned next action under a closed action vocabulary) is genuinely under-served by existing UAV benchmarks, and the paper's central methodological choice — reporting semantic scores, action-set agreement, and structural compliance side by side rather than collapsing them into one number — is well motivated and executed with unusual discipline. Specific strengths worth naming: a fixed, audited 3,024-sample strict set with a documented split contract (Tables VI–VII); 17 models under a uniform audit with normalization failures retained rather than silently dropped (Phi-4-Multimodal's 3,024 failures are reported, not removed); a 5% normalization-failure credibility rule that prevents action-only fallback numbers from being over-read; per-dimension, per-condition, and error-chain analyses (Tables XIII–XVII) that localize failures to the middle of the decision chain; and explicit, honest limitation statements about the injected-semantics premise. The multi-expert reference pilot, the conditional robustness matrix (Table XVI), and the modality-ablation study are all falsifiable, checka
major comments (5)
- [§III.B (Benchmark Generation, step four) and Table XI] Every metric in the paper (Eqs. 5, 10, 13–17) is scored against Γ_i = (A*_i, A+_i, A−_i) and the dimension projections g_id = π_d(x_i), yet the manuscript nowhere reports who produced these labels, under what protocol, with what adjudication, or with what inter-rater agreement. Many labels are judgment calls rather than facts: whether a dust-occluded frame warrants hover vs. reobserve vs. request-another-UAV; which constraints are 'active' under a given protocol; what the 'primary' degradation is when conditions co-occur (Table XIV shows samples carrying multiple condition labels). The Multi-Expert Reference row cannot substitute for label validation: it is a 600-sample pilot with round values (0.83/0.84/0.94/0.78), no rater count, no agreement statistic, and no statement of whether the reference annotators are disjoint from the original label authors — if they overlap, the row measures
- [§III.D / Fig. 5 and Eq. (10)] The controlled semantic score s_d(ŷ_im, g_id) is computed via a scoring prompt (Fig. 5), which implies an LLM-based judge, but the manuscript never states which model executes the judge, at what settings, or how the judge itself was validated against human scoring. Since all semantic metrics (S_md, P_m, the group scores in Table XII, and the conditional tables) derive from s_d, judge identity and judge–human agreement are load-bearing. Please report the judge model/version, decoding settings, and a judge-validation study (e.g., judge vs. human labels on a subsample, with agreement statistics), or clarify if scoring is rule-based.
- [Abstract / §IV.E.1 and Eq. (17)] The headline 'best strict mean scoring-dimension accuracy is only 0.1599' conflates two distinct failure modes the paper itself separates elsewhere: normalization/format compliance and decision competence. The manuscript documents an action-only fallback parser and a 5% credibility rule precisely because strict validity q_imd is partly a format-compliance measure. As reported in the abstract, 0.1599 reads as a decision-competence number and likely overstates the capability gap in that interpretation. Please decompose strict failures into (i) structural/normalization failures and (ii) semantically wrong but well-formed responses, at least for the top models, and temper the abstract framing accordingly (e.g., report strict accuracy alongside the share of failures attributable to format).
- [§IV.D and Table X] The modality-removal study supports an abstract-level claim ('confirming that both visual and textual inputs influence decisions'), but it rests on a 20-anchor subset with no stated selection procedure, no uncertainty quantification, and full results shown for only 2 of 17 models (Table X). 'Changes 4–15 of 20 action selections' on n=20 is compatible with a wide range of effect sizes. Either expand the anchor set (with a stated sampling scheme and confidence intervals or an exact test) and report all models, or downgrade the claim in the abstract to a pilot-scale sensitivity observation consistent with §IV.H's framing of the human-reference pilot.
- [Reproducibility (Abstract; §III–IV)] The abstract claims a 'reproducible benchmark,' but the manuscript contains no data/code release statement, no pointer to the full scoring prompt and normalization parser, and no per-model inference settings (prompts, temperature, max tokens) backing the 'uniformly audited' claim. For a benchmark paper the artifact is the contribution; please add an explicit release plan (samples, Γ_i labels, scoring code, parser, prompts) and an appendix with the evaluation prompt and per-model decoding configuration. The semantic-scoring prompt in Fig. 5 appears abbreviated; the full template should be available.
minor comments (8)
- [§III.A] Terminology overload: 'task families,' 'taxonomy nodes,' 'scoring dimensions,' 'dimension groups,' and 'evaluation links' are used near-interchangeably in places (e.g., 'four evaluation links' vs. the groups in Eq. 4). A single glossary or consistent naming would help readers track the taxonomy-vs-dimension distinction the paper (rightly) insists on.
- [§III.C, Eq. (9)] Normalized entropy H_d and the hard-case support n_hard_d are defined but never used in any subsequent table or figure. Either report them (they would be informative for class-balance auditing of Y_d) or remove the definitions.
- [Table XVII] Modality-trust mismatch (Gemma-4-E4B) and collaboration-decision mismatch (Qwen3-VL-4B) are both reported as 3,024/3,024. A 100% trigger rate is either a pipeline artifact (e.g., the diagnostic fires whenever the response does not exactly match the required field) or a substantive claim that needs discussion. As tabulated it is uninformative; please clarify.
- [Table XIV] The 'Clean baseline' row reports degradation recognition 0.0000. Presumably there is no degradation to recognize on clean samples, so the metric is undefined rather than zero; scoring it as 0 risks misleading readers. Consider a dash or an explicit note.
- [§IV.C, Eqs. (15)–(19)] The MAD 0.25 alternative penalty, PACS equal weights, and the 5% normalization cutoff are disclosed as conventions, which is good practice; a brief sensitivity check (e.g., do model orderings under PACS survive weight perturbation, or do rankings change if the alternative penalty ranges 0.1–0.5) would strengthen the claim that these choices are not driving conclusions.
- [Table I and §II.C] The coverage legend renders as garbled glyphs ('#', 'G #', blank) and 'α3-Bench' appears with a corrupted character; check symbol fonts. Also 'UA V' spacing artifacts appear throughout the extracted text — verify these are not in the source.
- [§IV.A] State whether model inference used zero-shot prompts identical across models, whether any model-specific chat templates were needed, and how ties/parse ambiguities in the action-only fallback parser were resolved.
- [§IV.E.2 / Table XII] SmolVLM2-2.2B scores 0.6255 on the Action group but 0.0992 on Context — a striking inversion worth one sentence of interpretation, since it bears on the paper's claim that action metrics alone do not establish protocol-grounded capability.
Circularity Check
Empirical benchmark paper with no derivation chain that reduces predictions or first-principles claims to their inputs by construction.
full rationale
MulRobBench is an offline evaluation benchmark: it constructs labeled samples (Oi, Ci, Ri, Γi), defines scoring dimensions D1–D12 and action metrics (SafeAcc, UnsafeRate, MAD, PACS, strict Q), and measures 17 models against those contracts. Reporting that the best semantic protocol-decision score is 0.5141 and the best strict mean dimension accuracy is 0.1599 is an empirical measurement against author-assigned ground truth, not a claimed derivation or prediction forced by the inputs. Metric definitions (Eqs. 10–19) are ordinary benchmark contracts—averages and indicators over labeled sets—not self-definitional reductions of a scientific claim. There is no fitted parameter re-presented as an out-of-sample prediction, no uniqueness theorem imported from overlapping authors to forbid alternatives, and no ansatz smuggled in via self-citation. Weaknesses in label provenance, inter-rater agreement, or semantic-scorer design (if any) are validity/correctness concerns outside this circularity pass. The paper is self-contained as an empirical leaderboard; steps is empty.
Axiom & Free-Parameter Ledger
free parameters (4)
- MAD acceptable-alternative penalty (0.25) =
0.25
- PACS equal weights over four components =
1/4 each
- Normalization-failure credibility cutoff (5%) =
5%
- 20-anchor modality-ablation subset size =
20 anchors
axioms (5)
- domain assumption Offline single-step next-action selection under a fixed observation and injected rule state is a valid diagnostic proxy for decision-level cyber-physical safety and security-policy compliance.
- ad hoc to paper Protocol semantics (restricted zones, privacy/media rules, sensitive-place norms, temporary mission rules) may be injected as benchmark mission context without pixel-level entity verification in source imagery.
- domain assumption A closed vocabulary of ten action semantics plus sample-level standard/alternative/forbidden sets is sufficient to audit safe vs unsafe UAV decisions.
- domain assumption Controlled semantic scoring under paraphrase equivalence plus strict structural diagnostics together measure decision quality without requiring real flight certification.
- domain assumption UAVScenes-derived multimodal observations plus readability/alignment filters yield a representative strict evaluation distribution for smart-city decision pressure.
invented entities (3)
-
MulRobBench decision contract (Oi, Ci, Ri, Γi) with 17 taxonomy nodes and D1–D12 scoring dimensions
no independent evidence
-
Protocol-decision semantic score P_m over dimensions {8,9,10,11,12}
no independent evidence
-
Protocol-Action Composite Score (PACS)
no independent evidence
read the original abstract
Smart-city airspace is transforming Uncrewed Aerial Vehicles (UAVs) from passive sensing platforms into cyber-physical decision makers that must follow operational rules under degraded observations and ambiguous language. Existing UAV and multimodal benchmarks evaluate perception, navigation, collaboration, and reasoning, but few assess whether physical evidence, protocol constraints, and action risk remain coupled during critical decisions. We introduce MulRobBench, an offline, protocol-conditioned benchmark for Vision-Language-Action (VLA) UAV agents in smart-city environments. MulRobBench integrates real UAV multimodal observations, protocol-level security policies, and action-level cyber-physical safety into a unified evaluation framework. The benchmark contains 3,024 samples spanning 17 task taxonomy nodes and 12 scoring dimensions across four stages: operational context understanding, multimodal evidence arbitration, degradation-aware reasoning, and risk-aware action planning. Evaluation combines semantic scoring with structural diagnostics, including policy compliance, format compliance, unsafe actions, parsing failures, and dimension-level validity. Across 17 multimodal models, the best semantic protocol-decision score reaches only 0.5141, while the best strict mean scoring-dimension accuracy is 0.1599. A controlled 20-anchor modality-ablation study changes 4-15 action selections per model, confirming that both visual and textual inputs influence decisions. Analysis identifies modality-trust selection, constraint extraction, glare, missing data, and operator shorthand as the primary causes of decision instability. MulRobBench provides a reproducible benchmark for trustworthy multimodal UAV decision making under realistic operational constraints.
Figures
Reference graph
Works this paper leans on
-
[1]
M. A. Ferrag, A. Lakas, and M. Debbah, “UA VBench: An open benchmark dataset for autonomous and agentic AI UA V systems via LLM-generated flight scenarios,” 2025, arXiv:2511.11252. [Online]. Available: https://arxiv.org/abs/2511.11252
arXiv 2025
-
[2]
Drones as a service (DaaS) for 5G networks and blockchain-assisted IoT-based smart city infrastructure,
T. Garg, S. Gupta, M. S. Obaidat, and M. Raj, “Drones as a service (DaaS) for 5G networks and blockchain-assisted IoT-based smart city infrastructure,”Cluster Computing, vol. 27, pp. 8725–8788, 2024
2024
-
[3]
Advancing UA V security with artificial intelligence: A comprehensive survey of techniques and future directions,
F. Tlili, S. Ayed, and L. C. Fourati, “Advancing UA V security with artificial intelligence: A comprehensive survey of techniques and future directions,”Internet of Things, vol. 27, p. 101281, 2024
2024
-
[4]
Toward secure complex UA V cyber-physical systems: A unified threat taxonomy and cross-layer survey of cybersecurity challenges,
M. I. Umrani, B. Butler, A. O’ Driscoll, and S. Davy, “Toward secure complex UA V cyber-physical systems: A unified threat taxonomy and cross-layer survey of cybersecurity challenges,”Internet of Things, vol. 37, p. 101902, 2026
2026
-
[5]
Cyber physical systems: Design challenges,
E. A. Lee, “Cyber physical systems: Design challenges,” in2008 11th IEEE International Symposium on Object and Component-Oriented Real-Time Distributed Computing, 2008, pp. 363–369
2008
-
[6]
Cyber-physical systems security: A survey,
A. Humayed, J. Lin, F. Li, and B. Luo, “Cyber-physical systems security: A survey,”IEEE Internet of Things Journal, vol. 4, no. 6, pp. 1802–1831, 2017
2017
-
[7]
AirCopBench: A benchmark for multi-drone collaborative embodied perception and reasoning,
J. Zhaet al., “AirCopBench: A benchmark for multi-drone collaborative embodied perception and reasoning,”Proceedings of the AAAI Confer- ence on Artificial Intelligence, vol. 40, no. 2, pp. 1507–1515, 2026
2026
-
[8]
Benchmarking neural network robustness to common corruptions and perturbations,
D. Hendrycks and T. Dietterich, “Benchmarking neural network robustness to common corruptions and perturbations,” inInternational Conference on Learning Representations, 2019. [Online]. Available: https://arxiv.org/abs/1903.12261
Pith/arXiv arXiv 2019
-
[9]
EmbodiedCity: A benchmark platform for embodied agent in real-world city environment,
C. Gaoet al., “EmbodiedCity: A benchmark platform for embodied agent in real-world city environment,” 2024, arXiv:2410.09604. [Online]. Available: https://arxiv.org/abs/2410.09604
Pith/arXiv arXiv 2024
-
[10]
UrbanVideo-bench: Benchmarking vision-language models on embodied intelligence with video data in urban spaces,
B. Zhaoet al., “UrbanVideo-bench: Benchmarking vision-language models on embodied intelligence with video data in urban spaces,” inProceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). Vienna, Austria: Association for Computational Linguistics, Jul. 2025, pp. 32 400–32 423. [Online]. Available: ht...
2025
-
[11]
CityNav: A large-scale dataset for real-world aerial navigation,
J. Leeet al., “CityNav: A large-scale dataset for real-world aerial navigation,” inProceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), Oct. 2025, pp. 5912–5922. [Online]. Available: https://openaccess.thecvf.com/content/ICCV2025/html/Lee CityNav A Large-Scale Dataset for Real-World Aerial Navigation ICCV 2025 paper.html
2025
-
[12]
CityNavAgent: Aerial vision-and-language navigation with hierarchical semantic planning and global memory,
W. Zhanget al., “CityNavAgent: Aerial vision-and-language navigation with hierarchical semantic planning and global memory,” inProceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). Vienna, Austria: Association for Computational Linguistics, Jul. 2025, pp. 31 292–31 309. [Online]. Available: https:...
2025
-
[13]
S. Daiet al., “MM-UA VBench: How well do multimodal large language models see, think, and plan in low-altitude UA V scenarios?” 2025, arXiv:2512.23219. [Online]. Available: https://arxiv.org/abs/2512.23219
arXiv 2025
-
[14]
ESARBench: A benchmark for agentic UA V embodied search and rescue,
D. Zhang, P. Chen, J. Zhou, and S. Yang, “ESARBench: A benchmark for agentic UA V embodied search and rescue,” 2026, arXiv:2605.01371. [Online]. Available: https://arxiv.org/abs/2605.01371
Pith/arXiv arXiv 2026
-
[15]
UA V-ON: A benchmark for open-world object goal navigation with aerial agents,
J. Xiaoet al., “UA V-ON: A benchmark for open-world object goal navigation with aerial agents,” inProceedings of the 33rd ACM International Conference on Multimedia. Association for Computing Machinery, Oct. 2025, pp. 13 023–13 029. [Online]. Available: https://dl.acm.org/doi/10.1145/3746027.3758251
arXiv 2025
-
[16]
RT-2: Vision-language-action models transfer web knowledge to robotic control,
B. Zitkovichet al., “RT-2: Vision-language-action models transfer web knowledge to robotic control,” inProceedings of The 7th Conference on Robot Learning, ser. Proceedings of Machine Learning Research, vol. 229. PMLR, 2023, pp. 2165–2183. [Online]. Available: https://proceedings.mlr.press/v229/zitkovich23a.html
2023
-
[17]
OpenVLA: An open-source vision-language- action model,
M. J. Kimet al., “OpenVLA: An open-source vision-language- action model,” inProceedings of The 8th Conference on Robot Learning, ser. Proceedings of Machine Learning Research, vol
-
[18]
ARViP: Adversarial regularization in visuomotor policies for robotic VLA purpose,
H. Wang, B. Wu, and S. Zheng, “ARViP: Adversarial regularization in visuomotor policies for robotic VLA purpose,”IEEE Transactions on Industrial Informatics, vol. 22, no. 4, pp. 3275–3285, Apr. 2026
2026
-
[19]
A survey on vision– language–action models for embodied AI,
Y . Ma, Z. Song, Y . Zhuang, J. Hao, and I. King, “A survey on vision– language–action models for embodied AI,”IEEE Transactions on Neural Networks and Learning Systems, vol. 37, no. 7, pp. 3031–3051, Jul. 2026
2026
-
[20]
UA VScenes: A multi-modal dataset for UA Vs,
S. Wanget al., “UA VScenes: A multi-modal dataset for UA Vs,” in Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), Oct. 2025, pp. 28 946–28 958. [Online]. Available: https: //openaccess.thecvf.com/content/ICCV2025/html/Wang UA VScenes A Multi-Modal Dataset for UA VsICCV 2025 paper.html
2025
-
[21]
Replanning-oriented framework for efficient real-time decision-making in multi-UA V sys- tems,
X. Hai, L. Tan, Q. Feng, H. Duan, and C. Wen, “Replanning-oriented framework for efficient real-time decision-making in multi-UA V sys- tems,”IEEE Transactions on Industrial Informatics, vol. 21, no. 7, pp. 5127–5137, Jul. 2025
2025
-
[22]
Task offloading for multi-UA V asset edge computing with deep reinforcement learning,
S. A. Zakaryia, M. A. Mead, T. Nabil, and M. K. Hussein, “Task offloading for multi-UA V asset edge computing with deep reinforcement learning,”Cluster Computing, vol. 28, 2025, article 462
2025
-
[23]
Resilient event-triggered formation control and secure estimation of multi-UA V systems,
Z. Gu, T. Yin, Q. Lu, and J. H. Park, “Resilient event-triggered formation control and secure estimation of multi-UA V systems,”IEEE Transactions on Industrial Informatics, vol. 21, no. 6, pp. 4915–4923, Jun. 2025
2025
-
[24]
Authentica- tion framework for secure smart farming system deployed for sustainable development of smart cities: A review,
A. Patwal, M. Wazid, D. P. Singh, A. K. Das, and V . B. K, “Authentica- tion framework for secure smart farming system deployed for sustainable development of smart cities: A review,”Cluster Computing, vol. 29, 2026, article 433
2026
-
[25]
AerialVLN: Vision-and-language navigation for UA Vs,
S. Liu, H. Zhang, Y . Qi, P. Wang, Y . Zhang, and Q. Wu, “AerialVLN: Vision-and-language navigation for UA Vs,” inProceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), Oct. 2023, pp. 15 384–15 394. [Online]. Available: https: //openaccess.thecvf.com/content/ICCV2023/html/Liu AerialVLN Vision-and-Language Navigation for UA VsICCV ...
2023
-
[26]
MMMU: A massive multi-discipline multimodal understanding and reasoning benchmark for expert AGI,
X. Yueet al., “MMMU: A massive multi-discipline multimodal understanding and reasoning benchmark for expert AGI,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Jun. 2024, pp. 9556–9567. [Online]. Available: https://openaccess.thecvf.com/content/CVPR2024/html/Yue MMMU A Massive Multi-discipline Multimodal Unde...
2024
-
[27]
MMBench: Is your multi-modal model an all-around player?
Y . Liuet al., “MMBench: Is your multi-modal model an all-around player?” inComputer Vision – ECCV 2024, ser. Lecture Notes in Computer Science, vol. 15064. Cham: Springer, 2025, pp. 216–233. 22
2024
-
[28]
Video-MME: The first-ever comprehensive evaluation benchmark of multi-modal LLMs in video analysis,
C. Fuet al., “Video-MME: The first-ever comprehensive evaluation benchmark of multi-modal LLMs in video analysis,” inProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Jun. 2025, pp. 24 108–24 118. [Online]. Available: https://openaccess.thecvf.com/content/CVPR2025/html/Fu Video-MME The First-Ever Comprehensive Evalu...
2025
-
[29]
Holistic evaluation of language models,
P. Lianget al., “Holistic evaluation of language models,”Transactions on Machine Learning Research, 2023. [Online]. Available: https: //arxiv.org/abs/2211.09110
Pith/arXiv arXiv 2023
-
[30]
DecodingTrust: A comprehensive assessment of trustworthiness in GPT models,
B. Wanget al., “DecodingTrust: A comprehensive assessment of trustworthiness in GPT models,” inAdvances in Neural Information Processing Systems, 2023. [Online]. Available: https://proceedings.neurips.cc/paper files/paper/2023/ hash/63cb9921eecf51bfad27a99b2c53dd6d-Abstract-Datasets and Benchmarks.html
2023
-
[31]
Vision-based learning for drones: A survey,
J. Xiao, R. Zhang, Y . Zhang, and M. Feroskhan, “Vision-based learning for drones: A survey,”IEEE Transactions on Neural Networks and Learning Systems, vol. 36, no. 9, pp. 15 601–15 621, Sep. 2025
2025
-
[32]
M. A. Ferrag, A. Lakas, and M. Debbah, “α 3-Bench: A unified benchmark of safety, robustness, and efficiency for LLM-based UA V agents over 6G networks,” 2026, arXiv:2601.03281. [Online]. Available: https://arxiv.org/abs/2601.03281
arXiv 2026
-
[33]
HUGE-Bench: A benchmark for high-level UA V vision- language-action tasks,
J. Guoet al., “HUGE-Bench: A benchmark for high-level UA V vision- language-action tasks,” 2026, arXiv:2603.19822. [Online]. Available: https://arxiv.org/abs/2603.19822
Pith/arXiv arXiv 2026
-
[34]
A mathematical theory of communication,
C. E. Shannon, “A mathematical theory of communication,”The Bell System Technical Journal, vol. 27, no. 3–4, pp. 379–423, 623–656,
-
[35]
SelectiveNet: A deep neural network with an integrated reject option,
Y . Geifman and R. El-Yaniv, “SelectiveNet: A deep neural network with an integrated reject option,” inProceedings of the 36th International Conference on Machine Learning, ser. Proceedings of Machine Learning Research, vol. 97. PMLR, 2019, pp. 2151–2159. [Online]. Available: https://proceedings.mlr.press/v97/geifman19a.html
2019
-
[36]
On calibration of modern neural networks,
C. Guo, G. Pleiss, Y . Sun, and K. Q. Weinberger, “On calibration of modern neural networks,” inProceedings of the 34th International Conference on Machine Learning, ser. Proceedings of Machine Learning Research, vol. 70. PMLR, 2017, pp. 1321–1330. [Online]. Available: https://proceedings.mlr.press/v70/guo17a.html
2017
-
[37]
S. Baiet al., “Qwen3-VL technical report,” 2025, arXiv:2511.21631. [Online]. Available: https://arxiv.org/abs/2511.21631
Pith/arXiv arXiv 2025
-
[38]
Qwen3.5: Towards native multimodal agents,
Qwen Team, “Qwen3.5: Towards native multimodal agents,” 2026, official model-family release page. [Online]. Available: https://qwen.ai/ blog?id=qwen3.5
2026
-
[39]
SmolVLM2-2.2B-Instruct model card,
Hugging Face, “SmolVLM2-2.2B-Instruct model card,” 2025, official model card. [Online]. Available: https://huggingface.co/HuggingFaceTB/ SmolVLM2-2.2B-Instruct
2025
-
[40]
Phi-4-Mini technical report: Compact yet powerful multimodal language models via mixture-of-LoRAs,
Microsoft, “Phi-4-Mini technical report: Compact yet powerful multimodal language models via mixture-of-LoRAs,” 2025, arXiv:2503.01743; includes Phi-4-Multimodal. [Online]. Available: https://arxiv.org/abs/2503.01743
Pith/arXiv arXiv 2025
-
[41]
Gemma 4 model overview,
Google, “Gemma 4 model overview,” 2026, official model documenta- tion. [Online]. Available: https://ai.google.dev/gemma/docs/core
2026
-
[42]
InternVL3.5: Advancing open-source multimodal models in versatility, reasoning, and efficiency,
W. Wanget al., “InternVL3.5: Advancing open-source multimodal models in versatility, reasoning, and efficiency,” 2025, arXiv:2508.18265. [Online]. Available: https://arxiv.org/abs/2508.18265
Pith/arXiv arXiv 2025
-
[43]
GLM-4.6V overview,
Z.AI, “GLM-4.6V overview,” 2026, official developer documentation for the GLM-4.6V series, including GLM-4.6V-Flash. [Online]. Available: https://docs.z.ai/guides/vlm/glm-4.6v
2026
-
[44]
S. Baiet al., “Qwen2.5-VL technical report,” 2025, arXiv:2502.13923. [Online]. Available: https://arxiv.org/abs/2502.13923
Pith/arXiv arXiv 2025
-
[45]
MiniCPM-V 4.5: Cooking efficient MLLMs via architecture, data, and training recipe,
T. Yuet al., “MiniCPM-V 4.5: Cooking efficient MLLMs via architecture, data, and training recipe,” 2025, arXiv:2509.18154. [Online]. Available: https://arxiv.org/abs/2509.18154
Pith/arXiv arXiv 2025
-
[46]
Aya Vision: Advancing the frontier of multilingual multimodality,
S. Dashet al., “Aya Vision: Advancing the frontier of multilingual multimodality,” 2025, arXiv:2505.08751. [Online]. Available: https: //arxiv.org/abs/2505.08751
Pith/arXiv arXiv 2025
-
[47]
Building and better understanding vision-language models,
H. Laurenc ¸on, L. Tronchon, M. Cord, and V . Sanh, “Building and better understanding vision-language models,” 2024, arXiv:2408.12637; includes Idefics3-8B. [Online]. Available: https://arxiv.org/abs/2408. 12637
Pith/arXiv arXiv 2024
-
[48]
LLaV A-NeXT: Improved reasoning, OCR, and world knowledge,
LLaV A Team, “LLaV A-NeXT: Improved reasoning, OCR, and world knowledge,” 2024, official project release page. [Online]. Available: https://llava-vl.github.io/blog/2024-01-30-llava-next/
2024
-
[49]
Concrete problems in AI safety,
D. Amodei, C. Olah, J. Steinhardt, P. Christiano, J. Schulman, and D. Mane, “Concrete problems in AI safety,” 2016, arXiv:1606.06565. [Online]. Available: https://arxiv.org/abs/1606.06565
Pith/arXiv arXiv 2016
-
[50]
Advanced security frameworks for UA V and IoT: A deep learning approach,
N. Quadar, A. Chehri, and B. Debaque, “Advanced security frameworks for UA V and IoT: A deep learning approach,”Internet of Things, vol. 32, p. 101594, 2025
2025
-
[270]
2679–2713
PMLR, 2025, pp. 2679–2713. [Online]. Available: https: //proceedings.mlr.press/v270/kim25c.html
2025
-
[1948]
Available: https://people.math.harvard.edu/ ∼ctm/home/ text/others/shannon/entropy/entropy.pdf
[Online]. Available: https://people.math.harvard.edu/ ∼ctm/home/ text/others/shannon/entropy/entropy.pdf
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.