Pith. sign in

REVIEW 5 major objections 8 minor 52 references

MulRobBench: A Decision-Level Benchmark for Safe and Security-Policy-Compliant Multimodal UAV Agents

T0 review · 5 major / 8 minor · reviewed 2026-07-30 · grok-4.5

Pith's one-line read Current multimodal models cannot reliably keep degraded evidence, security rules, and safe UAV actions coupled at the decision point.

desk verdict Solid offline UAV decision benchmark with a real multi-model gap; the headline numbers are only as strong as unvalidated action/policy labels. read the letter →

arxiv 2607.23870 v1 pith:RBI5XKEM submitted 2026-07-26 cs.MA cs.AI

classification cs.MAcs.AI
keywords cyber-physicalsecuritysafeactiondecisionmakingsecurity-policycompliancesmart-cityUAVagentsbenchmarkvision-language-actionmultimodalevidencearbitrationdegradation-awarereasoning
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

Smart-city drones are no longer just cameras; they must choose a next action under bad visibility, messy operator language, and explicit security rules. This paper argues that existing UAV and vision-language benchmarks miss that coupling and offers MulRobBench, an offline test that forces models to recover mission context, arbitrate multimodal evidence, handle degradation, and pick a controlled safe action. On 3,024 strict samples across 17 models, the best semantic protocol-decision score is only 0.5141 and the best strict dimension accuracy is 0.1599. Failures concentrate in modality trust, constraint extraction, collaboration or abstention thresholds, and action consistency—not coarse scene recognition. A sympathetic reader should care because a fluent scene description can still violate a restricted zone, ignore missing data, or skip a required reobservation, which is exactly the cyber-physical risk the benchmark isolates.

What carries the argument

MulRobBench: an offline, protocol-conditioned Vision-Language-Action decision contract that binds real UAV multimodal observations, injected security-policy mission context, and a closed ten-action safety vocabulary, scored along four links (context, evidence arbitration, degradation-aware reasoning, risk-aware action) with 12 dimensions and both semantic scores and strict structural diagnostics.

What would settle it

Find a model that, on the same 3,024-sample strict split and closed action contract, simultaneously posts high semantic protocol-decision score, high strict mean dimension accuracy, low unsafe-action rate, and near-zero normalization failures, with modality removal no longer flipping many of the 20-anchor actions.

Watch

Extended reading notes

Core claim

Across 17 uniformly audited multimodal models on MulRobBench’s 3,024 strict samples, current systems remain far from reliable protocol-conditioned UAV decision making: the best semantic protocol-decision score is 0.5141 and the best strict mean scoring-dimension accuracy is 0.1599. Models handle coarse context and targets relatively well but break on modality-trust selection, constraint extraction, collaboration and abstention triggers, and risk-aware action planning. Modality-removal on a matched 20-anchor subset changes 4–15 action selections per model, showing both vision and text shape decisions while exposing unstable combination of those inputs.

Load-bearing premise

Treating airport boundaries, sensitive-place rules, privacy limits, and other policies as injected mission context—rather than labels verified in the source imagery—still fairly measures whether a model preserves security-policy compliance when it chooses an action.

Editorial extensions

If this is right

  • UAV VLA evaluation must report semantic validity, structural parseability, and action safety side by side; a single accuracy score hides decision-chain breaks.
  • Improving coarse scene or target recognition will not close the gap; gains must target modality-trust, constraint extraction, and abstention or collaboration triggers.
  • Glare, missing data, and operator shorthand are priority stress cases because they systematically decouple evidence quality from allowed actions.
  • Offline protocol-conditioned next-action audits become a necessary gate before claims of smart-city UAV cyber-physical safety.
  • Human multi-expert pilots on stratified subsets remain far above models, setting a concrete gap for future systems work.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • Training objectives that reward fluent captions or open-ended plans may actively work against the structured action contract this benchmark requires.
  • If injected policies are later grounded in verified maps and entity labels, the same four-link chain could become a live mission-audit layer rather than only an offline test.
  • The non-monotonic modality-ablation results suggest some models are already over-relying on text priors; denser paired vision-text counterfactuals would expose that shortcut more sharply.
  • Similar decision-contract thinking likely transfers to other rule-bound embodied settings (ground robots in restricted sites, inspection agents) where evidence degradation and forbidden actions coexist.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

5 major / 8 minor

Summary. The paper introduces MulRobBench, an offline benchmark for protocol-conditioned safe next-action selection by multimodal UAV agents in smart-city settings. Built on UAVScenes-derived observations with injected protocol semantics (restricted zones, privacy/media rules, temporary mission rules), the benchmark comprises 3,024 strict samples organized into 17 primary-attribution taxonomy nodes and 12 scoring dimensions (D1–D12) across four decision links. Seventeen multimodal models are evaluated under a shared decision contract with separated semantic scoring (Eq. 10–11) and strict structural diagnostics (Eq. 17, 20), plus action-level metrics (SafeAcc, UnsafeRate, MAD, PACS). Headline results: best semantic protocol-decision score 0.5141, best strict mean dimension accuracy 0.1599; a 20-anchor modality-removal study changes 4–15/20 actions per model. The authors attribute failures primarily to modality-trust selection, constraint extraction, collaboration/abstention thresholds, and action-rationale consistency rather than coarse context recognition.

Significance. If the results hold, this is a useful and timely contribution. The decision-level endpoint (protocol-conditioned next action under a closed action vocabulary) is genuinely under-served by existing UAV benchmarks, and the paper's central methodological choice — reporting semantic scores, action-set agreement, and structural compliance side by side rather than collapsing them into one number — is well motivated and executed with unusual discipline. Specific strengths worth naming: a fixed, audited 3,024-sample strict set with a documented split contract (Tables VI–VII); 17 models under a uniform audit with normalization failures retained rather than silently dropped (Phi-4-Multimodal's 3,024 failures are reported, not removed); a 5% normalization-failure credibility rule that prevents action-only fallback numbers from being over-read; per-dimension, per-condition, and error-chain analyses (Tables XIII–XVII) that localize failures to the middle of the decision chain; and explicit, honest limitation statements about the injected-semantics premise. The multi-expert reference pilot, the conditional robustness matrix (Table XVI), and the modality-ablation study are all falsifiable, checka

major comments (5)
  1. [§III.B (Benchmark Generation, step four) and Table XI] Every metric in the paper (Eqs. 5, 10, 13–17) is scored against Γ_i = (A*_i, A+_i, A−_i) and the dimension projections g_id = π_d(x_i), yet the manuscript nowhere reports who produced these labels, under what protocol, with what adjudication, or with what inter-rater agreement. Many labels are judgment calls rather than facts: whether a dust-occluded frame warrants hover vs. reobserve vs. request-another-UAV; which constraints are 'active' under a given protocol; what the 'primary' degradation is when conditions co-occur (Table XIV shows samples carrying multiple condition labels). The Multi-Expert Reference row cannot substitute for label validation: it is a 600-sample pilot with round values (0.83/0.84/0.94/0.78), no rater count, no agreement statistic, and no statement of whether the reference annotators are disjoint from the original label authors — if they overlap, the row measures
  2. [§III.D / Fig. 5 and Eq. (10)] The controlled semantic score s_d(ŷ_im, g_id) is computed via a scoring prompt (Fig. 5), which implies an LLM-based judge, but the manuscript never states which model executes the judge, at what settings, or how the judge itself was validated against human scoring. Since all semantic metrics (S_md, P_m, the group scores in Table XII, and the conditional tables) derive from s_d, judge identity and judge–human agreement are load-bearing. Please report the judge model/version, decoding settings, and a judge-validation study (e.g., judge vs. human labels on a subsample, with agreement statistics), or clarify if scoring is rule-based.
  3. [Abstract / §IV.E.1 and Eq. (17)] The headline 'best strict mean scoring-dimension accuracy is only 0.1599' conflates two distinct failure modes the paper itself separates elsewhere: normalization/format compliance and decision competence. The manuscript documents an action-only fallback parser and a 5% credibility rule precisely because strict validity q_imd is partly a format-compliance measure. As reported in the abstract, 0.1599 reads as a decision-competence number and likely overstates the capability gap in that interpretation. Please decompose strict failures into (i) structural/normalization failures and (ii) semantically wrong but well-formed responses, at least for the top models, and temper the abstract framing accordingly (e.g., report strict accuracy alongside the share of failures attributable to format).
  4. [§IV.D and Table X] The modality-removal study supports an abstract-level claim ('confirming that both visual and textual inputs influence decisions'), but it rests on a 20-anchor subset with no stated selection procedure, no uncertainty quantification, and full results shown for only 2 of 17 models (Table X). 'Changes 4–15 of 20 action selections' on n=20 is compatible with a wide range of effect sizes. Either expand the anchor set (with a stated sampling scheme and confidence intervals or an exact test) and report all models, or downgrade the claim in the abstract to a pilot-scale sensitivity observation consistent with §IV.H's framing of the human-reference pilot.
  5. [Reproducibility (Abstract; §III–IV)] The abstract claims a 'reproducible benchmark,' but the manuscript contains no data/code release statement, no pointer to the full scoring prompt and normalization parser, and no per-model inference settings (prompts, temperature, max tokens) backing the 'uniformly audited' claim. For a benchmark paper the artifact is the contribution; please add an explicit release plan (samples, Γ_i labels, scoring code, parser, prompts) and an appendix with the evaluation prompt and per-model decoding configuration. The semantic-scoring prompt in Fig. 5 appears abbreviated; the full template should be available.
minor comments (8)
  1. [§III.A] Terminology overload: 'task families,' 'taxonomy nodes,' 'scoring dimensions,' 'dimension groups,' and 'evaluation links' are used near-interchangeably in places (e.g., 'four evaluation links' vs. the groups in Eq. 4). A single glossary or consistent naming would help readers track the taxonomy-vs-dimension distinction the paper (rightly) insists on.
  2. [§III.C, Eq. (9)] Normalized entropy H_d and the hard-case support n_hard_d are defined but never used in any subsequent table or figure. Either report them (they would be informative for class-balance auditing of Y_d) or remove the definitions.
  3. [Table XVII] Modality-trust mismatch (Gemma-4-E4B) and collaboration-decision mismatch (Qwen3-VL-4B) are both reported as 3,024/3,024. A 100% trigger rate is either a pipeline artifact (e.g., the diagnostic fires whenever the response does not exactly match the required field) or a substantive claim that needs discussion. As tabulated it is uninformative; please clarify.
  4. [Table XIV] The 'Clean baseline' row reports degradation recognition 0.0000. Presumably there is no degradation to recognize on clean samples, so the metric is undefined rather than zero; scoring it as 0 risks misleading readers. Consider a dash or an explicit note.
  5. [§IV.C, Eqs. (15)–(19)] The MAD 0.25 alternative penalty, PACS equal weights, and the 5% normalization cutoff are disclosed as conventions, which is good practice; a brief sensitivity check (e.g., do model orderings under PACS survive weight perturbation, or do rankings change if the alternative penalty ranges 0.1–0.5) would strengthen the claim that these choices are not driving conclusions.
  6. [Table I and §II.C] The coverage legend renders as garbled glyphs ('#', 'G #', blank) and 'α3-Bench' appears with a corrupted character; check symbol fonts. Also 'UA V' spacing artifacts appear throughout the extracted text — verify these are not in the source.
  7. [§IV.A] State whether model inference used zero-shot prompts identical across models, whether any model-specific chat templates were needed, and how ties/parse ambiguities in the action-only fallback parser were resolved.
  8. [§IV.E.2 / Table XII] SmolVLM2-2.2B scores 0.6255 on the Action group but 0.0992 on Context — a striking inversion worth one sentence of interpretation, since it bears on the paper's claim that action metrics alone do not establish protocol-grounded capability.

Circularity Check

0 steps flagged · score 0.0 of 10

Empirical benchmark paper with no derivation chain that reduces predictions or first-principles claims to their inputs by construction.

full rationale

MulRobBench is an offline evaluation benchmark: it constructs labeled samples (Oi, Ci, Ri, Γi), defines scoring dimensions D1–D12 and action metrics (SafeAcc, UnsafeRate, MAD, PACS, strict Q), and measures 17 models against those contracts. Reporting that the best semantic protocol-decision score is 0.5141 and the best strict mean dimension accuracy is 0.1599 is an empirical measurement against author-assigned ground truth, not a claimed derivation or prediction forced by the inputs. Metric definitions (Eqs. 10–19) are ordinary benchmark contracts—averages and indicators over labeled sets—not self-definitional reductions of a scientific claim. There is no fitted parameter re-presented as an out-of-sample prediction, no uniqueness theorem imported from overlapping authors to forbid alternatives, and no ansatz smuggled in via self-citation. Weaknesses in label provenance, inter-rater agreement, or semantic-scorer design (if any) are validity/correctness concerns outside this circularity pass. The paper is self-contained as an empirical leaderboard; steps is empty.

Assumptions & free parameters 4 free parameters · 5 assumptions · 3 invented entities

Load-bearing commitments are benchmark-design choices and evaluation conventions rather than physical laws: offline single-step decision points stand in for cyber-physical safety; injected protocol overlays define ground-truth constraints; a closed ten-action vocabulary defines safe/forbidden sets; semantic paraphrase scoring plus strict parse/action diagnostics jointly define success; and composite indices (PACS, MAD) use hand-chosen weights/penalties. Free parameters are few and explicit reporting conventions. Invented entities are benchmark constructs (taxonomy, dimensions, composites), not new physical mediators.

free parameters (4)
  • MAD acceptable-alternative penalty (0.25) = 0.25
    Eq. 15 assigns deviation 0.25 to safe alternatives vs 0/1 for exact/other; authors call it a reporting convention, not calibrated physical risk, yet it enters MAD and multiplicatively scales PACS.
  • PACS equal weights over four components = 1/4 each
    Eq. 18 averages Q_m4, Q_m5, SafeAcc_m, and Q_m11 with equal weight by design choice; different weights would reorder composite diagnostics.
  • Normalization-failure credibility cutoff (5%) = 5%
    Models above 5% normalization failure are barred from best-action markings; cutoff is stated as interpretability policy, not empirically justified safety threshold.
  • 20-anchor modality-ablation subset size = 20 anchors
    Modality-sensitivity claim rests on a hand-chosen matched 20-sample anchor set rather than the full 3,024.
assumptions (5)
  • domain assumption Offline single-step next-action selection under a fixed observation and injected rule state is a valid diagnostic proxy for decision-level cyber-physical safety and security-policy compliance.
    Stated throughout §I, §III, and Limitations; closed-loop control, latency, and long-horizon missions are explicitly out of scope.
  • ad hoc to paper Protocol semantics (restricted zones, privacy/media rules, sensitive-place norms, temporary mission rules) may be injected as benchmark mission context without pixel-level entity verification in source imagery.
    §III.B and Limitations treat these as protocolised decision conditions, not visually verified labels; central compliance claims depend on this separation.
  • domain assumption A closed vocabulary of ten action semantics plus sample-level standard/alternative/forbidden sets is sufficient to audit safe vs unsafe UAV decisions.
    Table III and Eqs. 12–14; only seven standard actions appear in the strict set, with ten retained for alternatives/extensions.
  • domain assumption Controlled semantic scoring under paraphrase equivalence plus strict structural diagnostics together measure decision quality without requiring real flight certification.
    §III.D and §IV.C; authors explicitly say semantic scores are not direct flight-control certification.
  • domain assumption UAVScenes-derived multimodal observations plus readability/alignment filters yield a representative strict evaluation distribution for smart-city decision pressure.
    §III.B–C construction pipeline; long-tailed preset schedule rather than real incident frequencies (authors caution on this).
invented entities (3)
  • MulRobBench decision contract (Oi, Ci, Ri, Γi) with 17 taxonomy nodes and D1–D12 scoring dimensions
    purpose: Organize protocol-conditioned VLA evaluation into auditable context, evidence, degradation, and action links.
    Core benchmark ontology introduced in §III.A; not an external standard prior to this paper.
  • Protocol-decision semantic score P_m over dimensions {8,9,10,11,12}
    purpose: Aggregate constraint, target, safe action, abstention/reobservation, and collaboration semantics into one headline semantic metric.
    Eq. 11 defines a benchmark-local aggregate used for ranking claims.
  • Protocol-Action Composite Score (PACS)
    purpose: Summarize degradation judgment, modality trust, safe-set agreement, reobservation, and action deviation in one composite.
    Eqs. 18–19; authors state it is not an inherited standard metric and not a sole ranking.

how reviews work

0 comments
Cite this review

Pith. "Pith review of MulRobBench: A Decision-Level Benchmark for Safe and Security-Policy-Compliant Multimodal UAV Agents." pith.science (2026). https://pith.science/paper/RBI5XKEM

@misc{pith2026260723870,
  author       = {Pith},
  title        = {Pith review of: MulRobBench: A Decision-Level Benchmark for Safe and Security-Policy-Compliant Multimodal UAV Agents},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/RBI5XKEM}},
  note         = {Machine review of arXiv:2607.23870}
}
read the original abstract

Smart-city airspace is transforming Uncrewed Aerial Vehicles (UAVs) from passive sensing platforms into cyber-physical decision makers that must follow operational rules under degraded observations and ambiguous language. Existing UAV and multimodal benchmarks evaluate perception, navigation, collaboration, and reasoning, but few assess whether physical evidence, protocol constraints, and action risk remain coupled during critical decisions. We introduce MulRobBench, an offline, protocol-conditioned benchmark for Vision-Language-Action (VLA) UAV agents in smart-city environments. MulRobBench integrates real UAV multimodal observations, protocol-level security policies, and action-level cyber-physical safety into a unified evaluation framework. The benchmark contains 3,024 samples spanning 17 task taxonomy nodes and 12 scoring dimensions across four stages: operational context understanding, multimodal evidence arbitration, degradation-aware reasoning, and risk-aware action planning. Evaluation combines semantic scoring with structural diagnostics, including policy compliance, format compliance, unsafe actions, parsing failures, and dimension-level validity. Across 17 multimodal models, the best semantic protocol-decision score reaches only 0.5141, while the best strict mean scoring-dimension accuracy is 0.1599. A controlled 20-anchor modality-ablation study changes 4-15 action selections per model, confirming that both visual and textual inputs influence decisions. Analysis identifies modality-trust selection, constraint extraction, glare, missing data, and operator shorthand as the primary causes of decision instability. MulRobBench provides a reproducible benchmark for trustworthy multimodal UAV decision making under realistic operational constraints.

Figures

Figures reproduced from arXiv: 2607.23870 by the authors.

Figure 1
Figure 1. Two-level UAV-VLA task taxonomy. The inner ring contains four task families, and the outer ring contains 17 primary-attribution task-taxonomy [PITH_FULL_IMAGE:figures/full_fig_p006_1.png] view at source ↗
Figure 2
Figure 2. Primary-attribution taxonomy of the strict evaluation set. The figure [PITH_FULL_IMAGE:figures/full_fig_p007_2.png] view at source ↗
Figure 3
Figure 3. Protocol-context distribution by source-domain bucket. The figure [PITH_FULL_IMAGE:figures/full_fig_p008_3.png] view at source ↗
Figures from the paper (11 more)
Figure 4
Figure 4. Figure 4: Observation-condition family distribution. Labels include the clean [PITH_FULL_IMAGE:figures/full_fig_p008_4.png]
Figure 5
Figure 5. Figure 5: Semantic scoring prompt template. The template constrains physical observation, mission context, security-policy rules, degradation conditions, action [PITH_FULL_IMAGE:figures/full_fig_p010_5.png]
Figure 6
Figure 6. Figure 6: Semantic decision contrast under ordinary patrol and emergency hazard. A readable boulevard observation permits slow conservative progress, whereas [PITH_FULL_IMAGE:figures/full_fig_p010_6.png]
Figure 7
Figure 7. Figure 7: Boundary-restricted decision check near an airport perimeter. The case tests whether limited visual evidence and restricted-zone protocol semantics [PITH_FULL_IMAGE:figures/full_fig_p011_7.png]
Figure 8
Figure 8. Figure 8: Degradation-aware decision check under dust-obscured boulevard patrol. The case tests whether degraded RGB evidence shifts trust toward nonvisual [PITH_FULL_IMAGE:figures/full_fig_p011_8.png]
Figure 10
Figure 10. Figure 10: Per-dimension result structure over D1–D12. The figure localizes model differences to context, evidence, degradation, or action links; these scoring dimensions are distinct from the 17 primary-attribution task-taxonomy nodes. can identify protocol context reasonably w…
Figure 11
Figure 11. Figure 11: Result structure after scoring-dimension group compression. Context [PITH_FULL_IMAGE:figures/full_fig_p015_11.png]
Figure 12
Figure 12. Figure 12: Security-policy semantics and unsafe-action tradeoff. The plot [PITH_FULL_IMAGE:figures/full_fig_p015_12.png]
Figure 15
Figure 15. Figure 15: Degradation-specific robustness for Qwen3-VL 8B. The plot relates [PITH_FULL_IMAGE:figures/full_fig_p016_15.png]
Figure 16
Figure 16. Figure 16: Protocol-context action pressure for Qwen3-VL 8B. The plot [PITH_FULL_IMAGE:figures/full_fig_p016_16.png]
Figure 18
Figure 18. Figure 18: Unsafe action replacement failure under a fire response scenario. The [PITH_FULL_IMAGE:figures/full_fig_p019_18.png]

Discussion (0). Sign in to comment.

Reference graph

Works this paper leans on

52 extracted references · 13 linked inside Pith

  1. [1]

    UA VBench: An open benchmark dataset for autonomous and agentic AI UA V systems via LLM-generated flight scenarios,

    M. A. Ferrag, A. Lakas, and M. Debbah, “UA VBench: An open benchmark dataset for autonomous and agentic AI UA V systems via LLM-generated flight scenarios,” 2025, arXiv:2511.11252. [Online]. Available: https://arxiv.org/abs/2511.11252

  2. [2]

    Drones as a service (DaaS) for 5G networks and blockchain-assisted IoT-based smart city infrastructure,

    T. Garg, S. Gupta, M. S. Obaidat, and M. Raj, “Drones as a service (DaaS) for 5G networks and blockchain-assisted IoT-based smart city infrastructure,”Cluster Computing, vol. 27, pp. 8725–8788, 2024

  3. [3]

    Advancing UA V security with artificial intelligence: A comprehensive survey of techniques and future directions,

    F. Tlili, S. Ayed, and L. C. Fourati, “Advancing UA V security with artificial intelligence: A comprehensive survey of techniques and future directions,”Internet of Things, vol. 27, p. 101281, 2024

  4. [4]

    Toward secure complex UA V cyber-physical systems: A unified threat taxonomy and cross-layer survey of cybersecurity challenges,

    M. I. Umrani, B. Butler, A. O’ Driscoll, and S. Davy, “Toward secure complex UA V cyber-physical systems: A unified threat taxonomy and cross-layer survey of cybersecurity challenges,”Internet of Things, vol. 37, p. 101902, 2026

  5. [5]

    Cyber physical systems: Design challenges,

    E. A. Lee, “Cyber physical systems: Design challenges,” in2008 11th IEEE International Symposium on Object and Component-Oriented Real-Time Distributed Computing, 2008, pp. 363–369

  6. [6]

    Cyber-physical systems security: A survey,

    A. Humayed, J. Lin, F. Li, and B. Luo, “Cyber-physical systems security: A survey,”IEEE Internet of Things Journal, vol. 4, no. 6, pp. 1802–1831, 2017

  7. [7]

    AirCopBench: A benchmark for multi-drone collaborative embodied perception and reasoning,

    J. Zhaet al., “AirCopBench: A benchmark for multi-drone collaborative embodied perception and reasoning,”Proceedings of the AAAI Confer- ence on Artificial Intelligence, vol. 40, no. 2, pp. 1507–1515, 2026

  8. [8]

    Benchmarking neural network robustness to common corruptions and perturbations,

    D. Hendrycks and T. Dietterich, “Benchmarking neural network robustness to common corruptions and perturbations,” inInternational Conference on Learning Representations, 2019. [Online]. Available: https://arxiv.org/abs/1903.12261

Show all 52 references
  1. [9]

    EmbodiedCity: A benchmark platform for embodied agent in real-world city environment,

    C. Gaoet al., “EmbodiedCity: A benchmark platform for embodied agent in real-world city environment,” 2024, arXiv:2410.09604. [Online]. Available: https://arxiv.org/abs/2410.09604

  2. [10]

    UrbanVideo-bench: Benchmarking vision-language models on embodied intelligence with video data in urban spaces,

    B. Zhaoet al., “UrbanVideo-bench: Benchmarking vision-language models on embodied intelligence with video data in urban spaces,” inProceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). Vienna, Austria: Association for ...

  3. [11]

    CityNav: A large-scale dataset for real-world aerial navigation,

    J. Leeet al., “CityNav: A large-scale dataset for real-world aerial navigation,” inProceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), Oct. 2025, pp. 5912–5922. [Online]. Available: https://openaccess.thecvf.com/content/ICCV2025/html/Lee CityNav A L...

  4. [12]

    CityNavAgent: Aerial vision-and-language navigation with hierarchical semantic planning and global memory,

    W. Zhanget al., “CityNavAgent: Aerial vision-and-language navigation with hierarchical semantic planning and global memory,” inProceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). Vienna, Austria: Association for Comp...

  5. [13]

    MM-UA VBench: How well do multimodal large language models see, think, and plan in low-altitude UA V scenarios?

    S. Daiet al., “MM-UA VBench: How well do multimodal large language models see, think, and plan in low-altitude UA V scenarios?” 2025, arXiv:2512.23219. [Online]. Available: https://arxiv.org/abs/2512.23219

  6. [14]

    ESARBench: A benchmark for agentic UA V embodied search and rescue,

    D. Zhang, P. Chen, J. Zhou, and S. Yang, “ESARBench: A benchmark for agentic UA V embodied search and rescue,” 2026, arXiv:2605.01371. [Online]. Available: https://arxiv.org/abs/2605.01371

  7. [15]

    UA V-ON: A benchmark for open-world object goal navigation with aerial agents,

    J. Xiaoet al., “UA V-ON: A benchmark for open-world object goal navigation with aerial agents,” inProceedings of the 33rd ACM International Conference on Multimedia. Association for Computing Machinery, Oct. 2025, pp. 13 023–13 029. [Online]. Available: https://dl.acm.org/doi/...

  8. [16]

    RT-2: Vision-language-action models transfer web knowledge to robotic control,

    B. Zitkovichet al., “RT-2: Vision-language-action models transfer web knowledge to robotic control,” inProceedings of The 7th Conference on Robot Learning, ser. Proceedings of Machine Learning Research, vol. 229. PMLR, 2023, pp. 2165–2183. [Online]. Available: https://proceedi...

  9. [17]

    OpenVLA: An open-source vision-language- action model,

    M. J. Kimet al., “OpenVLA: An open-source vision-language- action model,” inProceedings of The 8th Conference on Robot Learning, ser. Proceedings of Machine Learning Research, vol

  10. [18]

    ARViP: Adversarial regularization in visuomotor policies for robotic VLA purpose,

    H. Wang, B. Wu, and S. Zheng, “ARViP: Adversarial regularization in visuomotor policies for robotic VLA purpose,”IEEE Transactions on Industrial Informatics, vol. 22, no. 4, pp. 3275–3285, Apr. 2026

  11. [19]

    A survey on vision– language–action models for embodied AI,

    Y . Ma, Z. Song, Y . Zhuang, J. Hao, and I. King, “A survey on vision– language–action models for embodied AI,”IEEE Transactions on Neural Networks and Learning Systems, vol. 37, no. 7, pp. 3031–3051, Jul. 2026

  12. [20]

    UA VScenes: A multi-modal dataset for UA Vs,

    S. Wanget al., “UA VScenes: A multi-modal dataset for UA Vs,” in Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), Oct. 2025, pp. 28 946–28 958. [Online]. Available: https: //openaccess.thecvf.com/content/ICCV2025/html/Wang UA VScenes A Multi-Moda...

  13. [21]

    Replanning-oriented framework for efficient real-time decision-making in multi-UA V sys- tems,

    X. Hai, L. Tan, Q. Feng, H. Duan, and C. Wen, “Replanning-oriented framework for efficient real-time decision-making in multi-UA V sys- tems,”IEEE Transactions on Industrial Informatics, vol. 21, no. 7, pp. 5127–5137, Jul. 2025

  14. [22]

    Task offloading for multi-UA V asset edge computing with deep reinforcement learning,

    S. A. Zakaryia, M. A. Mead, T. Nabil, and M. K. Hussein, “Task offloading for multi-UA V asset edge computing with deep reinforcement learning,”Cluster Computing, vol. 28, 2025, article 462

  15. [23]

    Resilient event-triggered formation control and secure estimation of multi-UA V systems,

    Z. Gu, T. Yin, Q. Lu, and J. H. Park, “Resilient event-triggered formation control and secure estimation of multi-UA V systems,”IEEE Transactions on Industrial Informatics, vol. 21, no. 6, pp. 4915–4923, Jun. 2025

  16. [24]

    Authentica- tion framework for secure smart farming system deployed for sustainable development of smart cities: A review,

    A. Patwal, M. Wazid, D. P. Singh, A. K. Das, and V . B. K, “Authentica- tion framework for secure smart farming system deployed for sustainable development of smart cities: A review,”Cluster Computing, vol. 29, 2026, article 433

  17. [25]

    AerialVLN: Vision-and-language navigation for UA Vs,

    S. Liu, H. Zhang, Y . Qi, P. Wang, Y . Zhang, and Q. Wu, “AerialVLN: Vision-and-language navigation for UA Vs,” inProceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), Oct. 2023, pp. 15 384–15 394. [Online]. Available: https: //openaccess.thecvf.com/c...

  18. [26]

    MMMU: A massive multi-discipline multimodal understanding and reasoning benchmark for expert AGI,

    X. Yueet al., “MMMU: A massive multi-discipline multimodal understanding and reasoning benchmark for expert AGI,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Jun. 2024, pp. 9556–9567. [Online]. Available: https://openaccess.thec...

  19. [27]

    MMBench: Is your multi-modal model an all-around player?

    Y . Liuet al., “MMBench: Is your multi-modal model an all-around player?” inComputer Vision – ECCV 2024, ser. Lecture Notes in Computer Science, vol. 15064. Cham: Springer, 2025, pp. 216–233. 22

  20. [28]

    Video-MME: The first-ever comprehensive evaluation benchmark of multi-modal LLMs in video analysis,

    C. Fuet al., “Video-MME: The first-ever comprehensive evaluation benchmark of multi-modal LLMs in video analysis,” inProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Jun. 2025, pp. 24 108–24 118. [Online]. Available: https://openaccess....

  21. [29]

    Holistic evaluation of language models,

    P. Lianget al., “Holistic evaluation of language models,”Transactions on Machine Learning Research, 2023. [Online]. Available: https: //arxiv.org/abs/2211.09110

  22. [30]

    DecodingTrust: A comprehensive assessment of trustworthiness in GPT models,

    B. Wanget al., “DecodingTrust: A comprehensive assessment of trustworthiness in GPT models,” inAdvances in Neural Information Processing Systems, 2023. [Online]. Available: https://proceedings.neurips.cc/paper files/paper/2023/ hash/63cb9921eecf51bfad27a99b2c53dd6d-Abstract-Da...

  23. [31]

    Vision-based learning for drones: A survey,

    J. Xiao, R. Zhang, Y . Zhang, and M. Feroskhan, “Vision-based learning for drones: A survey,”IEEE Transactions on Neural Networks and Learning Systems, vol. 36, no. 9, pp. 15 601–15 621, Sep. 2025

  24. [32]

    α 3-Bench: A unified benchmark of safety, robustness, and efficiency for LLM-based UA V agents over 6G networks,

    M. A. Ferrag, A. Lakas, and M. Debbah, “α 3-Bench: A unified benchmark of safety, robustness, and efficiency for LLM-based UA V agents over 6G networks,” 2026, arXiv:2601.03281. [Online]. Available: https://arxiv.org/abs/2601.03281

  25. [33]

    HUGE-Bench: A benchmark for high-level UA V vision- language-action tasks,

    J. Guoet al., “HUGE-Bench: A benchmark for high-level UA V vision- language-action tasks,” 2026, arXiv:2603.19822. [Online]. Available: https://arxiv.org/abs/2603.19822

  26. [34]

    A mathematical theory of communication,

    C. E. Shannon, “A mathematical theory of communication,”The Bell System Technical Journal, vol. 27, no. 3–4, pp. 379–423, 623–656,

  27. [35]

    SelectiveNet: A deep neural network with an integrated reject option,

    Y . Geifman and R. El-Yaniv, “SelectiveNet: A deep neural network with an integrated reject option,” inProceedings of the 36th International Conference on Machine Learning, ser. Proceedings of Machine Learning Research, vol. 97. PMLR, 2019, pp. 2151–2159. [Online]. Available: ...

  28. [36]

    On calibration of modern neural networks,

    C. Guo, G. Pleiss, Y . Sun, and K. Q. Weinberger, “On calibration of modern neural networks,” inProceedings of the 34th International Conference on Machine Learning, ser. Proceedings of Machine Learning Research, vol. 70. PMLR, 2017, pp. 1321–1330. [Online]. Available: https:/...

  29. [37]

    Qwen3-VL technical report,

    S. Baiet al., “Qwen3-VL technical report,” 2025, arXiv:2511.21631. [Online]. Available: https://arxiv.org/abs/2511.21631

  30. [38]

    Qwen3.5: Towards native multimodal agents,

    Qwen Team, “Qwen3.5: Towards native multimodal agents,” 2026, official model-family release page. [Online]. Available: https://qwen.ai/ blog?id=qwen3.5

  31. [39]

    SmolVLM2-2.2B-Instruct model card,

    Hugging Face, “SmolVLM2-2.2B-Instruct model card,” 2025, official model card. [Online]. Available: https://huggingface.co/HuggingFaceTB/ SmolVLM2-2.2B-Instruct

  32. [40]

    Phi-4-Mini technical report: Compact yet powerful multimodal language models via mixture-of-LoRAs,

    Microsoft, “Phi-4-Mini technical report: Compact yet powerful multimodal language models via mixture-of-LoRAs,” 2025, arXiv:2503.01743; includes Phi-4-Multimodal. [Online]. Available: https://arxiv.org/abs/2503.01743

  33. [41]

    Gemma 4 model overview,

    Google, “Gemma 4 model overview,” 2026, official model documenta- tion. [Online]. Available: https://ai.google.dev/gemma/docs/core

  34. [42]

    InternVL3.5: Advancing open-source multimodal models in versatility, reasoning, and efficiency,

    W. Wanget al., “InternVL3.5: Advancing open-source multimodal models in versatility, reasoning, and efficiency,” 2025, arXiv:2508.18265. [Online]. Available: https://arxiv.org/abs/2508.18265

  35. [43]

    GLM-4.6V overview,

    Z.AI, “GLM-4.6V overview,” 2026, official developer documentation for the GLM-4.6V series, including GLM-4.6V-Flash. [Online]. Available: https://docs.z.ai/guides/vlm/glm-4.6v

  36. [44]

    Qwen2.5-VL technical report,

    S. Baiet al., “Qwen2.5-VL technical report,” 2025, arXiv:2502.13923. [Online]. Available: https://arxiv.org/abs/2502.13923

  37. [45]

    MiniCPM-V 4.5: Cooking efficient MLLMs via architecture, data, and training recipe,

    T. Yuet al., “MiniCPM-V 4.5: Cooking efficient MLLMs via architecture, data, and training recipe,” 2025, arXiv:2509.18154. [Online]. Available: https://arxiv.org/abs/2509.18154

  38. [46]

    Aya Vision: Advancing the frontier of multilingual multimodality,

    S. Dashet al., “Aya Vision: Advancing the frontier of multilingual multimodality,” 2025, arXiv:2505.08751. [Online]. Available: https: //arxiv.org/abs/2505.08751

  39. [47]

    Building and better understanding vision-language models,

    H. Laurenc ¸on, L. Tronchon, M. Cord, and V . Sanh, “Building and better understanding vision-language models,” 2024, arXiv:2408.12637; includes Idefics3-8B. [Online]. Available: https://arxiv.org/abs/2408. 12637

  40. [48]

    LLaV A-NeXT: Improved reasoning, OCR, and world knowledge,

    LLaV A Team, “LLaV A-NeXT: Improved reasoning, OCR, and world knowledge,” 2024, official project release page. [Online]. Available: https://llava-vl.github.io/blog/2024-01-30-llava-next/

  41. [49]

    Concrete problems in AI safety,

    D. Amodei, C. Olah, J. Steinhardt, P. Christiano, J. Schulman, and D. Mane, “Concrete problems in AI safety,” 2016, arXiv:1606.06565. [Online]. Available: https://arxiv.org/abs/1606.06565

  42. [50]

    Advanced security frameworks for UA V and IoT: A deep learning approach,

    N. Quadar, A. Chehri, and B. Debaque, “Advanced security frameworks for UA V and IoT: A deep learning approach,”Internet of Things, vol. 32, p. 101594, 2025

  43. [270]

    2679–2713

    PMLR, 2025, pp. 2679–2713. [Online]. Available: https: //proceedings.mlr.press/v270/kim25c.html

  44. [1948]

    Available: https://people.math.harvard.edu/ ∼ctm/home/ text/others/shannon/entropy/entropy.pdf

    [Online]. Available: https://people.math.harvard.edu/ ∼ctm/home/ text/others/shannon/entropy/entropy.pdf

Pith tools

Reviewed July 30, 2026 · model on record in the stance chip above.