An agentic mRAG framework uses GRPO-trained visual reranking and active rejection to verify retrieved candidate entities, achieving state-of-the-art on three KB-VQA benchmarks.
IEEE Transactions on Big Data (2025)
6 Pith papers cite this work. Polarity classification is still indexing.
years
2026 6representative citing papers
Cross-sectional patches with near-identical intensity but inconsistent masks are flagged as annotation noise, revealing systematic orientation-dependent bias in single-rater vascular CT labels.
Audit of KB-VQA benchmarks reveals systematic violations of answer derivability, question clarity, and visual disambiguation assumptions, with new repair and multi-entity augmentation protocols producing different model performance trends.
Semantic-level UI Element Injection distracts GUI agents by overlaying safety-aligned UI elements, achieving up to 4.4x higher attack success rates that transfer across models and create persistent attractors.
Training-free WSSS via decoupled mask proposals and offline semantic feature retrieval with boundary purification yields competitive benchmark performance without fine-tuning.
U-CESE integrates three CESE modules into a unified clip-based pipeline with DAKE keyframe extraction and ReCap captioning to support consistent multimodal event retrieval across video sources.
citing papers explorer
-
MMAgent-R$^2$: Learning to Rerank and Reject for Agentic mRAG
An agentic mRAG framework uses GRPO-trained visual reranking and active rejection to verify retrieved candidate entities, achieving state-of-the-art on three KB-VQA benchmarks.
-
Decoupled Single-Mask Annotation Noise Detection via Cross-Sectional Patch Self-Consistency
Cross-sectional patches with near-identical intensity but inconsistent masks are flagged as annotation noise, revealing systematic orientation-dependent bias in single-rater vascular CT labels.
-
Identifying and Resolving Pitfalls of Knowledge-Based VQA Benchmarks: Auditing, Repairing, and Augmenting
Audit of KB-VQA benchmarks reveals systematic violations of answer derivability, question clarity, and visual disambiguation assumptions, with new repair and multi-entity augmentation protocols producing different model performance trends.
-
Are GUI Agents Focused Enough? Automated Distraction via Semantic-level UI Element Injection
Semantic-level UI Element Injection distracts GUI agents by overlaying safety-aligned UI elements, achieving up to 4.4x higher attack success rates that transfer across models and create persistent attractors.
-
ModuSeg: Decoupling Object Discovery and Semantic Retrieval for Training-Free Weakly Supervised Segmentation
Training-free WSSS via decoupled mask proposals and offline semantic feature retrieval with boundary purification yields competitive benchmark performance without fine-tuning.
-
U-CESE: Unified Clip-based Event Search Engine for AI Challenge HCMC 2025
U-CESE integrates three CESE modules into a unified clip-based pipeline with DAKE keyframe extraction and ReCap captioning to support consistent multimodal event retrieval across video sources.