Linear and spherical interpolation of ANCE and QRACDR parameters yields a single dense retriever that recovers ad-hoc effectiveness while retaining conversational skill and improving zero-shot generalization.
InProceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 3: Industry Track)
12 Pith papers cite this work. Polarity classification is still indexing.
citation-role summary
citation-polarity summary
years
2026 12roles
background 2polarities
background 2representative citing papers
Training-free mechanistic interpretability locates lexical-encoding attention heads in ViT and shows that targeted interventions on them improve robustness to typographic attacks in CLIP and downstream LVLMs.
Centralized matching mechanisms outperform free negotiation in stability and efficiency with LLM agents, who also report preferences truthfully more often than humans, though not always in line with strategy-proofness predictions.
IPO-Mine releases a toolkit and large multimodal dataset for structured analysis of IPO filings and shows state-of-the-art models diverge from human judgments on chart quality and misleadingness.
H²MT uses offline semantic hierarchy construction, bottom-up memory aggregation, and coarse-to-fine query routing to achieve competitive QA quality with lower memory and latency than flat or retrieval baselines on LongBench tasks.
ReflectCAP distills model-specific hallucination and oversight patterns into Structured Reflection Notes that steer LVLMs toward more factual and complete image captions, reaching the Pareto frontier on factuality-coverage trade-offs.
AudioGuard combines waveform-level and policy-based detection to improve safety in audio AI systems, outperforming baselines on a new benchmark spanning languages, voices, and threat types.
MAESTRO adds a shared preference memory plus GUI-adaptation and workflow-navigation mechanisms to conversational agents with GUIs and tests them in a 33-person movie-booking study.
An agentic multi-source grounding system for marketplace query intent achieves 90.7% accuracy on long-tail queries at DoorDash by combining catalog grounding, web search, and deterministic disambiguation, outperforming baselines by up to 13pp.
A multi-agent framework decomposes multimodal empathetic response generation into structured reasoning steps and uses global reflection to reduce emotional biases, outperforming prior methods on IEMOCAP and MELD benchmarks.
AgentStop uses execution signals to early-terminate failing local LLM agent trajectories, cutting energy use 15-20% with minimal utility loss.
Reproducing GAR on BRIGHT shows it boosts reasoning-intensive retrieval effectiveness with low overhead when the reranker's signal quality is strong.
citing papers explorer
-
Improving Ad-hoc Search Effectiveness for Conversational Information Retrieval via Model Merging
Linear and spherical interpolation of ANCE and QRACDR parameters yields a single dense retriever that recovers ad-hoc effectiveness while retaining conversational skill and improving zero-shot generalization.
-
Towards Robustness against Typographic Attack with Training-free Concept Localization
Training-free mechanistic interpretability locates lexical-encoding attention heads in ViT and shows that targeted interventions on them improve robustness to typographic attacks in CLIP and downstream LVLMs.
-
Do Matching Mechanisms Work with LLM Agents?
Centralized matching mechanisms outperform free negotiation in stability and efficiency with LLM agents, who also report preferences truthfully more often than humans, though not always in line with strategy-proofness predictions.
-
IPO-Mine: A Toolkit and Dataset for Section-Structured Analysis of Long, Multimodal IPO Documents
IPO-Mine releases a toolkit and large multimodal dataset for structured analysis of IPO filings and shows state-of-the-art models diverge from human judgments on chart quality and misleadingness.
-
H$^{2}$MT: Semantic Hierarchy-Aware Hierarchical Memory Transformer
H²MT uses offline semantic hierarchy construction, bottom-up memory aggregation, and coarse-to-fine query routing to achieve competitive QA quality with lower memory and latency than flat or retrieval baselines on LongBench tasks.
-
ReflectCAP: Detailed Image Captioning with Reflective Memory
ReflectCAP distills model-specific hallucination and oversight patterns into Structured Reflection Notes that steer LVLMs toward more factual and complete image captions, reaching the Pareto frontier on factuality-coverage trade-offs.
-
AudioGuard: Toward Comprehensive Audio Safety Protection Across Diverse Threat Models
AudioGuard combines waveform-level and policy-based detection to improve safety in audio AI systems, outperforming baselines on a new benchmark spanning languages, voices, and threat types.
-
MAESTRO: Adapting GUIs and Guiding Navigation with User Preferences in Conversational Agents with GUIs
MAESTRO adds a shared preference memory plus GUI-adaptation and workflow-navigation mechanisms to conversational agents with GUIs and tests them in a 33-person movie-booking study.
-
Agentic Multi-Source Grounding for Enhanced Query Intent Understanding: A DoorDash Case Study
An agentic multi-source grounding system for marketplace query intent achieves 90.7% accuracy on long-tail queries at DoorDash by combining catalog grounding, web search, and deterministic disambiguation, outperforming baselines by up to 13pp.
-
A Multi-Agent Framework with Structured Reasoning and Reflective Refinement for Multimodal Empathetic Response Generation
A multi-agent framework decomposes multimodal empathetic response generation into structured reasoning steps and uses global reflection to reduce emotional biases, outperforming prior methods on IEMOCAP and MELD benchmarks.
-
AgentStop: Terminating Local AI Agents Early to Save Energy in Consumer Devices
AgentStop uses execution signals to early-terminate failing local LLM agent trajectories, cutting energy use 15-20% with minimal utility loss.
-
Reproducing Adaptive Reranking for Reasoning-Intensive IR
Reproducing GAR on BRIGHT shows it boosts reasoning-intensive retrieval effectiveness with low overhead when the reranker's signal quality is strong.