Index recall predicted before any index is built
Closed forms and synthetic twin corpora forecast PQ, IVF, HNSW, and FDE recall to within 0.03 on a million-document corpus.
Information Retrieval
Covers indexing, dictionaries, retrieval, content and analysis. Roughly includes material in ACM Subject Classes H.3.0, H.3.1, H.3.2, H.3.3, and H.3.4.
sort pith recommended most recent
Closed forms and synthetic twin corpora forecast PQ, IVF, HNSW, and FDE recall to within 0.03 on a million-document corpus.
GLIE stores 1KB per page, retains 79% of uncompressed retrieval quality, and traines in 3 GPU minutes.
· “Generative Late-Interaction Embeddings For Visual Document Retrieval”
RAG-Safety-Bench provides a controlled evaluation of LLM safety across non-RAG, oracle, on-topic, and random conditions, showing that…
· “RAG-Safety-Bench: Reliable Evaluation of Retrieval-Augmented LLM Safety”
24 non-lexical features reach 0.856 AUROC on ViDoRe, +0.207 over LLM, without reading document content.
· “Your Retriever Already Knows: Distribution-Shape QPP for RAG Retrieval Sufficiency”
A new framework retrieves helpful clients by their predicted utility rather than predefined similarity, validated on five real-world…
TimelyRAG reranks candidates using clause-level validity intervals, not just timestamps or keywords.
Top models reach 21% recall; evidence type inference is the key bottleneck.
· “ReGround: Grounding Reviewer Comments in Multimodal Evidence”
New method selects the best LLM per query while avoiding context loss and confusion, outperforming single-model baselines.
· “SWRouter: Similarity-Contractive Window Routing for Multi-Turn Large Language Model Conversations”
A directory-aware semantic storage and trace reuse system cuts token consumption dramatically while preserving high accuracy on structured…
· “VikingRAG: Accurate and Token-efficient Retrieval-augmented Generation over Structured Documents”
Joint optimization of pre-ranking and ranking fusion modules with dual-axis alignment and attribute-group regularization improves…
· “UniRec: Cross-stage Multi-Task Fusion with Preference Alignment for Cascaded Recommender Systems”
Glyph fuses three strategies to label 3.4M columns without reading data, reaching 99.8% steward acceptance.
A separate-context Critic and a structural commit rule produce per-claim audit traces with F1 0.84
· “GANDR: Claim Auditing for Verifiable Legal Answer Generation”
On two multi-hop QA benchmarks, LiteRAG matches or slightly beats stronger graph-RAG baselines in quality while cutting per-query token use…
· “LiteRAG: Cost-Efficient Graph-Based Retrieval-Augmented Generation”
A 30,841-trial study finds precision is flat once the answer path is present; recall is the lever that changes answers.
· “The Answer Path and the Grounding Instruction in LLM Question Answering over Knowledge Graphs”
One compact adapter per person: the model predicts its owner's text best — yet gains knowledge, not personal answers.
CHyD ensures every bracketed span matches the source, meeting safety-critical compliance needs.
· “Guaranteeing Faithful Evidence Extraction in Speculative Retrieval-Augmented Generation”
An audit of 67 real purchase conversations finds options in 77.6% of replies and zero recorded commitments.
· “Purchase Advice and Observable Buyer Responses in Real AI Conversations”
The central claim is that the comparison-level flip risk for quantized vector search can be decomposed into a boundary-mass term and a…
· “When Does Low-Bit Quantization Preserve the Decisions of Vector Search?”
A severity operating-point shift explains why rude or polite prompts change calibration far more than ranking, reconciling prior…
· “Should I Be Polite to My LLM Relevance Judge? Tone as a Severity Operating-Point Shift”
Extracting age, sex, and production modifiers from real-world datasets enables search without metadata standards.
· “Extracting Semantics from Cattle Reporting Categories for Data Interoperability and Findability”
New evaluation and training framework reveals and fixes the gap between accuracy and true evidence grounding.
· “Evaluating and Improving Evidence-Grounded Fact-Checking in LLMs via Multi-Round Evidence Ablation”
Diverse and non-diverse strategies produce expected score orderings across ten configurations, all constraints satisfied
· “FINALLY: A Dataset Recommender System for Recommender-Systems Research”
A private benchmark with 190 million documents, deep relevance labels, and a subcorpus that preserves rankings cuts evaluation cost by…
· “Q2D-Web: A Large-Scale Benchmark for Retrieval in Agentic RAG Systems”
A 44M-parameter model using atomic identifiers outranks a 220M-parameter model using naive or semantic IDs.
PDMR replaces single DocIDs with passage-grained targets, improving R@1 and MRR on NQ320K and MS MARCO over strong identifier-based…
Proof of concept for individualized knowledge simulation, but calibration and corpus size remain bottlenecks.
End-to-end model retains 83% of direct-scaling gain at 50× lower compute cost.
· “SequenceO1: End-to-End Ultra-Long (100K) Sequence Modeling in Recommendation with Low-Rank Caching”
Expanding ground truth to 5.3 combinations per query raises NDCG by 5–7 pp and reveals fine-tuning gains are partly artifact.
· “Tool Retrievers Are Underestimated: Annotation Expansion Reveals True Capability”
Hierarchical clustering from the ground up produces identifiers that are both unique and locally faithful, outperforming residual…
· “Exploring Bottom-Up Clustering for Creating Semantic IDs”
Multi-dimensional verification and stability analysis beat TSV by 3.3 points across three benchmarks.
· “Evidence-Aligned Entity Verification for Hallucination Detection in Retrieval-Augmented Generation”
A frozen VLM labels which documents lead to a correct answer, then trains a lightweight reranker that outperforms standard methods.
· “Bridging the Semantic-Utility Gap in Multimodal RAG via Generator-in-the-Loop Alignment”
Graph-to-hybrid distillation preserves teacher accuracy while cutting inference from hours to milliseconds on large databases.
· “Cassette: Case-to-Case Structural Distillation for Efficient Legal Case Retrieval”
It fires at median round 8 of 500 while holding F1 at 0.73 and finishes the full evaluation in 86 minutes.
Formal five-stage pipeline uses cheap filters before expensive AI, turning deduplication into data deepening
Exact values precomputed at ingestion stop small LLMs from inventing numbers; prose loads only on demand.
The field's real gap: connecting a person's need to the right tables, not just running analysis.
· “Open Tabular Insight Extraction: Where Do We Stand, and Where Should We Go?”
A domain-stripped computational skeleton surfaces same-problem papers in unrelated fields, lifting retrieval precision from 0.22 to 0.56.
Training-free spectral compression lifts nDCG@10 by up to 15% over k-means++ on ColBERTv2; its single-vector form beats MUVERA by ~70%.
Formal proof and empirical study show trajectory-level uncertainty must replace single-turn confidence metrics
MHR trains a full-width code first, then adds residual adaptors so all prefixes are directly searchable.
· “Matryoshka Hash Representations for Model-Aware Compact Semantic Retrieval”
Multi-task information co-evolves with sequence and feature representations at every layer, reducing latency by 30%
· “Task-Blind No MORE: Multi-Task Information Flow in Unified Ranking Backbones”
A compressed index shortlists pages, then a greedy token selector restores almost all of the uncompressed score, versus 93.93% for token…
· “Query-Aware Token Budgeting for Efficient Late-Interaction Visual Document Retrieval”
A two-stage LLM framework enriches intent diversity and aligns with business rules via GRPO, deployed at a major platform.
· “EAGER: Enrich-and-Align Generative Query Recommendation from Clicked Items in E-commerce Search”
A retrieve–localize–generate pipeline with self-reflective RL training beats all prior methods on four long-term memory QA benchmarks.
The prompt corpus and weights define the score's answer market; same rates yield 39.7% or 70.5%.
· “Measuring GEO Visibility: Prompt Corpora Define the Answer Market”
Audit of 1,920 queries across 12 countries and 4 languages finds language shift boosts local citations up to 13.5-fold.
· “Who Anchors AI Overviews in Health? Baidu, Google, and the Geography of Authority”
A three-level hierarchy extracted from evidence phrases, linking each node to specific source spans for auditing.
· “EviMap: Evidence-Grounded Hierarchical Topic Maps for Exploring Unlabeled Corpora”
Replacing flat paper similarities with a conference-specific taxonomy matches human session curation and survives live use.
· “TaxoConf: Taxonomy-Guided Automatic Conference Program Organization”
A communication-oriented RAG system reduces context by 25x, showing completeness is a distinct, human-centered objective.
Replacing dot-product scoring and tuning only bias and LayerNorm keeps short-window quality, with zero per-user storage.
· “Closing the Long-Short View Gap in Sequential Recommendation without Cached History”
A 100-person survey finds the data-quality layer only enables, never directly drives engagement.
· “Customer Relationship Intelligence: Integrating CRM and MDM for Enhanced Customer Engagement”
A new visualization traces the effect to expert-specific subspaces and reveals which experts specialize.
· “ExpertLens: Visualizing Embedding Spaces for Post-Hoc Explainability in MoE Enhanced Retrievers”
Learned components may rank and propose, but cannot promote a hypothesis to certified ground truth
Interactive RAG example selection in a visual analytics pipeline turns 69% accuracy into near-perfect extraction for MOF synthesis…
· “Visual Analysis of LLM-based Entity Resolution from Scientific Papers”
Benchmark of 10 LLM agents shows they check more but don’t change their minds under increasingly credible false claims
· “Evaluating Deep-Search Agents under Hierarchical Web Evidence Poisoning”
Multi‑agent system retrieves from multiple databases, re‑ranks by impact, and uses real peer‑review comments to revise drafts, achieving…
A method that injects knowledge only for nodes whose CF signals are unstable outperforms uniform KG fusion.
AtomCite verifies supplied page citations at about 93 percent accuracy and repairs them at 87-90 percent precision.
· “AtomCite: Verification and Correction of Supplied Page-Level Citations in Multi-Page Documents”
RAGMark enables fine-grained, reproducible benchmarking of RAG pipelines across retrievers, vector databases, reranking, compression, and…
· “RAGMark: A Comprehensive Framework for Benchmarking Retrieval-Augmented Generation Systems”
Re-ranking for fairness can look cheap once but adds per-query energy; training-time fixes vary by model and hardware.
· “What Price Fairness? Evaluating Energy - Fairness - Accuracy Trade-off in Recommender Systems”
Item-based collaborative filtering and graph neural networks together recover more alternative properties than either alone, especially at…
A fully black-box, zero-training method achieves ordering of texts by date and location using only API embeddings.
· “Recovering Temporal and Geographic Signals from Language Model Embeddings”
ROMCIR 2026 overview reveals focus on explainable AI, LLM reliability, and crowdsourced verification as key thrusts.
Confidence-gated routing captures half the gain at 40% cost and auto-disables where merging hurts.
· “Better Together: Complementary Query Rewriting Under a Strong RAG Baseline”
Notes lose 13 points, embedding migrations capture only half the gain, and store-only repair never recovers.
· “Does Your Agent's Memory Survive a Model Upgrade? A Controlled Study of Memory Portability”
Survey of 42 Master's students shows career motivation and ethical concerns differ by background.
A new pipeline extends multimodal RAG to return hand and special tools alongside procedures, reducing trips to the tool crib.
· “Beyond Maintenance Manual Multimodal RAG: Suggesting What Tool”
A lightweight convex optimization method improves retrieval effectiveness across benchmarks without retraining or index rebuilds.
· “Embedding Surgery: Localized Updates for Adaptive Ranking Correction in Dense Retrieval”
TeQHallu converts reference documents into relational databases and uses SQL queries to verify model responses, matching fine-tuned…
· “Leveraging Low-Level Symbolic Competences for Unsupervised Grounding in Hallucination Detection”
Adaptive expansion with sibling and child evidence reuse lifts answer F1 and retrieval recall on two benchmarks.
Image-grounded, business-scored expansion also cuts zero-result searches 17% on the same traffic.
· “SAM-D2Q: Aligning Multimodal Doc2Query with Search Demand and Conversion for E-commerce”
Fine-grained notes, semantic links, and field-level evolution preserve preference stages for evidence-path ranking.
· “AtomRec: Evolving Atomic Memory for Agentic Recommendation”
PTDG uses low-rank factorization to learn item-specific task dependency graphs and adaptive parameter masking to improve multi-task…
· “Personalized Task Dependency Graphs for Mitigating Signal Erosion in Multi-Task Recommendation”
IGPO boosts CTR 3.17% and cuts bad cases 39% in a commercial AI search with a changing catalog.
· “Inventory-Grounded Policy-Level Optimization for Training-Free AI Search”
VizIt wraps single-cell, spatial, and genetic data in one navigable interface, built on Parkinson's brain datasets.
· “VizIt: A multi-view framework for exploring single-cell, spatial, and genetic data online”
Scoring passages by how they fit the whole set, not just the query, rescues the bridge facts multi-hop questions rely on.
· “CAGE: Coherence-Aware Graph Encoding for Retrieval-Augmented Generation”
A two-stage VLM framework with dual alignment outperforms all baselines on four datasets and three downstream architectures.
The paper establishes that treating graph construction as a differentiable, retrieval-augmented learning task, combined with explicit…
· “MURAL: Multimodal Uncertainty-aware Recommendation via Adaptive edge Learning”
BioSync produces a continuous digital biomarker from HRV, EEG, actigraphy, and speech that correlates with severity and maintains…
· “BioSync: Transformer-Based Cross-Modal Fusion for a Multimodal Physiological Digital Biomarker”
A novel evaluation scheme using the Hullermeier-Rifqi Index with normalized edit distance is proposed to assess phonetic encoding…
· “Evaluation of Phonetic Encoding Algorithms on Transcription Datasets”