HAKARI-Bench reconstructs 35 benchmarks into 551 tasks across 43 languages, reproducing full MTEB, MMTEB, and BEIR rankings with Spearman correlation above 0.97 while supporting efficiency variant comparisons.
hub
C-Pack: Packed Resources For General Chinese Embeddings , booktitle =
15 Pith papers cite this work, alongside 3 external citations. Polarity classification is still indexing.
hub tools
citation-role summary
citation-polarity summary
years
2026 15roles
background 2polarities
background 2representative citing papers
Introduces P-CHR AUC and CRR metrics to demonstrate that semantic caching model selection is limited by calibration quality rather than ranking performance.
NL2SQLBench is a new modular benchmarking framework that evaluates LLM NL2SQL methods across three core modules on existing datasets, exposing large accuracy gaps and computational inefficiency.
PlanRAG models natural language exploratory reasoning problems as logical query trees, optimizes them via dynamic programming with a multi-dimensional cost model, and executes iterative retrieval-generation over the trees to outperform prior RAG methods on a new dataset.
SKIM is an adaptive multi-resolution soft-token framework that compresses procedural skills while aiming to preserve logical dependencies and task performance better than prior compression methods.
Occupational AI exposure is bipolar — physical execution and planning/design are opposite poles around a low-contrast middle — and the pole identities inverted between the 2013 and LLM eras.
StructuredSemanticSearch uses table discovery operators and orientation-aware integration on model-card tables to improve evidence coverage and diversity in model recommendation queries over a semantic baseline.
ClusterRAG applies density-based clustering to user profiles for collaborative retrieval in personalized RAG and reports best performance on LaMP tasks by combining target and similar-user profiles.
PETRA is a curated 1.36M-chunk petroleum-engineering retrieval dataset and pipeline that raises in-domain nDCG from 0.703 to 0.763 via score fusion and delivers 44% relative gain on an Earth Science benchmark through reranker adaptation on synthetic supervision.
D2D adaptively prioritizes informative attribute queries and times recommendations in conversational search, yielding 22-30% higher target accuracy and shorter conversations than baselines in simulations.
Credence replaces Jaccard-F1 with Semantic-F1 for claim decomposition quality and proves convergence properties for rule-based and LLM-based repair under stated assumptions, reporting +15-32pp gains on three domain benchmarks.
OMAGR decomposes queries into ontology-aligned anchors for parallel multi-dimensional graph retrieval, outperforming baselines on Context Precision and Faithfulness in the new TrafficLaw-QA dataset of 200 questions.
Full-horizon planning with on-demand replanning achieves accuracy parity with single-step planning in tool-calling agents for knowledge base and multi-hop question answering while consuming 2-3 times fewer tokens.
fMRI during naturalistic movies shows context (action target vs passive) remaps object representational geometry via double dissociation in affordance vs semantic dimensions within selective networks.
SciAtlas builds a large-scale multi-disciplinary academic knowledge graph and a neuro-symbolic retrieval system to support automated scientific research tasks such as literature review and idea positioning.
citing papers explorer
-
HAKARI-Bench: A Lightweight Benchmark for Comparing Retrieval Architectures and Efficiency Settings under Unified Conditions
HAKARI-Bench reconstructs 35 benchmarks into 551 tasks across 43 languages, reproducing full MTEB, MMTEB, and BEIR rankings with Spearman correlation above 0.97 while supporting efficiency variant comparisons.
-
Closing the Calibration Gap in Semantic Caching
Introduces P-CHR AUC and CRR metrics to demonstrate that semantic caching model selection is limited by calibration quality rather than ranking performance.
-
NL2SQLBench: A Modular Benchmarking Framework for LLM-Enabled NL2SQL Solutions
NL2SQLBench is a new modular benchmarking framework that evaluates LLM NL2SQL methods across three core modules on existing datasets, exposing large accuracy gaps and computational inefficiency.
-
When RAG Meets Query Planning: Logical Query Trees for Resolving Exploratory Reasoning Problems
PlanRAG models natural language exploratory reasoning problems as logical query trees, optimizes them via dynamic programming with a multi-dimensional cost model, and executes iterative retrieval-generation over the trees to outperform prior RAG methods on a new dataset.
-
Adaptive Multi-Resolution Procedural Knowledge Compression for Large Language Models
SKIM is an adaptive multi-resolution soft-token framework that compresses procedural skills while aiming to preserve logical dependencies and task performance better than prior compression methods.
-
The atomic structure of work: a micro-action instrument reveals two-pole AI occupational exposure and its decade-scale polar inversion
Occupational AI exposure is bipolar — physical execution and planning/design are opposite poles around a low-contrast middle — and the pole identities inverted between the 2013 and LLM eras.
-
Diversed Model Discovery via Structured Table Discovery
StructuredSemanticSearch uses table discovery operators and orientation-aware integration on model-card tables to improve evidence coverage and diversity in model recommendation queries over a semantic baseline.
-
ClusterRAG: Cluster-Based Collaborative Filtering for Personalized Retrieval-Augmented Generation
ClusterRAG applies density-based clustering to user profiles for collaborative retrieval in personalized RAG and reports best performance on LaMP tasks by combining target and similar-user profiles.
-
PETRA: Transforming Web Text for Petroleum-Engineering Domain Adaptation
PETRA is a curated 1.36M-chunk petroleum-engineering retrieval dataset and pipeline that raises in-domain nDCG from 0.703 to 0.763 via score fusion and delivers 44% relative gain on an Earth Science benchmark through reranker adaptation on synthetic supervision.
-
Dialogue to Discovery: Attribute-Aware Preference Elicitation for Conversational Product Search Assistants
D2D adaptively prioritizes informative attribute queries and times recommendations in conversational search, yielding 22-30% higher target accuracy and shorter conversations than baselines in simulations.
-
CREDENCE: Claim Reduction for Decomposition & Enhanced Credibility -- Semantic Metrics and Convergence Analysis
Credence replaces Jaccard-F1 with Semantic-F1 for claim decomposition quality and proves convergence properties for rule-based and LLM-based repair under stated assumptions, reporting +15-32pp gains on three domain benchmarks.
-
An Ontology-Guided Multi-Anchor Graph Retrieval Framework for Traffic Legal Liability Determination
OMAGR decomposes queries into ontology-aligned anchors for parallel multi-dimensional graph retrieval, outperforming baselines on Context Precision and Faithfulness in the new TrafficLaw-QA dataset of 200 questions.
-
Do Agents Need to Plan Step-by-Step? Rethinking Planning Horizon in Data-Centric Tool Calling
Full-horizon planning with on-demand replanning achieves accuracy parity with single-step planning in tool-calling agents for knowledge base and multi-hop question answering while consuming 2-3 times fewer tokens.
-
Contextual Role Modulates Object Representational Geometry in the Human Brain
fMRI during naturalistic movies shows context (action target vs passive) remaps object representational geometry via double dissociation in affordance vs semantic dimensions within selective networks.
-
SciAtlas: A Large-Scale Knowledge Graph for Automated Scientific Research
SciAtlas builds a large-scale multi-disciplinary academic knowledge graph and a neuro-symbolic retrieval system to support automated scientific research tasks such as literature review and idea positioning.