DataComp-VLM benchmark shows instruction-heavy data mixing outperforms filtering for VLM training, with DCVLM-Baseline achieving 63.6% on 33 tasks for 8B models (+5.4pp over FineVision).
Title resolution pending
11 Pith papers cite this work, alongside 14 external citations. Polarity classification is still indexing.
years
2026 11representative citing papers
HAKARI-Bench reconstructs 35 benchmarks into 551 tasks across 43 languages, reproducing full MTEB, MMTEB, and BEIR rankings with Spearman correlation above 0.97 while supporting efficiency variant comparisons.
DyNACO uses periodic observation of pheromone and incumbent solution for dynamic neural guidance in large-scale ACO, improving TSP and CVRP solvers with low overhead.
DiffMath is a symbol- and graph-aware latent diffusion framework that converts LaTeX structure into RelAST triplets and uses MathVAE plus MathDiT to generate structurally coherent handwritten expressions without explicit positional supervision.
IPO-Mine releases a toolkit and large multimodal dataset for structured analysis of IPO filings and shows state-of-the-art models diverge from human judgments on chart quality and misleadingness.
MVR-cache raises semantic cache hit rates by up to 37% on benchmarks using learned prompt segmentation, multi-vector retrieval, and RL training while preserving correctness.
A survey that categorizes RIR benchmarks by domain and modality, proposes a taxonomy for integrating reasoning into retrieval pipelines, and outlines key challenges.
SelPE introduces a selection-guided progressive evolution method for private structured text synthesis that decouples abstraction from schema realization and claims better validity and utility under tight DP budgets in low-data settings.
Reinforcement learning paired with a geometry-aware Polygons Transformer achieves area utilization competitive with the Sparrow heuristic solver for 2D irregular nesting.
The paper releases two adversarial malware datasets (44k family-labelled, 33k type-labelled) with high evasion rates and demonstrates that 0.5% poisoning injection raises evasion from 26.1% to 92.8%.
Cybersecurity's scale, adversaries, labeling issues, and operational demands make it the superior test-case for general AI progress over NLP or computer vision.
citing papers explorer
-
DataComp-VLM: Improved Open Datasets for Vision-Language Models
DataComp-VLM benchmark shows instruction-heavy data mixing outperforms filtering for VLM training, with DCVLM-Baseline achieving 63.6% on 33 tasks for 8B models (+5.4pp over FineVision).
-
HAKARI-Bench: A Lightweight Benchmark for Comparing Retrieval Architectures and Efficiency Settings under Unified Conditions
HAKARI-Bench reconstructs 35 benchmarks into 551 tasks across 43 languages, reproducing full MTEB, MMTEB, and BEIR rankings with Spearman correlation above 0.97 while supporting efficiency variant comparisons.
-
Beyond Static Priors: Dynamic Neural Guidance for Large-Scale Ant Colony Optimization
DyNACO uses periodic observation of pheromone and incumbent solution for dynamic neural guidance in large-scale ACO, improving TSP and CVRP solvers with low overhead.
-
DiffMath: Symbol- and Graph-Aware Latent Diffusion Transformer for Handwritten Mathematical Expression Generation
DiffMath is a symbol- and graph-aware latent diffusion framework that converts LaTeX structure into RelAST triplets and uses MathVAE plus MathDiT to generate structurally coherent handwritten expressions without explicit positional supervision.
-
IPO-Mine: A Toolkit and Dataset for Section-Structured Analysis of Long, Multimodal IPO Documents
IPO-Mine releases a toolkit and large multimodal dataset for structured analysis of IPO filings and shows state-of-the-art models diverge from human judgments on chart quality and misleadingness.
-
MVR-cache: Optimizing Semantic Caching via Multi-Vector Retrieval and Learned Prompt Segmentation
MVR-cache raises semantic cache hit rates by up to 37% on benchmarks using learned prompt segmentation, multi-vector retrieval, and RL training while preserving correctness.
-
A Survey of Reasoning-Intensive Retrieval: Progress and Challenges
A survey that categorizes RIR benchmarks by domain and modality, proposes a taxonomy for integrating reasoning into retrieval pipelines, and outlines key challenges.
-
SelPE: Progressive Selection for Private Structured Text Synthesis
SelPE introduces a selection-guided progressive evolution method for private structured text synthesis that decouples abstraction from schema realization and claims better validity and utility under tight DP budgets in low-data settings.
-
Geometry-Aware Reinforcement Learning for 2D Irregular Nesting
Reinforcement learning paired with a geometry-aware Polygons Transformer achieves area utilization competitive with the Sparrow heuristic solver for 2D irregular nesting.
-
Building an Adversarial Malware Dataset by Family and Type: Generation, Evasion, and Poisoning Evaluation
The paper releases two adversarial malware datasets (44k family-labelled, 33k type-labelled) with high evasion rates and demonstrates that 0.5% poisoning injection raises evasion from 26.1% to 92.8%.
-
Cybersecurity is the True Frontier for Generative AI Success or Failure
Cybersecurity's scale, adversaries, labeling issues, and operational demands make it the superior test-case for general AI progress over NLP or computer vision.