Thinking-mode VLMs collapse answer-token entropy, but thinking-chain entropy and length serve as robust, zero-cost hallucination predictors.
Title resolution pending
13 Pith papers cite this work. Polarity classification is still indexing.
citation-role summary
citation-polarity summary
years
2026 13roles
background 2polarities
background 2representative citing papers
HyphaeDB introduces an agent-native memory system using HNSW topology for gossip-based knowledge propagation, enabling emergent behaviors in multi-agent AI.
Audited olympiad corpus and Physics-R1 recipe improve 8B VLM by up to 18 points on held-out physics problems while exposing contamination in prior evals.
ProactBench measures LLM conversational proactivity in three phases using 198 multi-agent dialogues and finds recovery behavior hard to predict from existing benchmarks.
MULTITEXTEDIT benchmark reveals that all tested text-in-image editing models show pronounced degradation on non-English languages, especially Hebrew and Arabic, mainly in text accuracy and script fidelity.
FourTune matches full-precision LoRA quality on diffusion post-training via native W4A4G4 with a frozen SVD stabilizer, block-wise quant, and fused kernels, cutting memory 2.25× and speeding training 2.27× on FLUX.1-dev.
A single set of trained weights can be decoded either with MLA's compact latent cache or with a GQA-style expanded cache, letting the runtime match GPU rooflines without retraining.
Pre-trained MoE models exhibit deep-layer routing collapse for low-resource languages like Hebrew, largely corrected by continual pre-training on balanced bilingual data, with consistent patterns observed in Japanese.
RGAO combines retrieval-based complexity assessment with a formal budget algebra to enable dynamic topology selection in multi-agent code generation with provable conservation.
A frozen average of the last two cycles matches or exceeds eight shape-learning alternatives on 97 GIFT-Eval configurations for periodic time series forecasting.
citing papers explorer
-
When Thinking Hurts: Epistemic Signals in the Reasoning Chains of Visual Language Models
Thinking-mode VLMs collapse answer-token entropy, but thinking-chain entropy and length serve as robust, zero-cost hallucination predictors.
-
HyphaeDB: A Living Knowledge Topology for Agent-First Memory
HyphaeDB introduces an agent-native memory system using HNSW topology for gossip-based knowledge propagation, enabling emergent behaviors in multi-agent AI.
-
Physics-R1: An Audited Olympiad Corpus and Recipe for Visual Physics Reasoning
Audited olympiad corpus and Physics-R1 recipe improve 8B VLM by up to 18 points on held-out physics problems while exposing contamination in prior evals.
-
ProactBench: Beyond What The User Asked For
ProactBench measures LLM conversational proactivity in three phases using 198 multi-agent dialogues and finds recovery behavior hard to predict from existing benchmarks.
-
MULTITEXTEDIT: Benchmarking Cross-Lingual Degradation in Text-in-Image Editing
MULTITEXTEDIT benchmark reveals that all tested text-in-image editing models show pronounced degradation on non-English languages, especially Hebrew and Arabic, mainly in text accuracy and script fidelity.
-
FourTune: Towards Fully 4-Bit Efficient Post-Training for Diffusion Models
FourTune matches full-precision LoRA quality on diffusion post-training via native W4A4G4 with a frozen SVD stabilizer, block-wise quant, and fused kernels, cutting memory 2.25× and speeding training 2.27× on FLUX.1-dev.
-
GQLA: Group-Query Latent Attention for Hardware-Adaptive Large Language Model Decoding
A single set of trained weights can be decoded either with MLA's compact latent cache or with a GQA-style expanded cache, letting the runtime match GPU rooflines without retraining.
-
Mixture of Experts for Low-Resource LLMs
Pre-trained MoE models exhibit deep-layer routing collapse for low-resource languages like Hebrew, largely corrected by continual pre-training on balanced bilingual data, with consistent patterns observed in Japanese.
-
Retrieval-Conditioned Topology Selection with Provable Budget Conservation for Multi-Agent Code Generation
RGAO combines retrieval-based complexity assessment with a formal budget algebra to enable dynamic topology selection in multi-agent code generation with provable conservation.
-
Don't Learn the Shape: Forecasting Periodic Time Series by Rank-1 Decomposition
A frozen average of the last two cycles matches or exceeds eight shape-learning alternatives on 97 GIFT-Eval configurations for periodic time series forecasting.
- Know When to Stop: Segment-Level Credit Assignment for Reducing Overthinking
- Injection-Execution Dissociation: A Mechanistic Evaluation of Persistent Memory Attacks and Defenses in Stateful LLM Agents
- Training Non-Differentiable Networks via Optimal Transport