INDEX
Every paper Pith has read and judged, in one searchable index.
-
cs.CV arXiv submitted 2026-08-05Full-context model wins multi-turn image retrieval
Shengcao Cao +8 · “CoCo-IR: Contextual Composed Image Retrieval”
2608.05149 -
cs.CL arXiv submitted 2026-08-0550 procedural generators beat three rival reasoning datasets
Damien Sileo +2 · “Reasoning Core: Designing Broad Procedural Data for Completion-Supervised Reasoning Training”
2608.05148 -
cs.CV arXiv submitted 2026-08-05Few taps plus photos reconstruct an object's sound
Zisen Shao +3 · “Objects as Audio-Visual Modal Sound Fields”
2608.05145 -
cs.AI arXiv submitted 2026-08-05Agent runtime hits 78% on SWE-Bench Pro without retraining
Boxiu Li +25 · “Argus: A General-Purpose Agentic Runtime for Long-Horizon Reasoning”
2608.05144 -
cs.AI arXiv submitted 2026-08-0512% cross-repo code in the training mix lifts long-context skills
Indraneil Paul +3 · “OctoLong: Mid-Training On Cross-Repository Code Contexts Enhances Long-Context Modeling”
2608.05141 -
cs.CL arXiv submitted 2026-08-05Skill entropy predicts and fixes LLM skill-switching drops
Yinghui He +8 · “Toward Skill-Native LLMs: Skill Entropy for Benchmarking and Training Long-Horizon Reasoning”
2608.05139 -
eess.AS arXiv submitted 2026-08-05Fine-tuned 1B retriever beats 8B embedders and BM25 on Greek
Ayoub Kirouane +1 · “Teaching Nemotron Greek: Mining a Corpus, Adapting Retrieval, and Grounding Generation for Modern Greek across Specialist Domains”
2608.05138 -
cs.CV arXiv submitted 2026-08-05Adaptive modality mixing beats fixed fusion on five 3D benchmarks
Yue Zhang +4 · “SmartMage: Dynamic Modality Orchestration for 3D Scene Understanding”
2608.05137 -
cs.LG arXiv submitted 2026-08-05Gauge symmetry decides which optimizers keep the low-rank bias
Devender Singh · “The Loss Does Not See the Basis, but Adam Does”
2608.05136 -
cs.HC arXiv submitted 2026-08-05Turn vague research goals into vetted collaborator shortlists
Yingchaojie Feng +4 · “DeepConnect: A Visual Analytics System for Bridging Interdisciplinary Research Collaborations”
2608.05134 -
cs.CV arXiv submitted 2026-08-05Brain shape forecast beats the 'no-change' baseline by 2.29%
Hao Ding +3 · “Predicting Brain Morphometry with MT-GNN: Mesh Evolution in Continuous Time with Graph-Based Metric Tensor Embeddings”
2608.05132 -
cs.CV arXiv submitted 2026-08-05Balance-aware distillation lifts a 4B MLLM to 80%
Aniri +7 · “OPD-V: Visual On-Policy Self-Distillation with Modality Balance”
2608.05131 -
cs.CL arXiv submitted 2026-08-05Function calls beat intent-and-slot parsing for spoken AI
Yuezhang Peng +5 · “Spoken Function Calling: A New Perspective on Spoken Language Understanding for Large Audio Language Models”
2608.05126 -
cs.CL arXiv submitted 2026-08-05Chained fresh LLM calls gain 13.75 points on long-context tasks
Purbesh Mitra +1 · “Chained Recursive Language Models for Multi-Iteration Reasoning”
2608.05124 -
cs.CV arXiv submitted 2026-08-05Frozen ViT orientation peaks pick the best fine-tuning depth
Vaishnavi B Mohan +3 · “IRIS: A Visual Cortex-Inspired Framework for Analyzing Orientation Selectivity in Vision Transformers”
2608.05122 -
cs.SE arXiv submitted 2026-08-05AI coding tools' visual barriers fall into three recurring types
Sabrina Haque +1 · “Characterizing Visual Accessibility Issues in AI Developer Tools: An Empirical Study”
2608.05116 -
cs.CV arXiv submitted 2026-08-05Pose-only student beats larger models on classroom incidents
Paritosh Parmar +4 · “Robust and Efficient Motion Reasoning for Privacy-Aware Classroom Incident Recognition”
2608.05115 -
stat.ML arXiv submitted 2026-08-05SCMS converges to a stable ridge
Wanli Qiao · “Stable Density Ridges: Consistency and Convergence of Subspace Constrained Mean Shift”
2608.05112 -
cs.LG arXiv submitted 2026-08-05Same bonus amplifies, equalizes, or nulls—reward structure decides
Jai Malegaonkar +2 · “Reward Structure Shapes the Interaction Between Episodic Exploration and Neural Memory in Reinforcement Learning”
2608.05111 -
quant-ph arXiv submitted 2026-08-05One shared random bit outpowers shallow unitary quantum models
Arunava Majumder +3 · “Representational separation between unitary and channel quantum generative models via shared classical randomness at shallow depth”
2608.05110 -
eess.IV arXiv submitted 2026-08-05Single-shot laparoscopic depth hits 26 Hz without projector sync
Wayne Wonseok Rodgers +7 · “AI-based single-shot structured-light depth reconstruction for real-time laparoscopic surgical guidance”
2608.05109 -
cs.CR arXiv submitted 2026-08-05Memory-augmented attacker closes prompt-injection red-teaming gap
Yanting Wang +3 · “Agent Against Agent: An Agentic System for Automatic Prompt Injection Red Teaming”
2608.05108 -
cs.AI arXiv submitted 2026-08-05CoPlan lets care planners edit AI reasoning before the final plan
Hung Truong Thanh Nguyen +6 · “CoPlan: A Trustworthy Co-Intelligence Interface for Care Planning through Role-Based Contestable Argument Graphs”
2608.05107 -
cs.LG arXiv submitted 2026-08-05Stripped to 10% of its weights
Sajib Hossain +4 · “BnBERT-iPET: Sparse Few-Shot Language Modeling for Bengali via Lottery Ticket Pruning”
2608.05104 -
cs.LG arXiv submitted 2026-08-05One frozen weather prior runs filtering
Dibyajyoti Chakraborty +1 · “Multimodal Spatiotemporal Atmospheric Data Assimilation with Latent Flow-matching”
2608.05103 -
cs.AI arXiv submitted 2026-08-054B agent hits 55% on BrowseComp via answer-backtracked step credit
Yijun Lu +6 · “ABSeeker: Training Long-Horizon Search Agents via Answer-Backtracked Credit Assignment”
2608.05102 -
cs.CV arXiv submitted 2026-08-05Mask-free detector beats CT deepfake baselines by 9 AUC
Orazio Pontorno +3 · “HexMIL: Hierarchical Attention MIL for Ante-Hoc Explainable Detection of AI-Manipulated CT Volumes”
2608.05101 -
cs.CV arXiv submitted 2026-08-05A 4x sharper lesion signal in pretraining lifts frozen CT detection
Mahmut S. Gokmen +5 · “Lesion Detection in CT with Frozen Self-Distilled Features: SALT, a Spatially Adaptive Label-Guided Temperature”
2608.05100 -
cs.CL arXiv submitted 2026-08-05Most LLMs ignore the modal logic they are told to use
R\'eemi Andrieu +1 · “Same Formulas, Different Semantics: Do Language Models Follow Modal Logic Specifications?”
2608.05097 -
cs.AI arXiv submitted 2026-08-05Localized memory paths lift LLM agents using 7% of tokens
Xiawei Yue +4 · “Hierarchical Graph Memory for LLM Agents with Path-level Localization and Rewrite”
2608.05095 -
cs.LG arXiv submitted 2026-08-05Diagonally preconditioned Muon beats Muon at GPT-2 pretraining
Tongle Wu +3 · “MALT: Lightweight Curvature-Aware Muon via Diagonal Preconditioning”
2608.05088 -
cs.PF arXiv submitted 2026-08-05Falcon-512 fits C-V2X, fails 90% delivery above lightest traffic
Akid Abrar +6 · “Deployment Feasibility Analysis of Post-Quantum Digital Signatures in Safety-Critical C-V2X Communication for Urban Mobility Scenario”
2608.05087 -
cs.AI arXiv submitted 2026-08-05AI safety benchmarks reduce to three hidden traits
Joshua Fonseca Rivera (1) +4 · “Item Response Theory for AI Safety”
2608.05086 -
cs.LG arXiv submitted 2026-08-05Every fixed-lookahead experiment rule can be arbitrarily suboptimal
Ahmed Hassoon +1 · “Capability-Gated Planning: Cost-to-Goal Discovery and the Limits of Myopic Experiment Selection”
2608.05085 -
cs.LG arXiv submitted 2026-08-05Diffusion policies run 2.7x faster by learning when to stop denoising
Rohit Kumar Salla +2 · “Learning When to Stop: Prefix-Optimal Dynamic Diffusion Policies for Continuous Control”
2608.05084 -
cs.LG arXiv submitted 2026-08-05A learned rollout controller beats uniform sampling and fixed rules
Zheyuan Zhang +10 · “Optimizing What Policies Learn From: Recoverability-aware Rollout Intervention Learning”
2608.05080 -
cs.RO arXiv submitted 2026-08-05SpikingNav lifts corrupted ObjectNav success from 8.45% to 13.71%
Jiahong Zhang +7 · “SpikingNav: Robust Embodied Navigation with Spiking Neural Policies”
2608.05078 -
cs.CV arXiv submitted 2026-08-04Shared-backbone branches beat single-branch multimodal models
Yang Yang +7 · “ParVL: Parallel Scaling and Expandable Compute Allocation for Multimodal LLMs”
2608.04010 -
cs.CL arXiv submitted 2026-08-04Best LLM scores 75/100 forecasting anonymized social events
Zhenran Wang +3 · “SocietyBench: Forecasting Counterfactual Social-World Evolution”
2608.04009 -
cs.CL arXiv submitted 2026-08-04Six LLMs matched the bookies on live World Cup picks
Zhenran Wang +3 · “WorldCup Arena: Prospective, Leakage-Free Evaluation of Frontier LLMs on a Live Tournament”
2608.04008 -
cs.CL arXiv submitted 2026-08-04TurnSight beats prior best tool-use RL by 7.7%
Changle Qu +6 · “TurnSight: Turn-Level Hindsight Self-Distillation for Tool-Integrated Reasoning”
2608.04007 -
cs.HC arXiv submitted 2026-08-04Trustworthiness metrics lift expert agreement on LLMs
Adam Coscia +4 · “Calibrating Trustworthiness: Co-Designing Metrics and Visualizations for Evaluating LLMs in Education”
2608.04006 -
cs.CL arXiv submitted 2026-08-04Personal agents do improve from memory—unevenly
Shuhan Xue +8 · “PAST-Bench: Benchmarking the Foundations of Recursive Self-Improvement in Personal Agents”
2608.04003 -
math.NA arXiv submitted 2026-08-04ROM misfit cuts clean layered-inversion error 14-fold
Konstantinos Alexopoulos +1 · “Reduced-order modeling for electromagnetic inverse problems: a layered medium benchmark”
2608.03996 -
cs.CL arXiv submitted 2026-08-04ALiBi's distance bias silently blinds attention heads at long contexts
Christopher Schr\"oder +4 · “When Attention Goes Blind: Numerical Failure in ALiBi Positional Encodings”
2608.03994 -
cs.CV arXiv submitted 2026-08-04Text calibration lifts training-free segmentation by up to 5.9 pts
Wanli Ma +3 · “Perceptual Anchoring: Prototype-Guided Text Calibration for Training-free Open-Vocabulary Semantic Segmentation”
2608.03991 -
cs.LG arXiv submitted 2026-08-04Pathology-trained metrics predict which synthetic images help
Seyed Kahaki +3 · “Assessment of Conditional Diffusion Model for Synthetic Histopathology Image Generation”
2608.03990 -
cs.CL arXiv submitted 2026-08-04Browser workbench runs string algorithms 2,500x faster
Mirac Suzgun +3 · “string2string Studio: An Interactive, In-Browser Platform for String-to-String Algorithms”
2608.03984 -
cs.PL arXiv submitted 2026-08-04Best LLM recovers compiler-missed speedups on 83% of attempts
Hailong Jiang +6 · “Can Large Language Models Recover Semantic Optimization Opportunities That Compilers Miss?”
2608.03983 -
cs.CV arXiv submitted 2026-08-04Video agent tops Claude-4.5, GPT-5 by grounding before searching
Zhen Fang +19 · “Video-DeepResearch: Towards the Next-Generation Multimodal Deepresearch Agent”
2608.03979