A new native-runtime benchmark reveals that current frontier AI agents succeed on at most 62 percent of realistic long-horizon CLI tasks.
Deepseek-v3.2: Pushing the frontier of open large language models
7 Pith papers cite this work. Polarity classification is still indexing.
citation-role summary
citation-polarity summary
years
2026 7representative citing papers
AutoResearchBench is a new benchmark showing top AI agents achieve under 10% success on complex scientific literature discovery tasks that demand deep comprehension and open-ended search.
Region4Web shows that shifting web agent observation to functional regions instead of element-level granularity produces shorter, more effective state representations and raises task success on WebArena across multiple LLMs and agent methods.
QuasiMoTTo uses quasi-Monte Carlo to produce correlated yet marginally correct samples from language models, matching i.i.d. pass@k with 25-47% fewer samples on reasoning benchmarks and 50% fewer RL training steps.
Frontier LRMs match human game-learning behavior and predict fMRI signals an order of magnitude better than RL or Bayesian agents because of their in-context game-state representations.
LERA is a retrieve-then-generate auction system that refines ad candidate ranking with LLM logits and applies a threshold-aware critical-value payment rule to maintain truthfulness in chatbot ad insertion.
JoyAI-LLM Flash delivers a 48B MoE LLM with 2.7B active parameters per token via FiberPO RL and dense multi-token prediction, released with checkpoints on Hugging Face.
citing papers explorer
-
WildClawBench: A Benchmark for Real-World, Long-Horizon Agent Evaluation
A new native-runtime benchmark reveals that current frontier AI agents succeed on at most 62 percent of realistic long-horizon CLI tasks.
-
AutoResearchBench: Benchmarking AI Agents on Complex Scientific Literature Discovery
AutoResearchBench is a new benchmark showing top AI agents achieve under 10% success on complex scientific literature discovery tasks that demand deep comprehension and open-ended search.
-
Region4Web: Rethinking Observation Space Granularity for Web Agents
Region4Web shows that shifting web agent observation to functional regions instead of element-level granularity produces shorter, more effective state representations and raises task success on WebArena across multiple LLMs and agent methods.
-
QuasiMoTTo: Quasi-Monte Carlo Test-Time Scaling
QuasiMoTTo uses quasi-Monte Carlo to produce correlated yet marginally correct samples from language models, matching i.i.d. pass@k with 25-47% fewer samples on reasoning benchmarks and 50% fewer RL training steps.
-
Reason to Play: Behavioral and Brain Alignment Between Frontier LRMs and Human Game Learners
Frontier LRMs match human game-learning behavior and predict fMRI signals an order of magnitude better than RL or Bayesian agents because of their in-context game-state representations.
-
LERA: LLM-Enhanced RAG for Ad Auction in Generative Chatbots
LERA is a retrieve-then-generate auction system that refines ad candidate ranking with LLM logits and applies a threshold-aware critical-value payment rule to maintain truthfulness in chatbot ad insertion.
-
JoyAI-LLM Flash: Advancing Mid-Scale LLMs with Token Efficiency
JoyAI-LLM Flash delivers a 48B MoE LLM with 2.7B active parameters per token via FiberPO RL and dense multi-token prediction, released with checkpoints on Hugging Face.