Agents finish only 21% of multi-hour desktop workflows
A 108-task suite shows they lose constraints and miss mid-task state, not basic clicks.
· “OSWorld 2.0: Benchmarking Computer Use Agents on Long-Horizon Real-World Tasks”
15,515 reviewed in the last 24 hours · 223 in review now 270,564 on the record · 9.3% of the arXiv record read Every review is machine-made and published openly.
Trending What readers are reading and passing along: page reads, share clicks, X engagement, and reader discussion over the last 7 days.
sort impact radar most recent
A 108-task suite shows they lose constraints and miss mid-task state, not basic clicks.
· “OSWorld 2.0: Benchmarking Computer Use Agents on Long-Horizon Real-World Tasks”
Thinking mode and cache tricks let Gemma 4 match far larger open systems on STEM and multimodal tasks
Latent actions, dual transformers, and dream forcing close the gaps that break world-action models
· “ABot-M0.5: Unified Mobility-and-Manipulation World Action Model”
MOPD distills domain teachers onto the student's own data, outperforming mix and cascade baselines while allowing independent teacher develo
· “MOPD: Multi-Teacher On-Policy Distillation for Capability Integration in LLM Post-Training”
Benchmark uses verified VM transitions to separate rule-following from copied persistent state.
· “ScratchWorld: Evaluating If World Models Compute Executable Consequences”
Unified sim-and-real tests of 30 models show single-digit success and a large gap to human teleoperation
A sharp projection estimate proves the 2-copy threshold matches the 1-copy threshold in every dimension.
Isolating question-conditioned corrections via reference-only teachers and PMI improves four models on two datasets without losing natural r
· “Purified OPSD: On-Policy Self-Distillation Without Losing How to Think”
Six research-level case studies show parallel workers can build arguments hundreds of facts deep
· “Danus: Orchestrating Mathematical Reasoning Agents with Fact-Graph Memory”
Semi-autoregressive drafts plus load-aware length scheduling shift the live throughput–speed Pareto frontier.
· “DSpark: Confidence-Scheduled Speculative Decoding with Semi-Autoregressive Generation”
Certified interval arithmetic proves a noncircular real-analytic domain with a boundary-constant Neumann eigenfunction.
· “A computer-assisted counterexample to the planar Pompeiu and Schiffer conjectures”
It proves Shannon-entropy Bregman iterates converge to critical points on nonconvex problems.
· “On the Iterate Convergence of Bregman Projected Gradient Method”
A new bootstrap turns sub-2^n list-coloring algorithms into faster k-coloring algorithms, closing a long-standing gap.
· “k-Coloring is Faster than Computing the Chromatic Number”
For binary length-8191 codes the gap already reaches 8, and it grows as n^(1/3), so no constant bound can hold.
The Gaussian guess is only the leading term; higher-order corrections are explicit Bernoulli polynomials of the fractional displacement of t
One geometric rule yields sharp bounds on diffusivity, viscosity and hydrodynamics in boosted frames
It predicted finite time-frequency shifts are independent; this paper constructs a genuine 12-term counterexample.
· “Linear dependence of time-frequency shifts of a Schwartz function”
Self-generated hindsight skills give token-level credit, lifting success up to 39 points and cutting data needs.
· “SEED: Self-Evolving On-Policy Distillation for Agentic Reinforcement Learning”
Chern conjecture holds for M^4 in S^5 when both scalar curvature and Gauss-Kronecker curvature are constant, using weighted 3-forms and χ(M)
At least d^{2−o(1)} copies are needed for constant-precision spectrum, entropy, and rank estimates.
New Jacobian lens reads, swaps, and ablates the representations that drive silent reasoning and strategic deliberation.
· “Verbalizable Representations Form a Global Workspace in Language Models”
Hand-written functional tests cut judge-verifier disagreement from 32% to 1.4% on matched audits.
· “DeepSWE: Measuring Frontier Coding Agents on Original, Long-Horizon Engineering Tasks”
A new algorithm isolates the soft+Glauber hidden region that creates the two-loop splitting amplitude's kinematic dependence.
Long 45K-token trajectories and domain teachers, not more parameters, drive the reported gains on search and science.
· “Scaling the Horizon, Not the Parameters: Reaching Trillion-Parameter Performance with a 35B Agent”
Frame-dressed observables make correlators and time evolution relative to the frame you choose.
· “Relational path integral, effective actions and quantum frame covariance in gravity”
DOPD assigns each token to teacher or student based on advantage gaps, separating capability transfer from information asymmetry.
Transfer the before–after policy log-ratio as a dense reward; skip sparse RL on the target.
· “Weak-to-Strong Generalization via Direct On-Policy Distillation”
A reduction from MAX CUT answers the 2001 open question and settles the complexity of Kemeny aggregation for every fixed voter count.
Across 134 day-long tasks and 38,000 hours, average performance tracks interaction time with R² near 0.998
· “EdgeBench: Unveiling Scaling Laws of Learning from Real-World Environments”
The Parisi measure for every β>1 has a smooth density and a single endpoint atom—full replica symmetry breaking follows.
· “Full replica symmetry breaking in the Sherrington-Kirkpatrick model”
A unified framework covers general separable kernels and mirror flows, closing a decades-old iterate convergence gap.
· “A Unified Framework for Iterate Convergence of Bregman Proximal Methods”
Valley, orbital, moiré symmetry and band labels map 600+ bilayers onto Hubbard, topological and 1D models
File tools, consensus mining and validation gating speed convergence and let smaller models outperform larger ones
· “SkillOpt-Lite: Better and Faster Agent Self-evolution via One Line of Vibe”
The least C in the polynomial bound is max of 1 and 2|q|/(1+sqrt(1-|q|^2)).
A sharp density threshold now guarantees a small subfamily of edges with every vertex appearing evenly.
Fixed intra-epoch criteria allow utility updates at boundaries, raising coding pass rates and acceptance rates with 1.35x-1.72x fewer tokens
· “The Red Queen G\"odel Machine: Co-Evolving Agents and Their Evaluators”
When amplified noise kills the signal, the method reduces to rescaling one noisy measurement, overshooting the ideal by up to 21%.
· “Benchmarking Error Mitigation: Artefactual Improvements in Zero-Noise Extrapolation”
The long-open equality between unrestricted and Gaussian entanglement of formation now holds, turning covariance data into exact values and
The paper restores linearity, erasure-vs-error equivalence, and the first asymptotically good deletion and Majorana code families.
· “Theory of approximate quantum error correction and the error-set model”
User-need evaluation guides fixes for long-tail knowledge and instruction following to serve broad user demands.
· “Seed2.0 Model Card: Towards Intelligence Frontier for Real-World Complexity”
Latest arXiv reviews