Completely independent Steiner trees are defined as a generalization of completely independent spanning trees and internally disjoint Steiner trees, accompanied by characterizations, bounds, algorithms, hardness results, and applications to planar graphs and bounded-treewidth graphs plus a directed-
hub Canonical reference
AceReason-Nemotron: Advancing math and code reasoning through reinforcement learning
Canonical reference. 85% of citing Pith papers cite this work as background.
hub tools
citation-role summary
citation-polarity summary
representative citing papers
On a bilingual NSPS benchmark, Korean language consistently lowers harmful compliance (~10pp TRS), Korean grounding often mitigates that drop, and open- vs closed-source models reverse under direct requests.
A new large-scale triplet dataset and diffusion transformer model using coarse human masks deliver improved video virtual try-on quality and generalization in challenging real-world conditions.
Tractable relational probabilistic hyperproperties for MDPs are identified with efficient algorithms for probability-equality queries on reachability and omega-regular events, plus hardness results and a fast implementation.
RoMathExam supplies a century-long collection of Romanian math exams together with a new intrinsic complexity metric that correlates across frontier models at r > 0.72.
SGNO achieves stable long-horizon PDE rollouts by organizing autoregressive steps as spectral evolution updates with a constrained diagonal generator and learned correction, delivering a median 74.8% reduction in GMean100 error across ten APEBench tasks.
Failure-driven self-improvement raises OpenCUA-72B success rate on OSWorld from 42.3% to 48.9% via LLM diagnosis and inference-time code patches, without retraining.
HealthAgentBench is a new benchmark of 54 healthcare agent tasks where even the strongest frontier AI agent reaches only about 42% success rate on end-to-end clinical workflows.
Next-token prediction on multi-modal tokenized sleep signals yields embeddings that match supervised performance with far less labels and generalize to daytime heart data.
SciVisAgentSkills provides reusable agent skills that raise mean task scores on a 108-task SciVis benchmark when paired with Codex and Claude Code agents.
TabPFN-3 scales tabular foundation models to 1M rows with synthetic pretraining, test-time compute, and benchmark-leading performance on tabular, relational, and tabular-text tasks while being up to 20x faster than TabPFN-2.5.
VISOR is a VLM-based automated test oracle that evaluates robot task correctness and quality from videos while reporting its own uncertainty, tested on GPT and Gemini across four tasks and over 1000 videos with Gemini showing higher recall and GPT higher precision but low uncertainty-correctness tie
Prefix Sampling replays self-generated trajectory prefixes to control rollout pass rates near 50% in binary-reward RL, delivering wall-clock speedups and modest performance gains on SWE-bench Verified and AIME tasks.
MedMosaic is a new large-scale medical audio question-answering benchmark showing that even Gemini-2.5-pro reaches only ~68% accuracy across diverse clinical audio scenarios.
Wi2SAR is a drone-based wireless system that locates wilderness victims by exploiting automatic Wi-Fi reconnection on their mobile devices, using a 3D-printed Luneburg Lens for direction finding and adaptive navigation.
TabPFN-2.5 scales tabular foundation models to 20x larger datasets, outperforms tuned tree models on TabArena, achieves near-perfect win rates against default XGBoost, and adds a distillation engine for fast production deployment.
A theory of time from wavefunction collapse in GR predicts emergent unitary tensor graviton dynamics and identifies long-wavelength scalar modes as a viable dark matter candidate in a cosmological constant-dominated universe.
This survey categorizes agentic environments for LLMs by eight attributes and domains, introduces symbolic and neural synthesis paradigms with evaluation, and outlines four agent evolution pathways plus three environment evolution paradigms.
A survey that maps safety risks in personalized LLMs, introduces a unified taxonomy, and highlights three structural inadequacies in existing research on user-invariant safety, isolated techniques, and short-term evaluations.
Empirical comparison of domain-specific, computer-use, and general-purpose LLM agents plus CLI/GUI modalities on SciVis tasks reveals general-purpose agents highest in success rate but costliest, domain-specific agents more efficient, and persistent memory beneficial depending on mode.
Eywa enables language-based agentic AI systems to collaborate with specialized scientific foundation models for improved performance on structured data tasks.
A new turbofan dataset with realistic maintenance patterns is used to benchmark Bayesian filters as strong baselines against self-supervised learning representations for component health estimation.
Analysis of 1,223 AI-HCI papers shows declining focus on human epistemic sovereignty and rising optimization of autonomous agents, leading to a proposal for scaffolded cognitive friction via multi-agent systems to preserve human cognitive agency.
Four classroom-time variables predict physics concept learning, and classes with 10–20% group worksheets, 20–40% group clickers, and ≥2 student questions per hour show effect sizes above 2.
citing papers explorer
-
Completely Independent Steiner Trees
Completely independent Steiner trees are defined as a generalization of completely independent spanning trees and internally disjoint Steiner trees, accompanied by characterizations, bounds, algorithms, hardness results, and applications to planar graphs and bounded-treewidth graphs plus a directed-
-
ROK-FORTRESS: Measuring the Effect of Geopolitical Transcreation for National Security and Public Safety
On a bilingual NSPS benchmark, Korean language consistently lowers harmful compliance (~10pp TRS), Korean grounding often mitigates that drop, and open- vs closed-source models reverse under direct requests.
-
TripVVT: A Large-Scale Triplet Dataset and a Coarse-Mask Baseline for In-the-Wild Video Virtual Try-On
A new large-scale triplet dataset and diffusion transformer model using coarse human masks deliver improved video virtual try-on quality and generalization in challenging real-world conditions.
-
Tractable Hyperproperties for MDPs
Tractable relational probabilistic hyperproperties for MDPs are identified with efficient algorithms for probability-equality queries on reachability and omega-regular events, plus hardness results and a fast implementation.
-
RoMathExam: A Longitudinal Dataset of Romanian Math Exams (1895-2025) with a Seven-Decade Core (1957-2025)
RoMathExam supplies a century-long collection of Romanian math exams together with a new intrinsic complexity metric that correlates across frontier models at r > 0.72.
-
SGNO: Spectral Generator Neural Operators for Stable Long Horizon PDE Rollouts
SGNO achieves stable long-horizon PDE rollouts by organizing autoregressive steps as spectral evolution updates with a constrained diagonal generator and learned correction, delivering a median 74.8% reduction in GMean100 error across ten APEBench tasks.
-
Learning from Failure: Inference-Time Self-Improvement for Computer-Use Agents
Failure-driven self-improvement raises OpenCUA-72B success rate on OSWorld from 42.3% to 48.9% via LLM diagnosis and inference-time code patches, without retraining.
-
HealthAgentBench: A Unified Benchmark Suite of Realistic Agentic Healthcare Environments for Challenging Frontier AI Agents
HealthAgentBench is a new benchmark of 54 healthcare agent tasks where even the strongest frontier AI agent reaches only about 42% success rate on end-to-end clinical workflows.
-
Next-Token Prediction Learns Generalisable Representations of Sleep Physiology
Next-token prediction on multi-modal tokenized sleep signals yields embeddings that match supervised performance with far less labels and generalize to daytime heart data.
-
SciVisAgentSkills: Design and Evaluation of Agent Skills for Scientific Data Analysis and Visualization
SciVisAgentSkills provides reusable agent skills that raise mean task scores on a 108-task SciVis benchmark when paired with Codex and Claude Code agents.
-
TabPFN-3: Technical Report
TabPFN-3 scales tabular foundation models to 1M rows with synthetic pretraining, test-time compute, and benchmark-leading performance on tabular, relational, and tabular-text tasks while being up to 20x faster than TabPFN-2.5.
-
VISOR: A Vision-Language Model-based Test Oracle for Testing Robots
VISOR is a VLM-based automated test oracle that evaluates robot task correctness and quality from videos while reporting its own uncertainty, tested on GPT and Gemini across four tasks and over 1000 videos with Gemini showing higher recall and GPT higher precision but low uncertainty-correctness tie
-
Rollout Pass-Rate Control: Steering Binary-Reward RL Toward Its Most Informative Regime
Prefix Sampling replays self-generated trajectory prefixes to control rollout pass rates near 50% in binary-reward RL, delivering wall-clock speedups and modest performance gains on SWE-bench Verified and AIME tasks.
-
MedMosaic: A Challenging Large Scale Benchmark of Diverse Medical Audio
MedMosaic is a new large-scale medical audio question-answering benchmark showing that even Gemini-2.5-pro reaches only ~68% accuracy across diverse clinical audio scenarios.
-
"Take Me Home, Wi-Fi Drone": A Drone-based Wireless System for Wilderness Search and Rescue
Wi2SAR is a drone-based wireless system that locates wilderness victims by exploiting automatic Wi-Fi reconnection on their mobile devices, using a 3D-printed Luneburg Lens for direction finding and adaptive navigation.
-
TabPFN-2.5: Advancing the State of the Art in Tabular Foundation Models
TabPFN-2.5 scales tabular foundation models to 20x larger datasets, outperforms tuned tree models on TabArena, achieves near-perfect win rates against default XGBoost, and adds a distillation engine for fast production deployment.
-
Emergent time and more from wavefunction collapse in general relativity
A theory of time from wavefunction collapse in GR predicts emergent unitary tensor graviton dynamics and identifies long-wavelength scalar modes as a viable dark matter candidate in a cosmological constant-dominated universe.
-
Agentic Environment Engineering for Large Language Models: A Survey of Environment Modeling, Synthesis, Evaluation, and Application
This survey categorizes agentic environments for LLMs by eight attributes and domains, introduces symbolic and neural synthesis paradigms with evaluation, and outlines four agent evolution pathways plus three environment evolution paradigms.
-
Personalization Meets Safety:Mechanisms,Risks,and Mitigations in Personalized LLMs
A survey that maps safety risks in personalized LLMs, introduces a unified taxonomy, and highlights three structural inadequacies in existing research on user-invariant safety, isolated techniques, and short-term evaluations.
-
Exploring LLM Agent Designs and Interaction Modalities for Scientific Visualization
Empirical comparison of domain-specific, computer-use, and general-purpose LLM agents plus CLI/GUI modalities on SciVis tasks reveals general-purpose agents highest in success rate but costliest, domain-specific agents more efficient, and persistent memory beneficial depending on mode.
-
Heterogeneous Scientific Foundation Model Collaboration
Eywa enables language-based agentic AI systems to collaborate with specialized scientific foundation models for improved performance on structured data tasks.
-
A Machine Learning Framework for Turbofan Health Estimation via Inverse Problem Formulation
A new turbofan dataset with realistic maintenance patterns is used to benchmark Bayesian filters as strong baselines against self-supervised learning representations for component health estimation.
-
Cognitive Agency Surrender: Defending Epistemic Sovereignty via Scaffolded AI Friction
Analysis of 1,223 AI-HCI papers shows declining focus on human epistemic sovereignty and rising optimization of autonomous agents, leading to a proposal for scaffolded cognitive friction via multi-agent systems to preserve human cognitive agency.
-
Predictive Modeling for High Impact Active Learning Classrooms
Four classroom-time variables predict physics concept learning, and classes with 10–20% group worksheets, 20–40% group clickers, and ≥2 student questions per hour show effect sizes above 2.
-
Aleena: Alignment Agent for Research Software Engineering Collaborations
Aleena is an open-source AI agent that ingests multi-modal research software collaboration artifacts and transforms them into structured GitHub records to maintain continuous stakeholder alignment across the project lifecycle.
-
Prompt Governance? On Governing Technologies Governed by Natural Language
Literature on system prompts for AI shows fragmented and contradictory claims that complicate policy efforts to use them as reliable governance mechanisms.
-
Opening new parameter space windows on galaxy/AGN co-evolution with SKA radio continuum surveys
Overview of SKAO radio surveys for galaxy/AGN co-evolution, including tiered surveys, multi-frequency imaging, and synergies with other observatories.
- GRM: Utility-Aware Jailbreak Attacks on Audio LLMs via Gradient-Ratio Masking