AutoMedBench evaluates AI agents on long-horizon medical workflows across five stages and finds validation and submission as dominant failure points based on thousands of runs.
Canonical reference
You only look once: Unified, real-time object detection
Canonical reference. 80% of citing Pith papers cite this work as background.
citation-role summary
citation-polarity summary
representative citing papers
FRTSearch reframes fast radio transient detection as instance segmentation on dynamic spectra and uses the segmented shapes to infer dispersion measure and time of arrival, achieving 98% recall with over 99.9% fewer false positives than traditional methods.
A new benchmark (IG-Bench) reveals that LLM-based scientists fail at compositional lineage reasoning, with the best system reaching only 27.3% exact accuracy.
TCG-AR is a real-time multi-view AR system for trading card games using only commodity RGB cameras and synthetic training data.
The work introduces a distributional view of visual mechanistic interpretability that casts the task as KL-minimal optimization and realizes it through a soft-constraint principle implemented with energy-guided diffusion posterior sampling on models such as DINOv3.
PIT uses a neural autoencoder with a differentiable physics module and a new Physics-Informed Landmark Loss to track single particles in video, achieving sub-pixel accuracy in supervised and unsupervised modes.
A PINN pretrained on mechanistic synthetic data and fine-tuned experimentally is deployed in an EKF-style filter to estimate separator phase heights from flow rates alone.
Holi-DETR improves fashion item detection by integrating co-occurrence probabilities, inter-item spatial arrangements, and body keypoint relationships into the DETR architecture.
A policy gradient method with differentiable quadrotor dynamics learns agile interception from direction vectors alone, outperforming point-mass baselines by 30% at speeds up to 10 m/s.
Hippocampus-DETR integrates a hippocampal memory network (HipNet) into DETR to simulate brain subregions for pattern separation, completion, and improved detection accuracy plus generalization.
CucumberVision compares five 3D length methods on 48 RGB-D captures of seven cucumbers and shows a novel medial-axis cubic spline with trapezoidal integration achieves the lowest 4.13% MAPE, outperforming baselines at corrected significance.
SCOPE introduces an edge-deployable natural-language PTZ camera agent, a simulation benchmark, and evaluations showing that stronger small language models reduce hallucinations while perception remains the main bottleneck.
A wavelet-guided neural pipeline recovers previously known narrowband radio events from FAST observations of 33 exoplanet systems and reduces 139,127 detections to 803 veto-ready candidates; its one new candidate, toward K2-155, is argued to be terrestrial RFI.
A visual-semantic spatiotemporal framework creates the Street Economic Vitality Index (SEVI) to diagnose urban street economic vitality by parsing streetscapes with AI, standardizing brands via VLM-LLM, and incorporating lagged LBS demand data with Gaussian spillover modeling.
A simulated pipeline combining CNN wildfire detection, Bayesian confidence updates, and reconfigurable satellite scheduling automates wildfire monitoring, but it is not validated on real satellite images.
MOBIUS is a multi-modal bipedal robot with hybrid reinforcement learning and force control plus an MIQCP planner that enables walking, crawling, climbing, and rolling on varied terrains.
A cross-verification strategy using three YOLO models trained on distinct views of a 2134-sample 3D GPR dataset detects road subsurface distress with over 98.6 percent recall on field data.
Deep learning models on standardized 2D CT projections of pelvis and skull from 141 cadavers reach 95.65% patient-level accuracy for biological sex determination.
Tile-based inference with topology-aware merging improves small PCB defect detection by preserving scale and resolving edge artifacts on two datasets.
GSA-YOLO modifies YOLOv8n with structured sparsity via Group Lasso and Sparse Structure Selection plus Adaptive Knowledge Distillation, reporting 189.62 FPS and mAP50:95 gains of 2.4% and 1.8% on HiXray and PIDray datasets.
XiYOLO uses iterative energy-aware neural architecture search and scaling to produce object detectors with stronger accuracy-energy tradeoffs than YOLO baselines on GPUs and NPUs.
A UAS with YOLO-based swimmer detection and DES simulations reduces drowning rescue response time by a factor of five versus standard operations in tested lake areas.
Detection-guided prompting raises small VLM hazard F1 from 34.5% to 50.6% and BERTScore from 0.61 to 0.82 on construction images with only 2.5 ms added latency.
Presents a distributed ROS 2 framework integrating local LLMs and VLMs for conversational human-robot manipulation tasks with operator confirmation and experimental evaluation on a Franka FR3 arm.
citing papers explorer
-
AutoMedBench: Towards Medical AutoResearch with Agentic AI Models
AutoMedBench evaluates AI agents on long-horizon medical workflows across five stages and finds validation and submission as dominant failure points based on thousands of runs.
-
FRTSearch: Unified Detection and Parameter Inference of Fast Radio Transients using Instance Segmentation
FRTSearch reframes fast radio transient detection as instance segmentation on dynamic spectra and uses the segmented shapes to infer dispersion measure and time of arrival, achieving 98% recall with over 99.9% fewer false positives than traditional methods.
-
Ideas Have Genomes: Benchmarking Scientific Lineage Reasoning and Lineage-Grounded Idea Generation
A new benchmark (IG-Bench) reveals that LLM-based scientists fail at compositional lineage reasoning, with the best system reaching only 27.3% exact accuracy.
-
TCG-AR: Real-Time Multi-View Augmented Reality for Trading Card Game Streaming
TCG-AR is a real-time multi-view AR system for trading card games using only commodity RGB cameras and synthetic training data.
-
A Distributional View for Visual Mechanistic Interpretability: KL-Minimal Soft-Constraint Principle
The work introduces a distributional view of visual mechanistic interpretability that casts the task as KL-minimal optimization and realizes it through a soft-constraint principle implemented with energy-guided diffusion posterior sampling on models such as DINOv3.
-
Physics-Informed Tracking (PIT)
PIT uses a neural autoencoder with a differentiable physics module and a new Physics-Informed Landmark Loss to track single particles in video, achieving sub-pixel accuracy in supervised and unsupervised modes.
-
Estimating Dense-Packed Zone Height in Liquid-Liquid Separation: A Physics-Informed Neural Network Approach
A PINN pretrained on mechanistic synthetic data and fine-tuned experimentally is deployed in an EKF-style filter to estimate separator phase heights from flow rates alone.
-
Holi-DETR: Holistic Fashion Item Detection Leveraging Contextual Information
Holi-DETR improves fashion item detection by integrating co-occurrence probabilities, inter-item spatial arrangements, and body keypoint relationships into the DETR architecture.
-
Learning Agile Intruder Interception using Differentiable Quadrotor Dynamics
A policy gradient method with differentiable quadrotor dynamics learns agile interception from direction vectors alone, outperforming point-mass baselines by 30% at speeds up to 10 m/s.
-
Hippocampus-DETR: An Explicit Memory Object Detection Framework Based on Hippocampus Modeling
Hippocampus-DETR integrates a hippocampal memory network (HipNet) into DETR to simulate brain subregions for pattern separation, completion, and improved detection accuracy plus generalization.
-
Curvature-aware 3D length estimation of greenhouse cucumbers using RGB-D imaging and cubic spline arc-length integration
CucumberVision compares five 3D length methods on 48 RGB-D captures of seven cucumbers and shows a novel medial-axis cubic spline with trapezoidal integration achieves the lowest 4.13% MAPE, outperforming baselines at corrected significance.
-
SCOPE: Real-Time Natural Language Camera Agent at the Edge
SCOPE introduces an edge-deployable natural-language PTZ camera agent, a simulation benchmark, and evaluations showing that stronger small language models reduce hallucinations while perception remains the main bottleneck.
-
A Wavelet-Integrated Search Pipeline for Narrowband Technosignatures in FAST Observations of 33 Exoplanet Systems
A wavelet-guided neural pipeline recovers previously known narrowband radio events from FAST observations of 33 exoplanet systems and reduces 139,127 detections to 803 veto-ready candidates; its one new candidate, toward K2-155, is argued to be terrestrial RFI.
-
Diagnosing Urban Street Vitality via a Visual-Semantic and Spatiotemporal Framework for Street-Level Economics
A visual-semantic spatiotemporal framework creates the Street Economic Vitality Index (SEVI) to diagnose urban street economic vitality by parsing streetscapes with AI, standardizing brands via VLM-LLM, and incorporating lagged LBS demand data with Gaussian spillover modeling.
-
Automating the Wildfire Detection and Scheduling Pipeline with Maneuverable Earth Observation Satellites
A simulated pipeline combining CNN wildfire detection, Bayesian confidence updates, and reconfigurable satellite scheduling automates wildfire monitoring, but it is not validated on real satellite images.
-
MOBIUS: A Multi-Modal Bipedal Robot that can Walk, Crawl, Climb, and Roll
MOBIUS is a multi-modal bipedal robot with hybrid reinforcement learning and force control plus an MIQCP planner that enables walking, crawling, climbing, and rolling on varied terrains.
-
Automatic Road Subsurface Distress Recognition from Ground Penetrating Radar Images using Deep Learning-based Cross-verification
A cross-verification strategy using three YOLO models trained on distinct views of a 2134-sample 3D GPR dataset detects road subsurface distress with over 98.6 percent recall on field data.
-
Biological Sex Determination in Cadavers Using Deep Learning Algorithms from Computed Tomography Images of Pelvis and Skull
Deep learning models on standardized 2D CT projections of pelvis and skull from 141 cadavers reach 95.65% patient-level accuracy for biological sex determination.
-
From Full Boards to Tiny Defects: Scale-Aware Tile Inference with Topology-Aware Merging for High-Resolution PCB Defect Detection
Tile-based inference with topology-aware merging improves small PCB defect detection by preserving scale and resolving edge artifacts on two datasets.
-
GSA-YOLO: A High-Efficiency Framework via Structured Sparsity and Adaptive Knowledge Distillation for Real-Time X-ray Security Inspection
GSA-YOLO modifies YOLOv8n with structured sparsity via Group Lasso and Sparse Structure Selection plus Adaptive Knowledge Distillation, reporting 189.62 FPS and mAP50:95 gains of 2.4% and 1.8% on HiXray and PIDray datasets.
-
XiYOLO: Energy-Aware Object Detection via Iterative Architecture Search and Scaling
XiYOLO uses iterative energy-aware neural architecture search and scaling to produce object detectors with stronger accuracy-energy tradeoffs than YOLO baselines on GPUs and NPUs.
-
Autonomous Unmanned Aircraft Systems for Enhanced Search and Rescue of Drowning Swimmers: Image-Based Localization and Mission Simulation
A UAS with YOLO-based swimmer detection and DES simulations reduces drowning rescue response time by a factor of five versus standard operations in tested lake areas.
-
Integration of Object Detection and Small VLMs for Construction Safety Hazard Identification
Detection-guided prompting raises small VLM hazard F1 from 34.5% to 50.6% and BERTScore from 0.61 to 0.82 on construction images with only 2.5 ms added latency.
-
A Conversational Framework for Human-Robot Collaborative Manipulation with Distributed Generative AI models
Presents a distributed ROS 2 framework integrating local LLMs and VLMs for conversational human-robot manipulation tasks with operator confirmation and experimental evaluation on a Franka FR3 arm.
-
Software Engineering for Self-Adaptive Robotics: A Research Agenda
This paper proposes a research agenda for software engineering of self-adaptive robotic systems along lifecycle stages and enabling technologies, identifying challenges and a roadmap to 2030.
-
SoccerNet 2026 Challenges Results
The SoccerNet 2026 Challenges benchmarked 427 teams across five soccer video understanding tasks, with leading submissions improving over baselines on all tasks.