Pith. sign in

hub Canonical reference

Mlgym: A new framework and benchmark for advancing ai research agents

Canonical reference. 83% of citing Pith papers cite this work as background.

18 Pith papers citing it
4 external citations · external index
Background 83% of classified citations

hub tools

citation-role summary

background 4 dataset 2

citation-polarity summary

years

2026 16 2025 2

representative citing papers

EO-Gym: A Multimodal, Interactive Environment for Earth Observation Agents

cs.AI · 2026-05-02 · unverdicted · novelty 7.0

EO-Gym supplies an executable multimodal environment and 9k-trajectory benchmark that turns Earth Observation into a tool-using, multi-step reasoning task, revealing that current VLMs struggle on temporal and cross-sensor workflows while fine-tuning lifts Pass@3 from 0.49 to 0.74.

AI-Driven Research for Databases

cs.DB · 2026-04-08 · unverdicted · novelty 6.0

Co-evolving LLM-generated solutions with their evaluators enables discovery of novel database algorithms that outperform state-of-the-art baselines, including a query rewrite policy with up to 6.8x lower latency.

AI for Auto-Research: Roadmap & User Guide

cs.AI · 2026-05-18 · conditional · novelty 4.0

AI can generate research artifacts faster than it can verify them, so across all eight lifecycle stages the credible deployment mode is human-governed collaboration rather than full autonomy.

citing papers explorer

Showing 18 of 18 citing papers.