Pith. sign in

Paper Citation Record · LEDGER

SciAgentGym: Benchmarking Multi-Step Scientific Tool-use in LLM Agents

As of 21 August 2026, this Paper Citation Record lists 49 of 49 outbound references and 2 inbound Pith citation observations for arXiv:2602.12984.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2602.12984 v2

Coverage vector

measured 49 of 49 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-02T23:41:46.847782Z

measured 51 of 51 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-21T06:32:19.484+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-06-27T09:34:09.347912Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-07-03T11:28:04.326064Z

Reference resolution

49 of 49 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved48
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 74473242-8dee-46df-bf85-2a44e35342cf · outbound

This paper cites On Evaluation of Embodied Navigation Agents.

SciAgentGym: Benchmarking Multi-Step Scientific Tool-use in LLM Agents On Evaluation of Embodied Navigation Agents

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-02T23:41:41.542179Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T23:41:41.542179Z digest=sha256:06cd46520606eb23be100f0682884393677ec70291ca91014fce1ee693e6097b

Observation 950e1056-e214-47dd-8521-8f0845cb7eff · outbound

This paper cites System card: Claude opus 4 & claude sonnet 4.

SciAgentGym: Benchmarking Multi-Step Scientific Tool-use in LLM Agents System card: Claude opus 4 & claude sonnet 4

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-02T23:41:41.626293Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T23:41:41.626293Z digest=sha256:423a7e5032d0b286aeb3af36b8d44b3256cc098f066168dbbb4294be2eff05b3

Observation dc4b5753-9acd-4630-802b-d7c3fcf6e0a1 · outbound

This paper cites Claude sonnet 4.5 system card.https://www.anthropic.com/claude-sonne t-4-5-system-card, 2025.

SciAgentGym: Benchmarking Multi-Step Scientific Tool-use in LLM Agents Claude sonnet 4.5 system card.https://www.anthropic.com/claude-sonne t-4-5-system-card, 2025

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-02T23:41:41.760033Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T23:41:41.760033Z digest=sha256:a14e1c842a3bc00b300de59fb4287d19c9987823daea15a7f250c53c27a5143a

Observation 3e40eec4-fa03-4bb0-bf45-28bb5653ea85 · outbound

This paper cites SciMaster: Towards General-Purpose Scientific AI Agents, Part I. X-Master as Foundation: Can We Lead on Humanity's Last Exam?.

SciAgentGym: Benchmarking Multi-Step Scientific Tool-use in LLM Agents SciMaster: Towards General-Purpose Scientific AI Agents, Part I. X-Master as Foundation: Can We Lead on Humanity's Last Exam?

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-02T23:41:41.833775Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T23:41:41.833775Z digest=sha256:0f67204d0051287865b5dfce39d8ffb109eaab25d5eeea8b2577603c4a89d1d9

Observation 4e29a8cc-7b09-435b-94b2-a796665c3b17 · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

SciAgentGym: Benchmarking Multi-Step Scientific Tool-use in LLM Agents DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-02T23:41:41.937664Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T23:41:41.937664Z digest=sha256:0403d39beb836a6c555a2c1f2f179903578f7995b5cabb516c9d9ee78dda4d9d

Observation 3dc69148-fa01-465f-bef5-f6d5a7a2240b · outbound

This paper cites SuperGPQA: Scaling LLM Evaluation across 285 Graduate Disciplines.

SciAgentGym: Benchmarking Multi-Step Scientific Tool-use in LLM Agents SuperGPQA: Scaling LLM Evaluation across 285 Graduate Disciplines

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-02T23:41:42.011986Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T23:41:42.011986Z digest=sha256:347fd109fcbd0314347080c86f7f7fe88b282f983fb2f193bc28b3daf2de3024

Observation 1ca25276-d045-401e-a83c-c106fa18b3ee · outbound

This paper cites RBench-V: A Primary Assessment for Visual Reasoning Models with Multi-modal Outputs.

SciAgentGym: Benchmarking Multi-Step Scientific Tool-use in LLM Agents RBench-V: A Primary Assessment for Visual Reasoning Models with Multi-modal Outputs

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-02T23:41:42.127992Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T23:41:42.127992Z digest=sha256:93500b928b4bf81dd042ca8b314db567963b534b51d31da4fe648dd338b413a8

Observation b0663a81-39ac-4ee8-b7e8-d3f351d37055 · outbound

This paper cites R-Bench: Graduate-level Multi-disciplinary Benchmarks for LLM & MLLM Complex Reasoning Evaluation.

SciAgentGym: Benchmarking Multi-Step Scientific Tool-use in LLM Agents R-Bench: Graduate-level Multi-disciplinary Benchmarks for LLM & MLLM Complex Reasoning Evaluation

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-02T23:41:42.234154Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T23:41:42.234154Z digest=sha256:2314be2bd195239e6145c30aac99ab98faed54d41eea957a4a4eb00e319b4feb

Observation 4976358a-5bc9-445a-bfb7-55cb644d8af7 · outbound

This paper cites Discoveryworld: A virtual environ- ment for developing and evaluating automated scientific discovery agents.Advances in Neural Information Processing Systems, 37:10088–10116, 2024.

SciAgentGym: Benchmarking Multi-Step Scientific Tool-use in LLM Agents Discoveryworld: A virtual environ- ment for developing and evaluating automated scientific discovery agents.Advances in Neural Information Processing Systems, 37:10088–10116, 2024

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-02T23:41:42.555326Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T23:41:42.555326Z digest=sha256:41d9cb666791ebe2a6cee173680e54b9e780c62dd7143aea6973e4852f269546

Observation a84e9cb5-7ffb-4bf3-afd3-a6f9e5cf63c3 · outbound

This paper cites R-Bench: Graduate-level Multi-disciplinary Benchmarks for LLM & MLLM Complex Reasoning Evaluation.

SciAgentGym: Benchmarking Multi-Step Scientific Tool-use in LLM Agents R-Bench: Graduate-level Multi-disciplinary Benchmarks for LLM & MLLM Complex Reasoning Evaluation

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-02T23:41:42.420473Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T23:41:42.420473Z digest=sha256:b6e0195c5d0a73434102fa51ed740b7eee8f4660557d4f095a84b9f65b9c9903

Observation 2188affb-e4be-4da2-934d-62c4c193b34b · outbound

This paper cites AgentBench: Evaluating LLMs as Agents.

SciAgentGym: Benchmarking Multi-Step Scientific Tool-use in LLM Agents AgentBench: Evaluating LLMs as Agents

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-02T23:41:42.810611Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T23:41:42.810611Z digest=sha256:eb770dfaff39f71671858d1e5da8e1742d2240f2c34b103d026f42a16cc5503a

Observation 03a0ca64-dbeb-45e7-b0fd-21cb65174b20 · outbound

This paper cites Search-o1: Agentic Search-Enhanced Large Reasoning Models.

SciAgentGym: Benchmarking Multi-Step Scientific Tool-use in LLM Agents Search-o1: Agentic Search-Enhanced Large Reasoning Models

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-02T23:41:42.720642Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T23:41:42.720642Z digest=sha256:9032f02c32b1b4c0397b174726b7e5ae09e2110706e1b6716f6150388df3d058

Observation 2adc9e63-4b74-4682-9bcc-435936b1d7d0 · outbound

This paper cites SciAgent: Tool-augmented Language Models for Scientific Reasoning.

SciAgentGym: Benchmarking Multi-Step Scientific Tool-use in LLM Agents SciAgent: Tool-augmented Language Models for Scientific Reasoning

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-02T23:41:43.110460Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T23:41:43.110460Z digest=sha256:9b49d767c92078868715b9419f717517c723f163edb02142ff8b3b1290329c1d

Observation e2644418-bd9b-4aeb-916e-a46e51f3f435 · outbound

This paper cites Learn to explain: Multimodal reasoning via thought chains for sciencequestionanswering.

SciAgentGym: Benchmarking Multi-Step Scientific Tool-use in LLM Agents Learn to explain: Multimodal reasoning via thought chains for sciencequestionanswering

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-02T23:41:42.972355Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T23:41:42.972355Z digest=sha256:19e284c7e4ac43ea6e7f16070fee9c05add8cd8bdb4e22e50a1bde9561c2b998

Observation 9df05970-bd16-4978-bedd-dafad7ef3e9d · outbound

This paper cites Patil, Huanzhi Mao, Fanjia Yan, Charlie Cheng-Jie Ji, Vishnu Suresh, Ion Stoica, and Joseph E.

SciAgentGym: Benchmarking Multi-Step Scientific Tool-use in LLM Agents Patil, Huanzhi Mao, Fanjia Yan, Charlie Cheng-Jie Ji, Vishnu Suresh, Ion Stoica, and Joseph E

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-02T23:41:43.489952Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T23:41:43.489952Z digest=sha256:059ce1c615b3086df10c9518d54af00594e04d524ef6e65efc1024fa2a8e5be6

Observation 4c320aa6-47e9-46bd-84c8-a3e4cfb63e72 · outbound

This paper cites GPT-4 Technical Report.

SciAgentGym: Benchmarking Multi-Step Scientific Tool-use in LLM Agents GPT-4 Technical Report

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-02T23:41:43.244893Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T23:41:43.244893Z digest=sha256:91c49b7cadca31955b65cfe5adf4047fa1dad49473e208f6a39a6f48a28aba6a

Observation 9d12ad95-afbf-4ad3-8b9d-dc75de34e558 · outbound

This paper cites GPQA: A Graduate-Level Google-Proof Q&A Benchmark.

SciAgentGym: Benchmarking Multi-Step Scientific Tool-use in LLM Agents GPQA: A Graduate-Level Google-Proof Q&A Benchmark

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-02T23:41:43.849352Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T23:41:43.849352Z digest=sha256:7253d9df59becb75da53cfe8c1deff2a72cd4be9f62f7e7226fe85f77b193902

Observation 13636209-cf9b-4ac8-aa01-4d39739d02c0 · outbound

This paper cites OpenAI GPT-5 System Card.

SciAgentGym: Benchmarking Multi-Step Scientific Tool-use in LLM Agents OpenAI GPT-5 System Card

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-02T23:41:43.964010Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T23:41:43.964010Z digest=sha256:61a97d1b58496833c2ba0a50ecfc6a7b2b06c9e957925274bb01793442264edc

Observation 8b829228-cf1d-4d46-8ee0-fac425ef68e2 · outbound

This paper cites ToolLLM: Facilitating Large Language Models to Master 16000+ Real-world APIs.

SciAgentGym: Benchmarking Multi-Step Scientific Tool-use in LLM Agents ToolLLM: Facilitating Large Language Models to Master 16000+ Real-world APIs

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-02T23:41:43.668704Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T23:41:43.668704Z digest=sha256:cc8a5e36b7ecfbe4276036de201c1a95a9301dcaffe814483b14c8d50d826002

Observation 51aa645c-fae6-4fa5-a76c-97b7b2b589f8 · outbound

This paper cites A Survey of AI for Materials Science: Foundation Models, LLM Agents, Datasets, and Tools.

SciAgentGym: Benchmarking Multi-Step Scientific Tool-use in LLM Agents A Survey of AI for Materials Science: Foundation Models, LLM Agents, Datasets, and Tools

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-02T23:41:44.325186Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T23:41:44.325186Z digest=sha256:dbacc1b9aeba45126ac839033b16f5901b87b9b272df7b6092ad91a04a9cb66c

Observation 822f1f30-8980-4930-92de-533dc8e382ca · outbound

This paper cites SciPy 1.0--Fundamental Algorithms for Scientific Computing in Python.

SciAgentGym: Benchmarking Multi-Step Scientific Tool-use in LLM Agents SciPy 1.0--Fundamental Algorithms for Scientific Computing in Python

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-02T23:41:44.466323Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T23:41:44.466323Z digest=sha256:98eb654bd32ea2ba6f4111dffa367aa4dca0e16a6a8e1d4cbee0f3536dec32b6

Observation 06f3a84d-5ff8-4037-8741-988b642024a5 · outbound

This paper cites Gemini: A Family of Highly Capable Multimodal Models.

SciAgentGym: Benchmarking Multi-Step Scientific Tool-use in LLM Agents Gemini: A Family of Highly Capable Multimodal Models

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-02T23:41:44.143270Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T23:41:44.143270Z digest=sha256:d307416a4c54df536a8b449093421881b371133370d521806068ca5a7ecd8a8e

Observation 5da8c59e-ac14-4cad-88ec-70bf662a7fb7 · outbound

This paper cites From AI for science to agentic science: A survey on autonomous scientific discovery.CoRR, abs/2508.14111, 2025.

SciAgentGym: Benchmarking Multi-Step Scientific Tool-use in LLM Agents From AI for science to agentic science: A survey on autonomous scientific discovery.CoRR, abs/2508.14111, 2025

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-02T23:41:44.760711Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T23:41:44.760711Z digest=sha256:0744dd2bf0c7bdc3c6c9f6aed8d2c477190eee9ed3d4b33cf351d42ab7784027

Observation e2328cb1-f3cd-4bd8-88e7-2a7dd8f8ff5f · outbound

This paper cites AgentGym: Evaluating and training large language model-based agents across diverse environments.

SciAgentGym: Benchmarking Multi-Step Scientific Tool-use in LLM Agents AgentGym: Evaluating and training large language model-based agents across diverse environments

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-02T23:41:44.879166Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T23:41:44.879166Z digest=sha256:d31e40b42205478f044acb7f251df931637f77e407f5b5b15f08aec081847e42

Observation 18793b37-a88c-4d30-aeee-1d82db5fb1e6 · outbound

This paper cites SciBench: Evaluating College-Level Scientific Problem-Solving Abilities of Large Language Models.

SciAgentGym: Benchmarking Multi-Step Scientific Tool-use in LLM Agents SciBench: Evaluating College-Level Scientific Problem-Solving Abilities of Large Language Models

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-02T23:41:44.622336Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T23:41:44.622336Z digest=sha256:f88da8ca05f5078082710b3a65165711177ae5f8599012cb6c27b4bffcf27281

Observation 043f1a4b-b753-4a0b-9e06-afe4d804ee16 · outbound

This paper cites BMMR: A Large-Scale Bilingual Multimodal Multi-Discipline Reasoning Dataset.

SciAgentGym: Benchmarking Multi-Step Scientific Tool-use in LLM Agents BMMR: A Large-Scale Bilingual Multimodal Multi-Discipline Reasoning Dataset

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-02T23:41:45.067017Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T23:41:45.067017Z digest=sha256:b7eda4f1a21966d054428c14204963a8a9401de5943bb1dbf888e4339e076b12

Observation dd321c8a-80c4-48f9-8e85-7e8113c85983 · outbound

This paper cites On the Tool Manipulation Capability of Open-source Large Language Models.

SciAgentGym: Benchmarking Multi-Step Scientific Tool-use in LLM Agents On the Tool Manipulation Capability of Open-source Large Language Models

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-02T23:41:45.208041Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T23:41:45.208041Z digest=sha256:fa33f95bd035205956063b71f180226443a3b16ba8935b2c40e81341a6ecdd2e

Observation 521c8bd2-31d2-4ddf-bc5f-0bfa840077e2 · outbound

This paper cites AgentGym-RL: Training LLM Agents for Long-Horizon Decision Making through Multi-Turn Reinforcement Learning.

SciAgentGym: Benchmarking Multi-Step Scientific Tool-use in LLM Agents AgentGym-RL: Training LLM Agents for Long-Horizon Decision Making through Multi-Turn Reinforcement Learning

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-02T23:41:44.971735Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T23:41:44.971735Z digest=sha256:8f40c476dea0ea0cf31c252fae068810bfe7b1273c14837c9b2e890dca2705ef

Observation 2ac982ff-d9a9-42a2-b187-1956894abcec · outbound

This paper cites Qwen3 Technical Report.

SciAgentGym: Benchmarking Multi-Step Scientific Tool-use in LLM Agents Qwen3 Technical Report

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-02T23:41:45.360064Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T23:41:45.360064Z digest=sha256:5972112fb4e7ef14d96d8e2d749ce420afcffc7005eb8ab143600abc8d871277

Observation f315465b-7099-4a1d-8722-cbb6ef739bb4 · outbound

This paper cites ReAct: Synergizing Reasoning and Acting in Language Models.

SciAgentGym: Benchmarking Multi-Step Scientific Tool-use in LLM Agents ReAct: Synergizing Reasoning and Acting in Language Models

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-02T23:41:45.410108Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T23:41:45.410108Z digest=sha256:d9d807de3e86df8237527506705a551981874784dffb64fc6299c9c416d01b63

Observation 1ad184a2-420a-4a78-be22-8aa42cfb9da2 · outbound

This paper cites Wang, Peifeng Ruan, Donghan Yang, Tao Wang, Guanghua Xiao, Xin Liu, Carl Yang, Yang Xie, and Wenqi Shi.

SciAgentGym: Benchmarking Multi-Step Scientific Tool-use in LLM Agents Wang, Peifeng Ruan, Donghan Yang, Tao Wang, Guanghua Xiao, Xin Liu, Carl Yang, Yang Xie, and Wenqi Shi

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-02T23:41:45.299746Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T23:41:45.299746Z digest=sha256:f677d0701c64cfc8ce33094d43d96e41c16dc4b756ab2c978626c770739b7090

Observation a366daec-3d43-43b2-a6f9-54bebbd1f34a · outbound

This paper cites debug-gym: A Text-Based Environment for Interactive Debugging.

SciAgentGym: Benchmarking Multi-Step Scientific Tool-use in LLM Agents debug-gym: A Text-Based Environment for Interactive Debugging

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-02T23:41:45.552764Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T23:41:45.552764Z digest=sha256:2647aa9f4aabe1236fadd6de4e659d40b49fa83ad87cb403551024bb3c992ed0

Observation 3022e7e4-6289-4b75-b337-7e9140606382 · outbound

This paper cites MMMU-Pro: A More Robust Multi-discipline Multimodal Understanding Benchmark.

SciAgentGym: Benchmarking Multi-Step Scientific Tool-use in LLM Agents MMMU-Pro: A More Robust Multi-discipline Multimodal Understanding Benchmark

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-02T23:41:45.646682Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T23:41:45.646682Z digest=sha256:f64e58f8d8cb17b9cf88b423de93f87c2c11d79d6ab8a33ed8457c96e19026ea

Observation 7a2a993a-c2e0-4d3b-825b-f272da28d61e · outbound

This paper cites URLhttps://arxiv.org/abs/24 06.12045.

SciAgentGym: Benchmarking Multi-Step Scientific Tool-use in LLM Agents URLhttps://arxiv.org/abs/24 06.12045

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-02T23:41:45.464222Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T23:41:45.464222Z digest=sha256:67973047a0504d6ab93d4ff66c6d39d9b709152af68bfb0d52e39b29c1c9f1d5

Observation 20337007-c0ef-4811-a432-d37b4836ab93 · outbound

This paper cites Sciinstruct: a self-reflective instruction annotated dataset for training scientific language models.

SciAgentGym: Benchmarking Multi-Step Scientific Tool-use in LLM Agents Sciinstruct: a self-reflective instruction annotated dataset for training scientific language models

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-02T23:41:45.806856Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T23:41:45.806856Z digest=sha256:40b9b8835172023e8df4d3229156bb6515b2d7f5eb6830432029ec09731a3b2f

Observation a3a7ae48-876a-4096-9503-ade0391bc9f1 · outbound

This paper cites The Landscape of Agentic Reinforcement Learning for LLMs: A Survey.

SciAgentGym: Benchmarking Multi-Step Scientific Tool-use in LLM Agents The Landscape of Agentic Reinforcement Learning for LLMs: A Survey

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-02T23:41:45.930272Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T23:41:45.930272Z digest=sha256:69ffb0741d1232b7b55a8149b2c2786a94dd6d39186cf59c284356ef5a9e1e07

Observation 5312309a-044c-49d6-bfc6-5dc8604930c6 · outbound

This paper cites GLM-4.5: Agentic, Reasoning, and Coding (ARC) Foundation Models.

SciAgentGym: Benchmarking Multi-Step Scientific Tool-use in LLM Agents GLM-4.5: Agentic, Reasoning, and Coding (ARC) Foundation Models

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-02T23:41:45.766345Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T23:41:45.766345Z digest=sha256:9d098a427810b8c67445836ccda88fc0a03861ea4465fb1dcc4392dfc23c4857

Observation a4e71bd1-acc7-45b5-8874-68bee738df54 · outbound

This paper cites Scientists’ first exam: Probing cognitive abilities of mllm via perception, understanding, and reasoning.arXiv preprint arXiv:2506.10521, 2025.

SciAgentGym: Benchmarking Multi-Step Scientific Tool-use in LLM Agents Scientists’ first exam: Probing cognitive abilities of mllm via perception, understanding, and reasoning.arXiv preprint arXiv:2506.10521, 2025

Reference 38

Resolution
malformed identifier
no resolver link, observed 2026-08-02T23:41:46.142352Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T23:41:46.142352Z digest=sha256:ac2eab4930420b02a57a7ca099f97b23bff35a3932fcca9f2725482f30d009a5

Observation 7e2daef1-1a3a-45d7-9cdb-109302624488 · outbound

This paper cites WebArena: A Realistic Web Environment for Building Autonomous Agents.

SciAgentGym: Benchmarking Multi-Step Scientific Tool-use in LLM Agents WebArena: A Realistic Web Environment for Building Autonomous Agents

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-02T23:41:46.044618Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T23:41:46.044618Z digest=sha256:0a97d3027de6d447362de525a06ca7cfd496527d5cbb51bdd1d079c182b07181

Observation df752ca1-612f-4c9f-8f17-37053c441d98 · outbound

This paper cites an unresolved cited work.

SciAgentGym: Benchmarking Multi-Step Scientific Tool-use in LLM Agents Unresolved cited work

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-02T23:41:46.239829Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T23:41:46.239829Z digest=sha256:ae1a6845e00896a1979d28985bb7de7105628e0128e1831773b04694cf071f65

Observation b1a4de21-0e95-4915-a80c-9eeca4213bd3 · outbound

This paper cites •Standard Return:{’result’: main_value, ’metadata’: {...}}(e.g., units, status flags, data sources).

SciAgentGym: Benchmarking Multi-Step Scientific Tool-use in LLM Agents •Standard Return:{’result’: main_value, ’metadata’: {...}}(e.g., units, status flags, data sources)

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-02T23:41:46.339736Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T23:41:46.339736Z digest=sha256:afda6ed42e76a711bbe2c5944429b113254c42fa37fe1f21ea373c5708711665

Observation e5ff4846-6c47-45d5-a582-582625dc238e · outbound

This paper cites •Type Hints:All tools must provide complete Python type hints for parameters and return values.

SciAgentGym: Benchmarking Multi-Step Scientific Tool-use in LLM Agents •Type Hints:All tools must provide complete Python type hints for parameters and return values

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-02T23:41:46.446763Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T23:41:46.446763Z digest=sha256:888bee513278c94a92159b081b2f7c945025c559952223f43a1e077005abbb97

Observation a0a8b7f6-be47-432e-b27b-768bd2e50fac · outbound

This paper cites •Query:Retrieve hierarchical facts/records from external resources or local indices and return normalized fields for downstream steps.

SciAgentGym: Benchmarking Multi-Step Scientific Tool-use in LLM Agents •Query:Retrieve hierarchical facts/records from external resources or local indices and return normalized fields for downstream steps

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-02T23:41:46.552473Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T23:41:46.552473Z digest=sha256:07f812fdc5fe028104da73db18b844f95406185b134a923cfa3aa7c87d409fda

Observation 128b0d44-45b7-4905-b6de-2b5d57aed577 · outbound

This paper cites must include X.

SciAgentGym: Benchmarking Multi-Step Scientific Tool-use in LLM Agents must include X

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-02T23:41:46.629695Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T23:41:46.629695Z digest=sha256:08cd976cf9e9eba31cd3e79da770efa89205211ce5511da852e886e1ee22b542

Observation 62b6cd3d-2f28-46bd-8a08-6aed599e8cc9 · outbound

This paper cites an unresolved cited work.

SciAgentGym: Benchmarking Multi-Step Scientific Tool-use in LLM Agents Unresolved cited work

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-02T23:41:46.712280Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T23:41:46.712280Z digest=sha256:f1f83d80ee0d1a11b90960e33dd95ff4f9f42cc4d16241e43c89f947920e8a66

Observation 3848077d-eac6-431d-85a0-b01af5be0ee1 · outbound

This paper cites an unresolved cited work.

SciAgentGym: Benchmarking Multi-Step Scientific Tool-use in LLM Agents Unresolved cited work

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-02T23:41:46.758337Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T23:41:46.758337Z digest=sha256:0f09923c6709ebb7bd3f33db86d2762420b61a2bef10471fbcae9f69fbb95ed6

Observation 039c8a1a-3706-4739-84ad-c9d4779a0ae3 · outbound

This paper cites Tool-round limit exceeded; stopped automatically.

SciAgentGym: Benchmarking Multi-Step Scientific Tool-use in LLM Agents Tool-round limit exceeded; stopped automatically

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-02T23:41:46.847782Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T23:41:46.847782Z digest=sha256:ae6b75c10a0f8445c0caaa7415f79cabdcb962a4db0a4ec1aa5f00d79e162fe9

Observation 6ac35669-a6d2-4d1a-adbe-3287f9bab8c2 · outbound

This paper cites GPT-4 Technical Report.

SciAgentGym: Benchmarking Multi-Step Scientific Tool-use in LLM Agents GPT-4 Technical Report

Reference 774

Resolution
unresolved
no resolver link, observed 2026-08-02T23:41:43.364678Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T23:41:43.364678Z digest=sha256:1e1a203e9189e5ba6d465505fceb5f005c62bf9b370100f53b107eda63c060dc

Observation 23f30786-a418-42b2-961e-5a0b4ebefc36 · outbound

This paper cites an unresolved cited work.

SciAgentGym: Benchmarking Multi-Step Scientific Tool-use in LLM Agents Unresolved cited work

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-02T23:41:41.688635Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T23:41:41.688635Z digest=sha256:91cebf6dd63f768bf330eb60ea4125fe671359c1e0bb357f0078aa28f941f438

Pith citing papers

Observation 34e2cafd-a6ce-4171-a250-d665cac90c96 · inbound

AI scientists produce results without reasoning scientifically cites this paper.

AI scientists produce results without reasoning scientifically SciAgentGym: Benchmarking Multi-Step Scientific Tool-use in LLM Agents

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-06-02T02:03:38.947361Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-10T03:56:34.581133Z digest=sha256:ba2c3063989736ac6a67f1188079823db323511f918234ede879604dfb69a1c9

Observation 4a5d332b-b3ee-410f-93d2-9b9884877e8d · inbound

Benchmarking AI Agents for Addressing Scientific Challenges Across Scales cites this paper.

Benchmarking AI Agents for Addressing Scientific Challenges Across Scales SciAgentGym: Benchmarking Multi-Step Scientific Tool-use in LLM Agents

Reference 12

Resolution
verified exact
local_arxiv, observed 2026-07-03T11:28:04.327347Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-06-27T09:34:09.347912Z digest=sha256:b2fc7fe53e2ecd276a4c0af3675b69b49eb2f4981053be38d3cb5793cbe01acb