Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T04:23:50.683069Z
Paper Citation Record · LEDGER
As of 7 August 2026, this Paper Citation Record lists 45 of 45 outbound references and 8 inbound Pith citation observations for arXiv:2506.10764.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T04:23:50.683069Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-02T14:04:57.314444Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-07-04T20:20:07.360446Z
45 of 45 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 5b6495d8-7106-4e4e-9e68-c534f22f34d1 · outbound
OPT-BENCH: Evaluating LLM Agent on Large-Scale Search Spaces Optimization Problems GPT-4 Technical Report
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 692304e6-190e-4b51-affa-5494d3e16856 · outbound
OPT-BENCH: Evaluating LLM Agent on Large-Scale Search Spaces Optimization Problems MathQA: Towards Interpretable Math Word Problem Solving with Operation-Based Formalisms
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2e969ac0-cb4b-40eb-82a3-f510cb7984d6 · outbound
OPT-BENCH: Evaluating LLM Agent on Large-Scale Search Spaces Optimization Problems Program Synthesis with Large Language Models
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 39138021-4c0d-44eb-af22-74039c8e02bb · outbound
OPT-BENCH: Evaluating LLM Agent on Large-Scale Search Spaces Optimization Problems Language models are few-shot learners.Advances in neural information processing systems, 33:1877–1901, 2020
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 895e5a4d-eca7-4c88-ba1b-70ad41bb08db · outbound
OPT-BENCH: Evaluating LLM Agent on Large-Scale Search Spaces Optimization Problems Evaluating Large Language Models Trained on Code
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d36b7c62-4ba5-4106-b181-07d11b280735 · outbound
OPT-BENCH: Evaluating LLM Agent on Large-Scale Search Spaces Optimization Problems Palm: Scaling language modeling with pathways.Journal of Machine Learning Research, 24(240):1– 113, 2023
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4bbc9d8e-a657-4e40-8e9a-35c1b8918355 · outbound
OPT-BENCH: Evaluating LLM Agent on Large-Scale Search Spaces Optimization Problems Training Verifiers to Solve Math Word Problems
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ac041da5-8e76-4147-b5dc-be0ffe382590 · outbound
OPT-BENCH: Evaluating LLM Agent on Large-Scale Search Spaces Optimization Problems Selection-Inference: Exploiting Large Language Models for Interpretable Logical Reasoning
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 54c5a251-52c7-496e-97db-ec3ed7ebaf2d · outbound
OPT-BENCH: Evaluating LLM Agent on Large-Scale Search Spaces Optimization Problems NPHardEval: Dynamic Benchmark on Reasoning Ability of Large Language Models via Complexity Classes
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a424da53-d40a-46f5-af3b-ef1a0875be4c · outbound
OPT-BENCH: Evaluating LLM Agent on Large-Scale Search Spaces Optimization Problems Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bc03159f-8cdf-4858-901d-19719e22bfa7 · outbound
OPT-BENCH: Evaluating LLM Agent on Large-Scale Search Spaces Optimization Problems The Llama 3 Herd of Models
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7cbf9a37-12b5-4b76-aecd-cc4a7a198749 · outbound
OPT-BENCH: Evaluating LLM Agent on Large-Scale Search Spaces Optimization Problems DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4380e4f1-ea7e-40b0-b55e-c85b8b215dc0 · outbound
OPT-BENCH: Evaluating LLM Agent on Large-Scale Search Spaces Optimization Problems R-Bench: Graduate-level Multi-disciplinary Benchmarks for LLM & MLLM Complex Reasoning Evaluation
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 94684895-c743-40d3-87f1-6e6be6a6bbe9 · outbound
OPT-BENCH: Evaluating LLM Agent on Large-Scale Search Spaces Optimization Problems Measuring Massive Multitask Language Understanding
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5b1a572f-7661-4758-a832-1cfbdf5d1d36 · outbound
OPT-BENCH: Evaluating LLM Agent on Large-Scale Search Spaces Optimization Problems Measuring Mathematical Problem Solving With the MATH Dataset
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 59c6256c-9287-43d1-96eb-183abc89676b · outbound
OPT-BENCH: Evaluating LLM Agent on Large-Scale Search Spaces Optimization Problems Mlagentbench: Evaluating language agents on machine learning experimentation
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 141b4162-76ee-4f17-8f17-b187f4a13aa0 · outbound
OPT-BENCH: Evaluating LLM Agent on Large-Scale Search Spaces Optimization Problems Aide: Ai-driven exploration in the space of code, 2025
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 316f20a2-a8c8-4f4d-923d-612cab304516 · outbound
OPT-BENCH: Evaluating LLM Agent on Large-Scale Search Spaces Optimization Problems Let’s verify step by step
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4ae518c8-e72e-4df0-8705-bd7480778af1 · outbound
OPT-BENCH: Evaluating LLM Agent on Large-Scale Search Spaces Optimization Problems Truthfulqa: Measuring how models mimic human falsehoods
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation c22a6d6e-7175-4207-896c-7be959e993f2 · outbound
OPT-BENCH: Evaluating LLM Agent on Large-Scale Search Spaces Optimization Problems Criticbench: Benchmarking llms for critique-correct reasoning, 2024
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation cbe914bb-a89b-488e-9a4e-23270b03daf5 · outbound
OPT-BENCH: Evaluating LLM Agent on Large-Scale Search Spaces Optimization Problems AgentBench: Evaluating LLMs as Agents
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6fbbee3c-4310-4576-b5dd-86a596378fa4 · outbound
OPT-BENCH: Evaluating LLM Agent on Large-Scale Search Spaces Optimization Problems Self-Refine: Iterative Refinement with Self-Feedback
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b37b2185-8b21-4127-a86d-11fa9a8c23d5 · outbound
OPT-BENCH: Evaluating LLM Agent on Large-Scale Search Spaces Optimization Problems MLE-bench: Evaluating Machine Learning Agents on Machine Learning Engineering
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3ad8f589-dcb4-4b9e-bcbd-43b850d2aae7 · outbound
OPT-BENCH: Evaluating LLM Agent on Large-Scale Search Spaces Optimization Problems Openai o1 system card
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 70add6e7-969c-42c6-85f3-69a81e02bb88 · outbound
OPT-BENCH: Evaluating LLM Agent on Large-Scale Search Spaces Optimization Problems Toolbench: An open platform for training, serving, and evaluating large language models as tool agents.https://github.com/OpenBMB/ToolBench, 2023
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 88b56bbd-2842-4e35-9c4a-b5e20125ad00 · outbound
OPT-BENCH: Evaluating LLM Agent on Large-Scale Search Spaces Optimization Problems Training language models to follow instructions with human feedback.Advances in neural information processing systems, 35:27730–27744, 2022
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation aa9aceab-8e31-46f2-a87d-48e6db534039 · outbound
OPT-BENCH: Evaluating LLM Agent on Large-Scale Search Spaces Optimization Problems ART: Automatic multi-step reasoning and tool-use for large language models
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 04db7ec6-6883-4523-be2d-152247168773 · outbound
OPT-BENCH: Evaluating LLM Agent on Large-Scale Search Spaces Optimization Problems Toolformer: Language Models Can Teach Themselves to Use Tools
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b5a94c9a-ccbe-4db5-ba5d-22ef39cabf80 · outbound
OPT-BENCH: Evaluating LLM Agent on Large-Scale Search Spaces Optimization Problems Reflexion: Language Agents with Verbal Reinforcement Learning
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cde1fcaf-a9ed-4c4d-b2fc-0222ca212c1b · outbound
OPT-BENCH: Evaluating LLM Agent on Large-Scale Search Spaces Optimization Problems ALFWorld: Aligning Text and Embodied Environments for Interactive Learning
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5c483474-2b68-44a9-af0b-9fb159ef37a7 · outbound
OPT-BENCH: Evaluating LLM Agent on Large-Scale Search Spaces Optimization Problems Beyond the Imitation Game: Quantifying and extrapolating the capabilities of language models
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1cedfafc-06fb-410e-80c2-4218bf49f1f6 · outbound
OPT-BENCH: Evaluating LLM Agent on Large-Scale Search Spaces Optimization Problems Commonsenseqa: A question answering challenge targeting commonsense knowledge
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation e0742482-4557-4ffc-802a-7fe44fd554b8 · outbound
OPT-BENCH: Evaluating LLM Agent on Large-Scale Search Spaces Optimization Problems Llama 2: Open Foundation and Fine-Tuned Chat Models
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0ee32718-1a8f-4d23-80ac-5bf23969389c · outbound
OPT-BENCH: Evaluating LLM Agent on Large-Scale Search Spaces Optimization Problems Unresolved cited work
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 77c7a4e8-b25b-4364-bbba-1c434b0551d9 · outbound
OPT-BENCH: Evaluating LLM Agent on Large-Scale Search Spaces Optimization Problems Superglue: A stickier benchmark for general-purpose language understanding systems
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 30ab32ec-0e81-4648-bd65-b0f1709bc8ab · outbound
OPT-BENCH: Evaluating LLM Agent on Large-Scale Search Spaces Optimization Problems Glue: A multi-task benchmark and analysis platform for natural language understanding
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation a3adac98-9a74-4d92-a4ca-8cc85b9b86ed · outbound
OPT-BENCH: Evaluating LLM Agent on Large-Scale Search Spaces Optimization Problems Chain-of-thought prompting elicits reasoning in large language models
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d5d9f157-174c-4510-85da-0e779b3a2316 · outbound
OPT-BENCH: Evaluating LLM Agent on Large-Scale Search Spaces Optimization Problems Are large language models really good logical reasoners? a comprehensive evaluation and beyond.IEEE Transactions on Knowledge and Data Engineering, 2025
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation fe4b2975-5b95-4c5c-9a92-7fd4b40436d2 · outbound
OPT-BENCH: Evaluating LLM Agent on Large-Scale Search Spaces Optimization Problems InterCode: Standardizing and Benchmarking Interactive Coding with Execution Feedback
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 651be0e9-89ba-4c4c-92e1-47304ff082b0 · outbound
OPT-BENCH: Evaluating LLM Agent on Large-Scale Search Spaces Optimization Problems WebShop: Towards Scalable Real-World Web Interaction with Grounded Language Agents
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fea9e1be-8724-4124-8197-19d1b07a18d8 · outbound
OPT-BENCH: Evaluating LLM Agent on Large-Scale Search Spaces Optimization Problems Tree of Thoughts: Deliberate Problem Solving with Large Language Models
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 08e606c5-141b-4ebd-bbb3-d0b2510a6b24 · outbound
OPT-BENCH: Evaluating LLM Agent on Large-Scale Search Spaces Optimization Problems ReAct: Synergizing Reasoning and Acting in Language Models
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 33d97532-bf7c-44fe-8194-c5efb9b6b978 · outbound
OPT-BENCH: Evaluating LLM Agent on Large-Scale Search Spaces Optimization Problems Hellaswag: Can a machine really finish your sentence? InProceedings of the 57th Annual Meeting of the Association for Computational Linguistics, pages 4791–4800, 2019
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 88981e13-795f-4a16-931d-086f52242ac5 · outbound
OPT-BENCH: Evaluating LLM Agent on Large-Scale Search Spaces Optimization Problems Iolbench: Benchmarking llms on linguistic reasoning.arXiv preprint arXiv:2501.04249, 2024
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 50a6d4b1-5579-4ce8-ae5a-68fe93b66e54 · outbound
OPT-BENCH: Evaluating LLM Agent on Large-Scale Search Spaces Optimization Problems WebArena: A Realistic Web Environment for Building Autonomous Agents
Reference 46
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d1d85105-0d9e-478a-b365-51b8874b8d35 · inbound
What Makes an LLM a Good Optimizer? A Trajectory Analysis of LLM-Guided Evolutionary Search OPT-BENCH: Evaluating LLM Agent on Large-Scale Search Spaces Optimization Problems
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation adec540f-a8ac-440d-a6f2-7f273a7efedc · inbound
Forge: Quality-Aware Reinforcement Learning for NP-Hard Optimization in LLMs OPT-BENCH: Evaluating LLM Agent on Large-Scale Search Spaces Optimization Problems
Reference 51
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation d9b968af-82cf-4ae5-90d6-7438f9d985eb · inbound
What and When to Distill: Selective Hindsight Distillation for Multi-Turn Agents OPT-BENCH: Evaluating LLM Agent on Large-Scale Search Spaces Optimization Problems
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 3b67f314-ccd9-4705-bca2-1bc82a5318f9 · inbound
Large Language Models for Operations Research: A Comprehensive Survey OPT-BENCH: Evaluating LLM Agent on Large-Scale Search Spaces Optimization Problems
Reference 174
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 1e605860-bb3a-4446-bdb3-38aadd4f87eb · inbound
Beyond the Current Observation: Evaluating Multimodal Large Language Models in Controllable Non-Markov Games OPT-BENCH: Evaluating LLM Agent on Large-Scale Search Spaces Optimization Problems
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation a99942cd-2771-402e-bc91-67b96f67ca3e · inbound
MiniOpt: Reasoning to Model and Solve General Optimization Problems with Limited Resources OPT-BENCH: Evaluating LLM Agent on Large-Scale Search Spaces Optimization Problems
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 0d3bf3f0-e0b5-4632-93cb-a063009b55b1 · inbound
MiniOpt: Reasoning to Model and Solve General Optimization Problems with Limited Resources OPT-BENCH: Evaluating LLM Agent on Large-Scale Search Spaces Optimization Problems
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation a9ffee4b-2b47-4112-8e2e-8d49e0cf63f5 · inbound
PEARL: Solver-in-the-Loop Interactive Optimization Modeling from Natural Language OPT-BENCH: Evaluating LLM Agent on Large-Scale Search Spaces Optimization Problems
Reference 77
Source-reported events for the cited work
Unavailable: canonical work link unavailable.