Pith. sign in

Paper Citation Record · LEDGER

OPT-BENCH: Evaluating LLM Agent on Large-Scale Search Spaces Optimization Problems

As of 7 August 2026, this Paper Citation Record lists 45 of 45 outbound references and 8 inbound Pith citation observations for arXiv:2506.10764.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.10764 v1

Coverage vector

measured 45 of 45 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T04:23:50.683069Z

measured 53 of 53 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 8 of 8 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-02T14:04:57.314444Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T20:20:07.360446Z

Reference resolution

45 of 45 outbound references displayed

  • verified exact0
  • verified fuzzy8
  • unresolved37
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 5b6495d8-7106-4e4e-9e68-c534f22f34d1 · outbound

This paper cites GPT-4 Technical Report.

OPT-BENCH: Evaluating LLM Agent on Large-Scale Search Spaces Optimization Problems GPT-4 Technical Report

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T04:23:50.543505Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:23:50.543505Z digest=sha256:52fe921d4e17ac3576c29ffce30e4b0be23052d14fa7325160eeea9801810c70

Observation 692304e6-190e-4b51-affa-5494d3e16856 · outbound

This paper cites MathQA: Towards Interpretable Math Word Problem Solving with Operation-Based Formalisms.

OPT-BENCH: Evaluating LLM Agent on Large-Scale Search Spaces Optimization Problems MathQA: Towards Interpretable Math Word Problem Solving with Operation-Based Formalisms

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T04:23:50.547360Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:23:50.547360Z digest=sha256:3bc9d08000893c93220bb5e9bd7e52f975573fbe9b6e503d5f2e0575229db23b

Observation 2e969ac0-cb4b-40eb-82a3-f510cb7984d6 · outbound

This paper cites Program Synthesis with Large Language Models.

OPT-BENCH: Evaluating LLM Agent on Large-Scale Search Spaces Optimization Problems Program Synthesis with Large Language Models

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T04:23:50.551498Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:23:50.551498Z digest=sha256:d13904b60ed49e11a4ddaa6b76e8b8bde81b5e58ebc97457a7b9687761c9fff5

Observation 39138021-4c0d-44eb-af22-74039c8e02bb · outbound

This paper cites Language models are few-shot learners.Advances in neural information processing systems, 33:1877–1901, 2020.

OPT-BENCH: Evaluating LLM Agent on Large-Scale Search Spaces Optimization Problems Language models are few-shot learners.Advances in neural information processing systems, 33:1877–1901, 2020

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T04:23:50.554887Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:23:50.554887Z digest=sha256:54aa407ccb3a2649cb30a21a6cfe16fce7bba6981c231baf7b2386278cd9bcd9

Observation 895e5a4d-eca7-4c88-ba1b-70ad41bb08db · outbound

This paper cites Evaluating Large Language Models Trained on Code.

OPT-BENCH: Evaluating LLM Agent on Large-Scale Search Spaces Optimization Problems Evaluating Large Language Models Trained on Code

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T04:23:50.561890Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:23:50.561890Z digest=sha256:e3233b51dc51a7eafcf3d60f599c5d5c7845533c9fd6a0459c468b8fb54ca13e

Observation d36b7c62-4ba5-4106-b181-07d11b280735 · outbound

This paper cites Palm: Scaling language modeling with pathways.Journal of Machine Learning Research, 24(240):1– 113, 2023.

OPT-BENCH: Evaluating LLM Agent on Large-Scale Search Spaces Optimization Problems Palm: Scaling language modeling with pathways.Journal of Machine Learning Research, 24(240):1– 113, 2023

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T04:23:50.565288Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:23:50.565288Z digest=sha256:dbd87da8fc6770db88fa3d980204fd5d9df6c576f40bfd762b2f55f3719e8ead

Observation 4bbc9d8e-a657-4e40-8e9a-35c1b8918355 · outbound

This paper cites Training Verifiers to Solve Math Word Problems.

OPT-BENCH: Evaluating LLM Agent on Large-Scale Search Spaces Optimization Problems Training Verifiers to Solve Math Word Problems

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T04:23:50.568987Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:23:50.568987Z digest=sha256:65e99df08bc145d033c39c121183124a3435c64d5eedd5d9c3802becd5ec886d

Observation ac041da5-8e76-4147-b5dc-be0ffe382590 · outbound

This paper cites Selection-Inference: Exploiting Large Language Models for Interpretable Logical Reasoning.

OPT-BENCH: Evaluating LLM Agent on Large-Scale Search Spaces Optimization Problems Selection-Inference: Exploiting Large Language Models for Interpretable Logical Reasoning

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T04:23:50.571999Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:23:50.571999Z digest=sha256:cf809d64309f730b5036393225dba615b99f6d0760d2576b4ee05615edb0a331

Observation 54c5a251-52c7-496e-97db-ec3ed7ebaf2d · outbound

This paper cites NPHardEval: Dynamic Benchmark on Reasoning Ability of Large Language Models via Complexity Classes.

OPT-BENCH: Evaluating LLM Agent on Large-Scale Search Spaces Optimization Problems NPHardEval: Dynamic Benchmark on Reasoning Ability of Large Language Models via Complexity Classes

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T04:23:50.574907Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:23:50.574907Z digest=sha256:f17866fe6c032a09ef3cc0da91e617f5128d9e749768d2380e6f8a971ea3f52b

Observation a424da53-d40a-46f5-af3b-ef1a0875be4c · outbound

This paper cites Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context.

OPT-BENCH: Evaluating LLM Agent on Large-Scale Search Spaces Optimization Problems Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T04:23:50.577993Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:23:50.577993Z digest=sha256:d5be56b860764f79d5ca2c67bd6a38e69d8f4759b747c206b1047ee88231f0b4

Observation bc03159f-8cdf-4858-901d-19719e22bfa7 · outbound

This paper cites The Llama 3 Herd of Models.

OPT-BENCH: Evaluating LLM Agent on Large-Scale Search Spaces Optimization Problems The Llama 3 Herd of Models

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T04:23:50.580934Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:23:50.580934Z digest=sha256:902689bc77c5a8683e252d86f4eeffda8b033f93c98fa72a46e8799073edaca9

Observation 7cbf9a37-12b5-4b76-aecd-cc4a7a198749 · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

OPT-BENCH: Evaluating LLM Agent on Large-Scale Search Spaces Optimization Problems DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T04:23:50.584251Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:23:50.584251Z digest=sha256:7d8cd64489dea77b2e56da51bbe4c4787a11c98a7487e37430738234aa0c65dd

Observation 4380e4f1-ea7e-40b0-b55e-c85b8b215dc0 · outbound

This paper cites R-Bench: Graduate-level Multi-disciplinary Benchmarks for LLM & MLLM Complex Reasoning Evaluation.

OPT-BENCH: Evaluating LLM Agent on Large-Scale Search Spaces Optimization Problems R-Bench: Graduate-level Multi-disciplinary Benchmarks for LLM & MLLM Complex Reasoning Evaluation

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T04:23:50.587173Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:23:50.587173Z digest=sha256:9565efd27738e2ab535a4f2ec7310a9b04270525d45beda7ec6842c48fc1074c

Observation 94684895-c743-40d3-87f1-6e6be6a6bbe9 · outbound

This paper cites Measuring Massive Multitask Language Understanding.

OPT-BENCH: Evaluating LLM Agent on Large-Scale Search Spaces Optimization Problems Measuring Massive Multitask Language Understanding

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T04:23:50.590207Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:23:50.590207Z digest=sha256:af7e9629870284e301f4c896ef27a095f42365a23433aac17ad5e30714d47942

Observation 5b1a572f-7661-4758-a832-1cfbdf5d1d36 · outbound

This paper cites Measuring Mathematical Problem Solving With the MATH Dataset.

OPT-BENCH: Evaluating LLM Agent on Large-Scale Search Spaces Optimization Problems Measuring Mathematical Problem Solving With the MATH Dataset

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T04:23:50.593077Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:23:50.593077Z digest=sha256:6becde493d103f558313673d9cefe9a912b372c9055a93df2f70d328425b0d4d

Observation 59c6256c-9287-43d1-96eb-183abc89676b · outbound

This paper cites Mlagentbench: Evaluating language agents on machine learning experimentation.

OPT-BENCH: Evaluating LLM Agent on Large-Scale Search Spaces Optimization Problems Mlagentbench: Evaluating language agents on machine learning experimentation

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T04:23:50.596213Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:23:50.596213Z digest=sha256:cbd5cf1c0c9a4cf4ba4acecaaa40bab85a06fde75642d4b17d0ca8b28449943e

Observation 141b4162-76ee-4f17-8f17-b187f4a13aa0 · outbound

This paper cites Aide: Ai-driven exploration in the space of code, 2025.

OPT-BENCH: Evaluating LLM Agent on Large-Scale Search Spaces Optimization Problems Aide: Ai-driven exploration in the space of code, 2025

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T04:23:50.599339Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:23:50.599339Z digest=sha256:ee114aefece04efbfb136a6496e3564df2ab3fca3ceba3486c541624bb3c0cc3

Observation 316f20a2-a8c8-4f4d-923d-612cab304516 · outbound

This paper cites Let’s verify step by step.

OPT-BENCH: Evaluating LLM Agent on Large-Scale Search Spaces Optimization Problems Let’s verify step by step

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T04:23:50.602556Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:23:50.602556Z digest=sha256:3c39eaae863da5d8dddfa6a34ad86cc89ac5405016a6af5ab798870aaacaa3f3

Observation 4ae518c8-e72e-4df0-8705-bd7480778af1 · outbound

This paper cites Truthfulqa: Measuring how models mimic human falsehoods.

OPT-BENCH: Evaluating LLM Agent on Large-Scale Search Spaces Optimization Problems Truthfulqa: Measuring how models mimic human falsehoods

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:23:51.237836Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T04:23:50.605362Z digest=sha256:23e6a205b154f59f4763c35c45cceaa27609edd0d4cc39a653f25fe5dbfa53b2

Observation c22a6d6e-7175-4207-896c-7be959e993f2 · outbound

This paper cites Criticbench: Benchmarking llms for critique-correct reasoning, 2024.

OPT-BENCH: Evaluating LLM Agent on Large-Scale Search Spaces Optimization Problems Criticbench: Benchmarking llms for critique-correct reasoning, 2024

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:23:51.227353Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T04:23:50.608019Z digest=sha256:d3202107fc80e6cfddce3b2b15720f52e413a5a70ff1a1676a95ef9bbf0466bf

Observation cbe914bb-a89b-488e-9a4e-23270b03daf5 · outbound

This paper cites AgentBench: Evaluating LLMs as Agents.

OPT-BENCH: Evaluating LLM Agent on Large-Scale Search Spaces Optimization Problems AgentBench: Evaluating LLMs as Agents

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T04:23:50.610680Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:23:50.610680Z digest=sha256:4eb0faf3c2701bade9623dce452864051147ccea39b18f507061acbae0d314b4

Observation 6fbbee3c-4310-4576-b5dd-86a596378fa4 · outbound

This paper cites Self-Refine: Iterative Refinement with Self-Feedback.

OPT-BENCH: Evaluating LLM Agent on Large-Scale Search Spaces Optimization Problems Self-Refine: Iterative Refinement with Self-Feedback

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T04:23:50.613897Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:23:50.613897Z digest=sha256:8bcd9a60617f8164f050212bbf9b610fa45bce32220e09a64c3045a55ea828bb

Observation b37b2185-8b21-4127-a86d-11fa9a8c23d5 · outbound

This paper cites MLE-bench: Evaluating Machine Learning Agents on Machine Learning Engineering.

OPT-BENCH: Evaluating LLM Agent on Large-Scale Search Spaces Optimization Problems MLE-bench: Evaluating Machine Learning Agents on Machine Learning Engineering

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T04:23:50.616857Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:23:50.616857Z digest=sha256:c269ae78e8f1ba5c80be0e285089dcf9a6c5d531a04dc54b2f081ef2cb9026ea

Observation 3ad8f589-dcb4-4b9e-bcbd-43b850d2aae7 · outbound

This paper cites Openai o1 system card.

OPT-BENCH: Evaluating LLM Agent on Large-Scale Search Spaces Optimization Problems Openai o1 system card

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:23:51.217387Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T04:23:50.619574Z digest=sha256:dbc8280307cd7f1f2497ed9a191165c493823e07fa223db326b9af574d11d8b6

Observation 70add6e7-969c-42c6-85f3-69a81e02bb88 · outbound

This paper cites Toolbench: An open platform for training, serving, and evaluating large language models as tool agents.https://github.com/OpenBMB/ToolBench, 2023.

OPT-BENCH: Evaluating LLM Agent on Large-Scale Search Spaces Optimization Problems Toolbench: An open platform for training, serving, and evaluating large language models as tool agents.https://github.com/OpenBMB/ToolBench, 2023

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:23:51.207456Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T04:23:50.622398Z digest=sha256:38c848cf56f4c66b457baedc3bf1d8b069e13b2a103878d7a24d5370c5d6d3e4

Observation 88b56bbd-2842-4e35-9c4a-b5e20125ad00 · outbound

This paper cites Training language models to follow instructions with human feedback.Advances in neural information processing systems, 35:27730–27744, 2022.

OPT-BENCH: Evaluating LLM Agent on Large-Scale Search Spaces Optimization Problems Training language models to follow instructions with human feedback.Advances in neural information processing systems, 35:27730–27744, 2022

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T04:23:50.625582Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:23:50.625582Z digest=sha256:aaf0e39006307b33e55f5ffc011fab2efa61905d09aef061bd3189280604a762

Observation aa9aceab-8e31-46f2-a87d-48e6db534039 · outbound

This paper cites ART: Automatic multi-step reasoning and tool-use for large language models.

OPT-BENCH: Evaluating LLM Agent on Large-Scale Search Spaces Optimization Problems ART: Automatic multi-step reasoning and tool-use for large language models

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T04:23:50.628403Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:23:50.628403Z digest=sha256:8c304add29a56bda87fa3267ddbf0d2544fbb5b79ad91b71331c82f69d74f8d7

Observation 04db7ec6-6883-4523-be2d-152247168773 · outbound

This paper cites Toolformer: Language Models Can Teach Themselves to Use Tools.

OPT-BENCH: Evaluating LLM Agent on Large-Scale Search Spaces Optimization Problems Toolformer: Language Models Can Teach Themselves to Use Tools

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T04:23:50.631361Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:23:50.631361Z digest=sha256:8ad4c64f4bf7a1440631701e76da9c1620265c915a7d8f162038ceb1a1ff92fa

Observation b5a94c9a-ccbe-4db5-ba5d-22ef39cabf80 · outbound

This paper cites Reflexion: Language Agents with Verbal Reinforcement Learning.

OPT-BENCH: Evaluating LLM Agent on Large-Scale Search Spaces Optimization Problems Reflexion: Language Agents with Verbal Reinforcement Learning

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T04:23:50.634456Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:23:50.634456Z digest=sha256:68b5016b382a02668ccf6079cb9b426bb1ed3257551cf3cc0a1402e6d071cb81

Observation cde1fcaf-a9ed-4c4d-b2fc-0222ca212c1b · outbound

This paper cites ALFWorld: Aligning Text and Embodied Environments for Interactive Learning.

OPT-BENCH: Evaluating LLM Agent on Large-Scale Search Spaces Optimization Problems ALFWorld: Aligning Text and Embodied Environments for Interactive Learning

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T04:23:50.637541Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:23:50.637541Z digest=sha256:63df654f6c4c64a00c00cef884ee8e8a0a46c13893ed2f8d3dba0077d7f02df4

Observation 5c483474-2b68-44a9-af0b-9fb159ef37a7 · outbound

This paper cites Beyond the Imitation Game: Quantifying and extrapolating the capabilities of language models.

OPT-BENCH: Evaluating LLM Agent on Large-Scale Search Spaces Optimization Problems Beyond the Imitation Game: Quantifying and extrapolating the capabilities of language models

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T04:23:50.640693Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:23:50.640693Z digest=sha256:6d2dcab5e5c4fa6a92447962ba568cc575c684a1768b8dff4216405d1c6952d0

Observation 1cedfafc-06fb-410e-80c2-4218bf49f1f6 · outbound

This paper cites Commonsenseqa: A question answering challenge targeting commonsense knowledge.

OPT-BENCH: Evaluating LLM Agent on Large-Scale Search Spaces Optimization Problems Commonsenseqa: A question answering challenge targeting commonsense knowledge

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:23:51.192086Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T04:23:50.643779Z digest=sha256:72b62158af08090e0088ff3a9dd97e43478879e26eae00668b8a4e2fdb282767

Observation e0742482-4557-4ffc-802a-7fe44fd554b8 · outbound

This paper cites Llama 2: Open Foundation and Fine-Tuned Chat Models.

OPT-BENCH: Evaluating LLM Agent on Large-Scale Search Spaces Optimization Problems Llama 2: Open Foundation and Fine-Tuned Chat Models

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T04:23:50.646976Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:23:50.646976Z digest=sha256:e3540a2e3f7c1b9f381bcceb79cbfceeb561aba1ec52cb71ab3838711312d987

Observation 0ee32718-1a8f-4d23-80ac-5bf23969389c · outbound

This paper cites an unresolved cited work.

OPT-BENCH: Evaluating LLM Agent on Large-Scale Search Spaces Optimization Problems Unresolved cited work

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T04:23:50.650424Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:23:50.650424Z digest=sha256:15210f4ce1d53d24b22fa529d553c19eb51ed7501d454083e7358ea70ef9b2bd

Observation 77c7a4e8-b25b-4364-bbba-1c434b0551d9 · outbound

This paper cites Superglue: A stickier benchmark for general-purpose language understanding systems.

OPT-BENCH: Evaluating LLM Agent on Large-Scale Search Spaces Optimization Problems Superglue: A stickier benchmark for general-purpose language understanding systems

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:23:51.175816Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T04:23:50.653556Z digest=sha256:e5ac73c47afa9dfa0bd2ff2fb5128ef8a1413f6e237a2431c1489c8178a0af40

Observation 30ab32ec-0e81-4648-bd65-b0f1709bc8ab · outbound

This paper cites Glue: A multi-task benchmark and analysis platform for natural language understanding.

OPT-BENCH: Evaluating LLM Agent on Large-Scale Search Spaces Optimization Problems Glue: A multi-task benchmark and analysis platform for natural language understanding

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:23:51.165231Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T04:23:50.656382Z digest=sha256:1c11e565bcb4e7034820ee1c2fd929d0cb48643d693fb84231ab312d938fdc4b

Observation a3adac98-9a74-4d92-a4ca-8cc85b9b86ed · outbound

This paper cites Chain-of-thought prompting elicits reasoning in large language models.

OPT-BENCH: Evaluating LLM Agent on Large-Scale Search Spaces Optimization Problems Chain-of-thought prompting elicits reasoning in large language models

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-07T04:23:50.659110Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:23:50.659110Z digest=sha256:b9854e7bc9c1c09094fa93a5162476e2b6930418906fadf9f97de179318fe9dc

Observation d5d9f157-174c-4510-85da-0e779b3a2316 · outbound

This paper cites Are large language models really good logical reasoners? a comprehensive evaluation and beyond.IEEE Transactions on Knowledge and Data Engineering, 2025.

OPT-BENCH: Evaluating LLM Agent on Large-Scale Search Spaces Optimization Problems Are large language models really good logical reasoners? a comprehensive evaluation and beyond.IEEE Transactions on Knowledge and Data Engineering, 2025

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:23:51.147672Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T04:23:50.661863Z digest=sha256:30d06649c09b68e06bcb6e17a37e27065c1f5c8de131411d34d8110b812c2bec

Observation fe4b2975-5b95-4c5c-9a92-7fd4b40436d2 · outbound

This paper cites InterCode: Standardizing and Benchmarking Interactive Coding with Execution Feedback.

OPT-BENCH: Evaluating LLM Agent on Large-Scale Search Spaces Optimization Problems InterCode: Standardizing and Benchmarking Interactive Coding with Execution Feedback

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-07T04:23:50.664728Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:23:50.664728Z digest=sha256:8fc024dabd19ddcc6f66fa3ce61204ea41b0735018d9b8b92207d582e80289d8

Observation 651be0e9-89ba-4c4c-92e1-47304ff082b0 · outbound

This paper cites WebShop: Towards Scalable Real-World Web Interaction with Grounded Language Agents.

OPT-BENCH: Evaluating LLM Agent on Large-Scale Search Spaces Optimization Problems WebShop: Towards Scalable Real-World Web Interaction with Grounded Language Agents

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-07T04:23:50.667812Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:23:50.667812Z digest=sha256:095a5c37458918d4f037073ce67a1af6b1755dbef02a28cd98486a969803832e

Observation fea9e1be-8724-4124-8197-19d1b07a18d8 · outbound

This paper cites Tree of Thoughts: Deliberate Problem Solving with Large Language Models.

OPT-BENCH: Evaluating LLM Agent on Large-Scale Search Spaces Optimization Problems Tree of Thoughts: Deliberate Problem Solving with Large Language Models

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-07T04:23:50.671193Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:23:50.671193Z digest=sha256:591dac928712e823330e2e52446019b226d9c0f220c655606dceef65189816ea

Observation 08e606c5-141b-4ebd-bbb3-d0b2510a6b24 · outbound

This paper cites ReAct: Synergizing Reasoning and Acting in Language Models.

OPT-BENCH: Evaluating LLM Agent on Large-Scale Search Spaces Optimization Problems ReAct: Synergizing Reasoning and Acting in Language Models

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-07T04:23:50.674463Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:23:50.674463Z digest=sha256:c92bfb0242311168ee34248a7e4c44059dd7e2858750817ce68d42495cdcfe0a

Observation 33d97532-bf7c-44fe-8194-c5efb9b6b978 · outbound

This paper cites Hellaswag: Can a machine really finish your sentence? InProceedings of the 57th Annual Meeting of the Association for Computational Linguistics, pages 4791–4800, 2019.

OPT-BENCH: Evaluating LLM Agent on Large-Scale Search Spaces Optimization Problems Hellaswag: Can a machine really finish your sentence? InProceedings of the 57th Annual Meeting of the Association for Computational Linguistics, pages 4791–4800, 2019

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-07T04:23:50.677471Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:23:50.677471Z digest=sha256:fb51e58e7099277a1c727d789728ef943567f305a60434b863f2aed0900b9639

Observation 88981e13-795f-4a16-931d-086f52242ac5 · outbound

This paper cites Iolbench: Benchmarking llms on linguistic reasoning.arXiv preprint arXiv:2501.04249, 2024.

OPT-BENCH: Evaluating LLM Agent on Large-Scale Search Spaces Optimization Problems Iolbench: Benchmarking llms on linguistic reasoning.arXiv preprint arXiv:2501.04249, 2024

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-07T04:23:50.680320Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:23:50.680320Z digest=sha256:c3ed1b93894c8849b2887c439c703cef05ea5b3b56816ff812db7d050fe95b12

Observation 50a6d4b1-5579-4ce8-ae5a-68fe93b66e54 · outbound

This paper cites WebArena: A Realistic Web Environment for Building Autonomous Agents.

OPT-BENCH: Evaluating LLM Agent on Large-Scale Search Spaces Optimization Problems WebArena: A Realistic Web Environment for Building Autonomous Agents

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-07T04:23:50.683069Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:23:50.683069Z digest=sha256:9d91a03fa6596442eda13c19f14b1336b641e10a02ae81645fcdb707fa98e1a6

Pith citing papers

Observation d1d85105-0d9e-478a-b365-51b8874b8d35 · inbound

What Makes an LLM a Good Optimizer? A Trajectory Analysis of LLM-Guided Evolutionary Search cites this paper.

What Makes an LLM a Good Optimizer? A Trajectory Analysis of LLM-Guided Evolutionary Search OPT-BENCH: Evaluating LLM Agent on Large-Scale Search Spaces Optimization Problems

Reference 4

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T13:06:05.611346Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T02:19:18.121220Z digest=sha256:0fe45ad92a1b76fef313394c262a8c2e6a27f47af43a628cd47723f5b6932989

Observation adec540f-a8ac-440d-a6f2-7f273a7efedc · inbound

Forge: Quality-Aware Reinforcement Learning for NP-Hard Optimization in LLMs cites this paper.

Forge: Quality-Aware Reinforcement Learning for NP-Hard Optimization in LLMs OPT-BENCH: Evaluating LLM Agent on Large-Scale Search Spaces Optimization Problems

Reference 51

Resolution
verified exact
arxiv_id, observed 2026-05-12T02:46:18.873851Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-12T02:44:33.143247Z digest=sha256:1c8cfce763e9af93bb17ac78d5233b55a9439031dcb11b3fca3b432cfce531ed

Observation d9b968af-82cf-4ae5-90d6-7438f9d985eb · inbound

What and When to Distill: Selective Hindsight Distillation for Multi-Turn Agents cites this paper.

What and When to Distill: Selective Hindsight Distillation for Multi-Turn Agents OPT-BENCH: Evaluating LLM Agent on Large-Scale Search Spaces Optimization Problems

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-05-20T05:43:05.828295Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-20T05:41:23.712146Z digest=sha256:e1926183d2e0c702e5f38090609b284f262f739abff1dd7b9c9dc859a97de184

Observation 3b67f314-ccd9-4705-bca2-1bc82a5318f9 · inbound

Large Language Models for Operations Research: A Comprehensive Survey cites this paper.

Large Language Models for Operations Research: A Comprehensive Survey OPT-BENCH: Evaluating LLM Agent on Large-Scale Search Spaces Optimization Problems

Reference 174

Resolution
verified exact
arxiv_id, observed 2026-05-21T03:59:32.604852Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-21T03:56:29.983335Z digest=sha256:7b3dc0898d5b7b21e7b888b714416238510f85485ce4ddfbb424fa1284ec6180

Observation 1e605860-bb3a-4446-bdb3-38aadd4f87eb · inbound

Beyond the Current Observation: Evaluating Multimodal Large Language Models in Controllable Non-Markov Games cites this paper.

Beyond the Current Observation: Evaluating Multimodal Large Language Models in Controllable Non-Markov Games OPT-BENCH: Evaluating LLM Agent on Large-Scale Search Spaces Optimization Problems

Reference 38

Resolution
verified exact
arxiv_id, observed 2026-07-04T00:19:13.672489Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-26T21:17:02.332687Z digest=sha256:64ea1b1980ea89f32bd74f3c3911ff37c4962c8ef7171bcd559c79a0334ee9ad

Observation a99942cd-2771-402e-bc91-67b96f67ca3e · inbound

MiniOpt: Reasoning to Model and Solve General Optimization Problems with Limited Resources cites this paper.

MiniOpt: Reasoning to Model and Solve General Optimization Problems with Limited Resources OPT-BENCH: Evaluating LLM Agent on Large-Scale Search Spaces Optimization Problems

Reference 15

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T20:20:07.361970Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-25T20:19:30.720291Z digest=sha256:dd889c57418d7a4dc99687781eb11aa3cc5eddf594307c31cff0c602406ea6c1

Observation 0d3bf3f0-e0b5-4632-93cb-a063009b55b1 · inbound

MiniOpt: Reasoning to Model and Solve General Optimization Problems with Limited Resources cites this paper.

MiniOpt: Reasoning to Model and Solve General Optimization Problems with Limited Resources OPT-BENCH: Evaluating LLM Agent on Large-Scale Search Spaces Optimization Problems

Reference 15

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T13:19:50.751708Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-26T05:18:55.074710Z digest=sha256:c07cff030c4052faec50a116283debe15cbbcdad7b6b14e17dd3635de67d0a80

Observation a9ffee4b-2b47-4112-8e2e-8d49e0cf63f5 · inbound

PEARL: Solver-in-the-Loop Interactive Optimization Modeling from Natural Language cites this paper.

PEARL: Solver-in-the-Loop Interactive Optimization Modeling from Natural Language OPT-BENCH: Evaluating LLM Agent on Large-Scale Search Spaces Optimization Problems

Reference 77

Resolution
unresolved
no resolver link, observed 2026-08-02T14:04:57.314444Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T14:04:57.314444Z digest=sha256:50dc78f191863f6a4f61c84076e0e5a5f59a6e894e6a0e109324e1659e1f3145