Pith. sign in

Paper Citation Record · LEDGER

EdgeBench: Unveiling Scaling Laws of Learning from Real-World Environments

As of 23 August 2026, this Paper Citation Record lists 87 of 87 outbound references and 5 inbound Pith citation observations for arXiv:2607.05155.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2607.05155 v1

Coverage vector

measured 87 of 87 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-07-11T07:57:43.000834Z

measured 92 of 92 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-23T06:30:58.430688+00:00

measured 5 of 5 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-12T15:18:12.343176Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-05T13:09:10.342277Z

Reference resolution

87 of 87 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved86
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation a4636831-ba28-49ec-9420-9d9004204895 · outbound

This paper cites an unresolved cited work.

EdgeBench: Unveiling Scaling Laws of Learning from Real-World Environments Unresolved cited work

Reference 1

Resolution
unresolved
no resolver link, observed 2026-07-11T07:57:43.000834Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T07:57:43.000834Z digest=sha256:cc4e2f31965b5f00e2df7ed14f887aad0e126e882eec38550093c91bb2438152

Observation 732f6555-3981-44dd-a3c0-9458b0301f07 · outbound

This paper cites The Claude Model Family.

EdgeBench: Unveiling Scaling Laws of Learning from Real-World Environments The Claude Model Family

Reference 2

Resolution
unresolved
no resolver link, observed 2026-07-11T07:57:43.000834Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T07:57:43.000834Z digest=sha256:4eb6b7b5c78c39939f61795623029c8effe9185677f949647b0a5f6feeb84f90

Observation a0b3b172-b6e6-4bca-96a0-08bda14eb0cf · outbound

This paper cites Claude Opus 4.8 System Card.

EdgeBench: Unveiling Scaling Laws of Learning from Real-World Environments Claude Opus 4.8 System Card

Reference 3

Resolution
unresolved
no resolver link, observed 2026-07-11T07:57:43.000834Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T07:57:43.000834Z digest=sha256:56a08670c95ddb51802f5f1e6fd3cbc87ed289a8c713c47895315513de24cbab

Observation d24d591d-37b8-4184-a985-90f903ff20e6 · outbound

This paper cites Continual Learning Bench: Evaluating Frontier AI Systems in Real-World Stateful Environments.

EdgeBench: Unveiling Scaling Laws of Learning from Real-World Environments Continual Learning Bench: Evaluating Frontier AI Systems in Real-World Stateful Environments

Reference 4

Resolution
unresolved
no resolver link, observed 2026-07-11T07:57:43.000834Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T07:57:43.000834Z digest=sha256:143e7cbf440da6d76392115eef74e47b613c1a22f0a0e6cbdd5f1ca4adc2c01b

Observation 1bbf5100-911c-4bee-9da9-096f27ec977c · outbound

This paper cites Self-organized criticality: An explanation of 1/f noise.Physical Review Letters, 59(4):381–384, 1987.

EdgeBench: Unveiling Scaling Laws of Learning from Real-World Environments Self-organized criticality: An explanation of 1/f noise.Physical Review Letters, 59(4):381–384, 1987

Reference 5

Resolution
unresolved
no resolver link, observed 2026-07-11T07:57:43.000834Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T07:57:43.000834Z digest=sha256:6765fe354b6bedbf168cc8b76eb4141e42fe11477210e1a906ee0b243f26f8cc

Observation f01fb070-a75a-4bc0-bdcc-27117a7f6acd · outbound

This paper cites Application of the logistic function to bio-assay.Journal of the American Statistical Association, 39(227):357–365, 1944.

EdgeBench: Unveiling Scaling Laws of Learning from Real-World Environments Application of the logistic function to bio-assay.Journal of the American Statistical Association, 39(227):357–365, 1944

Reference 6

Resolution
unresolved
no resolver link, observed 2026-07-11T07:57:43.000834Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T07:57:43.000834Z digest=sha256:284c354748e592c5724db937e35ab6b0fcedc301416fe66328f6ca388e350d07

Observation b178965b-c5bf-4417-a2ba-00ad557826a1 · outbound

This paper cites Establishing Task Scaling Laws via Compute-Efficient Model Ladders.

EdgeBench: Unveiling Scaling Laws of Learning from Real-World Environments Establishing Task Scaling Laws via Compute-Efficient Model Ladders

Reference 7

Resolution
unresolved
no resolver link, observed 2026-07-11T07:57:43.000834Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T07:57:43.000834Z digest=sha256:320e48673b5c8be764ff5434c93b636da9cfffb585ae33c7d91a481fef8d8766

Observation 112643a7-dff6-405a-b733-501d50eedcaa · outbound

This paper cites an unresolved cited work.

EdgeBench: Unveiling Scaling Laws of Learning from Real-World Environments Unresolved cited work

Reference 8

Resolution
unresolved
no resolver link, observed 2026-07-11T07:57:43.000834Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T07:57:43.000834Z digest=sha256:61e6c5c8dbdbd0f1a55380a5da7f98e5ef88322caaa54d7225e0839f8f8934d3

Observation b01ab9ac-10cd-4952-82a4-cc692c1c6142 · outbound

This paper cites Large Language Monkeys: Scaling Inference Compute with Repeated Sampling.

EdgeBench: Unveiling Scaling Laws of Learning from Real-World Environments Large Language Monkeys: Scaling Inference Compute with Repeated Sampling

Reference 9

Resolution
unresolved
no resolver link, observed 2026-07-11T07:57:43.000834Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T07:57:43.000834Z digest=sha256:3f70fed3b05d8912bb7420b8c5999f3d6d622ca411b5cb1a55db776a8dcb5915

Observation 57fd2404-5238-4411-93f4-e2cb70846ef4 · outbound

This paper cites MLE-bench: Evaluating machine learning agents on machine learning engineering.

EdgeBench: Unveiling Scaling Laws of Learning from Real-World Environments MLE-bench: Evaluating machine learning agents on machine learning engineering

Reference 10

Resolution
unresolved
no resolver link, observed 2026-07-11T07:57:43.000834Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T07:57:43.000834Z digest=sha256:9d6b87baed5377afcbc374e0f1f5355203bd0def897d8589b48b5392b0b50b13

Observation 6cfb9ad2-ae60-4a8b-9e9a-3d0287324c71 · outbound

This paper cites Evaluating Large Language Models Trained on Code.

EdgeBench: Unveiling Scaling Laws of Learning from Real-World Environments Evaluating Large Language Models Trained on Code

Reference 11

Resolution
unresolved
no resolver link, observed 2026-07-11T07:57:43.000834Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T07:57:43.000834Z digest=sha256:630feddfeefe09591eae5f723a30368a85f99408b38fbc04901dde898195ad2b

Observation 69d17f94-8c52-4124-84ea-a44724649fb1 · outbound

This paper cites ScienceAgentBench: Toward rigorous assessment of language agents for data-driven scientific discovery.

EdgeBench: Unveiling Scaling Laws of Learning from Real-World Environments ScienceAgentBench: Toward rigorous assessment of language agents for data-driven scientific discovery

Reference 12

Resolution
unresolved
no resolver link, observed 2026-07-11T07:57:43.000834Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T07:57:43.000834Z digest=sha256:e8f79a056012fa26eba695a66609908ac9d2d97e9a649d25f62c0d00396793a3

Observation 865a234f-fb3e-4033-adbf-de6358e25840 · outbound

This paper cites LLF-Bench: Benchmark for Interactive Learning from Language Feedback.

EdgeBench: Unveiling Scaling Laws of Learning from Real-World Environments LLF-Bench: Benchmark for Interactive Learning from Language Feedback

Reference 13

Resolution
unresolved
no resolver link, observed 2026-07-11T07:57:43.000834Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T07:57:43.000834Z digest=sha256:dedc6b8ce22c7dec7da38a6eeda22b5c021509f03eaf2057a93f24946d29763e

Observation 6881fdaa-50e7-4907-be81-cb9d72c33837 · outbound

This paper cites Frontier-Eng: Benchmarking Self-Evolving Agents on Real-World Engineering Tasks with Generative Optimization.

EdgeBench: Unveiling Scaling Laws of Learning from Real-World Environments Frontier-Eng: Benchmarking Self-Evolving Agents on Real-World Engineering Tasks with Generative Optimization

Reference 14

Resolution
unresolved
no resolver link, observed 2026-07-11T07:57:43.000834Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T07:57:43.000834Z digest=sha256:4794f61325af1980ebf6176ec2f76ce041e11db12c7b988729be44bed3a51e9f

Observation fcda4993-f6c6-4743-80ac-dd446547cbe4 · outbound

This paper cites FrontierSWE: Benchmarking coding agents at the limits of human abilities.https://www.frontierswe.com/blog, 2026.

EdgeBench: Unveiling Scaling Laws of Learning from Real-World Environments FrontierSWE: Benchmarking coding agents at the limits of human abilities.https://www.frontierswe.com/blog, 2026

Reference 15

Resolution
unresolved
no resolver link, observed 2026-07-11T07:57:43.000834Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T07:57:43.000834Z digest=sha256:6b8dc71e646019013753cd45388f177e8775e87f1bec3032ecf11fddca201e28

Observation a3fa9e3c-b2da-46a4-b722-fe0430636231 · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

EdgeBench: Unveiling Scaling Laws of Learning from Real-World Environments DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 16

Resolution
unresolved
no resolver link, observed 2026-07-11T07:57:43.000834Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T07:57:43.000834Z digest=sha256:a4a4b5b234bd335965f610ee6ec56ea68cc13a15d1a131eeeca7f9e40d3e0b05

Observation 7bcc2faf-d59f-4810-bd0d-a8079a128692 · outbound

This paper cites DeepSeek-V4: Towards highly efficient million-token context intelligence.https://huggingface.

EdgeBench: Unveiling Scaling Laws of Learning from Real-World Environments DeepSeek-V4: Towards highly efficient million-token context intelligence.https://huggingface

Reference 17

Resolution
unresolved
no resolver link, observed 2026-07-11T07:57:43.000834Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T07:57:43.000834Z digest=sha256:a099bb5e520c187e4bff3c75a0191d9ab5cb60ec6b937f3fe511c9abb1b7b736

Observation 4a490dd8-b0bc-45f7-aef6-0a85af173efa · outbound

This paper cites NL2Repo-Bench: Towards long-horizon repository generation evaluation of coding agents.arXiv preprint arXiv:2512.12730, 2025.

EdgeBench: Unveiling Scaling Laws of Learning from Real-World Environments NL2Repo-Bench: Towards long-horizon repository generation evaluation of coding agents.arXiv preprint arXiv:2512.12730, 2025

Reference 18

Resolution
unresolved
no resolver link, observed 2026-07-11T07:57:43.000834Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T07:57:43.000834Z digest=sha256:92aa6e1674534a4410607061a390f1e6370ed90a3735d31a5f56f3e4ebda88d0

Observation 73000c7a-9880-4b6e-b1dc-6669b40dc8ab · outbound

This paper cites EvaLearn: Quantifying the learning capability and efficiency of LLMs via sequential problem solving.arXiv preprint arXiv:2506.02672, 2025.

EdgeBench: Unveiling Scaling Laws of Learning from Real-World Environments EvaLearn: Quantifying the learning capability and efficiency of LLMs via sequential problem solving.arXiv preprint arXiv:2506.02672, 2025

Reference 19

Resolution
unresolved
no resolver link, observed 2026-07-11T07:57:43.000834Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T07:57:43.000834Z digest=sha256:d8a53032fa4756ad0fc5b2d3dce3a4f5fb819f4995d36d114eedeff9254aa459

Observation e3760e36-17c8-4e2f-a37e-5eabc56e6cb5 · outbound

This paper cites CL-bench Life: Can Language Models Learn from Real-Life Context?.

EdgeBench: Unveiling Scaling Laws of Learning from Real-World Environments CL-bench Life: Can Language Models Learn from Real-Life Context?

Reference 20

Resolution
unresolved
no resolver link, observed 2026-07-11T07:57:43.000834Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T07:57:43.000834Z digest=sha256:9f8d438dd645e2333337a259200bd1708079b789ab734d61b7e9d6f886d794ac

Observation b43d2412-b158-4a29-998b-a5746a0433e2 · outbound

This paper cites CL-bench: A benchmark for context learning.arXiv preprint arXiv:2602.03587, 2026.

EdgeBench: Unveiling Scaling Laws of Learning from Real-World Environments CL-bench: A benchmark for context learning.arXiv preprint arXiv:2602.03587, 2026

Reference 21

Resolution
unresolved
no resolver link, observed 2026-07-11T07:57:43.000834Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T07:57:43.000834Z digest=sha256:4120b1f26c141c1c380838cdb9d760a846d515728a668ecc9c6d732589a52974

Observation 0c1af386-ae71-4832-b1c3-0ee67cf747db · outbound

This paper cites The Llama 3 Herd of Models.

EdgeBench: Unveiling Scaling Laws of Learning from Real-World Environments The Llama 3 Herd of Models

Reference 22

Resolution
unresolved
no resolver link, observed 2026-07-11T07:57:43.000834Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T07:57:43.000834Z digest=sha256:71ab03b59eacb4a0e91aa101fbcd5a2f7d0acdfca72e0f2f3b17634b78c53de1

Observation 40b8279d-1367-4817-a873-71b42c23df15 · outbound

This paper cites Collider-Bench: Benchmarking AI Agents with Particle Physics Analysis Reproduction.

EdgeBench: Unveiling Scaling Laws of Learning from Real-World Environments Collider-Bench: Benchmarking AI Agents with Particle Physics Analysis Reproduction

Reference 23

Resolution
unresolved
no resolver link, observed 2026-07-11T07:57:43.000834Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T07:57:43.000834Z digest=sha256:06cb6f30424b3a0499e46f267c8b02250d08b8fb18314bbca66bf9f84c1cb4c4

Observation 3b5a05dc-39f8-4fce-a93c-387a838617d2 · outbound

This paper cites Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context.

EdgeBench: Unveiling Scaling Laws of Learning from Real-World Environments Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context

Reference 24

Resolution
unresolved
no resolver link, observed 2026-07-11T07:57:43.000834Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T07:57:43.000834Z digest=sha256:cf713d558919c90c490229e72a8d1feaea6671ff0002cf8021c2cf69faad37fc

Observation cca37178-d682-42ea-be6f-84203d0c6f06 · outbound

This paper cites Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities.

EdgeBench: Unveiling Scaling Laws of Learning from Real-World Environments Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities

Reference 25

Resolution
unresolved
no resolver link, observed 2026-07-11T07:57:43.000834Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T07:57:43.000834Z digest=sha256:1649726e23affa66d167da03915fe4142306e4e8278ba977aa27dc98c083d270

Observation 8ff51093-4805-4598-a39d-68b265670f07 · outbound

This paper cites GLM-5: from Vibe Coding to Agentic Engineering.

EdgeBench: Unveiling Scaling Laws of Learning from Real-World Environments GLM-5: from Vibe Coding to Agentic Engineering

Reference 26

Resolution
unresolved
no resolver link, observed 2026-07-11T07:57:43.000834Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T07:57:43.000834Z digest=sha256:639b9ba4569242a85484685fc7bd55be29a93af47e480a1fed2c67533863c59b

Observation 73ff0522-3efa-4149-9dc5-e4bee394a260 · outbound

This paper cites an unresolved cited work.

EdgeBench: Unveiling Scaling Laws of Learning from Real-World Environments Unresolved cited work

Reference 27

Resolution
unresolved
no resolver link, observed 2026-07-11T07:57:43.000834Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T07:57:43.000834Z digest=sha256:7e3dfb9ba7512e02c6d07f2d87ef88ecedaa913ea2ac5b1e34f69d5e0f28984b

Observation 2523e60c-3c42-4315-a9af-b181b53a2d41 · outbound

This paper cites Measuring massive multitask language understanding.

EdgeBench: Unveiling Scaling Laws of Learning from Real-World Environments Measuring massive multitask language understanding

Reference 28

Resolution
unresolved
no resolver link, observed 2026-07-11T07:57:43.000834Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T07:57:43.000834Z digest=sha256:366b0ecf6826a397bde7df2d74260299c68ea222ebd9cf066a84ecb294307bea

Observation 54d22059-df86-4576-8f40-02cd8a4c65ae · outbound

This paper cites Measuring mathematical problem solving with the MATH dataset.

EdgeBench: Unveiling Scaling Laws of Learning from Real-World Environments Measuring mathematical problem solving with the MATH dataset

Reference 29

Resolution
unresolved
no resolver link, observed 2026-07-11T07:57:43.000834Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T07:57:43.000834Z digest=sha256:65401b6e303ca2094b0d0318de411914b7a2905387b1a9adf609758ad7046c46

Observation 7e2a44c4-6fe0-4a55-b6d8-c27c0d7967b2 · outbound

This paper cites Scaling laws for single-agent reinforcement learning.

EdgeBench: Unveiling Scaling Laws of Learning from Real-World Environments Scaling laws for single-agent reinforcement learning

Reference 30

Resolution
unresolved
no resolver link, observed 2026-07-11T07:57:43.000834Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T07:57:43.000834Z digest=sha256:1fd621de885926709a1c816fc816c52a956972677e237894e98aafa65196f6b4

Observation a7a7b1a7-3149-4201-83cb-bdc89358b7b1 · outbound

This paper cites Training Compute-Optimal Large Language Models.

EdgeBench: Unveiling Scaling Laws of Learning from Real-World Environments Training Compute-Optimal Large Language Models

Reference 31

Resolution
unresolved
no resolver link, observed 2026-07-11T07:57:43.000834Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T07:57:43.000834Z digest=sha256:8257de44b3d46469623c6edad10781356557d45cfd7d1ef9ace8a18498270feb

Observation 34df4243-24d7-4d25-a765-11b343cd2d55 · outbound

This paper cites Everything Is a Ralph Loop.https://ghuntley.com/loop/, January 2026.

EdgeBench: Unveiling Scaling Laws of Learning from Real-World Environments Everything Is a Ralph Loop.https://ghuntley.com/loop/, January 2026

Reference 32

Resolution
unresolved
no resolver link, observed 2026-07-11T07:57:43.000834Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T07:57:43.000834Z digest=sha256:b615b49592d3554f7a0ec800a47d6fb0b39ebdb9ede0f4989c3a920d5a6fa35e

Observation 328a289c-ed48-440e-89ba-c5b5d21cf6d2 · outbound

This paper cites ALE-bench: A benchmark for long-horizon objective-driven algorithm engineering.arXiv preprint arXiv:2506.09050, 2025.

EdgeBench: Unveiling Scaling Laws of Learning from Real-World Environments ALE-bench: A benchmark for long-horizon objective-driven algorithm engineering.arXiv preprint arXiv:2506.09050, 2025

Reference 33

Resolution
unresolved
no resolver link, observed 2026-07-11T07:57:43.000834Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T07:57:43.000834Z digest=sha256:407bf9db93e2b5901f228c4dd95879ee49e35d5d7b33622b7377de19915d8e3b

Observation 210837f9-5455-4f29-96ae-c6a968a9bb08 · outbound

This paper cites Jimenez, John Yang, Alexander Wettig, Shunyu Yao, Kexin Pei, Ofir Press, and Karthik R.

EdgeBench: Unveiling Scaling Laws of Learning from Real-World Environments Jimenez, John Yang, Alexander Wettig, Shunyu Yao, Kexin Pei, Ofir Press, and Karthik R

Reference 34

Resolution
unresolved
no resolver link, observed 2026-07-11T07:57:43.000834Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T07:57:43.000834Z digest=sha256:c6ec4b51a8c96d0a88c8b86e2ca796f8d50b9890a730095f1170ceb195483627

Observation 7da9ee76-5a45-45bf-8fcf-2408bf9bc4d1 · outbound

This paper cites Scaling Laws for Neural Language Models.

EdgeBench: Unveiling Scaling Laws of Learning from Real-World Environments Scaling Laws for Neural Language Models

Reference 35

Resolution
unresolved
no resolver link, observed 2026-07-11T07:57:43.000834Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T07:57:43.000834Z digest=sha256:b984ce5d89baed8d45f12ebc149578b8d7cffe94fcf7332ff1f628b4c427d615

Observation b886bcb1-417d-4c26-8913-2c0937be3a9f · outbound

This paper cites The Art of Scaling Reinforcement Learning Compute for LLMs.

EdgeBench: Unveiling Scaling Laws of Learning from Real-World Environments The Art of Scaling Reinforcement Learning Compute for LLMs

Reference 36

Resolution
unresolved
no resolver link, observed 2026-07-11T07:57:43.000834Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T07:57:43.000834Z digest=sha256:c04097e43c201c95e2caf2798663bb23b207800b766c86468f70a35fa02e3996

Observation 54bd7582-59b9-4158-b3e3-f583d7bbc9dc · outbound

This paper cites Measuring AI ability to complete long tasks.arXiv preprint arXiv:2503.14499, 2025.

EdgeBench: Unveiling Scaling Laws of Learning from Real-World Environments Measuring AI ability to complete long tasks.arXiv preprint arXiv:2503.14499, 2025

Reference 37

Resolution
unresolved
no resolver link, observed 2026-07-11T07:57:43.000834Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T07:57:43.000834Z digest=sha256:bb0c9724987d630a3b68874191f6710df35bdbd6a319ac918bb131d4643dd44e

Observation 3ecf02af-9595-477c-a464-f98d9e60071f · outbound

This paper cites Competition-Level Code Generation with AlphaCode.

EdgeBench: Unveiling Scaling Laws of Learning from Real-World Environments Competition-Level Code Generation with AlphaCode

Reference 38

Resolution
unresolved
no resolver link, observed 2026-07-11T07:57:43.000834Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T07:57:43.000834Z digest=sha256:26a453193ee5bd28e9ce9cd55f6e1c0c399dcee1b112115924e8ce03c9b96945

Observation cef4f759-52e3-404c-be62-fb76adab6c5a · outbound

This paper cites Introducing FrontierCode.https://cognition.ai/blog/frontier-code, 2026.

EdgeBench: Unveiling Scaling Laws of Learning from Real-World Environments Introducing FrontierCode.https://cognition.ai/blog/frontier-code, 2026

Reference 39

Resolution
unresolved
no resolver link, observed 2026-07-11T07:57:43.000834Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T07:57:43.000834Z digest=sha256:cd56e5ea89106582c774487a941a8f809191abda9d759546123fbdf46ae199b4

Observation 026adc16-e3ac-4ca3-bc46-cd519e262cde · outbound

This paper cites MLS-Bench: A Holistic and Rigorous Assessment of AI Systems on Building Better AI.

EdgeBench: Unveiling Scaling Laws of Learning from Real-World Environments MLS-Bench: A Holistic and Rigorous Assessment of AI Systems on Building Better AI

Reference 40

Resolution
unresolved
no resolver link, observed 2026-07-11T07:57:43.000834Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T07:57:43.000834Z digest=sha256:d2965abae7c82b9218afae7605281cec97159717ee564ed411c375927aef0a2d

Observation 17e54dcf-b41c-41b5-b9fb-2af29b9d9467 · outbound

This paper cites FrontierCS: Evolving challenges for evolving intelligence.arXiv preprint arXiv:2512.15699, 2025.

EdgeBench: Unveiling Scaling Laws of Learning from Real-World Environments FrontierCS: Evolving challenges for evolving intelligence.arXiv preprint arXiv:2512.15699, 2025

Reference 41

Resolution
unresolved
no resolver link, observed 2026-07-11T07:57:43.000834Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T07:57:43.000834Z digest=sha256:1314b8c1834607458b85b537de4541437aed1a2d4d0b48f030f24bac8bab6c46

Observation bb355dfe-4034-43c7-b985-a751c46b1cb2 · outbound

This paper cites Humans still beat AI in the long horizon: Revisiting test-time scaling in the agent era.https://joyemang33.github.io/blog/ 2026/humans-dont-just-sample/, 2026.

EdgeBench: Unveiling Scaling Laws of Learning from Real-World Environments Humans still beat AI in the long horizon: Revisiting test-time scaling in the agent era.https://joyemang33.github.io/blog/ 2026/humans-dont-just-sample/, 2026

Reference 42

Resolution
unresolved
no resolver link, observed 2026-07-11T07:57:43.000834Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T07:57:43.000834Z digest=sha256:8fc485292c7879c7447f0d78e90bdeb07d22f8cb2a846e1877bc34f67f8cdea9

Observation 881c1fe0-2cdf-48e5-b147-dacf0d347d52 · outbound

This paper cites MAA invitational competitions: American invitational mathematics examination (AIME).https://maa.org/maa-invitational-competitions/, 2026.

EdgeBench: Unveiling Scaling Laws of Learning from Real-World Environments MAA invitational competitions: American invitational mathematics examination (AIME).https://maa.org/maa-invitational-competitions/, 2026

Reference 43

Resolution
unresolved
no resolver link, observed 2026-07-11T07:57:43.000834Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T07:57:43.000834Z digest=sha256:31ef2e6bce68aec8018f35e25e3cb2767424ed4130b1795f2a863250f493d83f

Observation d8bf8a6e-7811-4d3a-b6e1-6b1aef51ea03 · outbound

This paper cites Merrill, Alexander G.

EdgeBench: Unveiling Scaling Laws of Learning from Real-World Environments Merrill, Alexander G

Reference 44

Resolution
unresolved
no resolver link, observed 2026-07-11T07:57:43.000834Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T07:57:43.000834Z digest=sha256:31351ff161cac0086a9d808cd6b6b6f3e7be95172372173c6193b770ea27939f

Observation 91e9f2a4-293f-49ca-adc2-bfec779246a4 · outbound

This paper cites an unresolved cited work.

EdgeBench: Unveiling Scaling Laws of Learning from Real-World Environments Unresolved cited work

Reference 45

Resolution
unresolved
no resolver link, observed 2026-07-11T07:57:43.000834Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T07:57:43.000834Z digest=sha256:17c25c18684281e617a6f8869fb4d550719f0ab10d2712adbbbc96d883c38bad

Observation d6b770bb-1edc-4a58-97ca-832b32585cf4 · outbound

This paper cites GPT-4 Technical Report.

EdgeBench: Unveiling Scaling Laws of Learning from Real-World Environments GPT-4 Technical Report

Reference 46

Resolution
unresolved
no resolver link, observed 2026-07-11T07:57:43.000834Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T07:57:43.000834Z digest=sha256:593fd31b2c601b7840f4aa556da136e26c6bf0ee64f9c56b8cbf0b83d0d374c7

Observation a851a6a5-d2c6-477f-bf96-244b56372cab · outbound

This paper cites Learning to reason with LLMs.https://openai.com/index/learning-to-reason-with-llms/, 2024.

EdgeBench: Unveiling Scaling Laws of Learning from Real-World Environments Learning to reason with LLMs.https://openai.com/index/learning-to-reason-with-llms/, 2024

Reference 47

Resolution
unresolved
no resolver link, observed 2026-07-11T07:57:43.000834Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T07:57:43.000834Z digest=sha256:9d0d952b8ca38700248121b47701934af08a171b648accef1d14af5e9028a854

Observation 0d27af57-53f9-429e-a886-97f8b5f06485 · outbound

This paper cites Introducing SWE-bench verified.

EdgeBench: Unveiling Scaling Laws of Learning from Real-World Environments Introducing SWE-bench verified

Reference 48

Resolution
unresolved
no resolver link, observed 2026-07-11T07:57:43.000834Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T07:57:43.000834Z digest=sha256:19928b57057acbab4e3ff09753e77d3efc58ea4124d1e4babfb23ccd0f1d4ce6

Observation be1cf6c6-bf4b-44a7-a72f-bd0d64a9af93 · outbound

This paper cites Computer-using agent.https://openai.com/index/computer-using-agent/, 2025.

EdgeBench: Unveiling Scaling Laws of Learning from Real-World Environments Computer-using agent.https://openai.com/index/computer-using-agent/, 2025

Reference 49

Resolution
unresolved
no resolver link, observed 2026-07-11T07:57:43.000834Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T07:57:43.000834Z digest=sha256:40c68fc666f27e0cdf88e1d3bbfede4fd566d067f98a1986eb3bb900d84ff41c

Observation 3a1aaa90-bc35-47e1-b1a8-2381540d2446 · outbound

This paper cites GPT-4.5 System Card.https://cdn.openai.com/gpt-4-5-system-card-2272025.pdf, 2025.

EdgeBench: Unveiling Scaling Laws of Learning from Real-World Environments GPT-4.5 System Card.https://cdn.openai.com/gpt-4-5-system-card-2272025.pdf, 2025

Reference 50

Resolution
unresolved
no resolver link, observed 2026-07-11T07:57:43.000834Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T07:57:43.000834Z digest=sha256:5932d42773e4b4533beb98230f00f005d5f66b97fc2c39b03705509059e105e6

Observation 78be564a-c477-4f0e-8910-22ba64918f9a · outbound

This paper cites Update to GPT-5 System Card: GPT-5.2.

EdgeBench: Unveiling Scaling Laws of Learning from Real-World Environments Update to GPT-5 System Card: GPT-5.2

Reference 51

Resolution
unresolved
no resolver link, observed 2026-07-11T07:57:43.000834Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T07:57:43.000834Z digest=sha256:943cb19026250d57a90d4d1371a384cc2fd6f25c387f3ccc6f6f087a30273bcf

Observation e47e857d-0523-466f-9746-0b9b56a260c6 · outbound

This paper cites OpenAI GPT-5 System Card.

EdgeBench: Unveiling Scaling Laws of Learning from Real-World Environments OpenAI GPT-5 System Card

Reference 52

Resolution
unresolved
no resolver link, observed 2026-07-11T07:57:43.000834Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T07:57:43.000834Z digest=sha256:46ae8ac56dc1e786d2c9b7171c62a8c612cc43331b5c6323b8fb5dd93d9f6897

Observation 70905830-041f-486b-b4cc-9343160de6a8 · outbound

This paper cites Follow a Goal.

EdgeBench: Unveiling Scaling Laws of Learning from Real-World Environments Follow a Goal

Reference 53

Resolution
unresolved
no resolver link, observed 2026-07-11T07:57:43.000834Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T07:57:43.000834Z digest=sha256:3a09bb2e668b571a6f81b057a20364378f93ba28fd2c0b5e20b877402832fd34

Observation e16acc0c-60f2-4a96-8241-aebfbf01853c · outbound

This paper cites GPT-5.4 Thinking System Card.

EdgeBench: Unveiling Scaling Laws of Learning from Real-World Environments GPT-5.4 Thinking System Card

Reference 54

Resolution
unresolved
no resolver link, observed 2026-07-11T07:57:43.000834Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T07:57:43.000834Z digest=sha256:b5856385425ce3af570578c8a3ffa82c5a9170a874e0fa3bdc44419509261083

Observation fc34b859-fde2-4635-981c-05305836231f · outbound

This paper cites GPT-5.5 System Card.

EdgeBench: Unveiling Scaling Laws of Learning from Real-World Environments GPT-5.5 System Card

Reference 55

Resolution
unresolved
no resolver link, observed 2026-07-11T07:57:43.000834Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T07:57:43.000834Z digest=sha256:194b439e52bd6291e88becb4d0e12e6537202a1a117cd05f13becb5b4d8d4b61

Observation 55cb6e29-b47e-46a7-911f-8a59565b5b1f · outbound

This paper cites How predictable is language model benchmark performance?.

EdgeBench: Unveiling Scaling Laws of Learning from Real-World Environments How predictable is language model benchmark performance?

Reference 56

Resolution
unresolved
no resolver link, observed 2026-07-11T07:57:43.000834Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T07:57:43.000834Z digest=sha256:fd9729e67e19963f67716eabd67fbd24723ba7da55787214e1c5349e86ca5388

Observation 8afc3c7d-9e70-430e-bce6-d4eb4626ee1c · outbound

This paper cites GDPval: Evaluating AI Model Performance on Real-World Economically Valuable Tasks.

EdgeBench: Unveiling Scaling Laws of Learning from Real-World Environments GDPval: Evaluating AI Model Performance on Real-World Economically Valuable Tasks

Reference 57

Resolution
unresolved
no resolver link, observed 2026-07-11T07:57:43.000834Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T07:57:43.000834Z digest=sha256:f1f448b24db60deb9054924e6e3ccc0c74d5d8d378cc01e6572e9f6cc0f206f9

Observation f91086f6-b85b-41a4-bb86-9a70c300f0b3 · outbound

This paper cites PRBench: End-to-end paper reproduction in physics research.arXiv preprint arXiv:2603.27646, 2026.

EdgeBench: Unveiling Scaling Laws of Learning from Real-World Environments PRBench: End-to-end paper reproduction in physics research.arXiv preprint arXiv:2603.27646, 2026

Reference 58

Resolution
unresolved
no resolver link, observed 2026-07-11T07:57:43.000834Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T07:57:43.000834Z digest=sha256:5993d15fa1a4ae6f1b3bae1b35f07841ae557040c72cd01b704ecf9dbe54c906

Observation f1e1595b-1e44-4dd1-b005-eb1d9a436803 · outbound

This paper cites GPQA: A Graduate-Level Google-Proof Q&A Benchmark.

EdgeBench: Unveiling Scaling Laws of Learning from Real-World Environments GPQA: A Graduate-Level Google-Proof Q&A Benchmark

Reference 59

Resolution
unresolved
no resolver link, observed 2026-07-11T07:57:43.000834Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T07:57:43.000834Z digest=sha256:a87e6ce6d7486fdfc6e1f2126fae097e0a2f1de884c35a2059390d89b7f6886b

Observation ee29afa1-7422-4610-8ab1-71a255dac09e · outbound

This paper cites HCAST: Human-Calibrated Autonomy Software Tasks.

EdgeBench: Unveiling Scaling Laws of Learning from Real-World Environments HCAST: Human-Calibrated Autonomy Software Tasks

Reference 60

Resolution
unresolved
no resolver link, observed 2026-07-11T07:57:43.000834Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T07:57:43.000834Z digest=sha256:0ec3a174a64aa59eb6275d9670d264c3b091b4b68431aa84b4bd70d15e7c3c88

Observation c7c6e6ae-4d1e-41ba-9124-1a301b67e354 · outbound

This paper cites Observational Scaling Laws and the Predictability of Language Model Performance.

EdgeBench: Unveiling Scaling Laws of Learning from Real-World Environments Observational Scaling Laws and the Predictability of Language Model Performance

Reference 61

Resolution
unresolved
no resolver link, observed 2026-07-11T07:57:43.000834Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T07:57:43.000834Z digest=sha256:ec7fedcfd4f8d4ed3ffeb0de6ed99fb008a9cb685c97c6e4fb23bbc7617dc88f

Observation 29cf167a-bdf0-43fa-ae28-3af81c202ecb · outbound

This paper cites Thinking vs. Doing: Agents that Reason by Scaling Test-Time Interaction.

EdgeBench: Unveiling Scaling Laws of Learning from Real-World Environments Thinking vs. Doing: Agents that Reason by Scaling Test-Time Interaction

Reference 62

Resolution
unresolved
no resolver link, observed 2026-07-11T07:57:43.000834Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T07:57:43.000834Z digest=sha256:0555118b69d4a1c8b3a963dc748d5204ede3ae7b5121b3aaadb2152d3c94eeaa

Observation 0c2d66e0-eecb-4956-b3d5-71cb9d17f48c · outbound

This paper cites CORE-Bench: Fostering the Credibility of Published Research Through a Computational Reproducibility Agent Benchmark.

EdgeBench: Unveiling Scaling Laws of Learning from Real-World Environments CORE-Bench: Fostering the Credibility of Published Research Through a Computational Reproducibility Agent Benchmark

Reference 63

Resolution
unresolved
no resolver link, observed 2026-07-11T07:57:43.000834Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T07:57:43.000834Z digest=sha256:4d35608b763bb7f9e675278c556b5ef82960b8da6ae7e23736a741956d1db2a8

Observation 82694a33-b01e-4646-8691-20ad6d64bcdc · outbound

This paper cites Scaling LLM Test-Time Compute Optimally can be More Effective than Scaling Model Parameters.

EdgeBench: Unveiling Scaling Laws of Learning from Real-World Environments Scaling LLM Test-Time Compute Optimally can be More Effective than Scaling Model Parameters

Reference 64

Resolution
unresolved
no resolver link, observed 2026-07-11T07:57:43.000834Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T07:57:43.000834Z digest=sha256:1807b8e1e07b628299b607e7409a70854a1226e3e66ab3f828fe4deab02c57b4

Observation bfdd95cd-b5f7-4f98-b4d2-9fbad01feefb · outbound

This paper cites PaperBench: Evaluating AI’s ability to replicate AI research.

EdgeBench: Unveiling Scaling Laws of Learning from Real-World Environments PaperBench: Evaluating AI’s ability to replicate AI research

Reference 65

Resolution
unresolved
no resolver link, observed 2026-07-11T07:57:43.000834Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T07:57:43.000834Z digest=sha256:8b39d2ce21696379c0449bd0e368c23ec796ed4bfa46f2bc5ea4198eab4f7510

Observation 712d4198-2fad-4628-9742-ca7899573f11 · outbound

This paper cites Agents’ last exam.arXiv preprint arXiv:2606.05405, 2026.

EdgeBench: Unveiling Scaling Laws of Learning from Real-World Environments Agents’ last exam.arXiv preprint arXiv:2606.05405, 2026

Reference 66

Resolution
unresolved
no resolver link, observed 2026-07-11T07:57:43.000834Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T07:57:43.000834Z digest=sha256:13c1049042242c01843894cb98ddb3966a0c82ec9acb8a26f4fedb60ef40f93c

Observation 1defa02b-5608-43f3-a65d-70e52db69747 · outbound

This paper cites SWE-EVO: Benchmarking Coding Agents in Long-Horizon Software Evolution Scenarios.

EdgeBench: Unveiling Scaling Laws of Learning from Real-World Environments SWE-EVO: Benchmarking Coding Agents in Long-Horizon Software Evolution Scenarios

Reference 67

Resolution
unresolved
no resolver link, observed 2026-07-11T07:57:43.000834Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T07:57:43.000834Z digest=sha256:bb4363a26e3e47195089816402d09ef1c8b66133a125e97b9db5a128db774562

Observation d112d6e9-5606-40dc-a387-121b2e1da382 · outbound

This paper cites BrowseComp: A Simple Yet Challenging Benchmark for Browsing Agents.

EdgeBench: Unveiling Scaling Laws of Learning from Real-World Environments BrowseComp: A Simple Yet Challenging Benchmark for Browsing Agents

Reference 68

Resolution
unresolved
no resolver link, observed 2026-07-11T07:57:43.000834Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T07:57:43.000834Z digest=sha256:3b3d6655193fc0b10fcf3b46baaf7d09f7a87eb63d51b070634f9dfc8b5bb13d

Observation 252f20b3-6da8-46dd-b279-166d986c5e88 · outbound

This paper cites Evo-Memory: Benchmarking LLM Agent Test-time Learning with Self-Evolving Memory.

EdgeBench: Unveiling Scaling Laws of Learning from Real-World Environments Evo-Memory: Benchmarking LLM Agent Test-time Learning with Self-Evolving Memory

Reference 69

Resolution
unresolved
no resolver link, observed 2026-07-11T07:57:43.000834Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T07:57:43.000834Z digest=sha256:15797435765c86adca39100520634e57fb01b3a3e4051d18d152c9cf8fd5c334

Observation 22315c62-4bfa-47c5-a42d-840f1a112ec2 · outbound

This paper cites A statistical distribution function of wide applicability.Journal of Applied Mechanics, 18(3): 293–297, 1951.

EdgeBench: Unveiling Scaling Laws of Learning from Real-World Environments A statistical distribution function of wide applicability.Journal of Applied Mechanics, 18(3): 293–297, 1951

Reference 70

Resolution
unresolved
no resolver link, observed 2026-07-11T07:57:43.000834Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T07:57:43.000834Z digest=sha256:bf886717343a4a488c440b86854197f8ff8775640aead2331f8ab9f974699ffe

Observation f8cf512a-ebdb-4e42-8857-b0716d1e392e · outbound

This paper cites RE-bench: Evaluating frontier AI R&D capabilities of language model agents against human experts.

EdgeBench: Unveiling Scaling Laws of Learning from Real-World Environments RE-bench: Evaluating frontier AI R&D capabilities of language model agents against human experts

Reference 71

Resolution
unresolved
no resolver link, observed 2026-07-11T07:57:43.000834Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T07:57:43.000834Z digest=sha256:61279fb94562f5ab54cf77efb7b8a8599c6047772c121d890f3c8d3b8adbe525

Observation 7c0d626a-d78c-4aaa-b494-6941e0d25c4c · outbound

This paper cites Sigmoid function.https://en.wikipedia.org/wiki/Sigmoid_function, 2025.

EdgeBench: Unveiling Scaling Laws of Learning from Real-World Environments Sigmoid function.https://en.wikipedia.org/wiki/Sigmoid_function, 2025

Reference 72

Resolution
unresolved
no resolver link, observed 2026-07-11T07:57:43.000834Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T07:57:43.000834Z digest=sha256:01ae41b1f6b46baf24bb518f8e86c31b456e3c93718cb2b9f7a8a00dde57eded

Observation 2d6660aa-82ab-4b57-903c-a210987b8135 · outbound

This paper cites StreamBench: Towards benchmarking continuous improvement of language agents.

EdgeBench: Unveiling Scaling Laws of Learning from Real-World Environments StreamBench: Towards benchmarking continuous improvement of language agents

Reference 73

Resolution
unresolved
no resolver link, observed 2026-07-11T07:57:43.000834Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T07:57:43.000834Z digest=sha256:c6fa6efd5a11e69b356bf442a166168b5507aa9dfa17b60c3099c5346906ae5f

Observation 47937986-7f5f-4b95-b26e-823417d4f43e · outbound

This paper cites Inference Scaling Laws: An Empirical Analysis of Compute-Optimal Inference for Problem-Solving with Language Models.

EdgeBench: Unveiling Scaling Laws of Learning from Real-World Environments Inference Scaling Laws: An Empirical Analysis of Compute-Optimal Inference for Problem-Solving with Language Models

Reference 74

Resolution
unresolved
no resolver link, observed 2026-07-11T07:57:43.000834Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T07:57:43.000834Z digest=sha256:3475315574389c07b369fb26f6366fd99fe2ee1caf9f5fcec32751f44cff97bd

Observation 82f25be7-e4ef-4f8b-a507-683c691cfcf5 · outbound

This paper cites RoadmapBench: Evaluating Long-Horizon Agentic Software Development Across Version Upgrades.

EdgeBench: Unveiling Scaling Laws of Learning from Real-World Environments RoadmapBench: Evaluating Long-Horizon Agentic Software Development Across Version Upgrades

Reference 75

Resolution
unresolved
no resolver link, observed 2026-07-11T07:57:43.000834Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T07:57:43.000834Z digest=sha256:65bdbc61aed1612c667badb7af084b51d247f377e06da6759bfecbe78a6ffb1a

Observation 7b4ae5b9-58ca-4e14-bf39-7f8ab6ce8281 · outbound

This paper cites AutoLab: Can Frontier Models Solve Long-Horizon Auto Research and Engineering Tasks?.

EdgeBench: Unveiling Scaling Laws of Learning from Real-World Environments AutoLab: Can Frontier Models Solve Long-Horizon Auto Research and Engineering Tasks?

Reference 76

Resolution
unresolved
no resolver link, observed 2026-07-11T07:57:43.000834Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T07:57:43.000834Z digest=sha256:d11e25fa283bee0b2ca94e8e9a8f1a529605bfb26c0a50c94e6b1c2734a8a27f

Observation 83dd5da8-a8d8-400f-84f0-19ca19df90eb · outbound

This paper cites ProgramBench: Can Language Models Rebuild Programs From Scratch?.

EdgeBench: Unveiling Scaling Laws of Learning from Real-World Environments ProgramBench: Can Language Models Rebuild Programs From Scratch?

Reference 77

Resolution
unresolved
no resolver link, observed 2026-07-11T07:57:43.000834Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T07:57:43.000834Z digest=sha256:6141b49352061dc9812dab55f46182394afdb428bd241d2f197524abfb49df0b

Observation 3419edba-ac58-47af-9f84-e4b284203e24 · outbound

This paper cites an unresolved cited work.

EdgeBench: Unveiling Scaling Laws of Learning from Real-World Environments Unresolved cited work

Reference 78

Resolution
unresolved
no resolver link, observed 2026-07-11T07:57:43.000834Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T07:57:43.000834Z digest=sha256:d8a3e597a65071e7b81eea2ad1e18b957e6db62913bd6c6697eda77cb8a0a94a

Observation 2d2e2253-0e65-405a-9e00-2bec3401796b · outbound

This paper cites GLM-5.1: Towards long-horizon tasks.https://z.ai/blog/glm-5.1, 2026.

EdgeBench: Unveiling Scaling Laws of Learning from Real-World Environments GLM-5.1: Towards long-horizon tasks.https://z.ai/blog/glm-5.1, 2026

Reference 79

Resolution
unresolved
no resolver link, observed 2026-07-11T07:57:43.000834Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T07:57:43.000834Z digest=sha256:1b196df10c300ce355d83c147357dae742dc9eec883a9a46d991f4bf5291babc

Observation 8ffc4346-e530-4b02-b3e9-dc7f9b48d6ed · outbound

This paper cites Prescriptive Scaling Reveals the Evolution of Language Model Capabilities.

EdgeBench: Unveiling Scaling Laws of Learning from Real-World Environments Prescriptive Scaling Reveals the Evolution of Language Model Capabilities

Reference 80

Resolution
unresolved
no resolver link, observed 2026-07-11T07:57:43.000834Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T07:57:43.000834Z digest=sha256:9bc9cf6ac5f05f5e8ac22bec150730d0bd24b46cc311bd145d61806c82037a1e

Observation 00de3fbb-5693-4642-97aa-b58b811c76c6 · outbound

This paper cites SWE-AGI: Benchmarking specification-driven software construction with MoonBit in the era of autonomous agents.arXiv preprint arXiv:2602.09447, 2026.

EdgeBench: Unveiling Scaling Laws of Learning from Real-World Environments SWE-AGI: Benchmarking specification-driven software construction with MoonBit in the era of autonomous agents.arXiv preprint arXiv:2602.09447, 2026

Reference 81

Resolution
unresolved
no resolver link, observed 2026-07-11T07:57:43.000834Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T07:57:43.000834Z digest=sha256:85648a63b94e29c990d14d0f190ac4648e3f972ea1cf2d3bd592d38931553071

Observation a011244a-6d00-4119-a029-e0ee87c7f1a5 · outbound

This paper cites LifelongAgentBench: Evaluating LLM Agents as Lifelong Learners.

EdgeBench: Unveiling Scaling Laws of Learning from Real-World Environments LifelongAgentBench: Evaluating LLM Agents as Lifelong Learners

Reference 82

Resolution
unresolved
no resolver link, observed 2026-07-11T07:57:43.000834Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T07:57:43.000834Z digest=sha256:13b00febc07c1b56eee0338b190f5360630fb3e16b3ae1753c33b9c17d352d83

Observation 4f56a403-bb20-49d5-9cb9-a7ee627a9e52 · outbound

This paper cites an unresolved cited work.

EdgeBench: Unveiling Scaling Laws of Learning from Real-World Environments Unresolved cited work

Reference 83

Resolution
unresolved
no resolver link, observed 2026-07-11T07:57:43.000834Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T07:57:43.000834Z digest=sha256:bb65acc9f49e01ecf1d143674fb282e3b9f700d6b40968d6d7856fe3f05efdb9

Observation 10e6ca54-87a1-4472-b024-0179f90f2e44 · outbound

This paper cites Within one task, the conditional expected score-growth rate is a weighted cut from unlocked score nodes to locked score nodes of the task graph.

EdgeBench: Unveiling Scaling Laws of Learning from Real-World Environments Within one task, the conditional expected score-growth rate is a weighted cut from unlocked score nodes to locked score nodes of the task graph

Reference 84

Resolution
unresolved
no resolver link, observed 2026-07-11T07:57:43.000834Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T07:57:43.000834Z digest=sha256:a41d70aa7ea049ead5006e758432a131ce06630acac503cf2096066abe637545

Observation b30b25d8-f8f9-442b-9c10-ab5aff42b72b · outbound

This paper cites an unresolved cited work.

EdgeBench: Unveiling Scaling Laws of Learning from Real-World Environments Unresolved cited work

Reference 85

Resolution
unresolved
no resolver link, observed 2026-07-11T07:57:43.000834Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T07:57:43.000834Z digest=sha256:7f07935bc3ea4cdacd65bb925c57f54d2e0c3fb1b70c5cf6e30e9a21afea48db

Observation 2d9bd9f2-87a3-4a43-863b-97f3b7a59aea · outbound

This paper cites an unresolved cited work.

EdgeBench: Unveiling Scaling Laws of Learning from Real-World Environments Unresolved cited work

Reference 86

Resolution
unresolved
no resolver link, observed 2026-07-11T07:57:43.000834Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T07:57:43.000834Z digest=sha256:95667b630f8ac9ad9f89e09f6970be1d9a20ba2719e94e7993f85541fcc3c920

Observation 17b0d53d-3e14-4212-9ece-320167a3b5c4 · outbound

This paper cites Throughout the section, we denote raw interaction time byt > 0, and the raw task/benchmark score byS(t).

EdgeBench: Unveiling Scaling Laws of Learning from Real-World Environments Throughout the section, we denote raw interaction time byt > 0, and the raw task/benchmark score byS(t)

Reference 87

Resolution
malformed identifier
no resolver link, observed 2026-07-11T07:57:43.000834Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T07:57:43.000834Z digest=sha256:fde73b77bd206438224f450960e2b1ea084c5453fd3b0e111144ff6109bbcfd5

Pith citing papers

Observation c7ecb65b-43ab-4cc8-9de2-de63cc58e571 · inbound

Sample-Efficient Learning from Agent Experience cites this paper.

Sample-Efficient Learning from Agent Experience EdgeBench: Unveiling Scaling Laws of Learning from Real-World Environments

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-01T08:43:22.628158Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T08:43:22.628158Z digest=sha256:bd38eb225c8ea4aab8868d7115831cb0bdddeb525992c3a32760b11d5b493266

Observation 6bb0945f-10e2-4d11-910e-6d1883fab4cb · inbound

Do Agent Benchmarks Measure Capability? Protocol Validity in the Age of Agentic AI cites this paper.

Do Agent Benchmarks Measure Capability? Protocol Validity in the Age of Agentic AI EdgeBench: Unveiling Scaling Laws of Learning from Real-World Environments

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-01T05:03:45.411378Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T05:03:45.411378Z digest=sha256:188bd718640b8842e0b2546470306dd97273bd2654d0c91ddb55fa1c73e88976

Observation 4e34d1c5-958c-46d1-ae55-e7b29e03fa3f · inbound

The Physics of Multi-Turn Long-Horizon Planning: From Pre-training to Post-training via Single- and Multi-Teacher On-Policy Agentic Distillation cites this paper.

The Physics of Multi-Turn Long-Horizon Planning: From Pre-training to Post-training via Single- and Multi-Teacher On-Policy Agentic Distillation EdgeBench: Unveiling Scaling Laws of Learning from Real-World Environments

Reference 35

Resolution
unresolved
no resolver link, observed 2026-07-31T06:59:15.205579Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T06:59:15.205579Z digest=sha256:fcca409dee9bcfedaa2b35519647f6f2da5bcc9f7586032e06f177485fa508fd

Observation e58e4473-b390-4895-91ea-d33b71e320b1 · inbound

GDPevo: Evaluating Agent Self-Evolution on Real Business Tasks cites this paper.

GDPevo: Evaluating Agent Self-Evolution on Real Business Tasks EdgeBench: Unveiling Scaling Laws of Learning from Real-World Environments

Reference 7

Resolution
verified exact
local_arxiv, observed 2026-08-05T13:09:10.439914Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-05T13:09:06.494394Z digest=sha256:7d9d82f978476c08794af8cc69dde43f1a48df23711030bd213da6f0703051ca

Observation ce60c5f6-a37b-472c-b4e8-d654b87278e2 · inbound

VibeLifeBench: Can Your Life Agent Be Proactive and Persistent in a Living World? cites this paper.

VibeLifeBench: Can Your Life Agent Be Proactive and Persistent in a Living World? EdgeBench: Unveiling Scaling Laws of Learning from Real-World Environments

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-12T15:18:12.343176Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:18:12.343176Z digest=sha256:01191ead96fd802d03b87fd66e1cf45920ee5b7d771df2643745a48b1101dc24