Pith. sign in

Paper Citation Record · LEDGER

Predicting Empirical AI Research Outcomes with Language Models

As of 8 August 2026, this Paper Citation Record lists 19 of 19 outbound references and 3 inbound Pith citation observations for arXiv:2506.00794.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.00794 v1

Coverage vector

measured 19 of 19 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T12:05:56.920833Z

measured 22 of 22 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 3 of 3 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-06T22:46:32.625500Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

19 of 19 outbound references displayed

  • verified exact0
  • verified fuzzy4
  • unresolved15
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

0
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

Observation eeedf0a2-6028-41b2-88b6-a18fc3c612c4 · outbound

This paper cites Humans vs Large Language Models: Judgmental Forecasting in an Era of Advanced AI.

Predicting Empirical AI Research Outcomes with Language Models Humans vs Large Language Models: Judgmental Forecasting in an Era of Advanced AI

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T12:05:55.463704Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:05:55.463704Z digest=sha256:db90ae53b9c0a891f01585df44a9aab41efdc240f78433541b8f2aba4ee144ea

Observation 5afbfc62-d585-46c0-b450-f74f715fbc02 · outbound

This paper cites LitLLM: A Toolkit for Scientific Literature Review.

Predicting Empirical AI Research Outcomes with Language Models LitLLM: A Toolkit for Scientific Literature Review

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T12:05:55.548904Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:05:55.548904Z digest=sha256:a782bceac309e32e6c4b57a482bca6810442db583dd72e8147c712a97f284be1

Observation fc8f2d72-b221-469b-ae01-f7db11fb3458 · outbound

This paper cites Alpacafarm: A simulation framework for methods that learn from human feedback.

Predicting Empirical AI Research Outcomes with Language Models Alpacafarm: A simulation framework for methods that learn from human feedback

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T12:05:55.630510Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:05:55.630510Z digest=sha256:a6d7cd6cdbbb07eb1ec344d84ce8d4b1b7ef8d8b8805b081ba67c7dfc4a605ed

Observation 1aa981e8-d446-4ee9-b2e8-61bed3beb569 · outbound

This paper cites Towards an AI co-scientist.

Predicting Empirical AI Research Outcomes with Language Models Towards an AI co-scientist

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T12:05:55.709041Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:05:55.709041Z digest=sha256:84b76c4d5dc902bd3f5c1543b404223261ae318f91886ce9f7bd15aa9385a541

Observation 49b6c088-00af-4006-bb02-7765bf1c0c7a · outbound

This paper cites Approaching Human-Level Forecasting with Language Models.

Predicting Empirical AI Research Outcomes with Language Models Approaching Human-Level Forecasting with Language Models

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T12:05:55.787770Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:05:55.787770Z digest=sha256:91de88a87812bc428983d94dc27ccd23fa6ae4b5b924bb20d3eedff6c61e72bd

Observation cdc97199-e562-4ef9-abd7-04dff8b60230 · outbound

This paper cites MLAgentBench: Evaluating Language Agents on Machine Learning Experimentation.

Predicting Empirical AI Research Outcomes with Language Models MLAgentBench: Evaluating Language Agents on Machine Learning Experimentation

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T12:05:55.906410Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:05:55.906410Z digest=sha256:9cc98126511ab6fc687a58b596acbc6426b22987461c31b4f6b15d97641a8ee1

Observation 0cd05966-ff5e-464d-9e8a-51b3cec2e9f0 · outbound

This paper cites ForecastBench: A Dynamic Benchmark of AI Forecasting Capabilities.

Predicting Empirical AI Research Outcomes with Language Models ForecastBench: A Dynamic Benchmark of AI Forecasting Capabilities

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T12:05:55.960132Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:05:55.960132Z digest=sha256:f86236b1fc76a243f9b793ba1bf9746203379f6976781126c6b971d036094204

Observation 4296a1ca-9f15-4fae-a8ed-b57e1d4e8ce4 · outbound

This paper cites Can large language models provide useful feedback on research papers? a large-scale empirical analysis.

Predicting Empirical AI Research Outcomes with Language Models Can large language models provide useful feedback on research papers? a large-scale empirical analysis

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:05:57.932804Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T12:05:56.082833Z digest=sha256:d822f2f4ed80d85f6d9bfa00ada38057e41853db2147ed074a9e0cb47f0b02af

Observation 2c1e4a31-56d0-4603-a789-7efdf5865715 · outbound

This paper cites The AI Scientist: Towards Fully Automated Open-Ended Scientific Discovery.

Predicting Empirical AI Research Outcomes with Language Models The AI Scientist: Towards Fully Automated Open-Ended Scientific Discovery

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T12:05:56.133782Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:05:56.133782Z digest=sha256:ec5147e0517561436ed3f3e2159f70ec61f923162e0b283d087116ea38a3eeec

Observation 38538177-e28b-497e-a42a-d26786893a42 · outbound

This paper cites Large language models surpass human experts in predicting neuroscience results.

Predicting Empirical AI Research Outcomes with Language Models Large language models surpass human experts in predicting neuroscience results

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:05:57.711264Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T12:05:56.207802Z digest=sha256:158abb0f9353197809064a6963088663ed34470214eaaa656197db6e879306ca

Observation ffb9baa2-5f72-488d-a125-056f2bf2e482 · outbound

This paper cites Neurips 2021 summary, 2021.

Predicting Empirical AI Research Outcomes with Language Models Neurips 2021 summary, 2021

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:05:57.544494Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T12:05:56.278474Z digest=sha256:02a321d7f803ab7a990b665fb302af3c7ba6e1398d9176d9453a7075b86db883

Observation 0dde4865-f67f-48c4-b329-bb5f47e9ba75 · outbound

This paper cites Automatic Prompt Optimization with "Gradient Descent" and Beam Search.

Predicting Empirical AI Research Outcomes with Language Models Automatic Prompt Optimization with "Gradient Descent" and Beam Search

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T12:05:56.352320Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:05:56.352320Z digest=sha256:44b58953612c9be9914e1b45e7ee994f105fc115451a2df52937b2ced144f7ed

Observation eb511723-d2e8-495a-bfb2-835ceb4dc392 · outbound

This paper cites Large Language Model Prediction Capabilities: Evidence from a Real-World Forecasting Tournament.

Predicting Empirical AI Research Outcomes with Language Models Large Language Model Prediction Capabilities: Evidence from a Real-World Forecasting Tournament

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T12:05:56.438248Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:05:56.438248Z digest=sha256:5d8a3bf22a79c78d891744afb03a4cd2094a0e6d461f2d3e6d66cf6d8ccec705

Observation 4095f34d-131c-4688-80e5-493e69d74eb7 · outbound

This paper cites Can LLMs Generate Novel Research Ideas? A Large-Scale Human Study with 100+ NLP Researchers.

Predicting Empirical AI Research Outcomes with Language Models Can LLMs Generate Novel Research Ideas? A Large-Scale Human Study with 100+ NLP Researchers

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T12:05:56.506033Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:05:56.506033Z digest=sha256:c9e37f3e087e552a23be54967ddc645f913480f3140405638776db78910f3162

Observation 45f0081e-14e6-41e4-a1a7-3671dc72148b · outbound

This paper cites Chain-of-thought prompting elicits reasoning in large language models.

Predicting Empirical AI Research Outcomes with Language Models Chain-of-thought prompting elicits reasoning in large language models

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T12:05:56.611944Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:05:56.611944Z digest=sha256:25f4a2ed59e7bb37e219239f27a0b938bbaad9621646de70d3fbc47d8514cc00

Observation 7ad5eaaa-d310-4f25-8756-19b310dafb9b · outbound

This paper cites The AI Scientist-v2: Workshop-Level Automated Scientific Discovery via Agentic Tree Search.

Predicting Empirical AI Research Outcomes with Language Models The AI Scientist-v2: Workshop-Level Automated Scientific Discovery via Agentic Tree Search

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T12:05:56.687259Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:05:56.687259Z digest=sha256:c4cd08c75cad9cf832706d73702de6434502c768bb9a53794615629f5265ba21

Observation bcad7cf2-68d6-4dcf-9079-bf410cf2d921 · outbound

This paper cites Star: Self-taught reasoner bootstrapping reasoning with reasoning.

Predicting Empirical AI Research Outcomes with Language Models Star: Self-taught reasoner bootstrapping reasoning with reasoning

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T12:05:56.756008Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:05:56.756008Z digest=sha256:9b8068c1d935f38d50b34281c5a053a4a99233a1d4765d2fe37849b4c12b9620

Observation 25018129-528f-42b4-a652-008c26bb6588 · outbound

This paper cites Goal driven discovery of distributional differences via language descriptions.

Predicting Empirical AI Research Outcomes with Language Models Goal driven discovery of distributional differences via language descriptions

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:05:57.331730Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T12:05:56.806527Z digest=sha256:2b6ee94cca9f1bb8b33868d807e3dfde0f25bc1686bbad4eea0ac331ef28b5fa

Observation 8a627cc6-e1e6-4218-9782-d04bf14298e5 · outbound

This paper cites Forecasting Future World Events with Neural Networks.

Predicting Empirical AI Research Outcomes with Language Models Forecasting Future World Events with Neural Networks

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T12:05:56.920833Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:05:56.920833Z digest=sha256:eb29e11f78570ca42329802bc9e62c3fb81bec23999781671e3b5e6a252cf402

Pith citing papers

Observation 75729d29-f9d4-45d6-9258-680d865673ea · inbound

The Ideation-Execution Gap: Execution Outcomes of LLM-Generated versus Human Research Ideas cites this paper.

The Ideation-Execution Gap: Execution Outcomes of LLM-Generated versus Human Research Ideas Predicting Empirical AI Research Outcomes with Language Models

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-06T22:46:32.625500Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:46:32.625500Z digest=sha256:5a3670756ef75df9b05720952b856c9dbb5e04cf34d00bbe641d62e5db77378b

Observation ad6885f1-249d-4223-b8d8-ee4b259d3fc8 · inbound

Advancing Event Forecasting through Massive Training of Large Language Models: Challenges, Solutions, and Broader Impacts cites this paper.

Advancing Event Forecasting through Massive Training of Large Language Models: Challenges, Solutions, and Broader Impacts Predicting Empirical AI Research Outcomes with Language Models

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-06T14:22:31.403988Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T14:22:31.403988Z digest=sha256:493b473bb3b4e281c4de66f35b8348058a6c1368f8e926f1542520e3d337b10f

Observation 6f01eadf-26ff-4fc2-a488-672666bbc266 · inbound

When AI reviews science: Can we trust the referee? cites this paper.

When AI reviews science: Can we trust the referee? Predicting Empirical AI Research Outcomes with Language Models

Reference 102

Resolution
verified exact
arxiv_id, observed 2026-05-08T23:19:30.523123Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-08T06:19:54.727724Z digest=sha256:2be1f2e0f9e41cf1f72e3cd829984ae5d52d213a9baf9209cd067e85ef7131ef