Pith. sign in

Paper Citation Record · LEDGER

Predicting Empirical AI Research Outcomes with Language Models

As of 18 August 2026, this Paper Citation Record lists 19 of 19 outbound references and 3 inbound Pith citation observations for arXiv:2506.00794.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.00794 v1

Coverage vector

measured 19 of 19 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T12:05:56.920833Z

measured 22 of 22 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00

measured 3 of 3 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-06T22:46:32.625500Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

19 of 19 outbound references displayed

  • verified exact0
  • verified fuzzy4
  • unresolved15
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

0
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

Observation eeedf0a2-6028-41b2-88b6-a18fc3c612c4 · outbound

This paper cites Humans vs Large Language Models: Judgmental Forecasting in an Era of Advanced AI.

Predicting Empirical AI Research Outcomes with Language Models Humans vs Large Language Models: Judgmental Forecasting in an Era of Advanced AI

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T12:05:55.463704Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:05:55.463704Z digest=sha256:bca450d72d5283095bbf0b96099c610933977f5932e2ba4f26913487f92c69f4

Observation 5afbfc62-d585-46c0-b450-f74f715fbc02 · outbound

This paper cites LitLLM: A Toolkit for Scientific Literature Review.

Predicting Empirical AI Research Outcomes with Language Models LitLLM: A Toolkit for Scientific Literature Review

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T12:05:55.548904Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:05:55.548904Z digest=sha256:98b451fbfe6de3b7c6a30b233d5e56fc594bb099dabd08edca0097cf50602fbd

Observation fc8f2d72-b221-469b-ae01-f7db11fb3458 · outbound

This paper cites Alpacafarm: A simulation framework for methods that learn from human feedback.

Predicting Empirical AI Research Outcomes with Language Models Alpacafarm: A simulation framework for methods that learn from human feedback

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T12:05:55.630510Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:05:55.630510Z digest=sha256:55d9afd63c50d361b7295d2ff0133b59c3a04fd7686aa161f125153e8fd35437

Observation 1aa981e8-d446-4ee9-b2e8-61bed3beb569 · outbound

This paper cites Towards an AI co-scientist.

Predicting Empirical AI Research Outcomes with Language Models Towards an AI co-scientist

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T12:05:55.709041Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:05:55.709041Z digest=sha256:a8c163905053e659aee22657985780318f22abf4b66f9d055e9c14ac616503b0

Observation 49b6c088-00af-4006-bb02-7765bf1c0c7a · outbound

This paper cites Approaching Human-Level Forecasting with Language Models.

Predicting Empirical AI Research Outcomes with Language Models Approaching Human-Level Forecasting with Language Models

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T12:05:55.787770Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:05:55.787770Z digest=sha256:24b44cb0c56d9ce7070a457fca7f9f1184ad7d2feb2f6cd4b86c8ce9e5a22c7a

Observation cdc97199-e562-4ef9-abd7-04dff8b60230 · outbound

This paper cites MLAgentBench: Evaluating Language Agents on Machine Learning Experimentation.

Predicting Empirical AI Research Outcomes with Language Models MLAgentBench: Evaluating Language Agents on Machine Learning Experimentation

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T12:05:55.906410Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:05:55.906410Z digest=sha256:1e427204cadc76f3fb3e5c080707e6fede70d816259e7986679c50397842ed8e

Observation 0cd05966-ff5e-464d-9e8a-51b3cec2e9f0 · outbound

This paper cites ForecastBench: A Dynamic Benchmark of AI Forecasting Capabilities.

Predicting Empirical AI Research Outcomes with Language Models ForecastBench: A Dynamic Benchmark of AI Forecasting Capabilities

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T12:05:55.960132Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:05:55.960132Z digest=sha256:00140c87a006132b090f5fa24ff4f49c597ac72715765414b1949f69ea2ee627

Observation 4296a1ca-9f15-4fae-a8ed-b57e1d4e8ce4 · outbound

This paper cites Can large language models provide useful feedback on research papers? a large-scale empirical analysis.

Predicting Empirical AI Research Outcomes with Language Models Can large language models provide useful feedback on research papers? a large-scale empirical analysis

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:05:57.932804Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T12:05:56.082833Z digest=sha256:14cc5777ed41d2f7693f10ff754cfe69e8da2e36cc0fdb75a238414729881624

Observation 2c1e4a31-56d0-4603-a789-7efdf5865715 · outbound

This paper cites The AI Scientist: Towards Fully Automated Open-Ended Scientific Discovery.

Predicting Empirical AI Research Outcomes with Language Models The AI Scientist: Towards Fully Automated Open-Ended Scientific Discovery

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T12:05:56.133782Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:05:56.133782Z digest=sha256:226f268457db6bea4bdcafbacd18c59e089cf10f99590dffbbb37318ad349d00

Observation 38538177-e28b-497e-a42a-d26786893a42 · outbound

This paper cites Large language models surpass human experts in predicting neuroscience results.

Predicting Empirical AI Research Outcomes with Language Models Large language models surpass human experts in predicting neuroscience results

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:05:57.711264Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T12:05:56.207802Z digest=sha256:a924ff9ade1c3c1d047142464cffcd4f316bf8887eb0cda527189f2f7c2f55cb

Observation ffb9baa2-5f72-488d-a125-056f2bf2e482 · outbound

This paper cites Neurips 2021 summary, 2021.

Predicting Empirical AI Research Outcomes with Language Models Neurips 2021 summary, 2021

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:05:57.544494Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T12:05:56.278474Z digest=sha256:9f532e86ae6a42170dd37cabfc8f00d7a039b617cc80f45b583341b30b62013e

Observation 0dde4865-f67f-48c4-b329-bb5f47e9ba75 · outbound

This paper cites Automatic Prompt Optimization with "Gradient Descent" and Beam Search.

Predicting Empirical AI Research Outcomes with Language Models Automatic Prompt Optimization with "Gradient Descent" and Beam Search

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T12:05:56.352320Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:05:56.352320Z digest=sha256:4d632bb25233d9062a85809a3a38a0936336e99d1610ebb76905ddb620c79a8f

Observation eb511723-d2e8-495a-bfb2-835ceb4dc392 · outbound

This paper cites Large Language Model Prediction Capabilities: Evidence from a Real-World Forecasting Tournament.

Predicting Empirical AI Research Outcomes with Language Models Large Language Model Prediction Capabilities: Evidence from a Real-World Forecasting Tournament

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T12:05:56.438248Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:05:56.438248Z digest=sha256:7f7219b026494d6a61da229ad35b0291e9754db8ed592d897984f303ad66752a

Observation 4095f34d-131c-4688-80e5-493e69d74eb7 · outbound

This paper cites Can LLMs Generate Novel Research Ideas? A Large-Scale Human Study with 100+ NLP Researchers.

Predicting Empirical AI Research Outcomes with Language Models Can LLMs Generate Novel Research Ideas? A Large-Scale Human Study with 100+ NLP Researchers

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T12:05:56.506033Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:05:56.506033Z digest=sha256:28933f1afd9e6228dd8c3fccfa117d0bfd14e316ea009df4af878d98d53bb2fb

Observation 45f0081e-14e6-41e4-a1a7-3671dc72148b · outbound

This paper cites Chain-of-thought prompting elicits reasoning in large language models.

Predicting Empirical AI Research Outcomes with Language Models Chain-of-thought prompting elicits reasoning in large language models

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T12:05:56.611944Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:05:56.611944Z digest=sha256:74fd259511b5fb8681f9e14893815469cff7ec94635f5caaa02689c34bc09f81

Observation 7ad5eaaa-d310-4f25-8756-19b310dafb9b · outbound

This paper cites The AI Scientist-v2: Workshop-Level Automated Scientific Discovery via Agentic Tree Search.

Predicting Empirical AI Research Outcomes with Language Models The AI Scientist-v2: Workshop-Level Automated Scientific Discovery via Agentic Tree Search

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T12:05:56.687259Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:05:56.687259Z digest=sha256:2b78c45ca74550440e57d191ef8e31c4bdc353b0d993961d857decf4dad4ae1b

Observation bcad7cf2-68d6-4dcf-9079-bf410cf2d921 · outbound

This paper cites Star: Self-taught reasoner bootstrapping reasoning with reasoning.

Predicting Empirical AI Research Outcomes with Language Models Star: Self-taught reasoner bootstrapping reasoning with reasoning

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T12:05:56.756008Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:05:56.756008Z digest=sha256:ea2d1ae089889bbc63c31b2cc0473e14d6ec0cd20a8df8f4e31a0fe121ec1918

Observation 25018129-528f-42b4-a652-008c26bb6588 · outbound

This paper cites Goal driven discovery of distributional differences via language descriptions.

Predicting Empirical AI Research Outcomes with Language Models Goal driven discovery of distributional differences via language descriptions

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:05:57.331730Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T12:05:56.806527Z digest=sha256:ecea92b7186f17415acf5cfe0e448ccbf2da7f351d892be06850a0177eaab626

Observation 8a627cc6-e1e6-4218-9782-d04bf14298e5 · outbound

This paper cites Forecasting Future World Events with Neural Networks.

Predicting Empirical AI Research Outcomes with Language Models Forecasting Future World Events with Neural Networks

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T12:05:56.920833Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:05:56.920833Z digest=sha256:ee0e524c36fa522333f01d678ff5adeb1510b62cdab3f3df1427359db83993f1

Pith citing papers

Observation 75729d29-f9d4-45d6-9258-680d865673ea · inbound

The Ideation-Execution Gap: Execution Outcomes of LLM-Generated versus Human Research Ideas cites this paper.

The Ideation-Execution Gap: Execution Outcomes of LLM-Generated versus Human Research Ideas Predicting Empirical AI Research Outcomes with Language Models

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-06T22:46:32.625500Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:46:32.625500Z digest=sha256:a4540d49093aa851ae6f695ff51f057b29570f7e393cb0f620c124301c2495bd

Observation ad6885f1-249d-4223-b8d8-ee4b259d3fc8 · inbound

Advancing Event Forecasting through Massive Training of Large Language Models: Challenges, Solutions, and Broader Impacts cites this paper.

Advancing Event Forecasting through Massive Training of Large Language Models: Challenges, Solutions, and Broader Impacts Predicting Empirical AI Research Outcomes with Language Models

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-06T14:22:31.403988Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T14:22:31.403988Z digest=sha256:2f10c1bf02b183594f22cd393bdc647765026b8adf62ae6259b89bd2b57adb16

Observation 6f01eadf-26ff-4fc2-a488-672666bbc266 · inbound

When AI reviews science: Can we trust the referee? cites this paper.

When AI reviews science: Can we trust the referee? Predicting Empirical AI Research Outcomes with Language Models

Reference 102

Resolution
verified exact
arxiv_id, observed 2026-05-08T23:19:30.523123Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-08T06:19:54.727724Z digest=sha256:250e10d8311a1fb97e262a6a3383c23d29d093a61a2340e7bff8bdabfd04c2e4