Pith. sign in

Paper Citation Record · LEDGER

Are Large Language Models Reliable AI Scientists? Assessing Reverse-Engineering of Black-Box Systems

As of 15 August 2026, this Paper Citation Record lists 99 of 99 outbound references and 3 inbound Pith citation observations for arXiv:2505.17968.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.17968 v1

Coverage vector

measured 99 of 99 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T14:40:40.704557Z

measured 102 of 102 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-15T06:32:42.880941+00:00

measured 3 of 3 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-07-11T20:19:27.650696Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-07-10T00:06:37.955246Z

Reference resolution

99 of 99 outbound references displayed

  • verified exact0
  • verified fuzzy37
  • unresolved61
  • parse uncertain1
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 82402fab-ecce-4773-907f-cd8bb8afa647 · outbound

This paper cites Large-Scale Bandit Problems and KWIK Learning.

Are Large Language Models Reliable AI Scientists? Assessing Reverse-Engineering of Black-Box Systems Large-Scale Bandit Problems and KWIK Learning

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T14:40:32.381858Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:40:32.381858Z digest=sha256:88fb0222ef29d92ac504dd742892711cdc77e5dab85ea1ef4182b9e96f4d6694

Observation 50c39405-4d30-48ca-9495-442ea9c4d0ee · outbound

This paper cites Queries and Concept Learning.

Are Large Language Models Reliable AI Scientists? Assessing Reverse-Engineering of Black-Box Systems Queries and Concept Learning

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T14:40:32.448282Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:40:32.448282Z digest=sha256:90feb4b22b3751414ae56840c1372b58b3071620a6c82775aff072c5e885ed5a

Observation 57f7fe63-fce0-45b3-bd70-3443d251a9f7 · outbound

This paper cites Inductive Inference: Theory and Methods.

Are Large Language Models Reliable AI Scientists? Assessing Reverse-Engineering of Black-Box Systems Inductive Inference: Theory and Methods

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T14:40:32.491594Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:40:32.491594Z digest=sha256:0d502489bc6a63f7539b5c6e2c1554814b4f505faf41260557e7d4ba99d1d8ff

Observation d30d95ee-ceca-4221-a946-e5cbf8dc69cb · outbound

This paper cites Claude 3.5 sonnet.

Are Large Language Models Reliable AI Scientists? Assessing Reverse-Engineering of Black-Box Systems Claude 3.5 sonnet

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T14:40:32.547987Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:40:32.547987Z digest=sha256:ed986c5b77f29c0681fc56932e367bc92f953c6d44ae09fcd0c4b1dac8c0d952

Observation 9693341c-04d0-41a0-9cb9-4c3f94395333 · outbound

This paper cites Toward Efficient Exploration by Large Language Model Agents.

Are Large Language Models Reliable AI Scientists? Assessing Reverse-Engineering of Black-Box Systems Toward Efficient Exploration by Large Language Model Agents

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T14:40:32.601633Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:40:32.601633Z digest=sha256:d5f28aa57f1349bbc166347e55c7432e8b434465299ec5efda96510895390ed0

Observation fcefdc3a-515d-4d59-93b9-9c562e09f279 · outbound

This paper cites A Markovian Decision Process.

Are Large Language Models Reliable AI Scientists? Assessing Reverse-Engineering of Black-Box Systems A Markovian Decision Process

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T14:40:32.640312Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:40:32.640312Z digest=sha256:732ed13a1f29a692f2dda1c3c57cd8230773889ff2a2411601180562cc711927

Observation fc58a8ce-a78f-466e-a746-4cebf820daae · outbound

This paper cites Using cognitive psychology to understand GPT-3.

Are Large Language Models Reliable AI Scientists? Assessing Reverse-Engineering of Black-Box Systems Using cognitive psychology to understand GPT-3

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T14:40:32.682542Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:40:32.682542Z digest=sha256:f948dc640c3a6b7a7d43053782b7350694b40ce7a6b5a215185aebc3681efdbe

Observation d291f60f-f390-4fdc-8cb9-c8ac0b12fa8c · outbound

This paper cites Variational inference: A review for statisticians.

Are Large Language Models Reliable AI Scientists? Assessing Reverse-Engineering of Black-Box Systems Variational inference: A review for statisticians

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T14:40:32.763095Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:40:32.763095Z digest=sha256:84591379df0e07f3d45514f13155ac02e19fb3f2ef2579fbb7457b373afc037d

Observation d22695ee-a3ea-41f9-8dd8-bf2b9b39962b · outbound

This paper cites R-MAX – A General Polynomial Time Algorithm for Near-Optimal Reinforcement Learning.

Are Large Language Models Reliable AI Scientists? Assessing Reverse-Engineering of Black-Box Systems R-MAX – A General Polynomial Time Algorithm for Near-Optimal Reinforcement Learning

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T14:40:32.855269Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:40:32.855269Z digest=sha256:d624648128c70a43666151ba36a75f30d3f233c21d8dfbdcae25ee9a020aa963

Observation 6bc69ef7-5ca6-44d8-8eaf-0c4f7ce9e93c · outbound

This paper cites Language models are few-shot learners.

Are Large Language Models Reliable AI Scientists? Assessing Reverse-Engineering of Black-Box Systems Language models are few-shot learners

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T14:40:32.911448Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:40:32.911448Z digest=sha256:40a93a60fc9d0a8caf1b5f1936c6f5c6871c47c7cbac139de3faf7decddbc159

Observation 43339939-93b2-46e5-a19c-abb7de83b6e5 · outbound

This paper cites Why Do Multi-Agent LLM Systems Fail?.

Are Large Language Models Reliable AI Scientists? Assessing Reverse-Engineering of Black-Box Systems Why Do Multi-Agent LLM Systems Fail?

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T14:40:32.993539Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:40:32.993539Z digest=sha256:97ed20a50c0ac38c6115ba1282280e70f721811f85704608ed488ed279039d52

Observation 6ce3290c-0b2a-44f6-8b06-4dd45832900d · outbound

This paper cites Bayesian Experimental Design: A Review.Statistical Science, pp.

Are Large Language Models Reliable AI Scientists? Assessing Reverse-Engineering of Black-Box Systems Bayesian Experimental Design: A Review.Statistical Science, pp

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T14:40:33.070856Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:40:33.070856Z digest=sha256:e2979c323725259960acc67d40f0030fb96ef59c5c709e4a85762ecb10088ec5

Observation 3f6f2e9b-8ebe-4a0e-80b4-0d4cc2750171 · outbound

This paper cites Do NOT Think That Much for 2+3=? On the Overthinking of o1-Like LLMs.

Are Large Language Models Reliable AI Scientists? Assessing Reverse-Engineering of Black-Box Systems Do NOT Think That Much for 2+3=? On the Overthinking of o1-Like LLMs

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T14:40:33.147832Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:40:33.147832Z digest=sha256:3fa0c9f4ad14c5158ae53e35c7f3d2825746db3bf9f2205f28cd27d3fa96ccfa

Observation 51fed5c1-5bc9-4750-9138-b24393c37def · outbound

This paper cites The first crank of the cultural ratchet: Learning and transmitting concepts through language.

Are Large Language Models Reliable AI Scientists? Assessing Reverse-Engineering of Black-Box Systems The first crank of the cultural ratchet: Learning and transmitting concepts through language

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T14:40:33.213779Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:40:33.213779Z digest=sha256:321adbdee13393594ebc52e2a54b5bc8ce3cd48eea1595b1c0cf224aef319abf

Observation 573ec285-8443-4393-8478-1bca2ad57f2b · outbound

This paper cites CogBench: a large language model walks into a psychology lab.

Are Large Language Models Reliable AI Scientists? Assessing Reverse-Engineering of Black-Box Systems CogBench: a large language model walks into a psychology lab

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T14:40:33.302731Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:40:33.302731Z digest=sha256:21e558e58b7845c50bc0bb065cb7fc1b0a9ea37792a97b92dd148864871fc320

Observation a02f0fd1-c699-42df-a7ca-9513f1fda469 · outbound

This paper cites The Danger of Overthinking: Examining the Reasoning-Action Dilemma in Agentic Tasks.

Are Large Language Models Reliable AI Scientists? Assessing Reverse-Engineering of Black-Box Systems The Danger of Overthinking: Examining the Reasoning-Action Dilemma in Agentic Tasks

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T14:40:33.403417Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:40:33.403417Z digest=sha256:b35579ef53fec3b2e321bf4f14c9a51dad7f0bd84412b80f8abcfa77bf0df949

Observation c841fc75-f7d6-4854-b798-1586206b5ac9 · outbound

This paper cites Uncertainty, Information, and Sequential Experiments.

Are Large Language Models Reliable AI Scientists? Assessing Reverse-Engineering of Black-Box Systems Uncertainty, Information, and Sequential Experiments

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T14:40:33.481437Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:40:33.481437Z digest=sha256:e6063729b3f997639629376b705f3e5452e1f25c192079ec21d7dc8354fd8d70

Observation 65f83589-7dfb-43e6-8e21-a237ddcfcf42 · outbound

This paper cites PILCO: A Model-Based and Data-Efficient Approach to Policy Search.

Are Large Language Models Reliable AI Scientists? Assessing Reverse-Engineering of Black-Box Systems PILCO: A Model-Based and Data-Efficient Approach to Policy Search

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T14:40:33.569042Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:40:33.569042Z digest=sha256:83101f0ddb6121dae0bf781eab958726f814fd798f1fed2960cfc842c0501606

Observation 3e2b302c-9af2-4c54-9c27-023bc5314e69 · outbound

This paper cites Aleatory or Epistemic? Does it Matter? Structural Safety, 31(2):105–112, 2009.

Are Large Language Models Reliable AI Scientists? Assessing Reverse-Engineering of Black-Box Systems Aleatory or Epistemic? Does it Matter? Structural Safety, 31(2):105–112, 2009

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T14:40:33.675725Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:40:33.675725Z digest=sha256:cad9489c5e3a4d40dd065102368bf2f7e0c522b104b42f468f82d4ed0bee7626

Observation f610df2c-264d-4db8-a8fc-ce0f5c8e817a · outbound

This paper cites Is In-Context Learning in Large Language Models Bayesian? A Martingale Perspective.

Are Large Language Models Reliable AI Scientists? Assessing Reverse-Engineering of Black-Box Systems Is In-Context Learning in Large Language Models Bayesian? A Martingale Perspective

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T14:40:33.779818Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:40:33.779818Z digest=sha256:c679ba52f15073956ef4e23e6148a9e2986d0304474f7f11a9abb62f957ca397

Observation 377321d7-611e-4dd4-93f7-db14bbafc4ee · outbound

This paper cites Variational Bayesian optimal experimental design.

Are Large Language Models Reliable AI Scientists? Assessing Reverse-Engineering of Black-Box Systems Variational Bayesian optimal experimental design

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T14:40:33.861953Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:40:33.861953Z digest=sha256:f6023d38f82736844042ee4f71e6636f24a457caf9f7d8277233956a022e62fe

Observation 5a31fd6a-78f1-4d59-baae-dfc5895de5c9 · outbound

This paper cites Baby steps in evaluating the capacities of large language models.

Are Large Language Models Reliable AI Scientists? Assessing Reverse-Engineering of Black-Box Systems Baby steps in evaluating the capacities of large language models

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T14:40:33.954207Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:40:33.954207Z digest=sha256:12db0e40c5cfd5f1b65588d96ede62fa7f440ee9df1065b9fee4a33f7a1527bb

Observation d67b3eaf-3302-4617-a054-1155f531c76c · outbound

This paper cites BoxingGym: Benchmarking progress in automated experimental design and model discovery.

Are Large Language Models Reliable AI Scientists? Assessing Reverse-Engineering of Black-Box Systems BoxingGym: Benchmarking progress in automated experimental design and model discovery

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T14:40:34.042557Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:40:34.042557Z digest=sha256:20b53b7c2fc582029d5d23f33e7c4022dd2d0eb5869666eaf523a6c81189fa5e

Observation 9a177777-7675-4fca-8622-89383c8e4bff · outbound

This paper cites Amplify scientific discovery with artificial intelligence.

Are Large Language Models Reliable AI Scientists? Assessing Reverse-Engineering of Black-Box Systems Amplify scientific discovery with artificial intelligence

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T14:40:34.118874Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:40:34.118874Z digest=sha256:8ace5eb67ed4fb19636e4274bbb057444b4dea22346e71db0e1b83ff3b2f6c7b

Observation 2aaa365c-eb5e-49a1-963f-e3f9dac601e1 · outbound

This paper cites ANOVA: Repeated measures.

Are Large Language Models Reliable AI Scientists? Assessing Reverse-Engineering of Black-Box Systems ANOVA: Repeated measures

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T14:40:34.160962Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:40:34.160962Z digest=sha256:c0bc082fee31fa2ffef21097a18541ca1fd1d2b45d08cfb35e847a08f25ebdaf

Observation b96652ac-dc3f-48b4-b0cb-b936df006f2b · outbound

This paper cites Towards an AI co-scientist.

Are Large Language Models Reliable AI Scientists? Assessing Reverse-Engineering of Black-Box Systems Towards an AI co-scientist

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T14:40:34.213138Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:40:34.213138Z digest=sha256:0f3c8a10c439b429b8336d5e7db14941c81270062e7a5a1857bdbcaca8987f2e

Observation 7d509de6-b5d1-4fe6-a570-73c8714ad341 · outbound

This paper cites The Llama 3 Herd of Models.

Are Large Language Models Reliable AI Scientists? Assessing Reverse-Engineering of Black-Box Systems The Llama 3 Herd of Models

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T14:40:34.277604Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:40:34.277604Z digest=sha256:05c17fe0f782defebea4317bf1b53ce2539a9b2072765e8cabfc155f3993771c

Observation e94454b9-3bf0-4669-96f1-ec0efc5c3a10 · outbound

This paper cites Bayes in the Age of Intelligent Machines.

Are Large Language Models Reliable AI Scientists? Assessing Reverse-Engineering of Black-Box Systems Bayes in the Age of Intelligent Machines

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T14:40:34.344865Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:40:34.344865Z digest=sha256:4cbbeceda21d2fc3aa03cf9542f1001f74ec2cb33fca4f656f6d54cf89b049c2

Observation 796414a5-9386-45a3-9c79-e35aefff5510 · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

Are Large Language Models Reliable AI Scientists? Assessing Reverse-Engineering of Black-Box Systems DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T14:40:34.407975Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:40:34.407975Z digest=sha256:70130dc055109b3cc1b9630142c067738996d61e013005dc5be28050ed68bf88

Observation 997b1a60-ef79-45a8-81af-4733c6156182 · outbound

This paper cites Can Large Language Models Detect Errors in Long Chain-of-Thought Reasoning?.

Are Large Language Models Reliable AI Scientists? Assessing Reverse-Engineering of Black-Box Systems Can Large Language Models Detect Errors in Long Chain-of-Thought Reasoning?

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T14:40:34.473655Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:40:34.473655Z digest=sha256:7ec4dc5fd60b814f2414b589315a9992854ecf7149ca7308e221a46b2f512eef

Observation c1d76051-ac80-4a40-a808-7ee91b4823a7 · outbound

This paper cites GPT-4o System Card.

Are Large Language Models Reliable AI Scientists? Assessing Reverse-Engineering of Black-Box Systems GPT-4o System Card

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T14:40:34.532686Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:40:34.532686Z digest=sha256:13ccea45e2544be76ec364c40dae3d65f539f37d0be9cd441cf83dcdb9fdaecd

Observation 7974c83d-b72e-4146-a36e-208df11b3aa8 · outbound

This paper cites What Do Learning Dynamics Reveal About Generalization in LLM Reasoning?.

Are Large Language Models Reliable AI Scientists? Assessing Reverse-Engineering of Black-Box Systems What Do Learning Dynamics Reveal About Generalization in LLM Reasoning?

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T14:40:34.600789Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:40:34.600789Z digest=sha256:c0fef59f934d8015daa37c29c0d71afd848bc70c3d1d79de3fb0ab1411ffa2d0

Observation 81194d79-d3a4-47e4-882b-f0d6797a95f9 · outbound

This paper cites Using the tools of cognitive science to understand large language models at different levels of analysis.

Are Large Language Models Reliable AI Scientists? Assessing Reverse-Engineering of Black-Box Systems Using the tools of cognitive science to understand large language models at different levels of analysis

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T14:40:34.640066Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:40:34.640066Z digest=sha256:65967de7e3a8a9fcaf1574cc8db6aa4f0a39c5fb823300cfe0b8fb6319a19789

Observation 1a170178-e8a4-4ea6-adee-ed86d9bcc9b9 · outbound

This paper cites A robust class of context- sensitive languages.

Are Large Language Models Reliable AI Scientists? Assessing Reverse-Engineering of Black-Box Systems A robust class of context- sensitive languages

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:40:49.565656Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T14:40:34.679985Z digest=sha256:f29de8aadc77412f2441d868113d3cf372cd95df55400c41b192a05b4f60a48a

Observation 1d339c4b-a002-4688-9b16-c4a20db7b5e8 · outbound

This paper cites Passive learning of active causal strategies in agents and language models.

Are Large Language Models Reliable AI Scientists? Assessing Reverse-Engineering of Black-Box Systems Passive learning of active causal strategies in agents and language models

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:40:49.352850Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T14:40:34.740445Z digest=sha256:305d736d579b28ec3a0f88160b5151acdacee6b11f8ef1a31e2d1a1404b4a996

Observation 08d6ff58-e230-40a8-8229-997dc742e396 · outbound

This paper cites Structured chain-of-thought prompting for code generation.

Are Large Language Models Reliable AI Scientists? Assessing Reverse-Engineering of Black-Box Systems Structured chain-of-thought prompting for code generation

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:40:49.168710Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T14:40:34.788602Z digest=sha256:31ebdbd080750226fb7e15dcf13c44daf735a67d9ce98b35b6d1af509ac64680

Observation 2396efc4-279d-4b28-9905-b61235b5033b · outbound

This paper cites Reducing Reinforcement Learning to KWIK Online Regression.

Are Large Language Models Reliable AI Scientists? Assessing Reverse-Engineering of Black-Box Systems Reducing Reinforcement Learning to KWIK Online Regression

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:40:48.955565Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T14:40:34.860427Z digest=sha256:0d4f5f563dd716e016802c4ce70ce2e74ef2c9869cedf08321371be03fd7e85e

Observation 4d7ba1ae-a463-4956-8e4a-de594c96296b · outbound

This paper cites Knows What It Knows: A Framework for Self-Aware Learning.

Are Large Language Models Reliable AI Scientists? Assessing Reverse-Engineering of Black-Box Systems Knows What It Knows: A Framework for Self-Aware Learning

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:40:48.835093Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T14:40:34.926262Z digest=sha256:8dedcd65403909a56e951a911508fefe1bc42c4c5830ce19276a22ac4b06bb62

Observation 35ebf653-8272-48c8-8b50-723822d32714 · outbound

This paper cites On a measure of the information provided by an experiment.

Are Large Language Models Reliable AI Scientists? Assessing Reverse-Engineering of Black-Box Systems On a measure of the information provided by an experiment

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-07T14:40:34.999414Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:40:34.999414Z digest=sha256:e996958961b9dcd8a865d6acbf1f8f5404b0c0c421be9b99c1c0b29fa82be3f2

Observation 7e304cd7-3ff4-4750-91d1-b9ed954c52af · outbound

This paper cites Learning quickly when irrelevant attributes abound: A new linear-threshold algorithm.

Are Large Language Models Reliable AI Scientists? Assessing Reverse-Engineering of Black-Box Systems Learning quickly when irrelevant attributes abound: A new linear-threshold algorithm

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:40:48.661957Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T14:40:35.093317Z digest=sha256:93f4e46606d0a3444440e1d852271af17d109668f530b57243c500b37f052778

Observation 1ce2be8f-e194-42a5-87d0-115a6f41818f · outbound

This paper cites Decoupling Exploration and Exploitation for Meta-Reinforcement Learning Without Sacrifices.

Are Large Language Models Reliable AI Scientists? Assessing Reverse-Engineering of Black-Box Systems Decoupling Exploration and Exploitation for Meta-Reinforcement Learning Without Sacrifices

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:40:48.477838Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T14:40:35.176378Z digest=sha256:dd4772357dfd9f3d7446918fe1a2e7142a573cb7dde8a74a64b6dd3c48a27c6d

Observation e3fe3d14-33ed-45ca-8b75-7ec14e8b2bc8 · outbound

This paper cites Large Language Models Assume People are More Rational than We Really are.

Are Large Language Models Reliable AI Scientists? Assessing Reverse-Engineering of Black-Box Systems Large Language Models Assume People are More Rational than We Really are

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-07T14:40:35.241752Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:40:35.241752Z digest=sha256:f9e60990953bcb5c2cde06bab06f2811ac2f6c29bb8da537bf92112475e6272a

Observation d8884055-f106-44e7-a599-64123dd7f2b3 · outbound

This paper cites Mind Your Step (by Step): Chain-of-Thought can Reduce Performance on Tasks where Thinking Makes Humans Worse.

Are Large Language Models Reliable AI Scientists? Assessing Reverse-Engineering of Black-Box Systems Mind Your Step (by Step): Chain-of-Thought can Reduce Performance on Tasks where Thinking Makes Humans Worse

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-07T14:40:35.314587Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:40:35.314587Z digest=sha256:2b30c90c6b55fcaefb610a421544e1bcec29d4505381ea6c9a435600a0048cbf

Observation 0913bfcd-5527-4056-921c-763140c1f2b2 · outbound

This paper cites The AI Scientist: Towards Fully Automated Open-Ended Scientific Discovery.

Are Large Language Models Reliable AI Scientists? Assessing Reverse-Engineering of Black-Box Systems The AI Scientist: Towards Fully Automated Open-Ended Scientific Discovery

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-07T14:40:35.395340Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:40:35.395340Z digest=sha256:a8c27621c58cb01c2f9b55afc89b68d5ec020b90e43f0f615e1f8a4c321c12db

Observation c49abd94-62aa-43cb-8152-c385fdd9970a · outbound

This paper cites Deconstructing Long Chain-of-Thought: A Structured Reasoning Optimization Framework for Long CoT Distillation.

Are Large Language Models Reliable AI Scientists? Assessing Reverse-Engineering of Black-Box Systems Deconstructing Long Chain-of-Thought: A Structured Reasoning Optimization Framework for Long CoT Distillation

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-07T14:40:35.476624Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:40:35.476624Z digest=sha256:9cf9cf7e03d426b5cf6df1687b3b47cc18ba572e6c9dbfcb4bfc4f045b5248b7

Observation 676edd7b-4582-4f39-86de-04ef2f90fe85 · outbound

This paper cites Category learning through active sampling.

Are Large Language Models Reliable AI Scientists? Assessing Reverse-Engineering of Black-Box Systems Category learning through active sampling

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:40:48.302269Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T14:40:35.562091Z digest=sha256:0d6afa2b279ee1a8fa54c3025ae5e8e1055a6ca488abe03864e20618282a4a5c

Observation 6c60ac6e-d798-4f3c-8d5c-a7c133cbac82 · outbound

This paper cites Is it better to select or to receive? learning via active and passive hypothesis testing.

Are Large Language Models Reliable AI Scientists? Assessing Reverse-Engineering of Black-Box Systems Is it better to select or to receive? learning via active and passive hypothesis testing

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:40:48.147093Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T14:40:35.624883Z digest=sha256:d961e18a7d2ecc36bbe3bf26ec9670669ee521a73dadb9b4a4b9991602d62a59

Observation 1841677b-fbc5-4fad-ac05-f55ef2ea0065 · outbound

This paper cites Modeling rapid language learning by distilling Bayesian priors into artificial neural networks.

Are Large Language Models Reliable AI Scientists? Assessing Reverse-Engineering of Black-Box Systems Modeling rapid language learning by distilling Bayesian priors into artificial neural networks

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-07T14:40:35.712060Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:40:35.712060Z digest=sha256:3db0958e7b0f910d6bc4779dcb3eddb7ccce378f00fe8c66b1eb4d1b09675f4c

Observation 05f80f70-7c1a-4c6d-a7a7-c357f7027ba9 · outbound

This paper cites Embers of autoregression show how large language models are shaped by the problem they are trained to solve.

Are Large Language Models Reliable AI Scientists? Assessing Reverse-Engineering of Black-Box Systems Embers of autoregression show how large language models are shaped by the problem they are trained to solve

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:40:47.995708Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T14:40:35.818075Z digest=sha256:71d1edec15d7dafc65c00c742e112de98d4560adbacbb5390cd9915433eaa516

Observation e3be7c28-fb95-48a7-a078-4fd94a17d26a · outbound

This paper cites MatPilot: an LLM-enabled AI Materials Scientist under the Framework of Human-Machine Collaboration.

Are Large Language Models Reliable AI Scientists? Assessing Reverse-Engineering of Black-Box Systems MatPilot: an LLM-enabled AI Materials Scientist under the Framework of Human-Machine Collaboration

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-07T14:40:35.888952Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:40:35.888952Z digest=sha256:787f6604e2d6b5c49f6cf7a33bbc68143b5343a395272b9415beb09ca03ba4b7

Observation caa38cd3-48dd-4eef-81c8-b9f9bf08921f · outbound

This paper cites Sparks of Science: Hypothesis Generation Using Structured Paper Data.

Are Large Language Models Reliable AI Scientists? Assessing Reverse-Engineering of Black-Box Systems Sparks of Science: Hypothesis Generation Using Structured Paper Data

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-07T14:40:36.009484Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:40:36.009484Z digest=sha256:527fd78d322b004a82fdbec0418a79499fbf02c751210fcb543fcbbdffc4e2e5

Observation 3c36647d-194f-4806-981d-f738f5196c12 · outbound

This paper cites (More) Efficient Reinforcement Learning via Posterior Sampling.

Are Large Language Models Reliable AI Scientists? Assessing Reverse-Engineering of Black-Box Systems (More) Efficient Reinforcement Learning via Posterior Sampling

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:40:47.815242Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T14:40:36.102354Z digest=sha256:776df72059766daa51ce799509d1f319edc613a8ff07681b7a8e946f0da5e760

Observation 038b21dc-64a8-4b1e-978d-80f4a4284dc1 · outbound

This paper cites Puterman.

Are Large Language Models Reliable AI Scientists? Assessing Reverse-Engineering of Black-Box Systems Puterman

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:40:47.645625Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T14:40:36.190057Z digest=sha256:589ef48d18670e9069ba64017dfdd9077f3b3ce74b2f72f5280a2ff4769c6f6f

Observation 844a8a34-8b35-471e-8b6e-faea4f154c0c · outbound

This paper cites Towards scientific discovery with generative ai: Progress, opportunities, and challenges.

Are Large Language Models Reliable AI Scientists? Assessing Reverse-Engineering of Black-Box Systems Towards scientific discovery with generative ai: Progress, opportunities, and challenges

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:40:47.543996Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T14:40:36.295907Z digest=sha256:658ebe1ac191fcd35eacc71ffc77c2283c567d1482a71d3706b9674c9d105efb

Observation 61eec6b9-b181-4567-a92f-b729e4834598 · outbound

This paper cites Diversity-Based Inference of Finite Automata.

Are Large Language Models Reliable AI Scientists? Assessing Reverse-Engineering of Black-Box Systems Diversity-Based Inference of Finite Automata

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:40:47.335383Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T14:40:36.395591Z digest=sha256:466f18924d6db0ce8656fa66b6be563e6edbc860b2161f839c5be89361b4aa0d

Observation e096bed7-de24-4dad-a5fe-528ebfa59b43 · outbound

This paper cites Inference of Finite Automata Using Homing Sequences.

Are Large Language Models Reliable AI Scientists? Assessing Reverse-Engineering of Black-Box Systems Inference of Finite Automata Using Homing Sequences

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:40:47.116883Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T14:40:36.486680Z digest=sha256:c1e0ecc607f4ad39c6535ebafb21a60426818f86c490047ab06c3bb8442eca5e

Observation 391f4971-cca2-42ce-b42e-b7e909f1f4f5 · outbound

This paper cites Jagadish, Marvin Mathony, Tobias Ludwig, and Eric Schulz.

Are Large Language Models Reliable AI Scientists? Assessing Reverse-Engineering of Black-Box Systems Jagadish, Marvin Mathony, Tobias Ludwig, and Eric Schulz

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-07T14:40:36.544531Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:40:36.544531Z digest=sha256:66e4c8f1716ceef669a90943ab64b5d8b824068adae6d6c6477275417272ea67

Observation ac32d513-3360-4f9d-a0a6-e992cb0f6269 · outbound

This paper cites Symbolic metaprogram search improves learning efficiency and explains rule learning in humans.

Are Large Language Models Reliable AI Scientists? Assessing Reverse-Engineering of Black-Box Systems Symbolic metaprogram search improves learning efficiency and explains rule learning in humans

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:40:46.746513Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T14:40:36.636665Z digest=sha256:bf2dd2fb78f597987f0f513f292d747179c2cc3d8d171fab0c7bf2b034d23745

Observation b397f9fd-5cb7-4c33-a49c-73948ca38d19 · outbound

This paper cites Trading off Mistakes and Don’t- Know Predictions.

Are Large Language Models Reliable AI Scientists? Assessing Reverse-Engineering of Black-Box Systems Trading off Mistakes and Don’t- Know Predictions

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:40:46.416593Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T14:40:36.706230Z digest=sha256:72bcad800bdf95419e4b2f10113ba3839bf8587f0e7bd88bb5d2ffc3cab99fb7

Observation c0a039d9-d524-4528-b9c8-acf3dc5d5e92 · outbound

This paper cites Agent Laboratory: Using LLM Agents as Research Assistants.

Are Large Language Models Reliable AI Scientists? Assessing Reverse-Engineering of Black-Box Systems Agent Laboratory: Using LLM Agents as Research Assistants

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-07T14:40:36.829520Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:40:36.829520Z digest=sha256:a287f274f83cda6fe4f945fc8d62851f3c12e4f8f8fb7e510edb24e6766ff920

Observation 0c74deb9-8ff3-4bd5-8068-21cb2a079a07 · outbound

This paper cites Active Learning Literature Survey.

Are Large Language Models Reliable AI Scientists? Assessing Reverse-Engineering of Black-Box Systems Active Learning Literature Survey

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:40:46.070277Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T14:40:36.913290Z digest=sha256:24b812fb5bb6326cdc18aaf6eb3d7e316b661cc22b058989f1e38ae621f4fb28

Observation 439cc0a6-169c-4dfb-a18c-432cc02b15ec · outbound

This paper cites Llm-sr: Scientific equation discovery via programming with large language models.

Are Large Language Models Reliable AI Scientists? Assessing Reverse-Engineering of Black-Box Systems Llm-sr: Scientific equation discovery via programming with large language models

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:40:45.807125Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T14:40:37.006923Z digest=sha256:45d18376a3d22ada36dccd395e2b32aed6af5c7603efb42e1bea4621d9027106

Observation 28738de6-bb3c-4bfb-8ab9-454ce709f0e5 · outbound

This paper cites Can LLMs Generate Novel Research Ideas? A Large-Scale Human Study with 100+ NLP Researchers.

Are Large Language Models Reliable AI Scientists? Assessing Reverse-Engineering of Black-Box Systems Can LLMs Generate Novel Research Ideas? A Large-Scale Human Study with 100+ NLP Researchers

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-07T14:40:37.118461Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:40:37.118461Z digest=sha256:9af5765f32259f09b8277df8c31d3a881cb4c8f6da90ada30501512eb9e4b5e1

Observation 7290d399-1b29-42ab-bc4f-f3030c56f46e · outbound

This paper cites To CoT or not to CoT? Chain-of-thought helps mainly on math and symbolic reasoning.

Are Large Language Models Reliable AI Scientists? Assessing Reverse-Engineering of Black-Box Systems To CoT or not to CoT? Chain-of-thought helps mainly on math and symbolic reasoning

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-07T14:40:37.243258Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:40:37.243258Z digest=sha256:bb56d476dd1f1be3c8600c90eb8dcf0dde2920086204ac66cc8d77d4ae2b5f5b

Observation 971622c1-4ba9-46c2-ab12-73b9b13d43ba · outbound

This paper cites PaperBench: Evaluating AI's Ability to Replicate AI Research.

Are Large Language Models Reliable AI Scientists? Assessing Reverse-Engineering of Black-Box Systems PaperBench: Evaluating AI's Ability to Replicate AI Research

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-07T14:40:37.301272Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:40:37.301272Z digest=sha256:d818c827da113846104ac1d0b6b8a7915912b0a5609c948391fd2e3b4558f960

Observation ae21c6d6-7306-4574-a60b-247ea474c416 · outbound

This paper cites An Analysis of Model-based interval estimation for Markov Decision Processes.

Are Large Language Models Reliable AI Scientists? Assessing Reverse-Engineering of Black-Box Systems An Analysis of Model-based interval estimation for Markov Decision Processes

Reference 66

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:40:45.551239Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T14:40:37.397807Z digest=sha256:b393b0882ef1af6243ee42bdb8834ac4ebd2949ff1b7ab062b8b8bd563fd1fff

Observation 35f9fa0f-d275-4436-88e0-646ebed51a86 · outbound

This paper cites A Bayesian framework for reinforcement learning.

Are Large Language Models Reliable AI Scientists? Assessing Reverse-Engineering of Black-Box Systems A Bayesian framework for reinforcement learning

Reference 67

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:40:45.290247Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T14:40:37.465247Z digest=sha256:7fd63cc7817be844d91162796d1b77335e89698eeebc8d450e7bc62c8daf7bb7

Observation 87f047f3-9fbf-4cf9-b538-640ffa5b6c2a · outbound

This paper cites Stop Overthinking: A Survey on Efficient Reasoning for Large Language Models.

Are Large Language Models Reliable AI Scientists? Assessing Reverse-Engineering of Black-Box Systems Stop Overthinking: A Survey on Efficient Reasoning for Large Language Models

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-07T14:40:37.537997Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:40:37.537997Z digest=sha256:8e8963073e136fe6d0aef02ce10f62d5e7c34e6d5dd80909c303fe6975b4b11d

Observation 2c3ffde1-9c2f-435d-84b1-6a78dc4c7403 · outbound

This paper cites Integrated architectures for learning, planning, and reacting based on approximating dynamic programming.

Are Large Language Models Reliable AI Scientists? Assessing Reverse-Engineering of Black-Box Systems Integrated architectures for learning, planning, and reacting based on approximating dynamic programming

Reference 69

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:40:45.084955Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T14:40:37.666221Z digest=sha256:f758f3c629da2d0e19630f30ab94be83c9e041ab00e6b1a3fdcc17f432d633bb

Observation e35a777f-eb96-4192-b9e3-aacbbb0f762c · outbound

This paper cites Dyna, an integrated architecture for learning, planning, and reacting.

Are Large Language Models Reliable AI Scientists? Assessing Reverse-Engineering of Black-Box Systems Dyna, an integrated architecture for learning, planning, and reacting

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-07T14:40:37.794791Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:40:37.794791Z digest=sha256:018cbe462f6eea2140cf3b26051c68d1762fef5a4f1e6a8a9cebcff91c60236c

Observation 58d2fbfd-c0df-4dcc-94c2-ab0efdb77019 · outbound

This paper cites Introduction to Reinforcement Learning.

Are Large Language Models Reliable AI Scientists? Assessing Reverse-Engineering of Black-Box Systems Introduction to Reinforcement Learning

Reference 71

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:40:44.912222Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T14:40:37.907356Z digest=sha256:e7229b9577f245c71a73b65debc16e27b33b7ec7961d4bc8c6d72e9b76197c33

Observation 028339c6-93e4-4e83-8fe3-50410e086460 · outbound

This paper cites Agnostic KWIK learning and Efficient Approximate Reinforcement Learning.

Are Large Language Models Reliable AI Scientists? Assessing Reverse-Engineering of Black-Box Systems Agnostic KWIK learning and Efficient Approximate Reinforcement Learning

Reference 72

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:40:44.696029Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T14:40:38.027925Z digest=sha256:8689c183eeb80ed1ed26b807a9f75bcde7b502e437a104fad4b51bf985409bcd

Observation 9f2ec0af-3f1c-48dd-9e1d-f7c7ec585536 · outbound

This paper cites Active exploration in dynamic environments.

Are Large Language Models Reliable AI Scientists? Assessing Reverse-Engineering of Black-Box Systems Active exploration in dynamic environments

Reference 73

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:40:44.516965Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T14:40:38.110372Z digest=sha256:d32b242d9a05bb510a02257a4c848779b3d060b4716e8806afad0f1b762f557b

Observation 3af49d99-a977-495d-b799-48b19f383d99 · outbound

This paper cites Exploring Compact Reinforcement-Learning Representations with Linear Regression.

Are Large Language Models Reliable AI Scientists? Assessing Reverse-Engineering of Black-Box Systems Exploring Compact Reinforcement-Learning Representations with Linear Regression

Reference 74

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:40:44.353743Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T14:40:38.219036Z digest=sha256:bdeeb73944cb9b862cdcae1ce92f0fb93db19b15cf1e1d6e4d93344f120725d8

Observation 777acb38-9e0d-463b-9206-3e95660ebe3c · outbound

This paper cites Scientific Discovery in the Age of Artificial Intelligence.

Are Large Language Models Reliable AI Scientists? Assessing Reverse-Engineering of Black-Box Systems Scientific Discovery in the Age of Artificial Intelligence

Reference 75

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:40:44.215739Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T14:40:38.326222Z digest=sha256:9f898df70b04395d50daedf18e2fab63dc64dc6abb0b1949cdf92d9e943cb7e3

Observation b6f8b0d4-7ab8-4cd1-8299-3f881701ac04 · outbound

This paper cites Thoughts Are All Over the Place: On the Underthinking of o1-Like LLMs.

Are Large Language Models Reliable AI Scientists? Assessing Reverse-Engineering of Black-Box Systems Thoughts Are All Over the Place: On the Underthinking of o1-Like LLMs

Reference 76

Resolution
unresolved
no resolver link, observed 2026-08-07T14:40:38.414916Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:40:38.414916Z digest=sha256:56930f2ce0c33ac659866904015cdd5d27ee1efbd82b99d53fbaef984c2f5ea0

Observation 92d3dbcd-db87-46df-84ab-0cd4783b1d68 · outbound

This paper cites Chain-of-thought prompting elicits reasoning in large language models.

Are Large Language Models Reliable AI Scientists? Assessing Reverse-Engineering of Black-Box Systems Chain-of-thought prompting elicits reasoning in large language models

Reference 77

Resolution
unresolved
no resolver link, observed 2026-08-07T14:40:38.504795Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:40:38.504795Z digest=sha256:201a794898bd2f3e77decd11636a007a40503d7bfc1c41298a75a025052b6a8a

Observation eabe149f-13cf-467b-9b42-111b84afa1ed · outbound

This paper cites An explanation of in-context learning as implicit Bayesian inference.

Are Large Language Models Reliable AI Scientists? Assessing Reverse-Engineering of Black-Box Systems An explanation of in-context learning as implicit Bayesian inference

Reference 78

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:40:44.071384Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T14:40:38.625895Z digest=sha256:b832875996d80af7665813b3ceb9a6c99bbbfaf783806440e373be5e69a19541

Observation ba0347c7-6571-4e20-9d66-1b705672e0b1 · outbound

This paper cites Piantadosi.

Are Large Language Models Reliable AI Scientists? Assessing Reverse-Engineering of Black-Box Systems Piantadosi

Reference 79

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:40:43.980017Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T14:40:38.740427Z digest=sha256:c35642ac56d89863a2d23cbb9f874e069caee169b7cfda6ed997700a928d345a

Observation c6fa9a29-8b5d-4f61-90c6-ee414cc17cf0 · outbound

This paper cites On Benchmarking Human-Like Intelligence in Machines.

Are Large Language Models Reliable AI Scientists? Assessing Reverse-Engineering of Black-Box Systems On Benchmarking Human-Like Intelligence in Machines

Reference 80

Resolution
unresolved
no resolver link, observed 2026-08-07T14:40:38.848439Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:40:38.848439Z digest=sha256:53a6658f2ad3875c53e45f1ecfc916b08b8c11d58eddb44e2e9de7f52e7fc0b3

Observation 1d8b21e7-8c45-47e4-8a79-bc16e3fb95a8 · outbound

This paper cites People use fast, goal-directed simulation to reason about novel games.

Are Large Language Models Reliable AI Scientists? Assessing Reverse-Engineering of Black-Box Systems People use fast, goal-directed simulation to reason about novel games

Reference 81

Resolution
unresolved
no resolver link, observed 2026-08-07T14:40:38.988889Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:40:38.988889Z digest=sha256:23c6dacd59d7ffc010a2e902335d5ebeb15cccd66e70a217f949b99721dd6d26

Observation 5d74f7bd-c2f4-43c7-943a-c49c357804a9 · outbound

This paper cites Eliciting the Priors of Large Language Models using Iterated In-Context Learning.

Are Large Language Models Reliable AI Scientists? Assessing Reverse-Engineering of Black-Box Systems Eliciting the Priors of Large Language Models using Iterated In-Context Learning

Reference 82

Resolution
unresolved
no resolver link, observed 2026-08-07T14:40:39.080062Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:40:39.080062Z digest=sha256:12a7e5a3d15a30e377edc13184e6dd44eb296ceb02484194ea0a2cdbeca54c98

Observation cd18fa2d-5a8c-4ce0-b329-e413ad30db57 · outbound

This paper cites Incoherent Probability Judgments in Large Language Models.

Are Large Language Models Reliable AI Scientists? Assessing Reverse-Engineering of Black-Box Systems Incoherent Probability Judgments in Large Language Models

Reference 83

Resolution
unresolved
no resolver link, observed 2026-08-07T14:40:39.192857Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:40:39.192857Z digest=sha256:506ac69181d0fad576b0ba10d4df0b70fc7493f86a2d0b3ee58d911f118614eb

Observation fdf9af22-6755-463d-9fae-66dd7bb25d63 · outbound

This paper cites Provide a *thorough reasoning* before performing the action.

Are Large Language Models Reliable AI Scientists? Assessing Reverse-Engineering of Black-Box Systems Provide a *thorough reasoning* before performing the action

Reference 84

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:40:43.903505Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T14:40:39.315535Z digest=sha256:694349669e700bf3c1f1243e9cec6caff28144c4137462b566118987d3827e63

Observation 1b26e766-d0ee-4bc4-9852-8428de509d9a · outbound

This paper cites [5 point].

Are Large Language Models Reliable AI Scientists? Assessing Reverse-Engineering of Black-Box Systems [5 point]

Reference 85

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:40:43.769540Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T14:40:39.414387Z digest=sha256:3ebb61ca5c9a0f8b9aeb39702ec38e65d2e645521a7b2f9ca2c5098ec3631580

Observation 402086e6-04fc-44ae-b9b7-d37854a3b984 · outbound

This paper cites You will then output a score based on a set of assessment criteria.

Are Large Language Models Reliable AI Scientists? Assessing Reverse-Engineering of Black-Box Systems You will then output a score based on a set of assessment criteria

Reference 86

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:40:43.565586Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T14:40:39.489262Z digest=sha256:eb5190b120bb4bac9446b1a3d3bc00f20a067f68e7b45418d410fca1f6fa7e2c

Observation 2bd083ca-20b8-4d2e-ab63-70b2b9c2bba2 · outbound

This paper cites [3 points].

Are Large Language Models Reliable AI Scientists? Assessing Reverse-Engineering of Black-Box Systems [3 points]

Reference 87

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:40:43.409550Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T14:40:39.578875Z digest=sha256:00be84b731a6f7239ee30875b28bf6fe8636f0136922df0e894f2257aff386e2

Observation 02f91f73-9274-43e4-bdab-e63008affb85 · outbound

This paper cites an unresolved cited work.

Are Large Language Models Reliable AI Scientists? Assessing Reverse-Engineering of Black-Box Systems Unresolved cited work

Reference 88

Resolution
unresolved
raw_fallback, observed 2026-08-07T14:40:43.238105Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T14:40:39.698105Z digest=sha256:fab0c693689ff2e71ca0d7576acc3409bc4e3823301ee6084403e9e8ca85448e

Observation e0ab5211-cae1-4849-b8df-a67c01c50722 · outbound

This paper cites an unresolved cited work.

Are Large Language Models Reliable AI Scientists? Assessing Reverse-Engineering of Black-Box Systems Unresolved cited work

Reference 89

Resolution
unresolved
raw_fallback, observed 2026-08-07T14:40:43.073924Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T14:40:39.790190Z digest=sha256:0114c75d45529ee697d67f69a777b088774215fa6131785182dfb133f21fe5ac

Observation f8de0d1e-0dd6-4bfa-82fc-2dc4efb0b8c7 · outbound

This paper cites (Note that there will be multiple a_i 's.).

Are Large Language Models Reliable AI Scientists? Assessing Reverse-Engineering of Black-Box Systems (Note that there will be multiple a_i 's.)

Reference 90

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:40:42.903921Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T14:40:39.864719Z digest=sha256:8dd08c883733d9912110d18b0c477122e1f78e8235fe30129d6fcdc0a8fd6b16

Observation 7dde8e8c-00b8-4453-9b44-b0a605a0a75b · outbound

This paper cites an unresolved cited work.

Are Large Language Models Reliable AI Scientists? Assessing Reverse-Engineering of Black-Box Systems Unresolved cited work

Reference 91

Resolution
unresolved
raw_fallback, observed 2026-08-07T14:40:42.815299Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T14:40:39.964515Z digest=sha256:2aef794dff530305695194d00f3af69431ad503f39445c94c569a623a7b88193

Observation 091783d3-1335-4dc9-b82b-376aa5d26901 · outbound

This paper cites an unresolved cited work.

Are Large Language Models Reliable AI Scientists? Assessing Reverse-Engineering of Black-Box Systems Unresolved cited work

Reference 92

Resolution
unresolved
raw_fallback, observed 2026-08-07T14:40:42.686738Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T14:40:40.059543Z digest=sha256:6fd581ce7df46b6d2fd2b24ef3c655d3b42ed5d68cd6ba69fcadce188c0c76d7

Observation 1e0f90f9-c980-41b2-b8ff-5de72fd2854d · outbound

This paper cites The score for this bullet should be the accuracy percentage times the total allocated 6 points [6 points].

Are Large Language Models Reliable AI Scientists? Assessing Reverse-Engineering of Black-Box Systems The score for this bullet should be the accuracy percentage times the total allocated 6 points [6 points]

Reference 93

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:40:42.535878Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T14:40:40.163623Z digest=sha256:7caae8957bdeb8900504bd4335aca2c07be5cc9c54b18db30ddcc996c463a1bd

Observation 5c06d16d-e10c-4406-b32f-7e9fbd3d8dd9 · outbound

This paper cites an unresolved cited work.

Are Large Language Models Reliable AI Scientists? Assessing Reverse-Engineering of Black-Box Systems Unresolved cited work

Reference 94

Resolution
unresolved
raw_fallback, observed 2026-08-07T14:40:42.440082Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T14:40:40.257316Z digest=sha256:76f4433534171e29b1c5a75968dc1877c022e445ab254aada9a69ed0454f3328

Observation 1161a586-9ef7-4cc8-9640-82a58d2a0058 · outbound

This paper cites Win by connecting 3 stones in a column.

Are Large Language Models Reliable AI Scientists? Assessing Reverse-Engineering of Black-Box Systems Win by connecting 3 stones in a column

Reference 95

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:40:42.304484Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T14:40:40.338730Z digest=sha256:ea7947a4ba3730ee8924de3e08bbcc372867e2a6bfb0c753e7253d6111993307

Observation a46dffa2-aae7-4f09-8573-a72c4beaa0a2 · outbound

This paper cites an unresolved cited work.

Are Large Language Models Reliable AI Scientists? Assessing Reverse-Engineering of Black-Box Systems Unresolved cited work

Reference 96

Resolution
parse uncertain
raw_fallback, observed 2026-08-07T14:40:42.165301Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T14:40:40.418830Z digest=sha256:ece4c72e9d29e1fda46eec2e02c219eaf85d043a53d40a7555a5d9e0a1966063

Observation aa59d39a-2c8a-4d6f-8c50-875550e5f662 · outbound

This paper cites an unresolved cited work.

Are Large Language Models Reliable AI Scientists? Assessing Reverse-Engineering of Black-Box Systems Unresolved cited work

Reference 97

Resolution
unresolved
raw_fallback, observed 2026-08-07T14:40:41.981797Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T14:40:40.525674Z digest=sha256:3620f92c6e6e17df9122c584c29edc67eea0273927f6ee06e8daa40d1861808a

Observation 93d5b542-c867-4482-b680-55190bd4d052 · outbound

This paper cites an unresolved cited work.

Are Large Language Models Reliable AI Scientists? Assessing Reverse-Engineering of Black-Box Systems Unresolved cited work

Reference 98

Resolution
unresolved
raw_fallback, observed 2026-08-07T14:40:41.848711Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T14:40:40.632213Z digest=sha256:6408e7ba43a264807123388599adf2a959b41ad559a472b04bbe5cbfd6937a4c

Observation 99ce0b0d-b0af-4510-9b0a-d475301ec7dc · outbound

This paper cites AAA", "BBB.

Are Large Language Models Reliable AI Scientists? Assessing Reverse-Engineering of Black-Box Systems AAA", "BBB

Reference 99

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:40:41.698993Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T14:40:40.704557Z digest=sha256:37bb9d7299310e8028dd369946a60123ba57a8cdb75768570ea22c7cde101f58

Pith citing papers

Observation 3a5a7a26-a078-46cd-b7f9-9315ee619eec · inbound

A Survey of Self-Evolving Agents: What, When, How, and Where to Evolve on the Path to Artificial Super Intelligence cites this paper.

A Survey of Self-Evolving Agents: What, When, How, and Where to Evolve on the Path to Artificial Super Intelligence Are Large Language Models Reliable AI Scientists? Assessing Reverse-Engineering of Black-Box Systems

Reference 293

Resolution
verified exact
arxiv_id, observed 2026-05-14T22:23:15.430452Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-05-14T22:23:14.621091Z digest=sha256:bb862520cc6a8a311762649f997321287a333baca7c30631aee2d021cca3a480

Observation f0ecc73d-9b18-4cec-9108-f6ee09342aa6 · inbound

CausalGame: Benchmarking Causal Thinking of LLM Agents in Games cites this paper.

CausalGame: Benchmarking Causal Thinking of LLM Agents in Games Are Large Language Models Reliable AI Scientists? Assessing Reverse-Engineering of Black-Box Systems

Reference 3

Resolution
unresolved
no resolver link, observed 2026-07-11T20:19:27.650696Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T20:19:27.650696Z digest=sha256:01baf4ac23c694ebbbdfe6d984393001c51e8f2c4870fcb79eaee5876842116e

Observation 5bddc95b-cf75-4d8b-bd2c-3374077b922a · inbound

Mining Workflow Graphs for Black-Box Boundary Testing of Conversational LLM Agents cites this paper.

Mining Workflow Graphs for Black-Box Boundary Testing of Conversational LLM Agents Are Large Language Models Reliable AI Scientists? Assessing Reverse-Engineering of Black-Box Systems

Reference 26

Resolution
metadata mismatch
local_arxiv, observed 2026-07-10T00:06:37.956842Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-07-10T00:02:58.383840Z digest=sha256:9f2dae854b1d49ad6ac24e8eeb56f0ee837bae634e6756f324900147f6d623cd