Pith. sign in

Paper Citation Record · LEDGER

Can AI Validate Science? Benchmarking LLMs for Accurate Scientific Claim $\rightarrow$ Evidence Reasoning

As of 19 August 2026, this Paper Citation Record lists 25 of 25 outbound references and 2 inbound Pith citation observations for arXiv:2506.08235.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.08235 v1

Coverage vector

measured 25 of 25 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T05:20:16.002827Z

measured 27 of 27 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-19T06:32:44.657259+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-01T16:42:03.917326Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-06-29T12:03:24.010456Z

Reference resolution

25 of 25 outbound references displayed

  • verified exact1
  • verified fuzzy13
  • unresolved10
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch1

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 4c32a2a7-4c99-4a1c-a9f4-0543c9b9b8eb · outbound

This paper cites an unresolved cited work.

Can AI Validate Science? Benchmarking LLMs for Accurate Scientific Claim $\rightarrow$ Evidence Reasoning Unresolved cited work

Reference 1

Resolution
unresolved
raw_fallback, observed 2026-08-07T05:20:16.447954Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T05:20:15.910167Z digest=sha256:72ed9c7b96c6c695286df3556548fd007489a5402fcb96d2e54b693bc571975e

Observation b2a3cc3e-b0d8-448b-86b7-2d2796554e80 · outbound

This paper cites analysis.

Can AI Validate Science? Benchmarking LLMs for Accurate Scientific Claim $\rightarrow$ Evidence Reasoning analysis

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:20:16.434564Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T05:20:15.914615Z digest=sha256:6756f61e02b8f4cf0dbb156db63647ad4519b36535d9863c75523b1d0eb0c13f

Observation 9f01c245-3caf-4bda-92f0-9b67e010b907 · outbound

This paper cites ArXiv:2412.08185.

Can AI Validate Science? Benchmarking LLMs for Accurate Scientific Claim $\rightarrow$ Evidence Reasoning ArXiv:2412.08185

Reference 3

Resolution
verified exact
raw_fallback, observed 2026-08-07T05:20:16.152177Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T05:20:15.896947Z digest=sha256:4e6ad3c19eeaf2a6ba8ec426aca4fba383e2aad3e7da08b3fb710b05d57465b3

Observation e2f86da8-feec-4c89-9a0b-5b561bffe1f2 · outbound

This paper cites XL$^2$Bench: A Benchmark for Extremely Long Context Understanding with Long-range Dependencies.

Can AI Validate Science? Benchmarking LLMs for Accurate Scientific Claim $\rightarrow$ Evidence Reasoning XL$^2$Bench: A Benchmark for Extremely Long Context Understanding with Long-range Dependencies

Reference 4

Resolution
metadata mismatch
local_arxiv, observed 2026-08-07T05:20:16.048050Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T05:20:15.900985Z digest=sha256:e9ef9c11d1cd78ad3651e7344d27c688809b46ad7c716c3d92e1d3b8b4bdf378

Observation 863244dd-c774-444e-82fd-6d821c0a3537 · outbound

This paper cites InThe Thirteenth International Conference on Learning Representa- tions.

Can AI Validate Science? Benchmarking LLMs for Accurate Scientific Claim $\rightarrow$ Evidence Reasoning InThe Thirteenth International Conference on Learning Representa- tions

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:20:16.461676Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T05:20:15.905791Z digest=sha256:44d99e90efcf986edbb5023e86a2c770257c495480187065eb90ff7bdbca8b5c

Observation 5279927f-3e2c-4bfe-8cbe-fa5000bdb776 · outbound

This paper cites claims": [ {.

Can AI Validate Science? Benchmarking LLMs for Accurate Scientific Claim $\rightarrow$ Evidence Reasoning claims": [ {

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:20:16.394401Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T05:20:15.926722Z digest=sha256:d5434cb4ddbf1728de081f1cd16f40a9fdca2784582ef35e7d4fb25a836e4991

Observation 89192eff-0186-41c1-874d-37047daf369d · outbound

This paper cites an unresolved cited work.

Can AI Validate Science? Benchmarking LLMs for Accurate Scientific Claim $\rightarrow$ Evidence Reasoning Unresolved cited work

Reference 12

Resolution
unresolved
raw_fallback, observed 2026-08-07T05:20:16.367515Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T05:20:15.935746Z digest=sha256:d48086894a2cc59e9c8acfc4c61d4e9ad314a13c217fa8fcb0afae378fc0a576

Observation 49a12ffb-5f0b-4fe3-be40-ed7b7cb583b4 · outbound

This paper cites evidence_sets.

Can AI Validate Science? Benchmarking LLMs for Accurate Scientific Claim $\rightarrow$ Evidence Reasoning evidence_sets

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:20:16.340199Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T05:20:15.943752Z digest=sha256:ef2cff2ab7b9d120df49a18c4e7241093e04275e82d0ba25b400b4b334e297c7

Observation ed4500ee-614e-4b7b-b60f-06147d45b0b1 · outbound

This paper cites an unresolved cited work.

Can AI Validate Science? Benchmarking LLMs for Accurate Scientific Claim $\rightarrow$ Evidence Reasoning Unresolved cited work

Reference 15

Resolution
unresolved
raw_fallback, observed 2026-08-07T05:20:16.327160Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T05:20:15.947762Z digest=sha256:a7820bfa86840d48bc04aa7d52026a44a431eeb98027d2e85938cb0a7c1fcf96

Observation eeb6bcc2-7d60-423e-8e3a-0fed56249ae7 · outbound

This paper cites an unresolved cited work.

Can AI Validate Science? Benchmarking LLMs for Accurate Scientific Claim $\rightarrow$ Evidence Reasoning Unresolved cited work

Reference 16

Resolution
unresolved
raw_fallback, observed 2026-08-07T05:20:16.313022Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T05:20:15.951837Z digest=sha256:9217db5e2ebf4bc070857345ad5404b708b7bd1c162210ddeacd84bf6ce0465d

Observation 680fef4b-1c4e-40a8-a3b6-73408549d669 · outbound

This paper cites conclusions.

Can AI Validate Science? Benchmarking LLMs for Accurate Scientific Claim $\rightarrow$ Evidence Reasoning conclusions

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:20:16.300044Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T05:20:15.955384Z digest=sha256:c2d8f3806c763555514b23e3e7b5a0a1d468ffb5b9a80c79da519bc6ebb6cfff

Observation 3d45f78c-dca3-49ed-8885-6cb4e36199f9 · outbound

This paper cites an unresolved cited work.

Can AI Validate Science? Benchmarking LLMs for Accurate Scientific Claim $\rightarrow$ Evidence Reasoning Unresolved cited work

Reference 18

Resolution
unresolved
raw_fallback, observed 2026-08-07T05:20:16.421143Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T05:20:15.959848Z digest=sha256:c929d2438983e8c6d166a583ea4b4eb14e8a4ad647294d082600aad8823dc837

Observation deafc3d6-3d50-4696-a829-2a41e34ac214 · outbound

This paper cites an unresolved cited work.

Can AI Validate Science? Benchmarking LLMs for Accurate Scientific Claim $\rightarrow$ Evidence Reasoning Unresolved cited work

Reference 19

Resolution
unresolved
raw_fallback, observed 2026-08-07T05:20:16.407445Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T05:20:15.963777Z digest=sha256:88cc4e3751ebb4686bbf6bb8b777d2c976de4621cd3d1c6d1ee00734a7b99e8b

Observation a3773792-9e70-44a0-987a-8af49e5b9583 · outbound

This paper cites claims": [ { 13.

Can AI Validate Science? Benchmarking LLMs for Accurate Scientific Claim $\rightarrow$ Evidence Reasoning claims": [ { 13

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:20:16.286527Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T05:20:15.967510Z digest=sha256:d9af24178e28b90adb8e54498b8178d16143937f26de8735f1399a3446eb80f6

Observation 5077d8d6-c97d-4f93-a807-3b18501cd66c · outbound

This paper cites an unresolved cited work.

Can AI Validate Science? Benchmarking LLMs for Accurate Scientific Claim $\rightarrow$ Evidence Reasoning Unresolved cited work

Reference 21

Resolution
unresolved
raw_fallback, observed 2026-08-07T05:20:16.380457Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T05:20:15.971758Z digest=sha256:0256188630cd76a273348e65c222dc782bfafc1d9a167a16c091de0c05527745

Observation fc792574-d942-4a26-9237-fcf602ddb8fb · outbound

This paper cites an unresolved cited work.

Can AI Validate Science? Benchmarking LLMs for Accurate Scientific Claim $\rightarrow$ Evidence Reasoning Unresolved cited work

Reference 22

Resolution
unresolved
raw_fallback, observed 2026-08-07T05:20:16.272785Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T05:20:15.975279Z digest=sha256:8ced3add7e3dca4a87e85714950a01a99a89682244c36ca6bcb57e901819c4c7

Observation 833ab0a3-b12f-4096-a324-38e9606baede · outbound

This paper cites an unresolved cited work.

Can AI Validate Science? Benchmarking LLMs for Accurate Scientific Claim $\rightarrow$ Evidence Reasoning Unresolved cited work

Reference 23

Resolution
unresolved
raw_fallback, observed 2026-08-07T05:20:16.353996Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T05:20:15.979115Z digest=sha256:d69ab0039c64237040b65e28a5dc25ca1c148e7e821ee51dcc6f05680d403eca

Observation d90a50a2-1fc2-47ad-8875-b55432db6df5 · outbound

This paper cites claim_id.

Can AI Validate Science? Benchmarking LLMs for Accurate Scientific Claim $\rightarrow$ Evidence Reasoning claim_id

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:20:16.256921Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T05:20:15.982689Z digest=sha256:4b7913502c6e2833b8aedf3ccfef2e3e8ff76515795d4df85e6eeef7b1ef0179

Observation 079b007c-1e96-4286-a262-b28cb93941e5 · outbound

This paper cites • Consider both supporting and contradicting evidence.

Can AI Validate Science? Benchmarking LLMs for Accurate Scientific Claim $\rightarrow$ Evidence Reasoning • Consider both supporting and contradicting evidence

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:20:16.241245Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T05:20:15.986240Z digest=sha256:9b052076310602b5ce44f233ecaa6df1d30b20742ee5a3550ebd0759d2dc5284

Observation 9af2ed57-83bf-4a0d-bd63-ca7450b939ad · outbound

This paper cites • Evaluate if the conclusion is justified by the evidence.

Can AI Validate Science? Benchmarking LLMs for Accurate Scientific Claim $\rightarrow$ Evidence Reasoning • Evaluate if the conclusion is justified by the evidence

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:20:16.226446Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T05:20:15.990226Z digest=sha256:84c1710051351e4b850529c0f212ccaa5ec4348381ddf70a4ccfaafa96ca105c

Observation 07ce018c-e886-47ed-917e-607f14a7321a · outbound

This paper cites • Consider methodological strengths and weaknesses.

Can AI Validate Science? Benchmarking LLMs for Accurate Scientific Claim $\rightarrow$ Evidence Reasoning • Consider methodological strengths and weaknesses

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:20:16.211659Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T05:20:15.994215Z digest=sha256:711ec9e160fa24c760e047604ee4638b17a5cd09900cd0fc95fd5aa940e7cdc9

Observation 97179451-bfee-4292-8c70-515db52230e8 · outbound

This paper cites conclusions.

Can AI Validate Science? Benchmarking LLMs for Accurate Scientific Claim $\rightarrow$ Evidence Reasoning conclusions

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:20:16.197673Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T05:20:15.998335Z digest=sha256:7dbda6fd2725377ed2534b1eb9abd58366ed15cdc8cc724bb532d47af61d578c

Observation f0299097-2e65-4ef4-8479-327fce363239 · outbound

This paper cites Claim_id.

Can AI Validate Science? Benchmarking LLMs for Accurate Scientific Claim $\rightarrow$ Evidence Reasoning Claim_id

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:20:16.184114Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T05:20:16.002827Z digest=sha256:f95806f19525866d2ccbeafb01401cd4a12d090e99c4b6d371f49bfab9f22386

Observation 98cf8914-c989-4342-b579-d19a9d66b28d · outbound

This paper cites Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone.

Can AI Validate Science? Benchmarking LLMs for Accurate Scientific Claim $\rightarrow$ Evidence Reasoning Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-07T05:20:15.886923Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:20:15.886923Z digest=sha256:ac0f3eafdb99ab059a0ca937b42ba7f08baca02181256f8590afce2f90f8436b

Observation 39ed1a4b-f5f4-4171-977a-0bb81d4fed10 · outbound

This paper cites In Proceedings of the 31st International Conference on Computational Linguistics, pages 3613–3630, Abu Dhabi, UAE.

Can AI Validate Science? Benchmarking LLMs for Accurate Scientific Claim $\rightarrow$ Evidence Reasoning In Proceedings of the 31st International Conference on Computational Linguistics, pages 3613–3630, Abu Dhabi, UAE

Reference 2025

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:20:16.475902Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T05:20:15.892674Z digest=sha256:5c5551dfbc2ce2d0c06d95cb4121ca99c4660c6a2e87d956ae3da1ca359e08bd

Pith citing papers

Observation cbb972ac-4212-4e08-a69d-b48f74ef6aac · inbound

ResearchLoop: An Evidence-Gated Control Plane for AI-Assisted Research cites this paper.

ResearchLoop: An Evidence-Gated Control Plane for AI-Assisted Research Can AI Validate Science? Benchmarking LLMs for Accurate Scientific Claim $\rightarrow$ Evidence Reasoning

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-06-29T12:03:24.012115Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-06-29T11:58:23.295488Z digest=sha256:60e705cf4c830c7c4ca778eabf137119f3b65ec9310cf0ab40d57fce5067e7ec

Observation cbf45071-fa48-4849-8fde-3e32c03cca93 · inbound

PEARL: Auditable Repair for Scientific Reasoning Graph Extraction cites this paper.

PEARL: Auditable Repair for Scientific Reasoning Graph Extraction Can AI Validate Science? Benchmarking LLMs for Accurate Scientific Claim $\rightarrow$ Evidence Reasoning

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-01T16:42:03.917326Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T16:42:03.917326Z digest=sha256:ee3ff38885b1a1690997053c9ee2fa4bede2f505478e95a8e914e66b8a32507c