Pith. sign in

Paper Citation Record · LEDGER

Can AI Validate Science? Benchmarking LLMs for Accurate Scientific Claim $\rightarrow$ Evidence Reasoning

As of 23 August 2026, this Paper Citation Record lists 25 of 25 outbound references and 2 inbound Pith citation observations for arXiv:2506.08235.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.08235 v1

Coverage vector

measured 25 of 25 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T05:20:16.002827Z

measured 27 of 27 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-23T06:30:58.430688+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-01T16:42:03.917326Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-06-29T12:03:24.010456Z

Reference resolution

25 of 25 outbound references displayed

  • verified exact1
  • verified fuzzy13
  • unresolved10
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch1

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 4c32a2a7-4c99-4a1c-a9f4-0543c9b9b8eb · outbound

This paper cites an unresolved cited work.

Can AI Validate Science? Benchmarking LLMs for Accurate Scientific Claim $\rightarrow$ Evidence Reasoning Unresolved cited work

Reference 1

Resolution
unresolved
raw_fallback, observed 2026-08-07T05:20:16.447954Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-07T05:20:15.910167Z digest=sha256:77e9d448824b3cc4fc227cccf3d3ad09052cce10162ff51ef1d647f93df76fc8

Observation b2a3cc3e-b0d8-448b-86b7-2d2796554e80 · outbound

This paper cites analysis.

Can AI Validate Science? Benchmarking LLMs for Accurate Scientific Claim $\rightarrow$ Evidence Reasoning analysis

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:20:16.434564Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-07T05:20:15.914615Z digest=sha256:7ed4ac8d11e4a5f61774c70783b9b4e5aaad4321b53ae7d223c11eba6d921c23

Observation 9f01c245-3caf-4bda-92f0-9b67e010b907 · outbound

This paper cites ArXiv:2412.08185.

Can AI Validate Science? Benchmarking LLMs for Accurate Scientific Claim $\rightarrow$ Evidence Reasoning ArXiv:2412.08185

Reference 3

Resolution
verified exact
raw_fallback, observed 2026-08-07T05:20:16.152177Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-07T05:20:15.896947Z digest=sha256:994b1328e0bf79481419d798b6244cfffd98e346b723c55309a53d651a5b47dc

Observation e2f86da8-feec-4c89-9a0b-5b561bffe1f2 · outbound

This paper cites XL$^2$Bench: A Benchmark for Extremely Long Context Understanding with Long-range Dependencies.

Can AI Validate Science? Benchmarking LLMs for Accurate Scientific Claim $\rightarrow$ Evidence Reasoning XL$^2$Bench: A Benchmark for Extremely Long Context Understanding with Long-range Dependencies

Reference 4

Resolution
metadata mismatch
local_arxiv, observed 2026-08-07T05:20:16.048050Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-07T05:20:15.900985Z digest=sha256:604c10391b5decd0b3b863ae6192483e2833e4f279e6bf9d97bda9b7297c802c

Observation 863244dd-c774-444e-82fd-6d821c0a3537 · outbound

This paper cites InThe Thirteenth International Conference on Learning Representa- tions.

Can AI Validate Science? Benchmarking LLMs for Accurate Scientific Claim $\rightarrow$ Evidence Reasoning InThe Thirteenth International Conference on Learning Representa- tions

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:20:16.461676Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-07T05:20:15.905791Z digest=sha256:3a356b8a6225d0786b929597ab55ca69af8f7aabcded565c72b12583086cccbd

Observation 5279927f-3e2c-4bfe-8cbe-fa5000bdb776 · outbound

This paper cites claims": [ {.

Can AI Validate Science? Benchmarking LLMs for Accurate Scientific Claim $\rightarrow$ Evidence Reasoning claims": [ {

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:20:16.394401Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-07T05:20:15.926722Z digest=sha256:2eaea9c084b1c0bbcf6ddd9d1dac55f86fa8b57402e01890ae00976500ce665a

Observation 89192eff-0186-41c1-874d-37047daf369d · outbound

This paper cites an unresolved cited work.

Can AI Validate Science? Benchmarking LLMs for Accurate Scientific Claim $\rightarrow$ Evidence Reasoning Unresolved cited work

Reference 12

Resolution
unresolved
raw_fallback, observed 2026-08-07T05:20:16.367515Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-07T05:20:15.935746Z digest=sha256:211d7fdaaf9d617da10606b00df1808c8715e92e6c4b98a18a9f38f16c25f6de

Observation 49a12ffb-5f0b-4fe3-be40-ed7b7cb583b4 · outbound

This paper cites evidence_sets.

Can AI Validate Science? Benchmarking LLMs for Accurate Scientific Claim $\rightarrow$ Evidence Reasoning evidence_sets

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:20:16.340199Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-07T05:20:15.943752Z digest=sha256:12af3269c01ee7ee6178a7f0a566619fb73d7c9c6c3780de5e79ebca958ca8cb

Observation ed4500ee-614e-4b7b-b60f-06147d45b0b1 · outbound

This paper cites an unresolved cited work.

Can AI Validate Science? Benchmarking LLMs for Accurate Scientific Claim $\rightarrow$ Evidence Reasoning Unresolved cited work

Reference 15

Resolution
unresolved
raw_fallback, observed 2026-08-07T05:20:16.327160Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-07T05:20:15.947762Z digest=sha256:06ba1b04e84d2b1348a8d1cccec40eedb1fb78cfefabf32775ca77ebb497c542

Observation eeb6bcc2-7d60-423e-8e3a-0fed56249ae7 · outbound

This paper cites an unresolved cited work.

Can AI Validate Science? Benchmarking LLMs for Accurate Scientific Claim $\rightarrow$ Evidence Reasoning Unresolved cited work

Reference 16

Resolution
unresolved
raw_fallback, observed 2026-08-07T05:20:16.313022Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-07T05:20:15.951837Z digest=sha256:f6868af396cb9e7a256f930b760623b1c12bd682bf543937e4f4ee238596a406

Observation 680fef4b-1c4e-40a8-a3b6-73408549d669 · outbound

This paper cites conclusions.

Can AI Validate Science? Benchmarking LLMs for Accurate Scientific Claim $\rightarrow$ Evidence Reasoning conclusions

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:20:16.300044Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-07T05:20:15.955384Z digest=sha256:1dd8440155f34ec87c7053fa2999bd00f7d544b463ecbdeca1d5fd1ac5854497

Observation 3d45f78c-dca3-49ed-8885-6cb4e36199f9 · outbound

This paper cites an unresolved cited work.

Can AI Validate Science? Benchmarking LLMs for Accurate Scientific Claim $\rightarrow$ Evidence Reasoning Unresolved cited work

Reference 18

Resolution
unresolved
raw_fallback, observed 2026-08-07T05:20:16.421143Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-07T05:20:15.959848Z digest=sha256:613c305038dfe761505d12e2104752353f01ced9d32fa4e43f6a58c0ed57544d

Observation deafc3d6-3d50-4696-a829-2a41e34ac214 · outbound

This paper cites an unresolved cited work.

Can AI Validate Science? Benchmarking LLMs for Accurate Scientific Claim $\rightarrow$ Evidence Reasoning Unresolved cited work

Reference 19

Resolution
unresolved
raw_fallback, observed 2026-08-07T05:20:16.407445Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-07T05:20:15.963777Z digest=sha256:c4440d6c945d44e14005df9d53cbb4b773b69345a20813380ea4a399b97bb7e8

Observation a3773792-9e70-44a0-987a-8af49e5b9583 · outbound

This paper cites claims": [ { 13.

Can AI Validate Science? Benchmarking LLMs for Accurate Scientific Claim $\rightarrow$ Evidence Reasoning claims": [ { 13

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:20:16.286527Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-07T05:20:15.967510Z digest=sha256:3e69ddb45cb1c90e6ff6880bb779da808557b7361c1a06214aec40965cf5d61f

Observation 5077d8d6-c97d-4f93-a807-3b18501cd66c · outbound

This paper cites an unresolved cited work.

Can AI Validate Science? Benchmarking LLMs for Accurate Scientific Claim $\rightarrow$ Evidence Reasoning Unresolved cited work

Reference 21

Resolution
unresolved
raw_fallback, observed 2026-08-07T05:20:16.380457Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-07T05:20:15.971758Z digest=sha256:f378248bda9f0afc9eed543d6220e7d515ca316590801c17cd8a19f0908cb7bd

Observation fc792574-d942-4a26-9237-fcf602ddb8fb · outbound

This paper cites an unresolved cited work.

Can AI Validate Science? Benchmarking LLMs for Accurate Scientific Claim $\rightarrow$ Evidence Reasoning Unresolved cited work

Reference 22

Resolution
unresolved
raw_fallback, observed 2026-08-07T05:20:16.272785Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-07T05:20:15.975279Z digest=sha256:1e4e68cadf58b1dea8177ac56e8143db1f1cc04dc01dab21cf7ebe8a15239db3

Observation 833ab0a3-b12f-4096-a324-38e9606baede · outbound

This paper cites an unresolved cited work.

Can AI Validate Science? Benchmarking LLMs for Accurate Scientific Claim $\rightarrow$ Evidence Reasoning Unresolved cited work

Reference 23

Resolution
unresolved
raw_fallback, observed 2026-08-07T05:20:16.353996Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-07T05:20:15.979115Z digest=sha256:cb45abd30f13359f974ea20273d44b4bcc6bc911c27ba20c49fb6ed4cbebda81

Observation d90a50a2-1fc2-47ad-8875-b55432db6df5 · outbound

This paper cites claim_id.

Can AI Validate Science? Benchmarking LLMs for Accurate Scientific Claim $\rightarrow$ Evidence Reasoning claim_id

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:20:16.256921Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-07T05:20:15.982689Z digest=sha256:ec14d6d35e2cbe12fff25168956f1ea4160ca833c1bd2d830d27fcc3446b86ec

Observation 079b007c-1e96-4286-a262-b28cb93941e5 · outbound

This paper cites • Consider both supporting and contradicting evidence.

Can AI Validate Science? Benchmarking LLMs for Accurate Scientific Claim $\rightarrow$ Evidence Reasoning • Consider both supporting and contradicting evidence

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:20:16.241245Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-07T05:20:15.986240Z digest=sha256:fe0ea931c2ee5f77ec851cbd5c37144c67fd4efc6229c9325912c2fef5035b54

Observation 9af2ed57-83bf-4a0d-bd63-ca7450b939ad · outbound

This paper cites • Evaluate if the conclusion is justified by the evidence.

Can AI Validate Science? Benchmarking LLMs for Accurate Scientific Claim $\rightarrow$ Evidence Reasoning • Evaluate if the conclusion is justified by the evidence

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:20:16.226446Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-07T05:20:15.990226Z digest=sha256:6c459a6775064082ab76b2f2ecabeb55ed7db8cc8368b94d7992ff86cc5534a9

Observation 07ce018c-e886-47ed-917e-607f14a7321a · outbound

This paper cites • Consider methodological strengths and weaknesses.

Can AI Validate Science? Benchmarking LLMs for Accurate Scientific Claim $\rightarrow$ Evidence Reasoning • Consider methodological strengths and weaknesses

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:20:16.211659Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-07T05:20:15.994215Z digest=sha256:9992f16c06daba730a7333fef320f9cbc7cf49639166934734057916fe68afe7

Observation 97179451-bfee-4292-8c70-515db52230e8 · outbound

This paper cites conclusions.

Can AI Validate Science? Benchmarking LLMs for Accurate Scientific Claim $\rightarrow$ Evidence Reasoning conclusions

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:20:16.197673Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-07T05:20:15.998335Z digest=sha256:dbce38257b43beb5d58cfab639c03289d11f5c2051b8861f852debaf4b6f525a

Observation f0299097-2e65-4ef4-8479-327fce363239 · outbound

This paper cites Claim_id.

Can AI Validate Science? Benchmarking LLMs for Accurate Scientific Claim $\rightarrow$ Evidence Reasoning Claim_id

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:20:16.184114Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-07T05:20:16.002827Z digest=sha256:4e51b2a205c38828c68a6f1f428cf394fbe590951d595758724fe577aea07e14

Observation 98cf8914-c989-4342-b579-d19a9d66b28d · outbound

This paper cites Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone.

Can AI Validate Science? Benchmarking LLMs for Accurate Scientific Claim $\rightarrow$ Evidence Reasoning Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-07T05:20:15.886923Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:20:15.886923Z digest=sha256:472e8a65f5e0a103fd2808eeb437911c25f2f6edcb05c79af81d9e3f859eae4e

Observation 39ed1a4b-f5f4-4171-977a-0bb81d4fed10 · outbound

This paper cites In Proceedings of the 31st International Conference on Computational Linguistics, pages 3613–3630, Abu Dhabi, UAE.

Can AI Validate Science? Benchmarking LLMs for Accurate Scientific Claim $\rightarrow$ Evidence Reasoning In Proceedings of the 31st International Conference on Computational Linguistics, pages 3613–3630, Abu Dhabi, UAE

Reference 2025

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:20:16.475902Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-07T05:20:15.892674Z digest=sha256:268035da9cacfffc7652bb57eedf3756a2a3728432554ba02864b56fb635c13f

Pith citing papers

Observation cbb972ac-4212-4e08-a69d-b48f74ef6aac · inbound

ResearchLoop: An Evidence-Gated Control Plane for AI-Assisted Research cites this paper.

ResearchLoop: An Evidence-Gated Control Plane for AI-Assisted Research Can AI Validate Science? Benchmarking LLMs for Accurate Scientific Claim $\rightarrow$ Evidence Reasoning

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-06-29T12:03:24.012115Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-06-29T11:58:23.295488Z digest=sha256:8067bd91996e872ee83d2f8c37e7630f53a637dce5132143b158c3c68f44f9b5

Observation cbf45071-fa48-4849-8fde-3e32c03cca93 · inbound

PEARL: Auditable Repair for Scientific Reasoning Graph Extraction cites this paper.

PEARL: Auditable Repair for Scientific Reasoning Graph Extraction Can AI Validate Science? Benchmarking LLMs for Accurate Scientific Claim $\rightarrow$ Evidence Reasoning

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-01T16:42:03.917326Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T16:42:03.917326Z digest=sha256:c442ca3a555f858c10d39ed5a8fecce0793d2eab9efb583e9fcd6464d7861ae6