Pith. sign in

Paper Citation Record · LEDGER

Evaluating the Unseen Capabilities: How Many Theorems Do LLMs Know?

As of 9 August 2026, this Paper Citation Record lists 65 of 65 outbound references and 1 inbound Pith citation observation for arXiv:2506.02058.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.02058 v1

Coverage vector

measured 65 of 65 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T11:56:34.350227Z

measured 66 of 66 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-05-10T15:40:18.133637Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-11T10:06:01.278099Z

Reference resolution

65 of 65 outbound references displayed

  • verified exact3
  • verified fuzzy29
  • unresolved31
  • parse uncertain0
  • malformed identifier2
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 72ab1322-29cf-4754-a082-ab21a584d42e · outbound

This paper cites GPT-4 Technical Report.

Evaluating the Unseen Capabilities: How Many Theorems Do LLMs Know? GPT-4 Technical Report

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T11:56:29.406168Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:56:29.406168Z digest=sha256:5e1e26dcdc0ff5acd8c563a42f603086e7d0382c81114d0fdae3a77839f0e70c

Observation 6856b69a-28b4-4b10-85a5-ef242ebea546 · outbound

This paper cites Physics of Language Models: Part 3.1, Knowledge Storage and Extraction.

Evaluating the Unseen Capabilities: How Many Theorems Do LLMs Know? Physics of Language Models: Part 3.1, Knowledge Storage and Extraction

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T11:56:29.482658Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:56:29.482658Z digest=sha256:feed9db94f0ebadb7b1b34428dfa22bcaa9c822d4713478a50e8205f80cc9114

Observation 483df34c-5857-4588-a5f8-e33033e45f75 · outbound

This paper cites Claude3.7 Sonnetsystemcard.

Evaluating the Unseen Capabilities: How Many Theorems Do LLMs Know? Claude3.7 Sonnetsystemcard

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:56:35.145924Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T11:56:29.576094Z digest=sha256:51c8d7c9b893d281568b20e7728a47ed46f71c0d28704d6ad1b428132885a4f0

Observation 310dca0d-c683-4f36-8c64-c3e4099af996 · outbound

This paper cites On the Opportunities and Risks of Foundation Models.

Evaluating the Unseen Capabilities: How Many Theorems Do LLMs Know? On the Opportunities and Risks of Foundation Models

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T11:56:29.637691Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:56:29.637691Z digest=sha256:9da0ef3a35498ca5c4f74b7c2773f428513cb6de653ad4f5d97573ccbfc69ecc

Observation 688d1427-ca52-42e9-8eeb-76a76e84da90 · outbound

This paper cites Eight things to know about large language models.Critical AI, 2(2), 2024.

Evaluating the Unseen Capabilities: How Many Theorems Do LLMs Know? Eight things to know about large language models.Critical AI, 2(2), 2024

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T11:56:29.701569Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:56:29.701569Z digest=sha256:6e1c5864d5657666ce92ffe0f9c6d74dce4bf99ffdc1874823d126da4646012f

Observation 52c15b9c-a07c-4c0d-8f30-eec444771e7d · outbound

This paper cites Language models are few-shot learners.

Evaluating the Unseen Capabilities: How Many Theorems Do LLMs Know? Language models are few-shot learners

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T11:56:29.768209Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:56:29.768209Z digest=sha256:9635c3290d56488d5791c5a69d29ee65aeb79fb30078a8ffdaa5e8e6837f38cc

Observation eb8c60f9-0d8a-4f75-b3b3-1ac1710b9d85 · outbound

This paper cites Sparks of Artificial General Intelligence: Early experiments with GPT-4.

Evaluating the Unseen Capabilities: How Many Theorems Do LLMs Know? Sparks of Artificial General Intelligence: Early experiments with GPT-4

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T11:56:29.866099Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:56:29.866099Z digest=sha256:74992a15a8ad4fcc05b4ff8e58c6602edfa8db01089149a77888016594f04557

Observation c7bf8963-f614-4700-8a85-8fd0267fb6ab · outbound

This paper cites Quantifying memorization across neural language models.

Evaluating the Unseen Capabilities: How Many Theorems Do LLMs Know? Quantifying memorization across neural language models

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:56:35.124339Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T11:56:29.927015Z digest=sha256:8672c8073c5aa0a081ec24e7c1afa90bca39e73db4115908bc63c0b830d20f56

Observation acbe91c0-5b91-48fb-a285-504003286703 · outbound

This paper cites How do large language models acquire factual knowledge during pretraining? In Neural Information Processing Systems, 2024.

Evaluating the Unseen Capabilities: How Many Theorems Do LLMs Know? How do large language models acquire factual knowledge during pretraining? In Neural Information Processing Systems, 2024

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:56:35.114386Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T11:56:30.009946Z digest=sha256:74458be56c111f7f7e2cd167af8c2977fb24202e5fd6a7c387834a3f2b395ff2

Observation 3dcb6d7e-dcd2-4478-b92e-73d7c8488979 · outbound

This paper cites A survey on evaluation of large language models.

Evaluating the Unseen Capabilities: How Many Theorems Do LLMs Know? A survey on evaluation of large language models

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T11:56:30.089145Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:56:30.089145Z digest=sha256:0860ba65aa085bd4957101b5d902880bd53e7c3cbdab15eea6f9b1127c701d12

Observation c6a89fa4-efc2-4f5e-abc6-b7a8f00ec216 · outbound

This paper cites Estimating the number of species in a stochastic abundance model.

Evaluating the Unseen Capabilities: How Many Theorems Do LLMs Know? Estimating the number of species in a stochastic abundance model

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:56:35.099450Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T11:56:30.173634Z digest=sha256:f96a96b1311efcb5062299add5a33bd759d3aafab5b953d6f983a760c916a2e7

Observation c10ddbc8-4a61-4e07-a370-281b9cbf1b2d · outbound

This paper cites A new statistical approach for assessing similarity of species composition with incidence and abundance data.

Evaluating the Unseen Capabilities: How Many Theorems Do LLMs Know? A new statistical approach for assessing similarity of species composition with incidence and abundance data

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:56:35.090536Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T11:56:30.264147Z digest=sha256:f415f46b59a7751c07d8db2e9518475c60864f8c2956f3fac182748ef33749be

Observation f6da8e46-f8f9-4d7d-9a97-54c0cf7f2fed · outbound

This paper cites Gonzalez, and Ion Stoica.

Evaluating the Unseen Capabilities: How Many Theorems Do LLMs Know? Gonzalez, and Ion Stoica

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:56:35.081757Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T11:56:30.373879Z digest=sha256:7cf341334568dce82b295c324196215d9b4b291bb2f19da811cf1cbcd96e43ce

Observation b03eeb4f-30c9-42f9-96c8-d6df7e737740 · outbound

This paper cites Training Verifiers to Solve Math Word Problems.

Evaluating the Unseen Capabilities: How Many Theorems Do LLMs Know? Training Verifiers to Solve Math Word Problems

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T11:56:30.466711Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:56:30.466711Z digest=sha256:f3d82d914214a56f5fa9290e76b62f09b42cd46459f1d08307444d8fb23ff45f

Observation 378f9079-eb0f-42d9-8c91-8c2c76031828 · outbound

This paper cites Forgetwhat you know about LLMs evaluations—LLMs are like a chameleon.arXiv preprint arXiv:2502.07445, 2025.

Evaluating the Unseen Capabilities: How Many Theorems Do LLMs Know? Forgetwhat you know about LLMs evaluations—LLMs are like a chameleon.arXiv preprint arXiv:2502.07445, 2025

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T11:56:30.563192Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:56:30.563192Z digest=sha256:e395036b512ae6833b5654580249c64ad0dd778404a1e2038332a79ea7bde5f9

Observation 48ea9d4c-b2af-4788-8e7a-be70489ca6f4 · outbound

This paper cites Springer US, Boston, MA, 2009.

Evaluating the Unseen Capabilities: How Many Theorems Do LLMs Know? Springer US, Boston, MA, 2009

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T11:56:30.691748Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:56:30.691748Z digest=sha256:342c8a79b1e38e55acb12d8578a45b368cfd7e1b2643e1db75ad20dd86dde850

Observation 1127a67f-eacf-47b3-8e4f-5f31f4a9b94c · outbound

This paper cites CURIE: Evaluating LLMs on multitask scientific long-context understanding and reasoning.

Evaluating the Unseen Capabilities: How Many Theorems Do LLMs Know? CURIE: Evaluating LLMs on multitask scientific long-context understanding and reasoning

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:56:35.072802Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T11:56:30.800214Z digest=sha256:3a86ac93befec95fc3fc29cf6f1534d8109516b4dd8e27f4f23ff2661f29c279

Observation 3dd36863-3d46-41a1-b651-f92f6d11ca24 · outbound

This paper cites Data science at the singularity.Harvard Data Science Review, 6(1), 2024.

Evaluating the Unseen Capabilities: How Many Theorems Do LLMs Know? Data science at the singularity.Harvard Data Science Review, 6(1), 2024

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T11:56:30.902979Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:56:30.902979Z digest=sha256:d8dd7ab8975cd05dfe2868b48a84020446e20cc81388e95e6060807077c142ad

Observation b88b454a-e470-4841-8662-022324c9d8d5 · outbound

This paper cites Estimating the number of unseen species: How many words did Shakespeare know?Biometrika, 63(3):435–447, 1976.

Evaluating the Unseen Capabilities: How Many Theorems Do LLMs Know? Estimating the number of unseen species: How many words did Shakespeare know?Biometrika, 63(3):435–447, 1976

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:56:35.057578Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T11:56:30.995955Z digest=sha256:858737e9a36b5b223f6fe8312957f6f0325a5ce4c63b5155eb580e5e3de44ee2

Observation fec3978d-2ff7-424c-a75c-b4eb4fbd38ac · outbound

This paper cites Near-optimal estimation of the unseen under regularly varying tail populations.Bernoulli, 29(4):3423–3442, 2023.

Evaluating the Unseen Capabilities: How Many Theorems Do LLMs Know? Near-optimal estimation of the unseen under regularly varying tail populations.Bernoulli, 29(4):3423–3442, 2023

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:56:35.047540Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T11:56:31.108281Z digest=sha256:e781dbbbf6cc75090652dff6308282307759ab25af05265021ed9d31f45262a4

Observation ca5772c1-a696-4352-a239-2cbaa2b9bfed · outbound

This paper cites Good-turing frequency estimation without tears.Journal of quantitative linguistics, 2(3):217–237, 1995.

Evaluating the Unseen Capabilities: How Many Theorems Do LLMs Know? Good-turing frequency estimation without tears.Journal of quantitative linguistics, 2(3):217–237, 1995

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T11:56:31.201554Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:56:31.201554Z digest=sha256:f23d92885cb33891750ed7193713d739e46be03e7957bce3a57884dbc7cf8fdb

Observation 2f94198c-03f9-46b7-83f1-1ac2be5258f3 · outbound

This paper cites The population frequencies of species and the estimation of population parameters.

Evaluating the Unseen Capabilities: How Many Theorems Do LLMs Know? The population frequencies of species and the estimation of population parameters

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:56:35.031401Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T11:56:31.316125Z digest=sha256:4d29e3e7e0072b77f0da46e0b27c7ee2865a51a6afda8f303f2f043e8e74cb5c

Observation f11c90b2-fe19-45a4-9762-1131b0048e9b · outbound

This paper cites Estimating knowledge in large language models without generating a single token.

Evaluating the Unseen Capabilities: How Many Theorems Do LLMs Know? Estimating knowledge in large language models without generating a single token

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:56:35.021476Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T11:56:31.399782Z digest=sha256:04d07cc32a8d333efdb092b20620f235457bbee7e8cd7af596b4c79255ff5e74

Observation f17a9559-781b-4bc9-8727-ba3f6e3ed67c · outbound

This paper cites The Llama 3 Herd of Models.

Evaluating the Unseen Capabilities: How Many Theorems Do LLMs Know? The Llama 3 Herd of Models

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T11:56:31.486703Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:56:31.486703Z digest=sha256:770ad358e14da09e5dfea8e5bda686040ea6ea87ffb48f3ac0f9de7738bfbb3d

Observation aae552b3-fbaa-4cbf-9c85-4b0dfad527ba · outbound

This paper cites Optimal prediction of the number of unseen species with multiplicity.

Evaluating the Unseen Capabilities: How Many Theorems Do LLMs Know? Optimal prediction of the number of unseen species with multiplicity

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:56:35.011026Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T11:56:31.589683Z digest=sha256:f5227207b9585120434dc7a9ae325c9742992c677a4a66b7e528a053d6a8e9db

Observation e1b4f6e4-a378-4ed5-83a4-01e165bbcbb1 · outbound

This paper cites Measuring massive multitask language understanding.

Evaluating the Unseen Capabilities: How Many Theorems Do LLMs Know? Measuring massive multitask language understanding

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T11:56:31.701733Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:56:31.701733Z digest=sha256:50ee484be592bd326beb454a155fbac77835ce3a4a9dcf8202224787de2e8770

Observation 2d78d870-0ef7-45a6-8018-4b91dc929097 · outbound

This paper cites The curious case of neural text degeneration.

Evaluating the Unseen Capabilities: How Many Theorems Do LLMs Know? The curious case of neural text degeneration

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T11:56:31.806850Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:56:31.806850Z digest=sha256:3c0a6bdecd33ab5a05cbbc063276ed6882643d5c9b1a6b5efc1d06d05d2872ec

Observation 3369593b-ec38-4dc2-93f2-35b1cbc63d97 · outbound

This paper cites GPT-4o System Card.

Evaluating the Unseen Capabilities: How Many Theorems Do LLMs Know? GPT-4o System Card

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T11:56:31.905024Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:56:31.905024Z digest=sha256:eb72e6a61567ed7bafb3de04bc2ac779a8ec6e6f025605e71d450bfeeaba4494

Observation 8b7ece32-5645-4610-a2fb-3ebc3a7c75dc · outbound

This paper cites Mistral 7B.

Evaluating the Unseen Capabilities: How Many Theorems Do LLMs Know? Mistral 7B

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T11:56:32.057825Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:56:32.057825Z digest=sha256:fbbe3f116bf4c13636e9e3fe0bf450daf00096e1338191143a12aedf3481b473

Observation e6655ea8-530c-4da3-ae41-b5589d52714f · outbound

This paper cites Calibrated language models must hallucinate.

Evaluating the Unseen Capabilities: How Many Theorems Do LLMs Know? Calibrated language models must hallucinate

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:56:34.989799Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T11:56:32.156600Z digest=sha256:0d3c1fad5399bbde13e58860b29ca0c9131e902d729191d7f8af7bba970d140f

Observation b677d7cb-2225-45af-b630-69fe5edaf29d · outbound

This paper cites Line of duty: Evaluating LLM self-knowledge via consistency in feasibility boundaries.arXiv preprint arXiv:2503.11256, 2025.

Evaluating the Unseen Capabilities: How Many Theorems Do LLMs Know? Line of duty: Evaluating LLM self-knowledge via consistency in feasibility boundaries.arXiv preprint arXiv:2503.11256, 2025

Reference 31

Resolution
verified exact
raw_fallback, observed 2026-08-07T11:56:34.654800Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T11:56:32.274696Z digest=sha256:6abb76327f5d64a78c4068b8d3f426918264baf6c3b87794710baffffcf99058

Observation aec70103-f245-46c7-a6cd-500493c24cc3 · outbound

This paper cites Too many AIs.https://dev.to/leeaao/too-many-ais-24nb, 2025.

Evaluating the Unseen Capabilities: How Many Theorems Do LLMs Know? Too many AIs.https://dev.to/leeaao/too-many-ais-24nb, 2025

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:56:34.979147Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T11:56:32.398334Z digest=sha256:028d97ea931eff55ab47e6eaf89198f09705cb29dc9f76283936379e6663d99d

Observation 57c732cf-9d26-4733-8cec-29c15d6a59d4 · outbound

This paper cites BioASQ-QA: A manually curated corpus for biomedical question answering.Scientific Data, 10 (1):170, 2023.

Evaluating the Unseen Capabilities: How Many Theorems Do LLMs Know? BioASQ-QA: A manually curated corpus for biomedical question answering.Scientific Data, 10 (1):170, 2023

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:56:34.969869Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T11:56:32.490401Z digest=sha256:38e99672c5a44f449d416684d5872794d57f87c9dd290e4d7b20236a7cfa171b

Observation 1fa5d269-dc34-47e8-bc34-dcdc8e918390 · outbound

This paper cites How pre-trained language models capture factual knowledge? A causal-inspired analysis.

Evaluating the Unseen Capabilities: How Many Theorems Do LLMs Know? How pre-trained language models capture factual knowledge? A causal-inspired analysis

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:56:34.960897Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T11:56:32.562356Z digest=sha256:f053074122a984f15f62925ac536efc90fcdce101badfb6ee1577a87c148e9f9

Observation 95f36e0b-4d3d-4373-b917-47032910befc · outbound

This paper cites ROUGE: A package for automatic evaluation of summaries.

Evaluating the Unseen Capabilities: How Many Theorems Do LLMs Know? ROUGE: A package for automatic evaluation of summaries

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:56:34.951469Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T11:56:32.665033Z digest=sha256:7391035e4d376a69513d25dc71a877ffbd6e4f67c83be31da54bbfb731453048

Observation a9b5a12a-3641-4fcb-9030-656528f8c9ac · outbound

This paper cites DeepSeek-V3 Technical Report.

Evaluating the Unseen Capabilities: How Many Theorems Do LLMs Know? DeepSeek-V3 Technical Report

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T11:56:32.742702Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:56:32.742702Z digest=sha256:36c14c494db93aa732cdf719ec95ddf228696d0b18faed2ede3d6d79e072400e

Observation 0ab7b408-c554-4834-8f88-2094453ab49a · outbound

This paper cites Feder Cooper, Daphne Ippolito, Christopher A.

Evaluating the Unseen Capabilities: How Many Theorems Do LLMs Know? Feder Cooper, Daphne Ippolito, Christopher A

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:56:34.942557Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T11:56:32.831463Z digest=sha256:703d423935e8111307e7abe83d727db1c4cae4665ae99331e8aa7a5afb17f0fb

Observation 9bdadb55-1361-4762-aa38-54a9fcac97a8 · outbound

This paper cites ChatGPT-3.5-turbo.

Evaluating the Unseen Capabilities: How Many Theorems Do LLMs Know? ChatGPT-3.5-turbo

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:56:34.933728Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T11:56:32.939632Z digest=sha256:a369491e5850ceee0a4ce02006fb5aca37d616d263d4a3854cbf1bcc04b3be5f

Observation 85ce83e6-eff3-4976-9842-f6f8c64f84f4 · outbound

This paper cites Competitive distribution estimation: Why is Good-Turing good.

Evaluating the Unseen Capabilities: How Many Theorems Do LLMs Know? Competitive distribution estimation: Why is Good-Turing good

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:56:34.925398Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T11:56:33.057660Z digest=sha256:6f3b5d3c22dfd11c021c89fccdcdd4a37892d9c56b6b5bf2da3a37b4bd0991b1

Observation f17f5189-d0c5-46f6-8727-8bcb583d3a86 · outbound

This paper cites Always Good Turing: Asymptotically optimal probability estimation.Science, 302(5644):427–431, 2003.

Evaluating the Unseen Capabilities: How Many Theorems Do LLMs Know? Always Good Turing: Asymptotically optimal probability estimation.Science, 302(5644):427–431, 2003

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:56:34.916852Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T11:56:33.170593Z digest=sha256:d6a89d3f888eae9c3c2f461a97702501c1b1a9e3a2a411e7c74d663b84d6766a

Observation f2c6e254-e81a-4a3b-8fbc-06a456d11fd1 · outbound

This paper cites Optimal prediction of the number of unseen species.

Evaluating the Unseen Capabilities: How Many Theorems Do LLMs Know? Optimal prediction of the number of unseen species

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:56:34.907642Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T11:56:33.261347Z digest=sha256:d9cf66cd46649ded5f0d73e70f32a0861eda2d8cab0af4102354bf21a4c4c83d

Observation ce2a9fe4-7ca7-4a92-98d4-b5a1a592c3c0 · outbound

This paper cites BLEU: A method for automatic evaluation of machine translation.

Evaluating the Unseen Capabilities: How Many Theorems Do LLMs Know? BLEU: A method for automatic evaluation of machine translation

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:56:34.897789Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T11:56:33.392631Z digest=sha256:9fe4586cd5c3691d16e264b466b69e729fbb4c8f2348f5c8b611be63acd2408c

Observation 0faef811-19d8-4ae0-9219-45c3b109bcb2 · outbound

This paper cites an unresolved cited work.

Evaluating the Unseen Capabilities: How Many Theorems Do LLMs Know? Unresolved cited work

Reference 43

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:56:34.888538Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T11:56:33.542099Z digest=sha256:98e5b52f8f76e4d780e839df48046135b237544d7f2e174a3793bd287d16dda5

Observation 68e09110-55f4-421b-9086-8ebd96634a3c · outbound

This paper cites Humanity's Last Exam.

Evaluating the Unseen Capabilities: How Many Theorems Do LLMs Know? Humanity's Last Exam

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-07T11:56:33.693512Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:56:33.693512Z digest=sha256:dbeedae9a196fc5eb34581ed1d7e32e7b1014550aa51747b9ce703f9c9e88ac0

Observation 900d1761-f576-45ef-b239-387891575e95 · outbound

This paper cites Do large language models know how much they know?arXiv preprint arXiv:2502.19573, 2025.

Evaluating the Unseen Capabilities: How Many Theorems Do LLMs Know? Do large language models know how much they know?arXiv preprint arXiv:2502.19573, 2025

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-07T11:56:33.840301Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:56:33.840301Z digest=sha256:96f1e797c8469596a5cb2fc9163f6c95fb9afde184f12c4526126e1c5160e734

Observation 96fb3030-8cd4-4bf9-90ff-ba9b27161418 · outbound

This paper cites AI and the everything in the whole wide world benchmark.

Evaluating the Unseen Capabilities: How Many Theorems Do LLMs Know? AI and the everything in the whole wide world benchmark

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:56:34.879529Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T11:56:33.927131Z digest=sha256:60bab6c36c9d4d9b6568251b8852e047a2cdefd1ea9a4e732997f0badfb7065d

Observation 4b1c0bfa-8e40-4db1-82b9-0aa6c63ca32c · outbound

This paper cites NLP evaluation in trouble: On the need to measure LLM data contamination for each benchmark.

Evaluating the Unseen Capabilities: How Many Theorems Do LLMs Know? NLP evaluation in trouble: On the need to measure LLM data contamination for each benchmark

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:56:34.870380Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T11:56:33.972092Z digest=sha256:17db4defd72c5cb67196983ead718152b019b564cfb439a7a2339a329966b1fd

Observation af1bb3c2-8c34-48f4-ba7e-f3450f84ede4 · outbound

This paper cites NeurIPS 2023 LLM Efficiency Fine-tuning Competition.

Evaluating the Unseen Capabilities: How Many Theorems Do LLMs Know? NeurIPS 2023 LLM Efficiency Fine-tuning Competition

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-07T11:56:34.073212Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:56:34.073212Z digest=sha256:00d2b1a74dc2ea72a070f150f2247c72ac789291fe66b66412a69b3b0fe96ab3

Observation fadec12e-09ec-48bd-bdad-4678fb372d40 · outbound

This paper cites Human disease ontology 2022 update.

Evaluating the Unseen Capabilities: How Many Theorems Do LLMs Know? Human disease ontology 2022 update

Reference 49

Resolution
verified exact
doi, observed 2026-08-07T11:56:34.393035Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T11:56:34.173184Z digest=sha256:6318c5289767f55c0454828c7086cbf45f602e6f82a86f1f33127a939b02e792

Observation bfa37a32-4556-4365-bde5-612215903355 · outbound

This paper cites Auto- Prompt: Eliciting knowledge from language models with automatically generated prompts.

Evaluating the Unseen Capabilities: How Many Theorems Do LLMs Know? Auto- Prompt: Eliciting knowledge from language models with automatically generated prompts

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:56:34.860670Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T11:56:34.298457Z digest=sha256:996011188530ccb1a8988ec089467fbcac4addabc5bef9df3ac92c38e56e4f9b

Observation 4278205c-6193-4b12-809a-f2a1b2779fd5 · outbound

This paper cites Welcome to the era of experience.Google AI, 2025.

Evaluating the Unseen Capabilities: How Many Theorems Do LLMs Know? Welcome to the era of experience.Google AI, 2025

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-07T11:56:34.304140Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:56:34.304140Z digest=sha256:43999fbc34426f51eecbf099bd2e5b718f4beb22de307e0afbe1a877816f1b6c

Observation cb6f8d71-aa20-4602-a8c0-b521e391c82c · outbound

This paper cites The Leaderboard Illusion.

Evaluating the Unseen Capabilities: How Many Theorems Do LLMs Know? The Leaderboard Illusion

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-07T11:56:34.307356Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:56:34.307356Z digest=sha256:a957f66c178263c60bd417909b57e893b5f2a907f614391814cee501c0cd84e7

Observation 0714b062-436d-42fa-8830-9680c3342321 · outbound

This paper cites Large language models encode clinical knowledge.Nature, 620(7972):172–180, 2023.

Evaluating the Unseen Capabilities: How Many Theorems Do LLMs Know? Large language models encode clinical knowledge.Nature, 620(7972):172–180, 2023

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-07T11:56:34.311072Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:56:34.311072Z digest=sha256:478b1aa6f49215b65506d820e3bf030880a13bb6efa4931542c4d2aec62525f3

Observation 63154cd9-7896-4b3b-bc9f-231e38059ce7 · outbound

This paper cites The bitter lesson.Incomplete Ideas (blog), 13(1):38, 2019.

Evaluating the Unseen Capabilities: How Many Theorems Do LLMs Know? The bitter lesson.Incomplete Ideas (blog), 13(1):38, 2019

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-07T11:56:34.314529Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:56:34.314529Z digest=sha256:d59094c9805f5f7afe79c2395995f8393ff910e4787323508fe0ffca3e938279

Observation cdf82a61-b4c1-4319-99ba-e27ea41fb45c · outbound

This paper cites Galactica: A Large Language Model for Science.

Evaluating the Unseen Capabilities: How Many Theorems Do LLMs Know? Galactica: A Large Language Model for Science

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-07T11:56:34.317573Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:56:34.317573Z digest=sha256:5240b0630670c0cf1d295b90c3c96f8ac669a2e0f8b365d726f9252ea88596e0

Observation bf256b5f-6e75-449f-bc4e-52fb3deaba4d · outbound

This paper cites Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context.

Evaluating the Unseen Capabilities: How Many Theorems Do LLMs Know? Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-07T11:56:34.320882Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:56:34.320882Z digest=sha256:71c27bc57376fe669340ba814bd0d92b5ad2b6272e8090ebc9c7cf36a490975d

Observation 18c493fd-f9b1-4e8b-9689-78de7603398e · outbound

This paper cites an unresolved cited work.

Evaluating the Unseen Capabilities: How Many Theorems Do LLMs Know? Unresolved cited work

Reference 57

Resolution
malformed identifier
doi_truncated, observed 2026-08-07T11:56:34.382001Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T11:56:34.324022Z digest=sha256:d009fb7e5690673ff7b607e6b7e6f8f7f8186c4943fc56b275c4c15975801ab4

Observation 64e7e780-3939-43a2-953a-1324e6ab1c40 · outbound

This paper cites Language Models are Open Knowledge Graphs.

Evaluating the Unseen Capabilities: How Many Theorems Do LLMs Know? Language Models are Open Knowledge Graphs

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-07T11:56:34.327247Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:56:34.327247Z digest=sha256:6ec1714bc2e528e0907b3a2abfc3583763abaf5d4e4a03ec5598d722a04d876e

Observation 60d4733b-45b1-4bef-880f-000b5a15117c · outbound

This paper cites Can AI Be as Creative as Humans?.

Evaluating the Unseen Capabilities: How Many Theorems Do LLMs Know? Can AI Be as Creative as Humans?

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-07T11:56:34.330717Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:56:34.330717Z digest=sha256:9090272d069af8cc33e63b34f384356666d6752aef8aa1edb8abfea3a2fd8225

Observation 5d323f19-07f3-41a9-9f95-676aaa7aae99 · outbound

This paper cites Chi, Quoc V.

Evaluating the Unseen Capabilities: How Many Theorems Do LLMs Know? Chi, Quoc V

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:56:34.830980Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T11:56:34.334226Z digest=sha256:dd4dcc36851a32559b45aa3507a2ee5b50cb3f46139958a51a8c9cb14f296adf

Observation 40608028-8344-4c9d-bbfe-42717f898ea3 · outbound

This paper cites Estimating the probabilities of rare outputs in language models.

Evaluating the Unseen Capabilities: How Many Theorems Do LLMs Know? Estimating the probabilities of rare outputs in language models

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:56:34.822545Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T11:56:34.337670Z digest=sha256:db9880ec654af53c1f3cbc165dd2ef9062e7cc9b44e8bbacc3f0986ce6c666c7

Observation 8f4f37ea-0c4b-4110-a534-f91081bc8a5e · outbound

This paper cites Qwen2.5 Technical Report.

Evaluating the Unseen Capabilities: How Many Theorems Do LLMs Know? Qwen2.5 Technical Report

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-07T11:56:34.340659Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:56:34.340659Z digest=sha256:48b749ba9e039820bf326cfdd536bc80a6d4e9cbc9bed582b1be56bea6b59195

Observation 893d72c8-1572-424b-9370-94f79a7c44b2 · outbound

This paper cites A careful examination of large language model performance on grade school arithmetic.

Evaluating the Unseen Capabilities: How Many Theorems Do LLMs Know? A careful examination of large language model performance on grade school arithmetic

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:56:34.812963Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T11:56:34.343378Z digest=sha256:a4f8b99fa4ba50bd6e1b9d523154fa95ac2b64614db580bc1b898d4d73e66bc6

Observation 88fcaddc-d668-4a60-87fd-a06a821f89a5 · outbound

This paper cites Using Pretrained Large Language Model with Prompt Engineering to Answer Biomedical Questions.

Evaluating the Unseen Capabilities: How Many Theorems Do LLMs Know? Using Pretrained Large Language Model with Prompt Engineering to Answer Biomedical Questions

Reference 64

Resolution
verified exact
local_arxiv, observed 2026-08-07T11:56:34.413661Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T11:56:34.346706Z digest=sha256:d249e22097187231d9c2c2d870a67730620d1b9b32116337410c2201bbd99dfe

Observation 29d200ca-a50b-460f-8047-70ddb01d209f · outbound

This paper cites Test your knowledge of mathematical theorems by listing 20 theorem names, separated by commas.

Evaluating the Unseen Capabilities: How Many Theorems Do LLMs Know? Test your knowledge of mathematical theorems by listing 20 theorem names, separated by commas

Reference 65

Resolution
malformed identifier
raw_fallback, observed 2026-08-07T11:56:34.803679Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T11:56:34.350227Z digest=sha256:f0dcfd5e247d7db92c8adb637196b2d4b0924a5bccfc7f004c6eb25f138b77d6

Pith citing papers

Observation 4f39711d-05df-4367-bdf0-2987ca8ae6db · inbound

UCS: Estimating Unseen Coverage for Improved In-Context Learning cites this paper.

UCS: Estimating Unseen Coverage for Improved In-Context Learning Evaluating the Unseen Capabilities: How Many Theorems Do LLMs Know?

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-05-11T10:06:01.283344Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-10T15:40:18.133637Z digest=sha256:024f770d1d77ad3ce454ba97adb276646add7fa1090bd3e1ee3379ded9fe903b