Pith. sign in

Paper Citation Record · LEDGER

Biomed-Enriched: A Biomedical Dataset Enriched with LLMs for Pretraining and Extracting Rare and Hidden Content

As of 8 August 2026, this Paper Citation Record lists 18 of 18 outbound references and 1 inbound Pith citation observation for arXiv:2506.20331.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.20331 v1

Coverage vector

measured 18 of 18 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T22:54:52.786307Z

measured 19 of 19 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-05-13T04:36:51.494629Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-13T04:37:15.679691Z

Reference resolution

18 of 18 outbound references displayed

  • verified exact0
  • verified fuzzy4
  • unresolved13
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation b0187271-a61d-4bc2-bdbd-bd82735af5d0 · outbound

This paper cites https://www.ncbi.nlm.nih.gov/pmc/tools/textmining/.

Biomed-Enriched: A Biomedical Dataset Enriched with LLMs for Pretraining and Extracting Rare and Hidden Content https://www.ncbi.nlm.nih.gov/pmc/tools/textmining/

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:54:53.597095Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-06T22:54:52.109879Z digest=sha256:95bf0f1b2c7aaf3289fe8e5f9e0edd1ee93a8b437215dbe5c154bbdaaf2306d6

Observation e9ab342e-88dc-4856-b953-5366ab8bb5f2 · outbound

This paper cites FrenchMedMCQA: A French Multiple-Choice Question Answering Dataset for Medical domain.

Biomed-Enriched: A Biomedical Dataset Enriched with LLMs for Pretraining and Extracting Rare and Hidden Content FrenchMedMCQA: A French Multiple-Choice Question Answering Dataset for Medical domain

Reference 2

Resolution
malformed identifier
doi_truncated, observed 2026-08-06T22:54:52.901277Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-06T22:54:52.161654Z digest=sha256:3a95bfcf18261263d07c5f634e30ac6551bbeec6008d1ee0a02534938fcb87c9

Observation bc8c8414-cb2a-4f61-95f5-42f572b1d0cb · outbound

This paper cites MEDITRON-70B: Scaling Medical Pretraining for Large Language Models.

Biomed-Enriched: A Biomedical Dataset Enriched with LLMs for Pretraining and Extracting Rare and Hidden Content MEDITRON-70B: Scaling Medical Pretraining for Large Language Models

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-06T22:54:52.195014Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:54:52.195014Z digest=sha256:8a8e95c63c4ed8cccc8fefe36a9f364f0f535cafe9716ea87d259ebb0cbdb0cf

Observation 5664f23b-450f-4be0-8316-a697247dbd4f · outbound

This paper cites Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge.

Biomed-Enriched: A Biomedical Dataset Enriched with LLMs for Pretraining and Extracting Rare and Hidden Content Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-06T22:54:52.224175Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:54:52.224175Z digest=sha256:f947efd3f9555eeab82a8ac84f045b8e11c44d924d3cd10e9037fd152eb021aa

Observation 56899dfb-c708-4a26-8cb2-e6d14518222a · outbound

This paper cites Unsupervised cross-lingual representation learning at scale.

Biomed-Enriched: A Biomedical Dataset Enriched with LLMs for Pretraining and Extracting Rare and Hidden Content Unsupervised cross-lingual representation learning at scale

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:54:53.458450Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-06T22:54:52.274238Z digest=sha256:bf387815bd0679270bdf735c326cd7befd7ce1918962e00e9caf9ef2d58c590d

Observation 3c44f1ff-5ad7-452f-a01b-fdf300683b73 · outbound

This paper cites Measuring massive multitask language understanding.

Biomed-Enriched: A Biomedical Dataset Enriched with LLMs for Pretraining and Extracting Rare and Hidden Content Measuring massive multitask language understanding

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:54:53.355408Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-06T22:54:52.310204Z digest=sha256:9746c054c7bd9b3377eb2307443e96bdabdec4eb62dfc51f44919a34af79f4ac

Observation 9d925cc3-a321-4c0c-91f0-8a28c79f7b66 · outbound

This paper cites B io M istral: A collection of open-source pretrained large language models for medical domains.

Biomed-Enriched: A Biomedical Dataset Enriched with LLMs for Pretraining and Extracting Rare and Hidden Content B io M istral: A collection of open-source pretrained large language models for medical domains

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-06T22:54:52.355144Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:54:52.355144Z digest=sha256:ff5b53a9e543561bc03ecc1863cbbba0c6a1a0066082e9267c40aff31b4714db

Observation b436d23e-fe70-4601-9e5f-40be2b00ed8d · outbound

This paper cites DataComp-LM: In search of the next generation of training sets for language models.

Biomed-Enriched: A Biomedical Dataset Enriched with LLMs for Pretraining and Extracting Rare and Hidden Content DataComp-LM: In search of the next generation of training sets for language models

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-06T22:54:52.388115Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:54:52.388115Z digest=sha256:e00cfde740cc1a3608347529b594f62a777c896c1fe3b6da2adf909db4d990ff

Observation 6031f621-4bfe-424f-bdf3-51443a3e43f0 · outbound

This paper cites Can a suit of armor conduct electricity? a new dataset for open book question answering.

Biomed-Enriched: A Biomedical Dataset Enriched with LLMs for Pretraining and Extracting Rare and Hidden Content Can a suit of armor conduct electricity? a new dataset for open book question answering

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:54:53.175327Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-06T22:54:52.405517Z digest=sha256:862bdad7e1d90d20d3247359f371e6a8d3a3234b311a7efa74e503ac9718e669

Observation 42893064-84b8-431c-8184-f1e2dafb45f7 · outbound

This paper cites an unresolved cited work.

Biomed-Enriched: A Biomedical Dataset Enriched with LLMs for Pretraining and Extracting Rare and Hidden Content Unresolved cited work

Reference 10

Resolution
unresolved
raw_fallback, observed 2026-08-06T22:54:53.042847Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-06T22:54:52.453916Z digest=sha256:c1860d0426e0ae33af4c62989bec4563103501c405687ed8acc3071612620502

Observation 53605707-6254-49f1-a56a-fe8c8660a6f1 · outbound

This paper cites The FineWeb Datasets: Decanting the Web for the Finest Text Data at Scale.

Biomed-Enriched: A Biomedical Dataset Enriched with LLMs for Pretraining and Extracting Rare and Hidden Content The FineWeb Datasets: Decanting the Web for the Finest Text Data at Scale

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-06T22:54:52.500273Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:54:52.500273Z digest=sha256:8cb5286f64898def94a9c3713edd617435361e6c8206a86b7f24a2c3da162943

Observation c6e9c7ff-7f1c-446e-adcc-a8792295a363 · outbound

This paper cites The Llama 3 Herd of Models.

Biomed-Enriched: A Biomedical Dataset Enriched with LLMs for Pretraining and Extracting Rare and Hidden Content The Llama 3 Herd of Models

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-06T22:54:52.539788Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:54:52.539788Z digest=sha256:cd31d772171423c2e8856ab1fea247176c94ed0fd16f43e454ec2a62b88e0f82

Observation 54b28227-abef-4277-ad59-344a68261de0 · outbound

This paper cites Organize the Web: Constructing Domains Enhances Pre-Training Data Curation.

Biomed-Enriched: A Biomedical Dataset Enriched with LLMs for Pretraining and Extracting Rare and Hidden Content Organize the Web: Constructing Domains Enhances Pre-Training Data Curation

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-06T22:54:52.567233Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:54:52.567233Z digest=sha256:b8a2a4b4dc78d5b744d6376b00bdb471e8eed599ef4c6ee3e9c90362e03a0686

Observation 19a7004c-f6ed-4ed2-b8cf-e03f42479ab3 · outbound

This paper cites PMC-LLaMA : toward building open-source language models for medicine.

Biomed-Enriched: A Biomedical Dataset Enriched with LLMs for Pretraining and Extracting Rare and Hidden Content PMC-LLaMA : toward building open-source language models for medicine

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-06T22:54:52.587988Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:54:52.587988Z digest=sha256:953c3d3b82b4556e909e54119e7941ab2fb5e37b8f0f8775ba63b0afb2731c3b

Observation e42d0f00-1133-4f3a-bdbd-5c64686223bf · outbound

This paper cites write newline.

Biomed-Enriched: A Biomedical Dataset Enriched with LLMs for Pretraining and Extracting Rare and Hidden Content write newline

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-06T22:54:52.652943Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:54:52.652943Z digest=sha256:1de96c92c4ed4a37c49b19157d0a9a46c49c7fb8d9a00e64751e8b0c0dbca99e

Observation 84e6a092-f833-4b0c-968a-0bd9568b6191 · outbound

This paper cites @esa (Ref.

Biomed-Enriched: A Biomedical Dataset Enriched with LLMs for Pretraining and Extracting Rare and Hidden Content @esa (Ref

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-06T22:54:52.724253Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:54:52.724253Z digest=sha256:02b1fdfa076d186ec8762fa6cb67c0e2b620afae8c4ec5490e5f31f2df8dbba3

Observation 0b4e5ab3-8fa0-4f30-812d-fde0b117a419 · outbound

This paper cites an unresolved cited work.

Biomed-Enriched: A Biomedical Dataset Enriched with LLMs for Pretraining and Extracting Rare and Hidden Content Unresolved cited work

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-06T22:54:52.748635Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:54:52.748635Z digest=sha256:f071937117166fe7a035192b32dde750427c90be4f59b508dca45076f8fcaac2

Observation 847a2b62-cbf4-4417-84be-4638a2a7bc6d · outbound

This paper cites an unresolved cited work.

Biomed-Enriched: A Biomedical Dataset Enriched with LLMs for Pretraining and Extracting Rare and Hidden Content Unresolved cited work

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-06T22:54:52.786307Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:54:52.786307Z digest=sha256:99310a21950d98e4b450ad25c3796ebe611ae24b0164f54d9b2048cc7f0de26f

Pith citing papers

Observation aa12e4d3-7b9b-4034-9572-0afc99ab8c96 · inbound

A Causal Language Modeling Detour Improves Encoder Continued Pretraining cites this paper.

A Causal Language Modeling Detour Improves Encoder Continued Pretraining Biomed-Enriched: A Biomedical Dataset Enriched with LLMs for Pretraining and Extracting Rare and Hidden Content

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-05-13T04:37:15.681688Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-13T04:36:51.494629Z digest=sha256:228f2c72a82d4df70b4e5e1f4422871f1a82a3ea26f77c81e81b04734607b600