Pith. sign in

Paper Citation Record · LEDGER

Biomed-Enriched: A Biomedical Dataset Enriched with LLMs for Pretraining and Extracting Rare and Hidden Content

As of 15 August 2026, this Paper Citation Record lists 18 of 18 outbound references and 1 inbound Pith citation observation for arXiv:2506.20331.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.20331 v1

Coverage vector

measured 18 of 18 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T22:54:52.786307Z

measured 19 of 19 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-15T06:32:42.880941+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-05-13T04:36:51.494629Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-13T04:37:15.679691Z

Reference resolution

18 of 18 outbound references displayed

  • verified exact0
  • verified fuzzy4
  • unresolved13
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation b0187271-a61d-4bc2-bdbd-bd82735af5d0 · outbound

This paper cites https://www.ncbi.nlm.nih.gov/pmc/tools/textmining/.

Biomed-Enriched: A Biomedical Dataset Enriched with LLMs for Pretraining and Extracting Rare and Hidden Content https://www.ncbi.nlm.nih.gov/pmc/tools/textmining/

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:54:53.597095Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-06T22:54:52.109879Z digest=sha256:de8b110d427d60208d0341ea8c84e2ea942c369aeec1a5c5196e657310dd4927

Observation e9ab342e-88dc-4856-b953-5366ab8bb5f2 · outbound

This paper cites FrenchMedMCQA: A French Multiple-Choice Question Answering Dataset for Medical domain.

Biomed-Enriched: A Biomedical Dataset Enriched with LLMs for Pretraining and Extracting Rare and Hidden Content FrenchMedMCQA: A French Multiple-Choice Question Answering Dataset for Medical domain

Reference 2

Resolution
malformed identifier
doi_truncated, observed 2026-08-06T22:54:52.901277Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-06T22:54:52.161654Z digest=sha256:c249c8a9edef53192a8155fa1c264af672dbe3b6f1866727f161496a1681d672

Observation bc8c8414-cb2a-4f61-95f5-42f572b1d0cb · outbound

This paper cites MEDITRON-70B: Scaling Medical Pretraining for Large Language Models.

Biomed-Enriched: A Biomedical Dataset Enriched with LLMs for Pretraining and Extracting Rare and Hidden Content MEDITRON-70B: Scaling Medical Pretraining for Large Language Models

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-06T22:54:52.195014Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:54:52.195014Z digest=sha256:de7ace621f431851f61e32fc32b11e444337bd0a12732cc5513b32b19484cc51

Observation 5664f23b-450f-4be0-8316-a697247dbd4f · outbound

This paper cites Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge.

Biomed-Enriched: A Biomedical Dataset Enriched with LLMs for Pretraining and Extracting Rare and Hidden Content Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-06T22:54:52.224175Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:54:52.224175Z digest=sha256:3f0fe03e3a7960b05fb3d2b1c221f615355cad2c1ab915ee1bbd9672ed410c45

Observation 56899dfb-c708-4a26-8cb2-e6d14518222a · outbound

This paper cites Unsupervised cross-lingual representation learning at scale.

Biomed-Enriched: A Biomedical Dataset Enriched with LLMs for Pretraining and Extracting Rare and Hidden Content Unsupervised cross-lingual representation learning at scale

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:54:53.458450Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-06T22:54:52.274238Z digest=sha256:e8a772bc077e13c9b24d36a0d34a8498c21c6311413527d44ae9beae4f578b4b

Observation 3c44f1ff-5ad7-452f-a01b-fdf300683b73 · outbound

This paper cites Measuring massive multitask language understanding.

Biomed-Enriched: A Biomedical Dataset Enriched with LLMs for Pretraining and Extracting Rare and Hidden Content Measuring massive multitask language understanding

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:54:53.355408Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-06T22:54:52.310204Z digest=sha256:65990749e2b02f56bd099e0946ebfcea35e051a4265526137e57ca1cb0dee9c0

Observation 9d925cc3-a321-4c0c-91f0-8a28c79f7b66 · outbound

This paper cites B io M istral: A collection of open-source pretrained large language models for medical domains.

Biomed-Enriched: A Biomedical Dataset Enriched with LLMs for Pretraining and Extracting Rare and Hidden Content B io M istral: A collection of open-source pretrained large language models for medical domains

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-06T22:54:52.355144Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:54:52.355144Z digest=sha256:d3010d8975aba05e511840dada8a7e25946fd88d235a586da4488a408f2f5fe0

Observation b436d23e-fe70-4601-9e5f-40be2b00ed8d · outbound

This paper cites DataComp-LM: In search of the next generation of training sets for language models.

Biomed-Enriched: A Biomedical Dataset Enriched with LLMs for Pretraining and Extracting Rare and Hidden Content DataComp-LM: In search of the next generation of training sets for language models

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-06T22:54:52.388115Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:54:52.388115Z digest=sha256:0d41f08d49538671c6864e84479d11ceb8fd37479070ecdcae0b9bb3ae10fcf2

Observation 6031f621-4bfe-424f-bdf3-51443a3e43f0 · outbound

This paper cites Can a suit of armor conduct electricity? a new dataset for open book question answering.

Biomed-Enriched: A Biomedical Dataset Enriched with LLMs for Pretraining and Extracting Rare and Hidden Content Can a suit of armor conduct electricity? a new dataset for open book question answering

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:54:53.175327Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-06T22:54:52.405517Z digest=sha256:ece0b34b4ccc912576395e5941d999ca99310d3de79769ef1ed5c316ed062e1a

Observation 42893064-84b8-431c-8184-f1e2dafb45f7 · outbound

This paper cites an unresolved cited work.

Biomed-Enriched: A Biomedical Dataset Enriched with LLMs for Pretraining and Extracting Rare and Hidden Content Unresolved cited work

Reference 10

Resolution
unresolved
raw_fallback, observed 2026-08-06T22:54:53.042847Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-06T22:54:52.453916Z digest=sha256:55737774521d968e023d9787591c79b09464e23b646054200e319902471bb003

Observation 53605707-6254-49f1-a56a-fe8c8660a6f1 · outbound

This paper cites The FineWeb Datasets: Decanting the Web for the Finest Text Data at Scale.

Biomed-Enriched: A Biomedical Dataset Enriched with LLMs for Pretraining and Extracting Rare and Hidden Content The FineWeb Datasets: Decanting the Web for the Finest Text Data at Scale

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-06T22:54:52.500273Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:54:52.500273Z digest=sha256:f647b6188b61fe972afe85785b534d310a12259202236a76908dee04fd7e31dd

Observation c6e9c7ff-7f1c-446e-adcc-a8792295a363 · outbound

This paper cites The Llama 3 Herd of Models.

Biomed-Enriched: A Biomedical Dataset Enriched with LLMs for Pretraining and Extracting Rare and Hidden Content The Llama 3 Herd of Models

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-06T22:54:52.539788Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:54:52.539788Z digest=sha256:ceee6541d9d2ef4b7eee681af2be10168a4f705061b43b4d575f7b9c5388295a

Observation 54b28227-abef-4277-ad59-344a68261de0 · outbound

This paper cites Organize the Web: Constructing Domains Enhances Pre-Training Data Curation.

Biomed-Enriched: A Biomedical Dataset Enriched with LLMs for Pretraining and Extracting Rare and Hidden Content Organize the Web: Constructing Domains Enhances Pre-Training Data Curation

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-06T22:54:52.567233Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:54:52.567233Z digest=sha256:a4863a9f6024517b6cdcfb1b89c65af7ef194e170e292609c1f1b23150a304f7

Observation 19a7004c-f6ed-4ed2-b8cf-e03f42479ab3 · outbound

This paper cites PMC-LLaMA : toward building open-source language models for medicine.

Biomed-Enriched: A Biomedical Dataset Enriched with LLMs for Pretraining and Extracting Rare and Hidden Content PMC-LLaMA : toward building open-source language models for medicine

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-06T22:54:52.587988Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:54:52.587988Z digest=sha256:7440d3200db15b65af96068018ec78de7f0052646ae558834085f2cd2259d9eb

Observation e42d0f00-1133-4f3a-bdbd-5c64686223bf · outbound

This paper cites write newline.

Biomed-Enriched: A Biomedical Dataset Enriched with LLMs for Pretraining and Extracting Rare and Hidden Content write newline

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-06T22:54:52.652943Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:54:52.652943Z digest=sha256:e215b51bbc033b8d12a0ec3a180f65036dfdae91b5b9b249e85ea2330a701495

Observation 84e6a092-f833-4b0c-968a-0bd9568b6191 · outbound

This paper cites @esa (Ref.

Biomed-Enriched: A Biomedical Dataset Enriched with LLMs for Pretraining and Extracting Rare and Hidden Content @esa (Ref

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-06T22:54:52.724253Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:54:52.724253Z digest=sha256:724e72b5a98021da05d9a9f2884cbe9642b9a275d278b2451336786d1bf8c637

Observation 0b4e5ab3-8fa0-4f30-812d-fde0b117a419 · outbound

This paper cites an unresolved cited work.

Biomed-Enriched: A Biomedical Dataset Enriched with LLMs for Pretraining and Extracting Rare and Hidden Content Unresolved cited work

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-06T22:54:52.748635Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:54:52.748635Z digest=sha256:9c4e43971612b9261ea7084c837d2d68e32cdee5e092ff3b1a419e4585cf6dc3

Observation 847a2b62-cbf4-4417-84be-4638a2a7bc6d · outbound

This paper cites an unresolved cited work.

Biomed-Enriched: A Biomedical Dataset Enriched with LLMs for Pretraining and Extracting Rare and Hidden Content Unresolved cited work

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-06T22:54:52.786307Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:54:52.786307Z digest=sha256:f34ef39c4ea7a1f51e810f364e6aca0dbbad0da45ebc525d9a7e14b7e92adb4b

Pith citing papers

Observation aa12e4d3-7b9b-4034-9572-0afc99ab8c96 · inbound

A Causal Language Modeling Detour Improves Encoder Continued Pretraining cites this paper.

A Causal Language Modeling Detour Improves Encoder Continued Pretraining Biomed-Enriched: A Biomedical Dataset Enriched with LLMs for Pretraining and Extracting Rare and Hidden Content

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-05-13T04:37:15.681688Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-13T04:36:51.494629Z digest=sha256:d289ae88009e289614ae3ec8e8834bb2258b1c58e25777098164081f0f8a9936