Pith. sign in

Paper Citation Record · LEDGER

Investigating the Impact of Data Selection Strategies on Language Model Performance

As of 22 August 2026, this Paper Citation Record lists 16 of 16 outbound references and 0 inbound Pith citation observations for arXiv:2501.03826.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2501.03826 v1

Coverage vector

measured 16 of 16 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-10T21:49:32.290699Z

measured 16 of 16 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-21T06:32:19.484+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

16 of 16 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved16
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 022ea559-4547-430a-9882-d87cf8325cd1 · outbound

This paper cites URL: " 'urlintro :=.

Investigating the Impact of Data Selection Strategies on Language Model Performance URL: " 'urlintro :=

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-10T21:49:32.219932Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T21:49:32.219932Z digest=sha256:3596cd044f0a0da20f8b116c6cc365cf472bd48b96d6bfcfcff393594b759fc2

Observation 5f32381a-64f2-49a7-81db-a394ff527064 · outbound

This paper cites write newline.

Investigating the Impact of Data Selection Strategies on Language Model Performance write newline

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-10T21:49:32.225622Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T21:49:32.225622Z digest=sha256:d669a89b3fa0b9f8b1a47e0a41a61e49d0ba334cf511bdc11ece1c222d6fa11d

Observation 21a5bb3e-e817-40a5-969e-8980509f53f2 · outbound

This paper cites A Survey on Data Selection for Language Models.

Investigating the Impact of Data Selection Strategies on Language Model Performance A Survey on Data Selection for Language Models

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-10T21:49:32.230472Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T21:49:32.230472Z digest=sha256:3e87df6db832bb4ab8973fc1957dc52c534bb1c4cfc35d84aa5b29b1d0965e26

Observation 3494af29-6327-40c2-bb2b-ed0d6a85f2fe · outbound

This paper cites Language Models are Few-Shot Learners.

Investigating the Impact of Data Selection Strategies on Language Model Performance Language Models are Few-Shot Learners

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-10T21:49:32.235627Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T21:49:32.235627Z digest=sha256:e7d79f04af0c773406bce346e98890fadd11be56345e1f9e181d4381b0d42cad

Observation 239a1845-ca03-4a8a-957f-88f548584989 · outbound

This paper cites Sparks of Artificial General Intelligence: Early experiments with GPT-4.

Investigating the Impact of Data Selection Strategies on Language Model Performance Sparks of Artificial General Intelligence: Early experiments with GPT-4

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-10T21:49:32.239877Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T21:49:32.239877Z digest=sha256:345e59ca6d478d96a6a7d4e417cc4f16c2c5f27aedea6eed02c031c199f1772e

Observation 7a4fe3c5-57ca-4a87-bcc6-1536ca43ad24 · outbound

This paper cites Gradient-based Bi-level Optimization for Deep Learning: A Survey.

Investigating the Impact of Data Selection Strategies on Language Model Performance Gradient-based Bi-level Optimization for Deep Learning: A Survey

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-10T21:49:32.245069Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T21:49:32.245069Z digest=sha256:d290fffddb7ea73f4b43c5c5d3a61f24abbbcbfa1e38448309ac9109eb1464d5

Observation 8d2d8e1a-85ae-44e4-b16a-1bde8b52f8f3 · outbound

This paper cites an unresolved cited work.

Investigating the Impact of Data Selection Strategies on Language Model Performance Unresolved cited work

Reference 7

Resolution
unresolved
raw_fallback, observed 2026-08-10T21:49:32.479451Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-10T21:49:32.250611Z digest=sha256:fb8dc29eff8d2095d7c8114e1f8d536be9ae363997663f031f998846da0bbc70

Observation 9d9a4bc8-fc65-44a1-acd1-3cbbaeda5c5e · outbound

This paper cites PaLM: Scaling Language Modeling with Pathways.

Investigating the Impact of Data Selection Strategies on Language Model Performance PaLM: Scaling Language Modeling with Pathways

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-10T21:49:32.255398Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T21:49:32.255398Z digest=sha256:9e2e9f001740a0e350252948b4cf257ce4286ca5db75d7cce41c50b234299160

Observation eab36d1d-707f-402c-bf70-e98bedcfaa88 · outbound

This paper cites Don't Stop Pretraining: Adapt Language Models to Domains and Tasks.

Investigating the Impact of Data Selection Strategies on Language Model Performance Don't Stop Pretraining: Adapt Language Models to Domains and Tasks

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-10T21:49:32.260531Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T21:49:32.260531Z digest=sha256:5f9c69267b5900608a9e156236f62df49059d94f195a3e4e5fa6b233b4acdae0

Observation 3135496d-87e3-480f-a679-a225496398c1 · outbound

This paper cites an unresolved cited work.

Investigating the Impact of Data Selection Strategies on Language Model Performance Unresolved cited work

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-10T21:49:32.266611Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T21:49:32.266611Z digest=sha256:8ad9510f89b13c2770417c49b6beba7de062c1745b22525fa5016ef523e7ec90

Observation 8c41f9fc-4dfc-4113-a92f-3022d9c015a0 · outbound

This paper cites Sentence-BERT: Sentence Embeddings using Siamese BERT-Networks.

Investigating the Impact of Data Selection Strategies on Language Model Performance Sentence-BERT: Sentence Embeddings using Siamese BERT-Networks

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-10T21:49:32.271065Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T21:49:32.271065Z digest=sha256:1e43cb3969a411c4f792e291396cac9740618ac4d7c04a8072ab3f33f6f36021

Observation 042a6e48-3a63-452b-be60-f6e89ccf26f6 · outbound

This paper cites an unresolved cited work.

Investigating the Impact of Data Selection Strategies on Language Model Performance Unresolved cited work

Reference 12

Resolution
unresolved
raw_fallback, observed 2026-08-10T21:49:32.459945Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-10T21:49:32.275387Z digest=sha256:a70d2a9a96cba976570d3bb036d440efd0119fad3aaf4a15a92b586d0eec74f4

Observation 326e4e8e-f9d4-4b96-948f-5081f9855e58 · outbound

This paper cites GLUE: A Multi-Task Benchmark and Analysis Platform for Natural Language Understanding.

Investigating the Impact of Data Selection Strategies on Language Model Performance GLUE: A Multi-Task Benchmark and Analysis Platform for Natural Language Understanding

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-10T21:49:32.278940Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T21:49:32.278940Z digest=sha256:a2d8f56ecf8a801becc7a31c662b9e7eca588450d2ae097762eaf47453a34b90

Observation f520713b-44c9-4956-994a-6a7ff46e5da7 · outbound

This paper cites an unresolved cited work.

Investigating the Impact of Data Selection Strategies on Language Model Performance Unresolved cited work

Reference 14

Resolution
unresolved
raw_fallback, observed 2026-08-10T21:49:32.446891Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-10T21:49:32.282966Z digest=sha256:57d4fd294045fae7396e420a71bf358d36e9098074e0ec623bae63bc1ecf8f37

Observation 0a93cd73-86ba-4d1e-9729-1e33082a474d · outbound

This paper cites Pretraining Data Detection for Large Language Models: A Divergence-based Calibration Method.

Investigating the Impact of Data Selection Strategies on Language Model Performance Pretraining Data Detection for Large Language Models: A Divergence-based Calibration Method

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-10T21:49:32.286703Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T21:49:32.286703Z digest=sha256:b623da23223b080f017dafdc4363fae2bae51800c8803246a6753a9f57c31907

Observation 3ee829c9-ff87-488b-9849-4fb2ba4547c7 · outbound

This paper cites Large Language Model (LLM) for Telecommunications: A Comprehensive Survey on Principles, Key Techniques, and Opportunities.

Investigating the Impact of Data Selection Strategies on Language Model Performance Large Language Model (LLM) for Telecommunications: A Comprehensive Survey on Principles, Key Techniques, and Opportunities

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-10T21:49:32.290699Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T21:49:32.290699Z digest=sha256:24cd1ec0a100efdb50b25dba733ff58d79419c3f51707652b3e712beee5de004

Pith citing papers

No inbound Pith citation observations are available.