Pith. sign in

Paper Citation Record · LEDGER

Maybe Only 0.5% Data is Needed: A Preliminary Exploration of Low Training Data Instruction Tuning

As of 22 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 20 inbound Pith citation observations for arXiv:2305.09246.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2305.09246 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 20 of 20 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-22T06:32:14.747728+00:00

measured 20 of 20 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-16T11:46:10.856665Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

8
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation a17b9d42-079a-4054-9af9-99bd1de6bf93 · inbound

A Comprehensive Overview of Large Language Models cites this paper.

A Comprehensive Overview of Large Language Models Maybe Only 0.5% Data is Needed: A Preliminary Exploration of Low Training Data Instruction Tuning

Reference 184

Resolution
verified exact
arxiv_id, observed 2026-05-19T20:28:39.584384Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-19T20:28:38.900026Z digest=sha256:3668d9d321b1054fb807a5dd68810987917c833d29da24e72caea1d13b7d599c

Observation 3b3810f5-0a64-4834-a2fa-fc7cec50e540 · inbound

From Cool Demos to Production-Ready FMware: Core Challenges and a Technology Roadmap cites this paper.

From Cool Demos to Production-Ready FMware: Core Challenges and a Technology Roadmap Maybe Only 0.5% Data is Needed: A Preliminary Exploration of Low Training Data Instruction Tuning

Reference 34

Resolution
verified exact
arxiv_id, observed 2026-05-23T19:08:20.935697Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-23T19:07:21.016824Z digest=sha256:9cbf54dbf97ab515034a6612a7fdabd9cde3f82cdcfee55361dd3e7825b70338

Observation 831d248c-202b-4444-bf6b-ede708da17cb · inbound

A Survey on Large Language Models with some Insights on their Capabilities and Limitations cites this paper.

A Survey on Large Language Models with some Insights on their Capabilities and Limitations Maybe Only 0.5% Data is Needed: A Preliminary Exploration of Low Training Data Instruction Tuning

Reference 265

Resolution
unresolved
no resolver link, observed 2026-08-10T22:17:56.419192Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:17:56.419192Z digest=sha256:40ba42ff62e61a353d49482649f4cb7ee1f59991760b5649940ea74c7e2834c5

Observation 18374ee6-0418-4c25-b4ac-749a5f658f05 · inbound

CDS: Knowledge Component-Driven Data Synthesis Guided by Cognitive Diagnosis Theory cites this paper.

CDS: Knowledge Component-Driven Data Synthesis Guided by Cognitive Diagnosis Theory Maybe Only 0.5% Data is Needed: A Preliminary Exploration of Low Training Data Instruction Tuning

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-10T20:44:16.879150Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T20:44:16.879150Z digest=sha256:bc91266f012eba550dc5faaac3e5bc0eac24f8af3d560ebf383e5cc7ee6142b5

Observation 8799ccc2-9085-41d3-9fbe-10ed33cfcfc9 · inbound

Can Large Language Models Be Query Optimizer for Relational Databases? cites this paper.

Can Large Language Models Be Query Optimizer for Relational Databases? Maybe Only 0.5% Data is Needed: A Preliminary Exploration of Low Training Data Instruction Tuning

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-08T18:54:32.353949Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T18:54:32.353949Z digest=sha256:58c5b472dc315623653369d823fbcad6a32efc57e3c76dbe89c5ea0adab77d0d

Observation bc008993-0af5-4596-b09c-5d0eda0255c9 · inbound

DONOD: Efficient and Generalizable Instruction Fine-Tuning for LLMs via Model-Intrinsic Dataset Pruning cites this paper.

DONOD: Efficient and Generalizable Instruction Fine-Tuning for LLMs via Model-Intrinsic Dataset Pruning Maybe Only 0.5% Data is Needed: A Preliminary Exploration of Low Training Data Instruction Tuning

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-16T11:46:10.856665Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:46:10.856665Z digest=sha256:88e4dbe9fd58e7938ccfd391508adff58b3dd51f116b5cdb869034c65144a985

Observation 3188442b-d488-4227-9b42-53d04e518808 · inbound

A Survey of Foundation Model-Powered Recommender Systems: From Feature-Based, Generative to Agentic Paradigms cites this paper.

A Survey of Foundation Model-Powered Recommender Systems: From Feature-Based, Generative to Agentic Paradigms Maybe Only 0.5% Data is Needed: A Preliminary Exploration of Low Training Data Instruction Tuning

Reference 277

Resolution
unresolved
no resolver link, observed 2026-08-16T11:07:59.542979Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:07:59.542979Z digest=sha256:af8ff7824b9895929c6f3a06c57e9a2fd7fdf5be8f67ff88ba741ed8b979bd14

Observation 3f31b49d-8cb8-4be8-a523-a6ea5f6d7bee · inbound

Text2Cypher: Data Pruning using Hard Example Selection cites this paper.

Text2Cypher: Data Pruning using Hard Example Selection Maybe Only 0.5% Data is Needed: A Preliminary Exploration of Low Training Data Instruction Tuning

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-15T23:16:58.093030Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T23:16:58.093030Z digest=sha256:b96cc8c6d5ccb88aba7f62bf18c304c0e79422a838602757640a6f4683cbcedc

Observation 4df64c3e-463d-4036-8ef3-d208509a3a8f · inbound

ReqBrain: Task-Specific Instruction Tuning of LLMs for AI-Assisted Requirements Generation cites this paper.

ReqBrain: Task-Specific Instruction Tuning of LLMs for AI-Assisted Requirements Generation Maybe Only 0.5% Data is Needed: A Preliminary Exploration of Low Training Data Instruction Tuning

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T14:48:02.545145Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:48:02.545145Z digest=sha256:ff7cb318686e7890077021ee9407caf96618af90cd9aad06b304856fcf9256a7

Observation 7fcd1e42-e9fc-4010-b60f-aae0e4afd8d4 · inbound

ALPS: Attention Localization and Pruning Strategy for Efficient Alignment of Large Language Models cites this paper.

ALPS: Attention Localization and Pruning Strategy for Efficient Alignment of Large Language Models Maybe Only 0.5% Data is Needed: A Preliminary Exploration of Low Training Data Instruction Tuning

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T14:31:28.279489Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:31:28.279489Z digest=sha256:6ee28f11871f5707faf7551ec007e3509863e6a32bf4cc367ea88301be056f18

Observation f5508cc0-8a67-49c5-9e72-5f3b9356a530 · inbound

Towards Efficient and Effective Alignment of Large Language Models cites this paper.

Towards Efficient and Effective Alignment of Large Language Models Maybe Only 0.5% Data is Needed: A Preliminary Exploration of Low Training Data Instruction Tuning

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T04:55:35.255414Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:55:35.255414Z digest=sha256:15f6fa9742c2b7b7e5fdd8c86f6b217eb85b56989e9dbc2dabbe57c1f43cca49

Observation a36c054a-8148-4258-b397-d8431c44bc66 · inbound

Minifinetuning: Low-Data Generation Domain Adaptation through Corrective Self-Distillation cites this paper.

Minifinetuning: Low-Data Generation Domain Adaptation through Corrective Self-Distillation Maybe Only 0.5% Data is Needed: A Preliminary Exploration of Low Training Data Instruction Tuning

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T12:40:25.256358Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:40:25.256358Z digest=sha256:8158185ae67c4f0c8943e43cea9234ec4a9dc1851623920dbd694272dc5bdfe0

Observation f35c73f3-39d1-45fb-807c-6ba3f365470f · inbound

Vision-G1: Towards General Vision Language Reasoning with Multi-Domain Data Curation cites this paper.

Vision-G1: Towards General Vision Language Reasoning with Multi-Domain Data Curation Maybe Only 0.5% Data is Needed: A Preliminary Exploration of Low Training Data Instruction Tuning

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-15T17:24:08.231109Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:24:08.231109Z digest=sha256:08ecb76d9fd5bd01f3e606569a6f61fe292fd6373d97d39bae18a21722637497

Observation 31f8ba62-0d74-4f2a-93fa-901086584fb9 · inbound

LAMDAS: LLM as an Implicit Classifier for Domain-specific Data Selection cites this paper.

LAMDAS: LLM as an Implicit Classifier for Domain-specific Data Selection Maybe Only 0.5% Data is Needed: A Preliminary Exploration of Low Training Data Instruction Tuning

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-04T23:33:39.307095Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T23:33:39.307095Z digest=sha256:d3af6cf18b0218f04089b8bb4529dc1367d78728ab22ab1cfd5388ae8cc14e40

Observation d40b9ff7-d536-4acf-9353-ca8e723307c4 · inbound

A Systematic Survey on Large Language Models for Evolutionary Optimization: From Modeling to Solving cites this paper.

A Systematic Survey on Large Language Models for Evolutionary Optimization: From Modeling to Solving Maybe Only 0.5% Data is Needed: A Preliminary Exploration of Low Training Data Instruction Tuning

Reference 110

Resolution
unresolved
no resolver link, observed 2026-08-04T20:55:41.527188Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:55:41.527188Z digest=sha256:e22584a5793043db37c64d9a9208b7e34e3002e2750b2800f9a15ffec6471030

Observation 0e20ff72-985d-4a07-8e15-ce9c79849849 · inbound

A Critical Look at Targeted Instruction Selection: Disentangling What Matters (and What Doesn't) cites this paper.

A Critical Look at Targeted Instruction Selection: Disentangling What Matters (and What Doesn't) Maybe Only 0.5% Data is Needed: A Preliminary Exploration of Low Training Data Instruction Tuning

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-02T23:10:59.419614Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T23:10:59.419614Z digest=sha256:9f0304fcaa878a28020310c487be6af80d70a8e7041f28d77c8298cb265b367b

Observation 3dc571b7-546d-4855-b965-92104cab70fc · inbound

GRACE: A Dynamic Coreset Selection Framework for Large Language Model Optimization cites this paper.

GRACE: A Dynamic Coreset Selection Framework for Large Language Model Optimization Maybe Only 0.5% Data is Needed: A Preliminary Exploration of Low Training Data Instruction Tuning

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-05-10T20:30:48.972081Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-10T18:06:46.131725Z digest=sha256:d22d1c03e8ce515746fe07594c0ad16f19485d4ad073bdb07f381c78a54defd7

Observation 0d257c10-9733-40ce-8402-a1a36c6abb3e · inbound

Measuring Distribution Shift in User Prompts and Its Effects on LLM Performance cites this paper.

Measuring Distribution Shift in User Prompts and Its Effects on LLM Performance Maybe Only 0.5% Data is Needed: A Preliminary Exploration of Low Training Data Instruction Tuning

Reference 71

Resolution
verified exact
arxiv_id, observed 2026-05-10T06:56:47.928366Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-05-10T05:27:50.049798Z digest=sha256:a6c5bd3e34d48d66388a27c27267b74ecca7efb44d7a8e30f7f13dccff5d1882

Observation b8741f3e-86b9-44f7-a692-e87c6dbefe66 · inbound

InvEvolve: Evolving White-Box Inventory Policies via Large Language Models with Performance Guarantees cites this paper.

InvEvolve: Evolving White-Box Inventory Policies via Large Language Models with Performance Guarantees Maybe Only 0.5% Data is Needed: A Preliminary Exploration of Low Training Data Instruction Tuning

Reference 166

Resolution
verified exact
arxiv_id, observed 2026-05-11T15:31:08.034102Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-05-09T19:50:39.734124Z digest=sha256:f2c488d563e1ec06af7a484b71ee11a64ab60ac5c91fb264a8aa53f053f4e3f5

Observation f5bd5196-42f2-4383-a605-f82d92ada5f5 · inbound

InvEvolve: Evolving White-Box Inventory Policies via Large Language Models with Performance Guarantees cites this paper.

InvEvolve: Evolving White-Box Inventory Policies via Large Language Models with Performance Guarantees Maybe Only 0.5% Data is Needed: A Preliminary Exploration of Low Training Data Instruction Tuning

Reference 166

Resolution
verified exact
arxiv_id, observed 2026-05-12T02:41:17.650267Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-05-12T02:38:45.322351Z digest=sha256:9d1a1ee9ee64e8bb73952c8eed972ba43220a71a61a3dac98299d60370538229