Pith. sign in

Paper Citation Record · LEDGER

MATES: Model-Aware Data Selection for Efficient Pretraining with Data Influence Models

As of 21 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 15 inbound Pith citation observations for arXiv:2406.06046.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2406.06046 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 15 of 15 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-21T06:32:19.484+00:00

measured 15 of 15 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-16T11:07:15.177393Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-03T16:48:39.405954Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation e6f4674b-7d2a-431d-938d-686dfdeb9be6 · inbound

Evaluating Sample Utility for Efficient Data Selection by Mimicking Model Weights cites this paper.

Evaluating Sample Utility for Efficient Data Selection by Mimicking Model Weights MATES: Model-Aware Data Selection for Efficient Pretraining with Data Influence Models

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-10T21:02:00.900926Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:02:00.900926Z digest=sha256:47dda342315504e4a3030835b51864fbded26bd2560168c36ec51aed4d9d504e

Observation d13fa1ac-469f-47e0-b95f-4199f6516d46 · inbound

Improving Influence-based Instruction Tuning Data Selection for Balanced Learning of Diverse Capabilities cites this paper.

Improving Influence-based Instruction Tuning Data Selection for Balanced Learning of Diverse Capabilities MATES: Model-Aware Data Selection for Efficient Pretraining with Data Influence Models

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-10T17:34:52.523831Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T17:34:52.523831Z digest=sha256:e8dd2e226806f39ae6a860596c6a30f0262a9a55577c98dfdfc4d02ddfbc6888

Observation 44a041b2-6238-4687-8ab0-71a56e4a2e29 · inbound

QuaDMix: Quality-Diversity Balanced Data Selection for Efficient LLM Pretraining cites this paper.

QuaDMix: Quality-Diversity Balanced Data Selection for Efficient LLM Pretraining MATES: Model-Aware Data Selection for Efficient Pretraining with Data Influence Models

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-16T11:07:15.177393Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:07:15.177393Z digest=sha256:b0b0d81f060adbfb946a2bab3874b02b10165bf700b3d8ef941fb743b186b35f

Observation 4780d297-bb0f-4226-b312-8b81f5c1bbb3 · inbound

Not All Documents Are What You Need for Extracting Instruction Tuning Data cites this paper.

Not All Documents Are What You Need for Extracting Instruction Tuning Data MATES: Model-Aware Data Selection for Efficient Pretraining with Data Influence Models

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-15T20:43:13.892154Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:43:13.892154Z digest=sha256:db78b5a3f937730f36e4c4a68257d3f7f971a45a5e486c338460738c45926ea5

Observation b747885b-54fb-4515-b014-9f40ca1aff07 · inbound

IDEAL: Data Equilibrium Adaptation for Multi-Capability Language Model Alignment cites this paper.

IDEAL: Data Equilibrium Adaptation for Multi-Capability Language Model Alignment MATES: Model-Aware Data Selection for Efficient Pretraining with Data Influence Models

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-15T20:32:06.125097Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:32:06.125097Z digest=sha256:81947e5a4cb5f122c4cbf1611af532c9f32a19da6e1d8fde18721e5a08c37adb

Observation f3e0ccfe-6668-4d09-b5f5-7a27ce0e6b6c · inbound

RefineX: Learning to Refine Pre-training Data at Scale from Expert-Guided Programs cites this paper.

RefineX: Learning to Refine Pre-training Data at Scale from Expert-Guided Programs MATES: Model-Aware Data Selection for Efficient Pretraining with Data Influence Models

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-06T20:21:56.521904Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:21:56.521904Z digest=sha256:10a08c07efbb6409972bc8c4e016f8532339f3e8d58b551e3808874939fa1426

Observation c91ef90a-3734-44ab-832d-846d511e3051 · inbound

LLM Data Selection and Utilization via Dynamic Bi-level Optimization cites this paper.

LLM Data Selection and Utilization via Dynamic Bi-level Optimization MATES: Model-Aware Data Selection for Efficient Pretraining with Data Influence Models

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-06T15:21:58.101569Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:21:58.101569Z digest=sha256:56a7a5ac80779a6c75dafba2b7a92559326dd15f7f46cc2cf001f9e118f282c0

Observation 0ac98bfb-6693-4547-848f-59cec9b21172 · inbound

BLISS: A Lightweight Bilevel Influence Scoring Method for Data Selection in Language Model Pretraining cites this paper.

BLISS: A Lightweight Bilevel Influence Scoring Method for Data Selection in Language Model Pretraining MATES: Model-Aware Data Selection for Efficient Pretraining with Data Influence Models

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-04T11:16:15.281996Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T11:16:15.281996Z digest=sha256:b36fa28b5ba0247ce3cb05f9e91462b49c228c9b1632064dd357774ca14e504d

Observation 668e404a-e66a-4f73-8486-a25ecaa60f98 · inbound

GradAlign: Gradient-Aligned Data Selection for LLM Reinforcement Learning cites this paper.

GradAlign: Gradient-Aligned Data Selection for LLM Reinforcement Learning MATES: Model-Aware Data Selection for Efficient Pretraining with Data Influence Models

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-02T21:05:26.804742Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T21:05:26.804742Z digest=sha256:40a9e6dd80a74822231eb48cd2b4c36c633580e6d3d896f25f912abe2ecadca9

Observation 05bae7ed-8f67-43e1-87cf-22af4b88def0 · inbound

Efficient Dataset Selection for Continual Adaptation of Generative Recommenders cites this paper.

Efficient Dataset Selection for Continual Adaptation of Generative Recommenders MATES: Model-Aware Data Selection for Efficient Pretraining with Data Influence Models

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-05-11T05:16:03.771646Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-10T18:13:46.622289Z digest=sha256:ed8ed5c91033c8524f5b9f5792ff7acc830af22d982a583319533fb826a1b690

Observation d6b18352-4ab4-49ea-87fd-1b67c952ec14 · inbound

An Empirical Study on Influence-Based Pretraining Data Selection for Code Large Language Models cites this paper.

An Empirical Study on Influence-Based Pretraining Data Selection for Code Large Language Models MATES: Model-Aware Data Selection for Efficient Pretraining with Data Influence Models

Reference 44

Resolution
verified exact
arxiv_id, observed 2026-05-11T05:16:04.521623Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-10T18:13:24.750244Z digest=sha256:e2026b53fce9caa8876a0a579d264dd2b8410994a15481c6f4452ee62e176c31

Observation c1e8c891-7e22-489a-89ba-d09753bbf189 · inbound

GRACE: A Dynamic Coreset Selection Framework for Large Language Model Optimization cites this paper.

GRACE: A Dynamic Coreset Selection Framework for Large Language Model Optimization MATES: Model-Aware Data Selection for Efficient Pretraining with Data Influence Models

Reference 80

Resolution
malformed identifier
arxiv_id, observed 2026-05-11T05:30:57.980667Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-10T18:06:46.131725Z digest=sha256:8e9a8ec356f5a5d5b8495c42bb0abb9d483b908d41091e448e7a216f5c656bfb

Observation 030ce241-3d36-4241-873b-1a6b7690c8b4 · inbound

Sketching the Readout of Large Language Models for Scalable Data Attribution and Valuation cites this paper.

Sketching the Readout of Large Language Models for Scalable Data Attribution and Valuation MATES: Model-Aware Data Selection for Efficient Pretraining with Data Influence Models

Reference 57

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T08:48:00.969572Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-10T08:47:36.122054Z digest=sha256:5e4648293bd1a2fe07d31cb34f2262aa33e25adc06a003b9f64b38e5c4ef14a2

Observation 1205b449-a336-4516-8c77-888553e55417 · inbound

Sketching the Readout of Large Language Models for Scalable Data Attribution and Valuation cites this paper.

Sketching the Readout of Large Language Models for Scalable Data Attribution and Valuation MATES: Model-Aware Data Selection for Efficient Pretraining with Data Influence Models

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-02T16:12:13.358123Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T16:12:13.358123Z digest=sha256:46a643d0c3e82f9422167c7b03e5c9856837d02ad126844203d636763d16571b

Observation b3ced776-83bb-45e9-a0c6-bb99289f2fde · inbound

HERMES: A Multi-Granularity Labeling Substrate for Pre-training Data Mixtures cites this paper.

HERMES: A Multi-Granularity Labeling Substrate for Pre-training Data Mixtures MATES: Model-Aware Data Selection for Efficient Pretraining with Data Influence Models

Reference 46

Resolution
verified exact
arxiv_id, observed 2026-07-03T16:48:39.407393Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-07-03T16:44:41.720388Z digest=sha256:7561d96dc80bdcc01c34daf63ca7542392f5ccac9ff86a30774390613a2bdd97