Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T14:25:22.360806Z
Paper Citation Record · LEDGER
As of 7 August 2026, this Paper Citation Record lists 73 of 73 outbound references and 1 inbound Pith citation observation for arXiv:2505.19051.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T14:25:22.360806Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-06-30T17:01:21.521025Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-06-30T17:04:56.636609Z
73 of 73 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 8fbc55e1-4d0d-4961-b324-4f89de4aa9fe · outbound
Efficient Data Selection at Scale via Influence Distillation LESS: Selecting Influential Data for Targeted Instruction Tuning
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3d6ef85e-3975-4139-be2b-80015db191cc · outbound
Efficient Data Selection at Scale via Influence Distillation Compute-Constrained Data Selection
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f876d2a2-41fd-4918-908b-e958932f7cad · outbound
Efficient Data Selection at Scale via Influence Distillation Selecting Informative Contexts Improves Language Model Finetuning
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation d619e2ab-6886-44f8-9a45-7f500f198bcc · outbound
Efficient Data Selection at Scale via Influence Distillation When Less is More: Investigating Data Pruning for Pretraining LLMs at Scale
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 19b78a8c-059c-4d4e-a570-b9b7eac4fa33 · outbound
Efficient Data Selection at Scale via Influence Distillation Perplexed by Perplexity: Perplexity-Based Data Pruning With Small Reference Models
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 21e7f305-db1c-408e-8639-e83c44e0de5f · outbound
Efficient Data Selection at Scale via Influence Distillation From Quantity to Quality: Boosting LLM Performance with Self-Guided Data Selection for Instruction Tuning
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1842ce34-e2e5-4790-8ad3-bdc80f01ace0 · outbound
Efficient Data Selection at Scale via Influence Distillation Large-Scale Data Selection for Instruction Tuning
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 79ee8196-a446-4fcb-9266-7f377ae89de4 · outbound
Efficient Data Selection at Scale via Influence Distillation Woodruff, and Michael Wunder
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 60b8b8ae-e7a5-4c03-bf9d-c1d4aa4238b1 · outbound
Efficient Data Selection at Scale via Influence Distillation Data selection for language models via importance resampling
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation bb7a900b-430a-47e6-9e52-c239a7a5ec99 · outbound
Efficient Data Selection at Scale via Influence Distillation DsDm: Model-Aware Dataset Selection with Datamodels
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b5eec917-9062-45a4-bfdc-781b790f61e1 · outbound
Efficient Data Selection at Scale via Influence Distillation Dynimpt: A dynamic data selection method for improving model training efficiency
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 11e1d17c-c31b-4f87-8cb3-e328be64f6f9 · outbound
Efficient Data Selection at Scale via Influence Distillation Llama 2: Open Foundation and Fine-Tuned Chat Models
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0463f96f-69bc-4ecb-a1b0-d5d34f434f28 · outbound
Efficient Data Selection at Scale via Influence Distillation Camels in a Changing Climate: Enhancing LM Adaptation with Tulu 2
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f4790058-ee55-4ed6-955b-52cba7aa5027 · outbound
Efficient Data Selection at Scale via Influence Distillation The Llama 3 Herd of Models
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation eb35e51a-e430-4238-b334-dcf3b50017d4 · outbound
Efficient Data Selection at Scale via Influence Distillation Qwen2.5: A party of foundation models, September 2024
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation befea17d-e1d9-4c00-a113-d58a9d298519 · outbound
Efficient Data Selection at Scale via Influence Distillation Measuring massive multitask language understanding
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fcdb9fb3-9c3f-40f1-b9db-e891da515fbe · outbound
Efficient Data Selection at Scale via Influence Distillation Aligning ai with shared human values
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation bb3c6ee3-5dc6-4280-99c3-324731caf66f · outbound
Efficient Data Selection at Scale via Influence Distillation Beyond neural scaling laws: beating power law scaling via data pruning
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d5a8c5d7-e961-4745-875f-389ced08038f · outbound
Efficient Data Selection at Scale via Influence Distillation SemDeDup: Data-efficient learning at web-scale through semantic deduplication
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 88d799a2-6e30-40e3-81b2-49fab40ddd1c · outbound
Efficient Data Selection at Scale via Influence Distillation Cross-lingual transfer learning with data selection for large-scale spoken language understanding
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation a41bb63b-5888-4361-84d8-82c8b50b1045 · outbound
Efficient Data Selection at Scale via Influence Distillation Smalltolarge (s2l): Scalable data selection for fine-tuning large language models by summarizing training loss trajectories of small models
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation f361fb2b-cf34-40e7-b9cd-9725d3e3ad4a · outbound
Efficient Data Selection at Scale via Influence Distillation Language Models are Few-Shot Learners
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 00ac99af-f005-4de7-b4c8-10b652dab6c2 · outbound
Efficient Data Selection at Scale via Influence Distillation The Pile: An 800GB Dataset of Diverse Text for Language Modeling
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a73e737e-1a7e-436f-821b-970c7579f6ed · outbound
Efficient Data Selection at Scale via Influence Distillation Palm: Scaling language modeling with pathways
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 82471f86-8c9e-4aea-8eba-95d0d6fb6a6a · outbound
Efficient Data Selection at Scale via Influence Distillation Glam: Efficient scaling of language models with mixture-of-experts
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 2bfb858c-ab44-4093-b67f-c1f624d15e86 · outbound
Efficient Data Selection at Scale via Influence Distillation Intelligent selection of language model training data
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation c452aa61-e1e1-4d88-b0e2-234b48fdb01e · outbound
Efficient Data Selection at Scale via Influence Distillation Cynical Selection of Language Model Training Data
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 5a7ade20-0521-4473-bff6-3c1bb2bab8a8 · outbound
Efficient Data Selection at Scale via Influence Distillation Automatic Document Selection for Efficient Encoder Pretraining
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7690d485-e810-4b73-9afc-bf3a14ecad51 · outbound
Efficient Data Selection at Scale via Influence Distillation Exploring the limits of transfer learning with a unified text-to-text transformer
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a9e3bc3b-a137-4acd-8a29-5540a37335c8 · outbound
Efficient Data Selection at Scale via Influence Distillation Skill-it! a data-driven skills framework for understanding and training language models
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 0f8af9a2-3458-4764-aeb1-601c3b23e6f8 · outbound
Efficient Data Selection at Scale via Influence Distillation Efficient online data mixing for language model pre-training
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 9d30b2e6-7b81-498c-8c6b-2ded60bc6076 · outbound
Efficient Data Selection at Scale via Influence Distillation DavIR: Data Selection via Implicit Reward for Large Language Models
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 29860640-ff21-4f63-ab10-007a95bad28e · outbound
Efficient Data Selection at Scale via Influence Distillation Dataset cartography: Mapping and diagnosing datasets with training dynamics
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 17ea860b-5e66-4673-a0b5-82171a40dc6a · outbound
Efficient Data Selection at Scale via Influence Distillation An Experimental Design Framework for Label-Efficient Supervised Finetuning of Large Language Models
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e924ecee-3e7c-4541-92d8-d6c7929fd640 · outbound
Efficient Data Selection at Scale via Influence Distillation D4: improving LLM pretraining via document de-duplication and diversification
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 474777d3-2037-4e4a-b396-c0df17e86390 · outbound
Efficient Data Selection at Scale via Influence Distillation Dsdm: Model-aware dataset selection with datamodels
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation d87a0e1b-bd04-4182-a1aa-25df599736f9 · outbound
Efficient Data Selection at Scale via Influence Distillation Learning from less data: A unified data subset selection and active learning framework for computer vision
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 8c036948-b5b7-4b01-80f2-b7226ffdb001 · outbound
Efficient Data Selection at Scale via Influence Distillation Retrieve: Coreset selection for efficient and robust semi-supervised learning
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 5533b5e1-4487-4a27-967c-91930106f009 · outbound
Efficient Data Selection at Scale via Influence Distillation Submodularity in data subset selection and active learning
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 99c9d33d-1de1-48fa-897f-013f850ee489 · outbound
Efficient Data Selection at Scale via Influence Distillation AlpaGasus: Training A Better Alpaca with Fewer Data
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4553c525-29a0-4827-8eaf-0364a95e7b86 · outbound
Efficient Data Selection at Scale via Influence Distillation Instruction Mining: Instruction Data Selection for Tuning Large Language Models
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9427bd11-8fc2-4a97-a465-12efbb58d1fd · outbound
Efficient Data Selection at Scale via Influence Distillation Active Learning for Convolutional Neural Networks: A Core-Set Approach
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 065bb8c1-8daf-4f54-84ed-cbfac5366862 · outbound
Efficient Data Selection at Scale via Influence Distillation Learning multiple layers of features from tiny images
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9df2c883-9405-4900-8f21-f08980d7e081 · outbound
Efficient Data Selection at Scale via Influence Distillation A software package for sequential quadratic programming
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation c5fb7ade-f1d3-47db-8a18-55d965c6d369 · outbound
Efficient Data Selection at Scale via Influence Distillation Fundamental algorithms for scientific computing in python and scipy 1.0 contributors
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation ac74536e-43cd-47ae-a685-e1c856b6fd0b · outbound
Efficient Data Selection at Scale via Influence Distillation Adam: A Method for Stochastic Optimization
Reference 46
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0cb541b8-9d92-4626-a4e5-89ebb8074ab1 · outbound
Efficient Data Selection at Scale via Influence Distillation Training Verifiers to Solve Math Word Problems
Reference 47
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6fe5a871-a5ab-4ff0-94f1-476304d1d5d2 · outbound
Efficient Data Selection at Scale via Influence Distillation Challenging BIG-Bench Tasks and Whether Chain-of-Thought Can Solve Them
Reference 48
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 32472503-e5aa-4826-985e-446ce7b0347a · outbound
Efficient Data Selection at Scale via Influence Distillation Clark, Eunsol Choi, Michael Collins, Dan Garrette, Tom Kwiatkowski, Vitaly Nikolaev, and Jennimaria Palomaki
Reference 49
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation afb0ea76-2f23-4533-8c7d-5f4df29d6fcc · outbound
Efficient Data Selection at Scale via Influence Distillation Evaluating Large Language Models Trained on Code
Reference 50
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a31d4afe-dcec-4ff7-a1db-d94e91c7d41f · outbound
Efficient Data Selection at Scale via Influence Distillation SQ u AD : 100,000+ questions for machine comprehension of text
Reference 51
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9245a6b5-e785-4207-b157-b69a084dad85 · outbound
Efficient Data Selection at Scale via Influence Distillation Alpacaeval: An automatic evaluator of instruction-following models, 2023 b
Reference 52
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ebb0f529-46d0-4e5c-93d0-4a61e2cd81dc · outbound
Efficient Data Selection at Scale via Influence Distillation Scaling Laws for Neural Language Models
Reference 53
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 02a068f5-cbb0-4130-847b-dd55170a42e4 · outbound
Efficient Data Selection at Scale via Influence Distillation Second-Order Forward-Mode Automatic Differentiation for Optimization
Reference 54
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6677dad7-ed6b-4286-9978-3af6fbc2b58e · outbound
Efficient Data Selection at Scale via Influence Distillation Lora: Low-rank adaptation of large language models
Reference 55
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8e7dcbd0-9d4b-4b92-a234-6347c225cabd · outbound
Efficient Data Selection at Scale via Influence Distillation TRAK: Attributing Model Behavior at Scale
Reference 56
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 36afd5ad-b3d3-4fa0-8887-ea30f67f55c6 · outbound
Efficient Data Selection at Scale via Influence Distillation Sharpness-Aware Minimization for Efficiently Improving Generalization
Reference 57
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6e3305e8-fbf8-45f9-bbb6-941bef4d9fb2 · outbound
Efficient Data Selection at Scale via Influence Distillation CrAM: A Compression-Aware Minimizer
Reference 58
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 47a69f1c-e9f3-4c64-bc1b-069c5604184e · outbound
Efficient Data Selection at Scale via Influence Distillation Scaling instruction-finetuned language models
Reference 59
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 47599f09-4443-44ff-b2e8-12eb9cd12737 · outbound
Efficient Data Selection at Scale via Influence Distillation o pf, Yannic Kilcher, Dimitri Von R \
Reference 60
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation e7186ae3-75ed-441f-8451-c12f5d1b4c9e · outbound
Efficient Data Selection at Scale via Influence Distillation Free dolly: Introducing the world’s first truly open instruction-tuned llm, 2023
Reference 61
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f75019a2-5318-439a-88c0-baa0a93723a8 · outbound
Efficient Data Selection at Scale via Influence Distillation Instruction Tuning with GPT-4
Reference 62
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c8474305-6044-4d0e-9fb3-7910ac924ebb · outbound
Efficient Data Selection at Scale via Influence Distillation Code alpaca: An instruction-following llama model for code generation, 2023
Reference 63
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation d7bd0df5-54eb-484d-98e4-91e61d49214c · outbound
Efficient Data Selection at Scale via Influence Distillation Lima: Less is more for alignment
Reference 64
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 7068f832-bd6c-4fc5-86ef-328f9326f2f6 · outbound
Efficient Data Selection at Scale via Influence Distillation WizardLM: Empowering large pre-trained language models to follow complex instructions
Reference 65
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 04a7edd7-6903-46dc-a012-cb0049da96f7 · outbound
Efficient Data Selection at Scale via Influence Distillation Openorca: An open dataset of gpt augmented flan reasoning traces, 2023
Reference 66
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 01ff910e-d46c-4c31-9199-9d7d391b3146 · outbound
Efficient Data Selection at Scale via Influence Distillation Sciriff: A resource to enhance language model instruction-following over scientific literature
Reference 67
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation de3062b4-e1da-4ee2-b833-0b457111a35e · outbound
Efficient Data Selection at Scale via Influence Distillation PaLM 2 Technical Report
Reference 68
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e984cdef-35ff-4ca2-abb9-a64fe1702d18 · outbound
Efficient Data Selection at Scale via Influence Distillation NV-Embed: Improved Techniques for Training LLMs as Generalist Embedding Models
Reference 69
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 261a37da-89c6-4cf4-83e0-86e9bdbd9157 · outbound
Efficient Data Selection at Scale via Influence Distillation Large Dual Encoders Are Generalizable Retrievers
Reference 70
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4d899f91-98c7-411e-a87d-069ad122b8f8 · outbound
Efficient Data Selection at Scale via Influence Distillation GaLore: Memory-Efficient LLM Training by Gradient Low-Rank Projection
Reference 71
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f5039b9f-99d5-4e32-955f-13a37e0dc572 · outbound
Efficient Data Selection at Scale via Influence Distillation HadaCore: Tensor Core Accelerated Hadamard Transform Kernel
Reference 72
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 206e9a1f-4905-400d-a309-729e1907d5fb · outbound
Efficient Data Selection at Scale via Influence Distillation Fast hadamard transform in cuda, with a pytorch interface, 2023
Reference 73
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation ec243924-b070-4a3b-9b1b-9e31b51376e3 · inbound
PRISM: Preference-Aware Influence Function Based Data Selection Method for Efficient Fine-Tuning Efficient Data Selection at Scale via Influence Distillation
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.