Pith. sign in

Paper Citation Record · LEDGER

When Does Metadata Conditioning (NOT) Work for Language Model Pre-Training? A Study with Context-Free Grammars

As of 18 August 2026, this Paper Citation Record lists 17 of 17 outbound references and 1 inbound Pith citation observation for arXiv:2504.17562.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2504.17562 v2

Coverage vector

measured 17 of 17 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-16T10:40:21.640422Z

measured 18 of 18 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-05-17T04:46:33.641714Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-17T04:49:02.953758Z

Reference resolution

17 of 17 outbound references displayed

  • verified exact0
  • verified fuzzy3
  • unresolved11
  • parse uncertain0
  • malformed identifier3
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 5f6a8e5f-4826-437a-bd56-fb07154f952c · outbound

This paper cites HTLM: Hyper-Text Pre-Training and Prompting of Language Models.

When Does Metadata Conditioning (NOT) Work for Language Model Pre-Training? A Study with Context-Free Grammars HTLM: Hyper-Text Pre-Training and Prompting of Language Models

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-16T10:40:21.569567Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T10:40:21.569567Z digest=sha256:4c758416918a37995258e2fda138f5d92e4eded060eed797f81fdeab8906e321

Observation 5ca024e4-4918-47e3-bf12-a9e51078cda4 · outbound

This paper cites Plug and Play Language Models: A Simple Approach to Controlled Text Generation.

When Does Metadata Conditioning (NOT) Work for Language Model Pre-Training? A Study with Context-Free Grammars Plug and Play Language Models: A Simple Approach to Controlled Text Generation

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-16T10:40:21.597843Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T10:40:21.597843Z digest=sha256:c15e5b0e75cb7564f3f6b4814291603a25c7b749f8e95055d9ffef055b360e8c

Observation 29d8bd5e-f1ac-48ef-b137-78c7c3bb4333 · outbound

This paper cites A Theory of Emergent In-Context Learning as Implicit Structure Induction.

When Does Metadata Conditioning (NOT) Work for Language Model Pre-Training? A Study with Context-Free Grammars A Theory of Emergent In-Context Learning as Implicit Structure Induction

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-16T10:40:21.607260Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T10:40:21.607260Z digest=sha256:eecdb33afce95b0efc6943eb7d845771c3708c8853ff125848912d7209ff7e3a

Observation 5da20810-687a-4991-a0a4-8bb7687e7b92 · outbound

This paper cites CTRL: A Conditional Transformer Language Model for Controllable Generation.

When Does Metadata Conditioning (NOT) Work for Language Model Pre-Training? A Study with Context-Free Grammars CTRL: A Conditional Transformer Language Model for Controllable Generation

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-16T10:40:21.611507Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T10:40:21.611507Z digest=sha256:7fc03ba2871f2e2b565f952e1841be4f71f02c7221d95c792b540ee8e9d12b47

Observation 5e9c4b23-e569-4400-9b61-08e4f262930b · outbound

This paper cites Source-Aware Training Enables Knowledge Attribution in Language Models.

When Does Metadata Conditioning (NOT) Work for Language Model Pre-Training? A Study with Context-Free Grammars Source-Aware Training Enables Knowledge Attribution in Language Models

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-16T10:40:21.615899Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T10:40:21.615899Z digest=sha256:c4cdc5e6d44d58ee95c274f3ba3b65d8b6f2acc784e8dc9f799297dd1505b902

Observation 0c8d3eb8-84ea-472b-a12a-d0bfcf67ede1 · outbound

This paper cites Implicit meta-learning may lead language models to trust more reliable sources.

When Does Metadata Conditioning (NOT) Work for Language Model Pre-Training? A Study with Context-Free Grammars Implicit meta-learning may lead language models to trust more reliable sources

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-16T10:40:21.620137Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T10:40:21.620137Z digest=sha256:890805eb1e4300f562a194ac44a566b2ab387260f0816c1c7a3feceecd3aa47b

Observation fad96e55-2756-4869-9d56-10ff2f9de7b7 · outbound

This paper cites Hugo Touvron, Thibaut Lavril, Gautier Izacard, Xavier Martinet, Marie-Anne Lachaux, Timoth´ee Lacroix, Baptiste Rozi `ere, Naman Goyal, Eric Hambro, Faisal Azhar, et al.

When Does Metadata Conditioning (NOT) Work for Language Model Pre-Training? A Study with Context-Free Grammars Hugo Touvron, Thibaut Lavril, Gautier Izacard, Xavier Martinet, Marie-Anne Lachaux, Timoth´ee Lacroix, Baptiste Rozi `ere, Naman Goyal, Eric Hambro, Faisal Azhar, et al

Reference 13

Resolution
malformed identifier
no resolver link, observed 2026-08-16T10:40:21.624364Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T10:40:21.624364Z digest=sha256:86c05cd92aadf4ee67926384f87303c4a152222aebadf54062ffd614c8b6c9c9

Observation 59c8a36f-fe50-424f-a446-ae69a9dcffd0 · outbound

This paper cites Conditional Language Learning with Context.

When Does Metadata Conditioning (NOT) Work for Language Model Pre-Training? A Study with Context-Free Grammars Conditional Language Learning with Context

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-16T10:40:21.628207Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T10:40:21.628207Z digest=sha256:d3753ab60cded2d730fa9ca482556d33f45fff659f47c33ef05cf4c87e232d49

Observation 196d73a5-9b48-4591-8ac3-6c4af658c168 · outbound

This paper cites Deciphering the impact of pretraining data on large language models through machine unlearning.

When Does Metadata Conditioning (NOT) Work for Language Model Pre-Training? A Study with Context-Free Grammars Deciphering the impact of pretraining data on large language models through machine unlearning

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T10:40:21.919692Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T10:40:21.632162Z digest=sha256:b1292bdc1e3dea27c8e0ddebe497d27083d05c61c21bd3a1a42bd03174f2ef59

Observation a849e908-be1c-4485-8680-ec13837d9ccb · outbound

This paper cites 12 Published as a conference paper at COLM 2025 A More Details of the Experimental Setting Training Setting.

When Does Metadata Conditioning (NOT) Work for Language Model Pre-Training? A Study with Context-Free Grammars 12 Published as a conference paper at COLM 2025 A More Details of the Experimental Setting Training Setting

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T10:40:21.907550Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T10:40:21.636242Z digest=sha256:2e96805769dfe79332874b24479d55336616d800a8f6f77d1faf497e226fff15

Observation cbefdf3e-f022-4134-8587-64e57b6af8f0 · outbound

This paper cites B The loss of next-token prediction for various token positions.

When Does Metadata Conditioning (NOT) Work for Language Model Pre-Training? A Study with Context-Free Grammars B The loss of next-token prediction for various token positions

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T10:40:21.895155Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T10:40:21.640422Z digest=sha256:efae1ef6d9f5c62aa72a6b84a88d40b4e1b1d375e71f57cdbffb948c1013b337

Observation 88cdee63-919e-4970-a1ce-fae8ab7d020f · outbound

This paper cites URL https: //aclanthology.org/P18-1198/.

When Does Metadata Conditioning (NOT) Work for Language Model Pre-Training? A Study with Context-Free Grammars URL https: //aclanthology.org/P18-1198/

Reference 2018

Resolution
unresolved
no resolver link, observed 2026-08-16T10:40:21.592365Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T10:40:21.592365Z digest=sha256:a69eca95a5ce37ea660231f75c29e21c2a9df29a5555052ac42431328713ad30

Observation e3d15d29-0f1d-41c5-811b-1d81e8edf486 · outbound

This paper cites The Pile: An 800GB Dataset of Diverse Text for Language Modeling.

When Does Metadata Conditioning (NOT) Work for Language Model Pre-Training? A Study with Context-Free Grammars The Pile: An 800GB Dataset of Diverse Text for Language Modeling

Reference 2019

Resolution
unresolved
no resolver link, observed 2026-08-16T10:40:21.602496Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T10:40:21.602496Z digest=sha256:66e26ebe94686668415c830d4cd2a4af64cfa81936351d1a0bbd32fb7284d5fa

Observation 804b2c0f-49a5-4f8b-bb7f-bceea1a77a75 · outbound

This paper cites CoCon: A Self-Supervised Approach for Controlled Text Generation.

When Does Metadata Conditioning (NOT) Work for Language Model Pre-Training? A Study with Context-Free Grammars CoCon: A Self-Supervised Approach for Controlled Text Generation

Reference 2022

Resolution
malformed identifier
no resolver link, observed 2026-08-16T10:40:21.587527Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T10:40:21.587527Z digest=sha256:72e5d03600646292163bb80d395a23a056509eea660ec4c453301d6afb270197

Observation 705191bd-db36-43d6-966a-6c590f03ffbf · outbound

This paper cites Physics of Language Models: Part 3.3, Knowledge Capacity Scaling Laws.

When Does Metadata Conditioning (NOT) Work for Language Model Pre-Training? A Study with Context-Free Grammars Physics of Language Models: Part 3.3, Knowledge Capacity Scaling Laws

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-16T10:40:21.578676Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T10:40:21.578676Z digest=sha256:a847bebc9b5c2a2d60c65a424a56b29754dc192aa489cae05de6254fb001e66d

Observation 6c110442-d4c6-442d-8a5d-30f0951b371a · outbound

This paper cites A Theory for Emergence of Complex Skills in Language Models.

When Does Metadata Conditioning (NOT) Work for Language Model Pre-Training? A Study with Context-Free Grammars A Theory for Emergence of Complex Skills in Language Models

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-16T10:40:21.582670Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T10:40:21.582670Z digest=sha256:89771deaaad398dbcfeb941fc2f874e1a15cc5fcbaf1be2abc618af3bfbbe90c

Observation 99c4e181-5b57-431c-97e9-36560abea958 · outbound

This paper cites Physics of Language Models: Part 1, Learning Hierarchical Language Structures.

When Does Metadata Conditioning (NOT) Work for Language Model Pre-Training? A Study with Context-Free Grammars Physics of Language Models: Part 1, Learning Hierarchical Language Structures

Reference 2025

Resolution
malformed identifier
no resolver link, observed 2026-08-16T10:40:21.574336Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T10:40:21.574336Z digest=sha256:50177964266c5b0f8239086a0f549db5d371cb1971b0f5fabd4c8087288c78df

Pith citing papers

Observation 3d3eda9a-55b6-4fcf-8478-9f2759dde147 · inbound

Beyond URLs: Metadata Diversity and Position for Efficient LLM Pretraining cites this paper.

Beyond URLs: Metadata Diversity and Position for Efficient LLM Pretraining When Does Metadata Conditioning (NOT) Work for Language Model Pre-Training? A Study with Context-Free Grammars

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-05-17T04:49:02.956416Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-17T04:46:33.641714Z digest=sha256:ea55b403c92528b3187a37cefaf9697f2897ee7085e49eda28efc6e1a130c891