Pith. sign in

Paper Citation Record · LEDGER

From 128K to 4M: Efficient Training of Ultra-Long Context Large Language Models

As of 8 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 7 inbound Pith citation observations for arXiv:2504.06214.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2504.06214 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 7 of 7 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 7 of 7 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T15:16:00.996148Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T18:40:03.180712Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation dfb3ad48-2a43-47d1-9fef-2f118a452604 · inbound

Beyond Hard and Soft: Hybrid Context Compression for Balancing Local and Global Information Retention cites this paper.

Beyond Hard and Soft: Hybrid Context Compression for Balancing Local and Global Information Retention From 128K to 4M: Efficient Training of Ultra-Long Context Large Language Models

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-07T15:16:00.996148Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:16:00.996148Z digest=sha256:c8b86ad672585be6e1f3d19e790d33de218d014519cee3b2a2558779dd5fedc0

Observation dd34b158-bdcf-4517-8278-3005e3f6e66a · inbound

Shuffle the Context: RoPE-Perturbed Self-Distillation for Long-Context Adaptation cites this paper.

Shuffle the Context: RoPE-Perturbed Self-Distillation for Long-Context Adaptation From 128K to 4M: Efficient Training of Ultra-Long Context Large Language Models

Reference 5

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T13:45:28.120415Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-10T13:42:00.440049Z digest=sha256:01859f93b8353221e230ffdbf1197c75873aa5275b0fab578b1a9565c3debe50

Observation 20ccfe7c-cfdc-409b-ad11-236c93f7e9b3 · inbound

ZAYA1-8B Technical Report cites this paper.

ZAYA1-8B Technical Report From 128K to 4M: Efficient Training of Ultra-Long Context Large Language Models

Reference 198

Resolution
verified exact
arxiv_id, observed 2026-05-11T17:26:05.154164Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-05-08T17:36:37.182196Z digest=sha256:d3eeb76815f7394a5899f93baeab1ab9ef8253ca55c38df341d3604261bdebf7

Observation 2c63f0d9-3753-42a0-b754-98863efae634 · inbound

How Many Different Outputs Can a Transformer Generate? cites this paper.

How Many Different Outputs Can a Transformer Generate? From 128K to 4M: Efficient Training of Ultra-Long Context Large Language Models

Reference 55

Resolution
verified exact
arxiv_id, observed 2026-05-22T07:11:12.744720Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-05-22T07:09:23.107309Z digest=sha256:8f04ac7e56b5485554c33e90d4ba224b90626b510f50733491f478e4d3c4a278

Observation ea8533d5-77ec-4638-b380-70359f522517 · inbound

ZONOS2 Technical Report cites this paper.

ZONOS2 Technical Report From 128K to 4M: Efficient Training of Ultra-Long Context Large Language Models

Reference 239

Resolution
verified exact
arxiv_id, observed 2026-07-04T18:40:03.182027Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-06-25T22:37:15.072758Z digest=sha256:09f276d36f1373ceedb43fff53732f477cfb8b14929443846560d28130240c37

Observation 62c3ea62-a6fd-4bec-9716-acb0b19113e2 · inbound

ZONOS2 Technical Report cites this paper.

ZONOS2 Technical Report From 128K to 4M: Efficient Training of Ultra-Long Context Large Language Models

Reference 239

Resolution
verified exact
arxiv_id, observed 2026-07-01T18:15:58.911311Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-06-29T02:07:31.791835Z digest=sha256:9e7087753951d35b2b29e85ba259b0437823664e0e57646831b0798b2dbcb58e

Observation 8532207a-9dd6-4a16-8ac1-feafc98ff1d3 · inbound

ZUNA1.1: A more flexible EEG foundation model for Denoising and Super-resolution cites this paper.

ZUNA1.1: A more flexible EEG foundation model for Denoising and Super-resolution From 128K to 4M: Efficient Training of Ultra-Long Context Large Language Models

Reference 208

Resolution
unresolved
no resolver link, observed 2026-08-01T09:52:09.146522Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T09:52:09.146522Z digest=sha256:c6fb12e727a59220a3271b01ccbb4432d2185bd7fd60d530c2142c97112b7aa7