Pith. sign in

Paper Citation Record · LEDGER

NExtLong: Toward Effective Long-Context Training without Long Documents

As of 18 August 2026, this Paper Citation Record lists 100 of 117 outbound references and 5 inbound Pith citation observations for arXiv:2501.12766.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2501.12766 v2

Coverage vector

measured 100 of 117 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-10T16:52:49.881585Z

measured 105 of 105 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00

measured 5 of 5 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-06T20:40:09.996269Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-15T11:17:24.563972Z

Reference resolution

100 of 117 outbound references displayed

  • verified exact2
  • verified fuzzy7
  • unresolved91
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation b8ae1682-3c40-4687-a630-b011c5e2c9a5 · outbound

This paper cites Many-Shot In-Context Learning.

NExtLong: Toward Effective Long-Context Training without Long Documents Many-Shot In-Context Learning

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-10T16:52:49.596261Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T16:52:49.596261Z digest=sha256:a1d10412846b432611d510ce59b6dbc022a84a992388d81e907c954a34664240

Observation ef92d5aa-ab39-4f31-8643-62951e46d8bf · outbound

This paper cites Yi: Open foundation models by 01.ai, 2024.

NExtLong: Toward Effective Long-Context Training without Long Documents Yi: Open foundation models by 01.ai, 2024

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-10T16:52:49.599116Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T16:52:49.599116Z digest=sha256:6296b4746301e5f3cc3716678a2337610a80c24218ea6dff24d28b6b2f872414

Observation ec82821e-59b9-4268-a165-a7757384fe65 · outbound

This paper cites M., Neyshabur, B., and Zhai, X.

NExtLong: Toward Effective Long-Context Training without Long Documents M., Neyshabur, B., and Zhai, X

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-10T16:52:49.601633Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T16:52:49.601633Z digest=sha256:23492e4cfdb1404672f78db140ec3677550f0fa7a6c0712e346a24c2cc9d2058

Observation 3783d258-7d66-4799-b8cb-242a19435443 · outbound

This paper cites Training-Free Long-Context Scaling of Large Language Models.

NExtLong: Toward Effective Long-Context Training without Long Documents Training-Free Long-Context Scaling of Large Language Models

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-10T16:52:49.604307Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T16:52:49.604307Z digest=sha256:7ffbfbe452e54dae7d6da32e60cc332be21135fdc31d590af8adaebf1b728d14

Observation b417ecd5-918b-45ef-a074-782e32c56f70 · outbound

This paper cites Many-shot jailbreaking.

NExtLong: Toward Effective Long-Context Training without Long Documents Many-shot jailbreaking

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-10T16:52:49.607236Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T16:52:49.607236Z digest=sha256:2d6bf8546764b6e3d6ddd406b9172d0c7c20867ef39c0cdf38739fdbf81b3bfd

Observation 07b44178-b8f0-465e-b75f-c2976d39b72d · outbound

This paper cites an unresolved cited work.

NExtLong: Toward Effective Long-Context Training without Long Documents Unresolved cited work

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-10T16:52:49.610063Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T16:52:49.610063Z digest=sha256:0d9df34d9aa2577266b845d406648013532b36864e3ed06b336bdd410b52dd55

Observation b3904a4a-e84d-4761-9527-a968441afa4c · outbound

This paper cites and Lempitsky, V.

NExtLong: Toward Effective Long-Context Training without Long Documents and Lempitsky, V

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-10T16:52:49.612786Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T16:52:49.612786Z digest=sha256:165f2d6f004cf563e7a2388c8400207f0965aff6329f56959e73d1ad84dd1254

Observation a766e8d3-1417-404c-b18c-dc00bde38027 · outbound

This paper cites LongBench: A Bilingual, Multitask Benchmark for Long Context Understanding.

NExtLong: Toward Effective Long-Context Training without Long Documents LongBench: A Bilingual, Multitask Benchmark for Long Context Understanding

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-10T16:52:49.615464Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T16:52:49.615464Z digest=sha256:1a03aba9ff247ae6f805f8b15c3fdde2f31ac2293bd3f4d3b62867fe1499fc66

Observation 68ebb1a4-f872-41d9-a115-c915e5d26d6a · outbound

This paper cites LongBench v2: Towards Deeper Understanding and Reasoning on Realistic Long-context Multitasks.

NExtLong: Toward Effective Long-Context Training without Long Documents LongBench v2: Towards Deeper Understanding and Reasoning on Realistic Long-context Multitasks

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-10T16:52:49.618678Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T16:52:49.618678Z digest=sha256:baf6d1cced44e2003433cc147345be333cf0204bc9b7e39e5fb144458d2e3374

Observation 8b76c459-c5ed-4732-afba-1b8ff64cbaa7 · outbound

This paper cites Codeplan: Repository-level coding using llms and planning.

NExtLong: Toward Effective Long-Context Training without Long Documents Codeplan: Repository-level coding using llms and planning

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-10T16:52:49.621755Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T16:52:49.621755Z digest=sha256:c5c29054f61bfcebe065e3efe12044a9e2a88bdafa0c9cc76df6b0482c9757bb

Observation a44e6632-6b02-4616-b9d7-6a5dcfe9324d · outbound

This paper cites Smollm-corpus, 2024.

NExtLong: Toward Effective Long-Context Training without Long Documents Smollm-corpus, 2024

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-10T16:52:49.624687Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T16:52:49.624687Z digest=sha256:b9c1da1deb8353ae2ff210d27834ef293eed6fd188a97031d7151b2d89936e26

Observation 1b70d8eb-f91c-48c7-a662-edec16b9cd5e · outbound

This paper cites and LeCun, Y.

NExtLong: Toward Effective Long-Context Training without Long Documents and LeCun, Y

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-10T16:52:49.627468Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T16:52:49.627468Z digest=sha256:30b170f39ea83cf2e2127ad4638948868a37eb593db6f8524cd82959ee513def

Observation 2da1e6e3-832c-4efe-8489-cb87fe6f1226 · outbound

This paper cites In-Context Learning with Long-Context Models: An In-Depth Exploration.

NExtLong: Toward Effective Long-Context Training without Long Documents In-Context Learning with Long-Context Models: An In-Depth Exploration

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-10T16:52:49.630273Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T16:52:49.630273Z digest=sha256:c6c038f6660498f3fcbb4e0acbfe56bae8967f1d12f9d6e822c6e7109320e09c

Observation fbcd98db-3ab5-4344-9f77-7e4fc31a0133 · outbound

This paper cites DeepSeek LLM: Scaling Open-Source Language Models with Longtermism.

NExtLong: Toward Effective Long-Context Training without Long Documents DeepSeek LLM: Scaling Open-Source Language Models with Longtermism

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-10T16:52:49.634603Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T16:52:49.634603Z digest=sha256:a2d7c1d435bf9c8cad4ebe7ddea0afc0bf245dc63c7276d618fd8c4a52e640f7

Observation d3f9f1bc-6061-4d09-b14b-0967a2b00c40 · outbound

This paper cites G., Bradley, H., O’Brien, K., Hallahan, E., Khan, M.

NExtLong: Toward Effective Long-Context Training without Long Documents G., Bradley, H., O’Brien, K., Hallahan, E., Khan, M

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-10T16:52:49.637843Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T16:52:49.637843Z digest=sha256:8cb39d6a72b9da3e0f6f0bb18983265525c9e0cd3307abce65cc6b4cc7271353

Observation 598f7e81-aec7-4508-83b7-ee3e75952e1a · outbound

This paper cites Piqa: Reasoning about physical commonsense in natural language.

NExtLong: Toward Effective Long-Context Training without Long Documents Piqa: Reasoning about physical commonsense in natural language

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-10T16:52:49.640604Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T16:52:49.640604Z digest=sha256:591d59106b9654e678f88232097d6538f654af15e97285d7276483439992df3b

Observation 5ad28309-b0af-4b6a-a967-9a8f025aa474 · outbound

This paper cites Peek across: Improving multi-document modeling via cross-document question-answering.

NExtLong: Toward Effective Long-Context Training without Long Documents Peek across: Improving multi-document modeling via cross-document question-answering

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-10T16:52:49.643163Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T16:52:49.643163Z digest=sha256:ff6b87472dfaade4d968c86a2c8ecd6c742db4e1675025c3a8e4a4df2fcf0846

Observation e88be9cc-dcf6-4e41-b2bd-2648398b65c6 · outbound

This paper cites Resolving the imbalance issue in hierarchical disciplinary topic inference via llm-based data augmentation.

NExtLong: Toward Effective Long-Context Training without Long Documents Resolving the imbalance issue in hierarchical disciplinary topic inference via llm-based data augmentation

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-10T16:52:49.645836Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T16:52:49.645836Z digest=sha256:b0f8f26c9b46b6fc7a7df98a47e529d21750cb06070a91dad3014e8093208d1d

Observation 7fdc6626-4477-4d76-b0dc-b27a5289afda · outbound

This paper cites CLEX: Continuous Length Extrapolation for Large Language Models.

NExtLong: Toward Effective Long-Context Training without Long Documents CLEX: Continuous Length Extrapolation for Large Language Models

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-10T16:52:49.648565Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T16:52:49.648565Z digest=sha256:76014d486d8d3d39b47cd3004bccc5bc7b925c13e82f8ba9bf43dee73031cdf6

Observation 7304842d-9c25-478e-8b0f-2d9e08e7aba0 · outbound

This paper cites Extending Context Window of Large Language Models via Positional Interpolation.

NExtLong: Toward Effective Long-Context Training without Long Documents Extending Context Window of Large Language Models via Positional Interpolation

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-10T16:52:49.651533Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T16:52:49.651533Z digest=sha256:05198635329b92e0afc0a2467e3e5787e67edf24ef7495007919ee79c80b3814

Observation e13da2df-cfe2-4191-ae01-62b044db974f · outbound

This paper cites Incremental False Negative Detection for Contrastive Learning.

NExtLong: Toward Effective Long-Context Training without Long Documents Incremental False Negative Detection for Contrastive Learning

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-10T16:52:49.654524Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T16:52:49.654524Z digest=sha256:338a0de256cd23365257a4ce34a65dfd0d67f7ad1b21671dc89c5f67cba69bdc

Observation d71abe2e-6f35-4f4b-8360-3188ae0c8af1 · outbound

This paper cites LongLoRA: Efficient Fine-tuning of Long-Context Large Language Models.

NExtLong: Toward Effective Long-Context Training without Long Documents LongLoRA: Efficient Fine-tuning of Long-Context Large Language Models

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-10T16:52:49.657576Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T16:52:49.657576Z digest=sha256:45bd98de6d37cb0d3a4000ccb762b44b7d1474b2fdcafceb2d348da8ade16a67

Observation aac55ebb-5eed-4fcc-b41b-b55ae99c9ecd · outbound

This paper cites Language Models as Science Tutors.

NExtLong: Toward Effective Long-Context Training without Long Documents Language Models as Science Tutors

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-10T16:52:49.660461Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T16:52:49.660461Z digest=sha256:a51361f027ed0309ab490046dac0471aa5576905948af024bbc4c3400d82566c

Observation 64fb3737-ecca-4717-8d7c-77917eebca21 · outbound

This paper cites Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge.

NExtLong: Toward Effective Long-Context Training without Long Documents Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-10T16:52:49.663362Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T16:52:49.663362Z digest=sha256:65e75ab92a6f557ee948a516f16141265258253d84009fdad7e2addf803fee15

Observation cb4f1c2c-c51a-4b11-9265-d25e99141dda · outbound

This paper cites FlashAttention-2: Faster Attention with Better Parallelism and Work Partitioning.

NExtLong: Toward Effective Long-Context Training without Long Documents FlashAttention-2: Faster Attention with Better Parallelism and Work Partitioning

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-10T16:52:49.666210Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T16:52:49.666210Z digest=sha256:8c9ae0a163e90bf1f2e56d3a75cdfb7f08fec15511c3816c0499127a6122e592

Observation 511ee546-b07c-4fd5-82c8-e0679dd3465e · outbound

This paper cites Deepseek-v2: A strong, economical, and efficient mixture-of-experts language model, 2024.

NExtLong: Toward Effective Long-Context Training without Long Documents Deepseek-v2: A strong, economical, and efficient mixture-of-experts language model, 2024

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-10T16:52:49.668962Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T16:52:49.668962Z digest=sha256:9e2397821f460ef3cad8b60e11eca981d9256abaf2abc5ba641be50978ae9042

Observation 8afd3d6d-faa3-4021-8457-4d6d2ac0588b · outbound

This paper cites Fewer truncations improve language modeling.

NExtLong: Toward Effective Long-Context Training without Long Documents Fewer truncations improve language modeling

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-10T16:52:49.671609Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T16:52:49.671609Z digest=sha256:42a30d9e3064f3d6478defb75a46a0850d71d590e5518b066499b0f22b0de774

Observation 108b24ad-0d6a-4d4e-90e1-15314906cff0 · outbound

This paper cites Enhancing Chat Language Models by Scaling High-quality Instructional Conversations.

NExtLong: Toward Effective Long-Context Training without Long Documents Enhancing Chat Language Models by Scaling High-quality Instructional Conversations

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-10T16:52:49.674195Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T16:52:49.674195Z digest=sha256:ad2c40e2fb7147f5254cec52053b6c7c952e280595948697b85d923689578d24

Observation 762099b0-450b-4b28-a9ca-52d373fc15ba · outbound

This paper cites The Llama 3 Herd of Models.

NExtLong: Toward Effective Long-Context Training without Long Documents The Llama 3 Herd of Models

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-10T16:52:49.676823Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T16:52:49.676823Z digest=sha256:06cb41b7e0f777d98eb6da00acb31cb2c42da870abedcf5c43a4a6ab09612dd3

Observation e2964de3-5c80-4fef-843a-4412e1b95952 · outbound

This paper cites O., Hart, P.

NExtLong: Toward Effective Long-Context Training without Long Documents O., Hart, P

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-10T16:52:49.679652Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T16:52:49.679652Z digest=sha256:f9634569522f7ad35923352e8f6dbf3ce1a8295fc17b391c45f035ec1d4efdd1

Observation 88e10fbb-9019-43e8-962a-ff2634938633 · outbound

This paper cites Data Engineering for Scaling Language Models to 128K Context.

NExtLong: Toward Effective Long-Context Training without Long Documents Data Engineering for Scaling Language Models to 128K Context

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-10T16:52:49.681924Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T16:52:49.681924Z digest=sha256:6de26cd358d06f27815ade3fd7ae853f54b490db26ab7ed6139fc25c778a5371

Observation 0fa4a2af-350f-4345-ae4f-8cece78bd100 · outbound

This paper cites Quest: Query-centric Data Synthesis Approach for Long-context Scaling of Large Language Model.

NExtLong: Toward Effective Long-Context Training without Long Documents Quest: Query-centric Data Synthesis Approach for Long-context Scaling of Large Language Model

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-10T16:52:49.684394Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T16:52:49.684394Z digest=sha256:c4b1dd0d8d118221a1eb2efb8e416e24dfb1f181f31cec3ad1545818ffc0609d

Observation e1f16e7b-eaed-4cdc-a00c-9ab2404259b7 · outbound

This paper cites The Pile: An 800GB Dataset of Diverse Text for Language Modeling.

NExtLong: Toward Effective Long-Context Training without Long Documents The Pile: An 800GB Dataset of Diverse Text for Language Modeling

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-10T16:52:49.687471Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T16:52:49.687471Z digest=sha256:83aee873dfb502cc47c9894e19d482f5b705aa98a46d4304452d8c657f814a89

Observation ecc9ea48-3fab-483a-9850-07d03379ff00 · outbound

This paper cites How to train long-context language models (effectively).

NExtLong: Toward Effective Long-Context Training without Long Documents How to train long-context language models (effectively)

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-10T16:52:49.690352Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T16:52:49.690352Z digest=sha256:bad229e1892f7317013210aed7bf170420ee3bd22a536aeb11ee6e84803e0ae3

Observation 6b9bfb5a-1d8f-48ae-aff6-c7f609c6f012 · outbound

This paper cites Deep learning, volume 1.

NExtLong: Toward Effective Long-Context Training without Long Documents Deep learning, volume 1

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-10T16:52:49.692952Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T16:52:49.692952Z digest=sha256:46d480b73dbadeb454eb00d0844444565889b9f9d19eb2f5c7d612582eaee83d

Observation 380909dd-b631-41ed-ac72-bc7edd35b795 · outbound

This paper cites Retrieval augmented language model pre-training.

NExtLong: Toward Effective Long-Context Training without Long Documents Retrieval augmented language model pre-training

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-10T16:52:49.696012Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T16:52:49.696012Z digest=sha256:3d9a7bf0bbcd9f2b71f18f23ca2eafd3e56d94b0063f5d2739887f6141b320df

Observation 1f90e41c-179f-4239-9296-8d1b5e1491c4 · outbound

This paper cites LM-Infinite: Zero-Shot Extreme Length Generalization for Large Language Models.

NExtLong: Toward Effective Long-Context Training without Long Documents LM-Infinite: Zero-Shot Extreme Length Generalization for Large Language Models

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-10T16:52:49.698721Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T16:52:49.698721Z digest=sha256:53cdab5d1be22f827502399d316fdde4a78c8941011bebc2884ba5bd7ed607fd

Observation f09505fc-3b04-49c7-8867-4413fb92b92d · outbound

This paper cites Measuring Massive Multitask Language Understanding.

NExtLong: Toward Effective Long-Context Training without Long Documents Measuring Massive Multitask Language Understanding

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-10T16:52:49.701550Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T16:52:49.701550Z digest=sha256:f7e48fa4c1955f9ff33d2b9d6d385a47dbdbfcef08464993698be1daa50d0ab1

Observation 7d4597f3-7373-4045-a5cc-3db81c221f26 · outbound

This paper cites Scaling Laws for Autoregressive Generative Modeling.

NExtLong: Toward Effective Long-Context Training without Long Documents Scaling Laws for Autoregressive Generative Modeling

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-10T16:52:49.704664Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T16:52:49.704664Z digest=sha256:4fb17b16aa136c2cff01be440cdd5e5234d775ef1cd4522b7fa0b9c0de346b63

Observation 3a2382c3-c54c-4c74-a721-12e032519888 · outbound

This paper cites E., Osindero, S., and Teh, Y.

NExtLong: Toward Effective Long-Context Training without Long Documents E., Osindero, S., and Teh, Y

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-10T16:52:49.707589Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T16:52:49.707589Z digest=sha256:e7636f6fefe04bd4f1c00186d333c3fc556652275631b3601b30ac349fc1b18d

Observation def041a3-b0ad-4389-b601-a299c4a07578 · outbound

This paper cites Training Compute-Optimal Large Language Models.

NExtLong: Toward Effective Long-Context Training without Long Documents Training Compute-Optimal Large Language Models

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-10T16:52:49.710294Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T16:52:49.710294Z digest=sha256:9a62970f0a7b34ed13218613a23cd80210fb50774588917122e8bd92a0139e2c

Observation ba128b55-5878-480b-ac9e-cbf96b1fec11 · outbound

This paper cites RULER: What's the Real Context Size of Your Long-Context Language Models?.

NExtLong: Toward Effective Long-Context Training without Long Documents RULER: What's the Real Context Size of Your Long-Context Language Models?

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-10T16:52:49.713153Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T16:52:49.713153Z digest=sha256:7acd8e8c2646d12a5e2694048680855d49b769a0ab5cd088d9f1982ceac9745d

Observation 7edebc2f-17b2-44bb-b871-e6875b46cd1a · outbound

This paper cites Camels in a Changing Climate: Enhancing LM Adaptation with Tulu 2.

NExtLong: Toward Effective Long-Context Training without Long Documents Camels in a Changing Climate: Enhancing LM Adaptation with Tulu 2

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-10T16:52:49.716088Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T16:52:49.716088Z digest=sha256:61964595a819c402a7e76d58cfea8908ed082bb56698e4ed2a79f80a6e8cc54d

Observation ce720dd9-6067-4225-9cf1-349f28ad3a1c · outbound

This paper cites Long-Context LLMs Meet RAG: Overcoming Challenges for Long Inputs in RAG.

NExtLong: Toward Effective Long-Context Training without Long Documents Long-Context LLMs Meet RAG: Overcoming Challenges for Long Inputs in RAG

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-10T16:52:49.718953Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T16:52:49.718953Z digest=sha256:444be1e09147c208b74bec24b296d5a0698865abc9a15c6353ce49469bce6fb2

Observation fb940a89-586e-4315-b614-922cfaec16d4 · outbound

This paper cites Llm maybe longlm: Self-extend llm context window without tuning, 2024 b.

NExtLong: Toward Effective Long-Context Training without Long Documents Llm maybe longlm: Self-extend llm context window without tuning, 2024 b

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-10T16:52:49.721963Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T16:52:49.721963Z digest=sha256:f37095ead5e3f6c29552be1cd97e57a70fbb934a39744c27569f087f7403a882

Observation a0fd8472-b10c-436f-901a-587dcfe507f6 · outbound

This paper cites Understanding the effects of language-specific class imbalance in multilingual fine-tuning.

NExtLong: Toward Effective Long-Context Training without Long Documents Understanding the effects of language-specific class imbalance in multilingual fine-tuning

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-10T16:52:49.724847Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T16:52:49.724847Z digest=sha256:028dc192ef80aeea49f3ae558f259b473ed57589ac68025f2c7fa801b5e783c9

Observation 48b86255-52ec-4763-b0cd-8ff13f1cd786 · outbound

This paper cites B., Pion, N., Weinzaepfel, P., and Larlus, D.

NExtLong: Toward Effective Long-Context Training without Long Documents B., Pion, N., Weinzaepfel, P., and Larlus, D

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-10T16:52:49.727703Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T16:52:49.727703Z digest=sha256:01679e2f5003fbe6232cba616fb81468b03ac6c5c06fc3247b6a1b4b66c9cb8a

Observation aaea3f9b-36d6-49b7-92dd-a7c9faf3cf45 · outbound

This paper cites Scaling Laws for Neural Language Models.

NExtLong: Toward Effective Long-Context Training without Long Documents Scaling Laws for Neural Language Models

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-10T16:52:49.730470Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T16:52:49.730470Z digest=sha256:fcb8641960280e5fee32dac1d1ab6ee3a50cab6d29f49549d1cb299b56308b6d

Observation a1f00387-35bd-45ac-a33e-387b2d27b890 · outbound

This paper cites F., and Ramakrishnan, R.

NExtLong: Toward Effective Long-Context Training without Long Documents F., and Ramakrishnan, R

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-10T16:52:49.733534Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T16:52:49.733534Z digest=sha256:5f7928e06e22e9dd13eae65c224979cda874e7703c256da82392009443eba106

Observation 75cbb18a-ae66-4dc5-9e8a-30e0546ebeaf · outbound

This paper cites an unresolved cited work.

NExtLong: Toward Effective Long-Context Training without Long Documents Unresolved cited work

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-10T16:52:49.736077Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T16:52:49.736077Z digest=sha256:cce7a6f64eab6e024a932fef40b73c18500a07137d93b180415db0fcef87c189

Observation c0315e77-57ff-496d-afa3-af0eb05494de · outbound

This paper cites The Stack: 3 TB of permissively licensed source code.

NExtLong: Toward Effective Long-Context Training without Long Documents The Stack: 3 TB of permissively licensed source code

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-10T16:52:49.738727Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T16:52:49.738727Z digest=sha256:8edf3411d6a456b1a0d5dd57ff02a8439126a8eb60978a5cb8e6f224154bf0a0

Observation 33455ace-2951-4073-9d23-21d5da625b67 · outbound

This paper cites Crafting papers on machine learning.

NExtLong: Toward Effective Long-Context Training without Long Documents Crafting papers on machine learning

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-10T16:52:49.741630Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T16:52:49.741630Z digest=sha256:ca4ea325d760da28e1f5372d3011b9a220719e04f91b4d5f1cad39e86723e270

Observation 67259b43-c288-4cd9-88f8-db1048e9fd6e · outbound

This paper cites S., Yvon, F., Gall \'e , M., et al.

NExtLong: Toward Effective Long-Context Training without Long Documents S., Yvon, F., Gall \'e , M., et al

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-10T16:52:49.744131Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T16:52:49.744131Z digest=sha256:1d8d41822626c3959dc803fc4228587e35f105f483484e8acb5d7b84b4cb8e55

Observation 60eacaa4-c011-4097-bc00-9c2edf7c902f · outbound

This paper cites The Inductive Bias of In-Context Learning: Rethinking Pretraining Example Design.

NExtLong: Toward Effective Long-Context Training without Long Documents The Inductive Bias of In-Context Learning: Rethinking Pretraining Example Design

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-10T16:52:49.746561Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T16:52:49.746561Z digest=sha256:58f75385b90862104031182b3ec625e1748ec6d599a1221f8fc1d7fa858e78b3

Observation 4efbd360-56d0-449b-9ed2-5ef69d0d0a80 · outbound

This paper cites u ttler, H., Lewis, M., Yih, W.-t., Rockt \.

NExtLong: Toward Effective Long-Context Training without Long Documents u ttler, H., Lewis, M., Yih, W.-t., Rockt \

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-10T16:52:49.749394Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T16:52:49.749394Z digest=sha256:2112caaa1b1b2968d045667ce903a0559118c003a8cd96e6c7738f12eabbd17b

Observation 0a902c6a-73af-46a8-9f56-d24a860efe74 · outbound

This paper cites Functional Interpolation for Relative Positions Improves Long Context Transformers.

NExtLong: Toward Effective Long-Context Training without Long Documents Functional Interpolation for Relative Positions Improves Long Context Transformers

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-10T16:52:49.751740Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T16:52:49.751740Z digest=sha256:a79cc9d8080a7f5661b5660e080b54c6966393fbc89c27b0a8c44f65e8de03d2

Observation d83144d2-1bf8-41ee-84a0-225e65017feb · outbound

This paper cites From System 1 to System 2: A Survey of Reasoning Large Language Models.

NExtLong: Toward Effective Long-Context Training without Long Documents From System 1 to System 2: A Survey of Reasoning Large Language Models

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-10T16:52:49.754619Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T16:52:49.754619Z digest=sha256:320964604c61719a0d868c4c60f409aec44138cfcde02d0e885f75edd9266008

Observation fdb8ded2-0c1d-4086-aca4-63dadbac0a58 · outbound

This paper cites World Model on Million-Length Video And Language With Blockwise RingAttention.

NExtLong: Toward Effective Long-Context Training without Long Documents World Model on Million-Length Video And Language With Blockwise RingAttention

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-10T16:52:49.757307Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T16:52:49.757307Z digest=sha256:5a5bcf98a7324526ea0e79be6505059fdb18c378b46a5cb80c39853b3bb47bd9

Observation b49441af-a9a2-41f1-aa27-c946f7615cd5 · outbound

This paper cites LogiQA: A Challenge Dataset for Machine Reading Comprehension with Logical Reasoning.

NExtLong: Toward Effective Long-Context Training without Long Documents LogiQA: A Challenge Dataset for Machine Reading Comprehension with Logical Reasoning

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-10T16:52:49.760124Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T16:52:49.760124Z digest=sha256:91f7a46459715183058c4c89d2d5fe6ae01c0227e4caf1add2b50102ff9ec713

Observation 7d814e29-aac4-4e04-8f28-f752da1f4228 · outbound

This paper cites Decoupled Weight Decay Regularization.

NExtLong: Toward Effective Long-Context Training without Long Documents Decoupled Weight Decay Regularization

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-10T16:52:49.763213Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T16:52:49.763213Z digest=sha256:f75d70592be94eda9a7f0a41df8997c6c8e93eb159a5387188740f051afa16dc

Observation a3236e0b-e274-4afb-8fc9-d7a3205fac85 · outbound

This paper cites Fineweb-edu, May 2024.

NExtLong: Toward Effective Long-Context Training without Long Documents Fineweb-edu, May 2024

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-10T16:52:49.766364Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T16:52:49.766364Z digest=sha256:f645372ec295a36e87d1f6ea40591d5c95a05a93c583cec3d19266dca7e5c59f

Observation fe6d6e35-b7e6-473a-89bb-72eb8eb2fa15 · outbound

This paper cites Learning passage impacts for inverted indexes.

NExtLong: Toward Effective Long-Context Training without Long Documents Learning passage impacts for inverted indexes

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-10T16:52:49.769219Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T16:52:49.769219Z digest=sha256:f006b8fb4ebc055f716238fe942840bb15503f2a211da302ae29712b353c2d71

Observation 873dbf4f-fad0-4c84-aed0-d1538d325e6d · outbound

This paper cites and Liu, B.

NExtLong: Toward Effective Long-Context Training without Long Documents and Liu, B

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-10T16:52:49.772089Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T16:52:49.772089Z digest=sha256:3fe78ad03fc411d487cd85aadbaa1b35157e10bdc66aa0543178ee6dc5e1df16

Observation d09e2b34-1770-448a-acd2-735c740246aa · outbound

This paper cites Introducing meta llama 3: The most capable openly available llm to date, 2024.

NExtLong: Toward Effective Long-Context Training without Long Documents Introducing meta llama 3: The most capable openly available llm to date, 2024

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-10T16:52:49.774963Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T16:52:49.774963Z digest=sha256:947189ecdd95110e6075fd33531d0d57ca21f7e40072cf89e4224f3ab05df402

Observation d19d97c5-34d2-4fca-bcce-b21f9202bb57 · outbound

This paper cites S., Carbonell, J.

NExtLong: Toward Effective Long-Context Training without Long Documents S., Carbonell, J

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-10T16:52:49.777844Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T16:52:49.777844Z digest=sha256:dcd1f4bad09800d96c2af4aaf2b89f73b9cbe13d417f5079919c2de885c7ffb2

Observation 7768ffa1-ad0c-49eb-a762-975947208f82 · outbound

This paper cites an unresolved cited work.

NExtLong: Toward Effective Long-Context Training without Long Documents Unresolved cited work

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-10T16:52:49.780519Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T16:52:49.780519Z digest=sha256:7086f189976f996ae2ee385f3e26b251bae59a3548a9518cdbdb6241c80625c6

Observation 46d83723-8b22-4624-baee-3717e079833a · outbound

This paper cites and Rosenbloom, P.

NExtLong: Toward Effective Long-Context Training without Long Documents and Rosenbloom, P

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-10T16:52:49.783405Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T16:52:49.783405Z digest=sha256:6ef6b13251bbe2254684f806deaa0b1411888d6e48c3721a6759a335d5fb6946

Observation ad86d884-2cb2-4678-8f49-0eab00173228 · outbound

This paper cites Document Expansion by Query Prediction.

NExtLong: Toward Effective Long-Context Training without Long Documents Document Expansion by Query Prediction

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-10T16:52:49.786336Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T16:52:49.786336Z digest=sha256:5a210dcd73c2093c6fe2b246241149908477a29a46d135871a85830e8a1d58b5

Observation 2792de7c-82df-4926-a951-62f47a88274e · outbound

This paper cites GPT-4 Technical Report.

NExtLong: Toward Effective Long-Context Training without Long Documents GPT-4 Technical Report

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-10T16:52:49.789274Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T16:52:49.789274Z digest=sha256:8da5c187872be18a00e45af1f1755e79f1743424bf65955c1a577d08ddabd2a7

Observation 20e8d39d-7f00-4235-87a0-b8935d354b5e · outbound

This paper cites Training language models to follow instructions with human feedback.

NExtLong: Toward Effective Long-Context Training without Long Documents Training language models to follow instructions with human feedback

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-10T16:52:49.793338Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T16:52:49.793338Z digest=sha256:54f1ee860f53948a9908cd05467c148b9d7d4db1cdfe2fd19201619cd9e2d551

Observation bd3fd6e0-61eb-4b64-86ab-c2ca03e46a07 · outbound

This paper cites The LAMBADA dataset: Word prediction requiring a broad discourse context.

NExtLong: Toward Effective Long-Context Training without Long Documents The LAMBADA dataset: Word prediction requiring a broad discourse context

Reference 72

Resolution
unresolved
no resolver link, observed 2026-08-10T16:52:49.796342Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T16:52:49.796342Z digest=sha256:6aed5127ef3680866239ed8430aac452c62f429e9b726e0d74292b66788d6f69

Observation 111ca4d8-84d9-4467-b4a8-7058be393f07 · outbound

This paper cites OpenWebMath: An Open Dataset of High-Quality Mathematical Web Text.

NExtLong: Toward Effective Long-Context Training without Long Documents OpenWebMath: An Open Dataset of High-Quality Mathematical Web Text

Reference 73

Resolution
unresolved
no resolver link, observed 2026-08-10T16:52:49.799535Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T16:52:49.799535Z digest=sha256:d1bcb42cdf5b6732d9ce334368dc1a0e42b47cf03a6f7f7dad3a80d2c7382cdf

Observation 4658dd33-4243-4d1e-92d6-03e97b3974ad · outbound

This paper cites YaRN: Efficient Context Window Extension of Large Language Models.

NExtLong: Toward Effective Long-Context Training without Long Documents YaRN: Efficient Context Window Extension of Large Language Models

Reference 74

Resolution
unresolved
no resolver link, observed 2026-08-10T16:52:49.802483Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T16:52:49.802483Z digest=sha256:1dd215de9be8d29e3f9b16cc394d7579ab86ec8b6990cdfaf37f666d0c7388fa

Observation 76c77de0-f723-4676-969d-872242b9c1df · outbound

This paper cites Improving language understanding by generative pre-training.

NExtLong: Toward Effective Long-Context Training without Long Documents Improving language understanding by generative pre-training

Reference 75

Resolution
unresolved
no resolver link, observed 2026-08-10T16:52:49.805557Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T16:52:49.805557Z digest=sha256:e9085631e38232be911e243f30537a5e6a98dde5b87aaff8d76379e18ca646f1

Observation bd943460-5dce-42b8-bb44-fe9dd660b12f · outbound

This paper cites an unresolved cited work.

NExtLong: Toward Effective Long-Context Training without Long Documents Unresolved cited work

Reference 76

Resolution
unresolved
no resolver link, observed 2026-08-10T16:52:49.808368Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T16:52:49.808368Z digest=sha256:eeda311c611809728e84a8b6e69f93f9017c2f7f84b43f0719f45796981954c5

Observation d48ddb90-ea9a-4d9a-9578-30827b4a637b · outbound

This paper cites Zero: Memory optimizations toward training trillion parameter models.

NExtLong: Toward Effective Long-Context Training without Long Documents Zero: Memory optimizations toward training trillion parameter models

Reference 77

Resolution
unresolved
no resolver link, observed 2026-08-10T16:52:49.811224Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T16:52:49.811224Z digest=sha256:3d8a85963944a49de7bde8dd7c55211e94edaf18dc24c672a367d30bee5ad998

Observation 6bedb4a9-9893-4519-ad16-5a5f5bf00d63 · outbound

This paper cites Contrastive Learning with Hard Negative Samples.

NExtLong: Toward Effective Long-Context Training without Long Documents Contrastive Learning with Hard Negative Samples

Reference 78

Resolution
unresolved
no resolver link, observed 2026-08-10T16:52:49.813895Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T16:52:49.813895Z digest=sha256:8e8ba2162dd941b53379cdd0b21a2559e77d1eb03181eb43669706b16f9bf6e3

Observation a210bede-9715-4da9-92de-6fdb2ef66552 · outbound

This paper cites Code Llama: Open Foundation Models for Code.

NExtLong: Toward Effective Long-Context Training without Long Documents Code Llama: Open Foundation Models for Code

Reference 79

Resolution
unresolved
no resolver link, observed 2026-08-10T16:52:49.817410Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T16:52:49.817410Z digest=sha256:6063759e6c925cf19c6ffc02cee699687b4502c8457d6ea9d9e6d3c0d23f91b1

Observation 7820db58-1300-4ff6-b1aa-261ef793e1d2 · outbound

This paper cites Distributionally Robust Neural Networks for Group Shifts: On the Importance of Regularization for Worst-Case Generalization.

NExtLong: Toward Effective Long-Context Training without Long Documents Distributionally Robust Neural Networks for Group Shifts: On the Importance of Regularization for Worst-Case Generalization

Reference 80

Resolution
unresolved
no resolver link, observed 2026-08-10T16:52:49.820648Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T16:52:49.820648Z digest=sha256:774f1f07d86b86c1f0a1f18f0c67e24ff65afb893af6f289b122cfd6fb8274df

Observation 6919cf99-db81-404c-a474-334b9235cffd · outbound

This paper cites L., Bhagavatula, C., and Choi, Y.

NExtLong: Toward Effective Long-Context Training without Long Documents L., Bhagavatula, C., and Choi, Y

Reference 81

Resolution
unresolved
no resolver link, observed 2026-08-10T16:52:49.824058Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T16:52:49.824058Z digest=sha256:c2fb114a98ad3440788da68cc3320834e377291ccd08c3d56c9628e87562b467

Observation e4ebea3b-00c9-43ae-a172-28c1c62a66c5 · outbound

This paper cites an unresolved cited work.

NExtLong: Toward Effective Long-Context Training without Long Documents Unresolved cited work

Reference 82

Resolution
unresolved
raw_fallback, observed 2026-08-10T16:52:50.601550Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-10T16:52:49.827211Z digest=sha256:3d7655ab5f1b393ea993cfafeb68136ed52b2f09cdad6e73c2978a56b7a67228

Observation f36bbba1-6f0c-45f2-81be-d497e46ffee8 · outbound

This paper cites H., Sch \"a rli, N., and Zhou, D.

NExtLong: Toward Effective Long-Context Training without Long Documents H., Sch \"a rli, N., and Zhou, D

Reference 83

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T16:52:50.592746Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-10T16:52:49.830080Z digest=sha256:cbc918f01d8a132d6ea446bb18843b70fdbd626f9a7113fc620f3fad5ee97ce0

Observation d8672f83-f6bf-4502-baf1-c61c0f9622df · outbound

This paper cites In-context Pretraining: Language Modeling Beyond Document Boundaries.

NExtLong: Toward Effective Long-Context Training without Long Documents In-context Pretraining: Language Modeling Beyond Document Boundaries

Reference 84

Resolution
unresolved
no resolver link, observed 2026-08-10T16:52:49.832757Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T16:52:49.832757Z digest=sha256:4aeab9ee4c30b4e323dcf5f8796bb7328e75c42cd262e4bbeb4be17dd0b0ac3f

Observation 9d00f2de-e8e2-440a-ba4c-2f944ad14301 · outbound

This paper cites R., Hestness, J., and Dey, N.

NExtLong: Toward Effective Long-Context Training without Long Documents R., Hestness, J., and Dey, N

Reference 85

Resolution
unresolved
no resolver link, observed 2026-08-10T16:52:49.835313Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T16:52:49.835313Z digest=sha256:7f9bc3e4415776176326c6dfeae4ac6504626f14272c5ddef609b200df3e6435

Observation c658b6fc-b89e-4a89-b928-bdbd76b383ac · outbound

This paper cites Dolma: an Open Corpus of Three Trillion Tokens for Language Model Pretraining Research.

NExtLong: Toward Effective Long-Context Training without Long Documents Dolma: an Open Corpus of Three Trillion Tokens for Language Model Pretraining Research

Reference 86

Resolution
unresolved
no resolver link, observed 2026-08-10T16:52:49.837695Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T16:52:49.837695Z digest=sha256:da1e6dd9280a125949834d9bd9c5323371ec8c8b3c1b6158fd01c139eb5bf6ff

Observation 95e98ca3-b90c-4735-919f-e68b48da28cd · outbound

This paper cites Unraveling the Mystery of Scaling Laws: Part I.

NExtLong: Toward Effective Long-Context Training without Long Documents Unraveling the Mystery of Scaling Laws: Part I

Reference 87

Resolution
unresolved
no resolver link, observed 2026-08-10T16:52:49.840436Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T16:52:49.840436Z digest=sha256:5dab1147017b6a7e3799f6d9d0097a0c008223449a7c88efc52484127ce868c0

Observation 6973dcef-bb11-43f6-a1f5-da94b8ecfd97 · outbound

This paper cites Rectified rotary position embeddings.

NExtLong: Toward Effective Long-Context Training without Long Documents Rectified rotary position embeddings

Reference 88

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T16:52:50.578353Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-10T16:52:49.843463Z digest=sha256:87408b8d7a254d27102acfa8eb7c15fa25ab53fdc057297d61111c10db732b74

Observation a0576dd2-76e7-4439-a759-1d46675e5bd1 · outbound

This paper cites Roformer: Enhanced transformer with rotary position embedding, 2021.

NExtLong: Toward Effective Long-Context Training without Long Documents Roformer: Enhanced transformer with rotary position embedding, 2021

Reference 89

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T16:52:50.569475Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-10T16:52:49.846089Z digest=sha256:2a8d45c831bfadaf31e00298002d6925ab60a1699425923642697ddcda14d455

Observation 0ccd1e5f-90a2-4157-8659-4c307700c6de · outbound

This paper cites Qwen2.5: A party of foundation models, September 2024.

NExtLong: Toward Effective Long-Context Training without Long Documents Qwen2.5: A party of foundation models, September 2024

Reference 90

Resolution
unresolved
no resolver link, observed 2026-08-10T16:52:49.849185Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T16:52:49.849185Z digest=sha256:927e4951085be2cac858bec0e8fd1b581a4a650afbcb854f94f4ce6dab4221e3

Observation 70cc8a88-2e52-4715-9330-6bdf65767ab9 · outbound

This paper cites Untie the Knots: An Efficient Data Augmentation Strategy for Long-Context Pre-Training in Language Models.

NExtLong: Toward Effective Long-Context Training without Long Documents Untie the Knots: An Efficient Data Augmentation Strategy for Long-Context Pre-Training in Language Models

Reference 91

Resolution
verified exact
local_arxiv, observed 2026-08-10T16:52:50.182205Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-10T16:52:49.852046Z digest=sha256:176e2bb09c18dfae80759cddc5ebeeb2296526cc23c16d6cc8f99cfe074b3753

Observation 35e519c5-c2f7-4a48-9c4a-6a3dd00a3fad · outbound

This paper cites LLaMA: Open and Efficient Foundation Language Models.

NExtLong: Toward Effective Long-Context Training without Long Documents LLaMA: Open and Efficient Foundation Language Models

Reference 92

Resolution
unresolved
no resolver link, observed 2026-08-10T16:52:49.854942Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T16:52:49.854942Z digest=sha256:140423699f2a8e08cf0fdabd1b9f3e8a4608a9f33c15ec32e9a9a702d350a665

Observation 4617071d-0497-4297-8e80-f3535433745a · outbound

This paper cites Llama 2: Open Foundation and Fine-Tuned Chat Models.

NExtLong: Toward Effective Long-Context Training without Long Documents Llama 2: Open Foundation and Fine-Tuned Chat Models

Reference 93

Resolution
unresolved
no resolver link, observed 2026-08-10T16:52:49.858072Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T16:52:49.858072Z digest=sha256:f7b7f3696fd32baed1b419f1134e634d27a3cea2f52f6f396cba7c44e8fe07d8

Observation 2cfe92fa-969c-4db3-a504-c19cd265f158 · outbound

This paper cites Focused transformer: Contrastive training for context scaling.

NExtLong: Toward Effective Long-Context Training without Long Documents Focused transformer: Contrastive training for context scaling

Reference 94

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T16:52:50.556095Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-10T16:52:49.861284Z digest=sha256:6a867b9edb1cce883c8b76d5c168a3a38fafed1a553bbdcacf94a2bddf990575

Observation e9016387-c687-4aa8-a474-8750aec6a482 · outbound

This paper cites and Hinton, G.

NExtLong: Toward Effective Long-Context Training without Long Documents and Hinton, G

Reference 95

Resolution
unresolved
no resolver link, observed 2026-08-10T16:52:49.864269Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T16:52:49.864269Z digest=sha256:da10786ca7d79f44fc0642330df0f71fbda7b8964436b94265ff49089dfaec7a

Observation 18cdf622-fb3e-4cee-8bbe-7ce20167c6de · outbound

This paper cites A survey on large language model based autonomous agents.

NExtLong: Toward Effective Long-Context Training without Long Documents A survey on large language model based autonomous agents

Reference 96

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T16:52:50.542667Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-10T16:52:49.867012Z digest=sha256:e0063f24f689172507c3ec3db7ad2403ebdf7cfcf2bba297177dad641b56e25d

Observation 57957cc3-6561-4abb-b3a6-b72cac500fed · outbound

This paper cites Query-as-context Pre-training for Dense Passage Retrieval.

NExtLong: Toward Effective Long-Context Training without Long Documents Query-as-context Pre-training for Dense Passage Retrieval

Reference 97

Resolution
verified exact
local_arxiv, observed 2026-08-10T16:52:50.156728Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-10T16:52:49.869706Z digest=sha256:1b74cb833d39133aba39267265448853fa236d8d6eecde70b9183a397842ed71

Observation a58715eb-3b72-417d-8e33-021e29c0753c · outbound

This paper cites Contextual masked auto-encoder for dense passage retrieval.

NExtLong: Toward Effective Long-Context Training without Long Documents Contextual masked auto-encoder for dense passage retrieval

Reference 98

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T16:52:50.534405Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-10T16:52:49.872833Z digest=sha256:8cb9aa98c69be2afdf4331d3e6c303106f0bbd39fa3ec7b07ceded0fe8013aa5

Observation 5133502f-36cc-4dc9-96d0-a39b4837f685 · outbound

This paper cites Less is More for Long Document Summary Evaluation by LLMs.

NExtLong: Toward Effective Long-Context Training without Long Documents Less is More for Long Document Summary Evaluation by LLMs

Reference 99

Resolution
unresolved
no resolver link, observed 2026-08-10T16:52:49.875532Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T16:52:49.875532Z digest=sha256:0b89c83c0db9b66e89ceed92dffac28e69bbbf78549787aa0fbcd11dd45a43ea

Observation 1960fb80-c0c4-426d-abb0-9d2023308396 · outbound

This paper cites Efficient Streaming Language Models with Attention Sinks.

NExtLong: Toward Effective Long-Context Training without Long Documents Efficient Streaming Language Models with Attention Sinks

Reference 100

Resolution
unresolved
no resolver link, observed 2026-08-10T16:52:49.878420Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T16:52:49.878420Z digest=sha256:d9d9ccf712cc1cc15fbd635b50375e996cc2c4ddda4cbf56659b9783b398c26b

Observation fcfb9f50-dc73-4812-aaad-74c207584e1f · outbound

This paper cites M., Pham, H., Dong, X., Du, N., Liu, H., Lu, Y., Liang, P.

NExtLong: Toward Effective Long-Context Training without Long Documents M., Pham, H., Dong, X., Du, N., Liu, H., Lu, Y., Liang, P

Reference 101

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T16:52:50.525607Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-10T16:52:49.881585Z digest=sha256:2e4d75475273e2243f5402619497b5c1f57135eb8bb799f48118b5e0c5ee8fb3

Pith citing papers

Observation 48b4e371-f0c5-4084-92ef-69e093ce72b9 · inbound

From System 1 to System 2: A Survey of Reasoning Large Language Models cites this paper.

From System 1 to System 2: A Survey of Reasoning Large Language Models NExtLong: Toward Effective Long-Context Training without Long Documents

Reference 253

Resolution
verified exact
arxiv_id, observed 2026-05-13T01:36:24.431950Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-13T01:36:23.845366Z digest=sha256:83348880cc0e19d269c80c79bc58092a708649cb69deb5e668291e84e94c5190

Observation b3a1a969-801d-4be4-91d7-b2e5e67ff30b · inbound

MemAgent: Reshaping Long-Context LLM with Multi-Conv RL-based Memory Agent cites this paper.

MemAgent: Reshaping Long-Context LLM with Multi-Conv RL-based Memory Agent NExtLong: Toward Effective Long-Context Training without Long Documents

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-05-15T11:17:24.567598Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-15T11:17:24.406028Z digest=sha256:8868e435b0d1e22736290101c197836ee0f9ce3c115ea4a891c64249c2fda5f5

Observation b5722cbe-bd39-4a3e-ae62-3e775887f637 · inbound

MemAgent: Reshaping Long-Context LLM with Multi-Conv RL-based Memory Agent cites this paper.

MemAgent: Reshaping Long-Context LLM with Multi-Conv RL-based Memory Agent NExtLong: Toward Effective Long-Context Training without Long Documents

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-06T20:40:09.996269Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:40:09.996269Z digest=sha256:7b9582c7079e5888b92e9e27087aa996c80d24dda70ee2960a190989a986c486

Observation 136bcd82-410a-4977-9243-58df5a5f1f00 · inbound

Libra: Large Chinese-based Safeguard for AI Content cites this paper.

Libra: Large Chinese-based Safeguard for AI Content NExtLong: Toward Effective Long-Context Training without Long Documents

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-06T12:17:38.656970Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:17:38.656970Z digest=sha256:7d4ab5e21a8ba2fbb4774f70b4e331c57da7cd3cfc21ff1384b23a04afe9b31e

Observation ae8b22b1-d0f9-41cb-8a37-1b0031b7fe89 · inbound

Masked Diffusion Language Models are Strong and Steerable Text-Based World Models for Agentic RL cites this paper.

Masked Diffusion Language Models are Strong and Steerable Text-Based World Models for Agentic RL NExtLong: Toward Effective Long-Context Training without Long Documents

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-02T14:50:10.939315Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T14:50:10.939315Z digest=sha256:c26abb6329015bdc730da92adbbfd3a48b39613c4a81e4b4acdbfccafe193870