Pith. sign in

Paper Citation Record · LEDGER

AdaMTP: An Adaptive Training Paradigm for Multi-Token Prediction

As of 16 August 2026, this Paper Citation Record lists 56 of 56 outbound references and 0 inbound Pith citation observations for arXiv:2608.00434.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2608.00434 v1

Coverage vector

measured 56 of 56 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-15T15:25:35.727421Z

measured 56 of 56 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-16T06:30:59.297886+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

56 of 56 outbound references displayed

  • verified exact0
  • verified fuzzy4
  • unresolved52
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 6ca8042f-fe68-43f0-95c0-6a4cae5f0e81 · outbound

This paper cites arXiv preprint arXiv:2505.17505 , year=.

AdaMTP: An Adaptive Training Paradigm for Multi-Token Prediction arXiv preprint arXiv:2505.17505 , year=

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-15T15:25:35.488365Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T15:25:35.488365Z digest=sha256:556dffbf238da62c620ebd8511ed619baaff0aac52c6c036c562209b5b4675dd

Observation b15b3e1e-1ce3-4422-a3d1-93d64aaf5815 · outbound

This paper cites Multi-Token Prediction Needs Registers.

AdaMTP: An Adaptive Training Paradigm for Multi-Token Prediction Multi-Token Prediction Needs Registers

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-15T15:25:35.492948Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T15:25:35.492948Z digest=sha256:bff512922e5151c3159c57a834bd98dcbffa9778da494326ce84c64c18fc50ad

Observation 4af58e76-e53f-4e04-abac-96ec96e0b318 · outbound

This paper cites Faster Language Models with Better Multi-Token Prediction Using Tensor Decomposition.

AdaMTP: An Adaptive Training Paradigm for Multi-Token Prediction Faster Language Models with Better Multi-Token Prediction Using Tensor Decomposition

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-15T15:25:35.497472Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T15:25:35.497472Z digest=sha256:4028a0e21456402f1010be49f9814f6219d1b75757b3ae847fbb539f0ce826cf

Observation fddb3267-0753-401d-a3bc-791dfcf6b3f7 · outbound

This paper cites arXiv preprint arXiv:2510.14751 , year=.

AdaMTP: An Adaptive Training Paradigm for Multi-Token Prediction arXiv preprint arXiv:2510.14751 , year=

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-15T15:25:35.502284Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T15:25:35.502284Z digest=sha256:517309d7812eef5150a3eaa0ab3643a927e317f7905c25bee136c8b2881170a1

Observation ab5214b2-cd6d-42a0-b2d6-dabf286e9e62 · outbound

This paper cites Better & Faster Large Language Models via Multi-token Prediction.

AdaMTP: An Adaptive Training Paradigm for Multi-Token Prediction Better & Faster Large Language Models via Multi-token Prediction

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-15T15:25:35.506809Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T15:25:35.506809Z digest=sha256:00a434ed32c6ebb735a63c2b7e28d5b61f444542df630c9a942f929d489c5f41

Observation cb42265e-5eeb-47ad-bdfb-2e05c2109b78 · outbound

This paper cites Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) , pages=.

AdaMTP: An Adaptive Training Paradigm for Multi-Token Prediction Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) , pages=

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-15T15:25:35.511445Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T15:25:35.511445Z digest=sha256:27be268c64d4fa733391cf830d25a32ee8faf9c6dc699fb21f1a9e78ff0b6ba7

Observation ff1c6df5-af5c-4786-b8cb-c9c05f5430d5 · outbound

This paper cites DeepSeek-V3 Technical Report.

AdaMTP: An Adaptive Training Paradigm for Multi-Token Prediction DeepSeek-V3 Technical Report

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-15T15:25:35.515460Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T15:25:35.515460Z digest=sha256:81b4bd19f15a0b339b42b7792b239b6012911ea2065142f02d071862734f2067

Observation 8c78a3ba-b27a-42dc-8fae-fe74713e6b4b · outbound

This paper cites arXiv preprint arXiv:2512.24617 , year=.

AdaMTP: An Adaptive Training Paradigm for Multi-Token Prediction arXiv preprint arXiv:2512.24617 , year=

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-15T15:25:35.520379Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T15:25:35.520379Z digest=sha256:460b6294908a41cdc2b843e9ef52392492bfc03aaaa1e28d223b4c727dd949ab

Observation 59b27054-2001-4abb-bfbd-55df76adbf8b · outbound

This paper cites Efficient Joint Prediction of Multiple Future Tokens.

AdaMTP: An Adaptive Training Paradigm for Multi-Token Prediction Efficient Joint Prediction of Multiple Future Tokens

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-15T15:25:35.525235Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T15:25:35.525235Z digest=sha256:4bd6a6f5e79d7bcb240c5a5cd300576daf8fd05218666b0f6ca39773b7199801

Observation 7901a405-6c62-489d-a70e-1292971f33c1 · outbound

This paper cites arXiv preprint arXiv:2509.18362 , year=.

AdaMTP: An Adaptive Training Paradigm for Multi-Token Prediction arXiv preprint arXiv:2509.18362 , year=

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-15T15:25:35.529828Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T15:25:35.529828Z digest=sha256:97acd2cf2f9e4fbafbbaf2a3492d5793635ace769cbba90a2b79ebce3ec61a01

Observation 9e6ae26b-e742-4625-9c4d-5d6be734e0cd · outbound

This paper cites Medusa: Simple LLM Inference Acceleration Framework with Multiple Decoding Heads.

AdaMTP: An Adaptive Training Paradigm for Multi-Token Prediction Medusa: Simple LLM Inference Acceleration Framework with Multiple Decoding Heads

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-15T15:25:35.533490Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T15:25:35.533490Z digest=sha256:78aa7de255decac7f6809c5d64ed0d18f737e02f96429c1dc0cd249bce6f0a53

Observation fae43951-a2c7-4818-9d31-503092c9033b · outbound

This paper cites arXiv preprint arXiv:2508.19228 , year=.

AdaMTP: An Adaptive Training Paradigm for Multi-Token Prediction arXiv preprint arXiv:2508.19228 , year=

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-15T15:25:35.540077Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T15:25:35.540077Z digest=sha256:167a4405cfca654dc8a72f6ed7c3699a582a315f28fb2828681f4f9fe41c395b

Observation 11058976-3e5b-4c53-aeca-19444d6933a2 · outbound

This paper cites Your LLM Knows the Future: Uncovering Its Multi-Token Prediction Potential.

AdaMTP: An Adaptive Training Paradigm for Multi-Token Prediction Your LLM Knows the Future: Uncovering Its Multi-Token Prediction Potential

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-15T15:25:35.544478Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T15:25:35.544478Z digest=sha256:1e669287c92e9ae45e05c05272873e392ca3579367e249ef23d6d1d2ab28b3e1

Observation a0425c5c-cd91-4de9-9fb7-6485cfb8a0b1 · outbound

This paper cites MiMo: Unlocking the Reasoning Potential of Language Model -- From Pretraining to Posttraining.

AdaMTP: An Adaptive Training Paradigm for Multi-Token Prediction MiMo: Unlocking the Reasoning Potential of Language Model -- From Pretraining to Posttraining

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-15T15:25:35.548459Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T15:25:35.548459Z digest=sha256:591f3af40943a04d8eebb674570a7637dff68f1fb78f9097f577ae357bfd9b1d

Observation 49e34df3-9ec8-4be1-bd75-ee6fc9069bf1 · outbound

This paper cites Findings of the Association for Computational Linguistics: EMNLP 2020 , pages=.

AdaMTP: An Adaptive Training Paradigm for Multi-Token Prediction Findings of the Association for Computational Linguistics: EMNLP 2020 , pages=

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T15:25:36.836408Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-15T15:25:35.552740Z digest=sha256:859fa4555040acf242ba018568dde20d26f53ea4e892e02b037e46cb56f512c3

Observation 8305692f-b842-403e-ae79-19833b7bd8f4 · outbound

This paper cites journal of machine learning research , volume=.

AdaMTP: An Adaptive Training Paradigm for Multi-Token Prediction journal of machine learning research , volume=

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-15T15:25:35.556471Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T15:25:35.556471Z digest=sha256:901c5dcd0728643cce2b8c597831f5fa8437cb8bae869d76c9f871ab3335a4ce

Observation 97eec831-accd-4d0d-81d9-b292476d3de8 · outbound

This paper cites SqueezeLLM: Dense-and-Sparse Quantization.

AdaMTP: An Adaptive Training Paradigm for Multi-Token Prediction SqueezeLLM: Dense-and-Sparse Quantization

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-15T15:25:35.561194Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T15:25:35.561194Z digest=sha256:7a5f2ee896bdc9be4c672c2c61b6a88f34fd54d9e0042f38d3d9b4c468feecda

Observation b7a82594-e9ac-4fc4-80ea-f36e7ceb11bb · outbound

This paper cites Proceedings of machine learning and systems , volume=.

AdaMTP: An Adaptive Training Paradigm for Multi-Token Prediction Proceedings of machine learning and systems , volume=

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-15T15:25:35.565402Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T15:25:35.565402Z digest=sha256:dcf1019fa1cd7f9a553521a0379c0c1b358ca174288467568aa4624e51457874

Observation 7216b7b8-fb1a-45ea-98d3-c39d1f76900a · outbound

This paper cites Advances in neural information processing systems , volume=.

AdaMTP: An Adaptive Training Paradigm for Multi-Token Prediction Advances in neural information processing systems , volume=

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-15T15:25:35.569367Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T15:25:35.569367Z digest=sha256:6acfcfe8cb95bbfe631b5202ba5a14da065050c14c29d5e868c6134aa5ff8927

Observation cdb2b6dc-7b8a-43fd-968b-b0030e0584f8 · outbound

This paper cites Advances in Neural Information Processing Systems , volume=.

AdaMTP: An Adaptive Training Paradigm for Multi-Token Prediction Advances in Neural Information Processing Systems , volume=

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T15:25:36.805981Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-15T15:25:35.573313Z digest=sha256:eba17c4e2b089d6460965e0131d75091e246ee8d4c5f59f259ca2ff877c1f30e

Observation 9197c0e0-8b5d-4333-9d4e-92b4ba92d94f · outbound

This paper cites International conference on machine learning , pages=.

AdaMTP: An Adaptive Training Paradigm for Multi-Token Prediction International conference on machine learning , pages=

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-15T15:25:35.577111Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T15:25:35.577111Z digest=sha256:b9cadcca41e317616d946b2bbc5d642fc62665bdb99813c57c5471dee01abd9c

Observation cf1ec5e0-6419-4e95-9a30-a9d11bec1ba8 · outbound

This paper cites A Simple and Effective Pruning Approach for Large Language Models.

AdaMTP: An Adaptive Training Paradigm for Multi-Token Prediction A Simple and Effective Pruning Approach for Large Language Models

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-15T15:25:35.581263Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T15:25:35.581263Z digest=sha256:f0cc76de865509d1b4cd9c7e42537dbf727ac9b643581b341e9b4d30521d560a

Observation 0b6e21e7-5dcb-461b-ba01-d2f755056a13 · outbound

This paper cites MiniLLM: On-Policy Distillation of Large Language Models.

AdaMTP: An Adaptive Training Paradigm for Multi-Token Prediction MiniLLM: On-Policy Distillation of Large Language Models

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-15T15:25:35.584918Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T15:25:35.584918Z digest=sha256:3c8ac5ff522b2dc664ac1e073af967b92864829a6d0d281751785325831a5a9f

Observation c0b25a40-0d95-4f6e-8990-46ce33cdf93d · outbound

This paper cites Distilling the Knowledge in a Neural Network.

AdaMTP: An Adaptive Training Paradigm for Multi-Token Prediction Distilling the Knowledge in a Neural Network

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-15T15:25:35.588779Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T15:25:35.588779Z digest=sha256:60c9893c2aee70c965c1800436afa24973c9b80ed99ab2b4ff141a131edb57e6

Observation 684808a8-7cb9-42a2-99f7-92cb59bca5ca · outbound

This paper cites Instruction Tuning with GPT-4.

AdaMTP: An Adaptive Training Paradigm for Multi-Token Prediction Instruction Tuning with GPT-4

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-15T15:25:35.593350Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T15:25:35.593350Z digest=sha256:a7b187fc84bb700e6a83c791a02bed133d333132e0e998a31a45f341b6e7e270

Observation 9352eff0-d483-4feb-b47b-8d8aa54f07d0 · outbound

This paper cites Findings of the Association for Computational Linguistics: ACL 2023 , pages=.

AdaMTP: An Adaptive Training Paradigm for Multi-Token Prediction Findings of the Association for Computational Linguistics: ACL 2023 , pages=

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-15T15:25:35.597216Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T15:25:35.597216Z digest=sha256:f76baee0e1f31a52bc26764cc396002654e558637278e57582b707f3dd7abdd8

Observation 47e84ab0-6e18-4f92-9d22-cee4dac4cd70 · outbound

This paper cites Proceedings of the 61st annual meeting of the association for computational linguistics (volume 1: long papers) , pages=.

AdaMTP: An Adaptive Training Paradigm for Multi-Token Prediction Proceedings of the 61st annual meeting of the association for computational linguistics (volume 1: long papers) , pages=

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-15T15:25:35.601190Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T15:25:35.601190Z digest=sha256:5d5d80b5102e308dc4963e5ba4d8efc1fa96bd5778706b41be90e09d3f360dbc

Observation 2754cfba-0c70-4c6b-b757-e3c30c72157f · outbound

This paper cites International conference on machine learning , pages=.

AdaMTP: An Adaptive Training Paradigm for Multi-Token Prediction International conference on machine learning , pages=

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-15T15:25:35.604775Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T15:25:35.604775Z digest=sha256:ae19c0751f72a053840a3d476d5d626992c4477212a92b5787f099960740635b

Observation 68b7cbdc-223d-4165-9389-3710a91fcd55 · outbound

This paper cites Advances in neural information processing systems , volume=.

AdaMTP: An Adaptive Training Paradigm for Multi-Token Prediction Advances in neural information processing systems , volume=

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-15T15:25:35.608730Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T15:25:35.608730Z digest=sha256:e9e8b8ae7fbcd3a51f21c5a810c412fce94588a9836ba7378e2f52fefb8d7ab3

Observation 81b89411-4885-4e9d-b25b-3f9bfe3fb54e · outbound

This paper cites Generating Long Sequences with Sparse Transformers.

AdaMTP: An Adaptive Training Paradigm for Multi-Token Prediction Generating Long Sequences with Sparse Transformers

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-15T15:25:35.613055Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T15:25:35.613055Z digest=sha256:ff26c914d9e67fbf8d5f67a2a47113da82adadee95805d65719203e8ad949e1b

Observation 07f1da7e-6c0b-4e28-a59a-6fc269ebb59f · outbound

This paper cites MoBA: Mixture of Block Attention for Long-Context LLMs.

AdaMTP: An Adaptive Training Paradigm for Multi-Token Prediction MoBA: Mixture of Block Attention for Long-Context LLMs

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-15T15:25:35.617482Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T15:25:35.617482Z digest=sha256:d95d79e1a2645d41cc9710b5fd39864325000c783f713722d5f18fd3f359ffd5

Observation 4341a82d-73e1-40df-a8f2-61b47b70508f · outbound

This paper cites DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model.

AdaMTP: An Adaptive Training Paradigm for Multi-Token Prediction DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-15T15:25:35.621851Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T15:25:35.621851Z digest=sha256:2b67bf9ad8e7de2b28cd445f5896e6a716f988c2999cea5933d8a248b542e229

Observation 4578af6f-bcc6-41b3-8bf3-a61977e00d1f · outbound

This paper cites Advances in neural information processing systems , volume=.

AdaMTP: An Adaptive Training Paradigm for Multi-Token Prediction Advances in neural information processing systems , volume=

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-15T15:25:35.626115Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T15:25:35.626115Z digest=sha256:e962f2fe2f1421a6a6b317ee9af344220a2c5c2496c93871955341399318bd9c

Observation 93a6d0a6-88c6-4d2c-a717-34668cd039cd · outbound

This paper cites FlashAttention-2: Faster Attention with Better Parallelism and Work Partitioning.

AdaMTP: An Adaptive Training Paradigm for Multi-Token Prediction FlashAttention-2: Faster Attention with Better Parallelism and Work Partitioning

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-15T15:25:35.629844Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T15:25:35.629844Z digest=sha256:cb688f0ebd367791452289fe40d22a3042c4d501b88afcf715af17ada4c2cf9c

Observation c4b3350b-d095-4e51-a2ee-e117f5dcb70f · outbound

This paper cites Proceedings of the 29th symposium on operating systems principles , pages=.

AdaMTP: An Adaptive Training Paradigm for Multi-Token Prediction Proceedings of the 29th symposium on operating systems principles , pages=

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-15T15:25:35.634067Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T15:25:35.634067Z digest=sha256:32e50ba36e017acb63c99872b642ecfcfb0fe33e6625294d9fba6ce88618b9a7

Observation 7f4fc9a8-475d-48af-bd96-b1dbe02f3a2f · outbound

This paper cites Advances in neural information processing systems , volume=.

AdaMTP: An Adaptive Training Paradigm for Multi-Token Prediction Advances in neural information processing systems , volume=

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-15T15:25:35.638626Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T15:25:35.638626Z digest=sha256:8e7ef80666204813291ecda84eb60e5191b9d5d765aa3498cac03e9a794ab299

Observation 417c9b87-6b3d-4572-9a70-c8128b0923aa · outbound

This paper cites International Conference on Machine Learning , pages=.

AdaMTP: An Adaptive Training Paradigm for Multi-Token Prediction International Conference on Machine Learning , pages=

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-15T15:25:35.642428Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T15:25:35.642428Z digest=sha256:b3f5bc91205e5e90eef8e7d2429539cbbb71acc647ab86bc9f4842e39f7b6330

Observation 9308dce7-57e6-4084-bbbc-044c91fdb25e · outbound

This paper cites Accelerating Large Language Model Decoding with Speculative Sampling.

AdaMTP: An Adaptive Training Paradigm for Multi-Token Prediction Accelerating Large Language Model Decoding with Speculative Sampling

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-15T15:25:35.646913Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T15:25:35.646913Z digest=sha256:4b25c39ba0052f4939c7d9f4d83d26d8d2c4c964227eaf7ce495f4d2d8623973

Observation affac131-8332-409d-85cf-12b8e3c40c94 · outbound

This paper cites CCF International Conference on Natural Language Processing and Chinese Computing , pages=.

AdaMTP: An Adaptive Training Paradigm for Multi-Token Prediction CCF International Conference on Natural Language Processing and Chinese Computing , pages=

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T15:25:36.733156Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-15T15:25:35.651547Z digest=sha256:31d3e951d4de7fba08461ca934c65db63555e918847f2ac8519889c464d874d1

Observation 0d9a0d3b-0c2d-4d8f-bcde-c13d645c6218 · outbound

This paper cites DistillSpec: Improving Speculative Decoding via Knowledge Distillation.

AdaMTP: An Adaptive Training Paradigm for Multi-Token Prediction DistillSpec: Improving Speculative Decoding via Knowledge Distillation

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-15T15:25:35.655628Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T15:25:35.655628Z digest=sha256:0b1ba5e4ef15180a91ed5d945e388249c118991563df8f9a34bcaf07ba4451f3

Observation eacf1cdc-5280-41d9-8bb3-33c2eec32f08 · outbound

This paper cites Advances in Neural Information Processing Systems , volume=.

AdaMTP: An Adaptive Training Paradigm for Multi-Token Prediction Advances in Neural Information Processing Systems , volume=

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-15T15:25:35.660295Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T15:25:35.660295Z digest=sha256:c9e88949910fc428a4a1ace6e583e451283bcdc6a50a5f3070946c0579ea0e5f

Observation c1e56746-f9a9-4d69-baeb-e4335031de2b · outbound

This paper cites EAGLE: Speculative Sampling Requires Rethinking Feature Uncertainty.

AdaMTP: An Adaptive Training Paradigm for Multi-Token Prediction EAGLE: Speculative Sampling Requires Rethinking Feature Uncertainty

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-15T15:25:35.664561Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T15:25:35.664561Z digest=sha256:df4e5839af2bec70ea37486043248a4add26f4be02afd8b3c175fd434d73b671

Observation 05d3129c-804f-49f8-a138-9d38188f3620 · outbound

This paper cites Proceedings of the 29th ACM International Conference on Architectural Support for Programming Languages and Operating Systems, Volume 3 , pages=.

AdaMTP: An Adaptive Training Paradigm for Multi-Token Prediction Proceedings of the 29th ACM International Conference on Architectural Support for Programming Languages and Operating Systems, Volume 3 , pages=

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-15T15:25:35.668809Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T15:25:35.668809Z digest=sha256:993f427df3d7f0c28bfb4a69023d010b1d49312b036322c98bf4b8c4167a56c1

Observation 09811252-cabc-45a1-9504-239f8a8f959f · outbound

This paper cites Advances in Neural Information Processing Systems , volume=.

AdaMTP: An Adaptive Training Paradigm for Multi-Token Prediction Advances in Neural Information Processing Systems , volume=

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T15:25:36.706553Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-15T15:25:35.672899Z digest=sha256:0655a52697667b2b864b16e299be77216281ea78ff02d0fbf64216bf3a7233ea

Observation 55ee1ba7-a582-4b59-8970-58ca9abebb03 · outbound

This paper cites Large Concept Models: Language Modeling in a Sentence Representation Space.

AdaMTP: An Adaptive Training Paradigm for Multi-Token Prediction Large Concept Models: Language Modeling in a Sentence Representation Space

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-15T15:25:35.677765Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T15:25:35.677765Z digest=sha256:0e16483d4ad262dd67828ff9a2dc8e9d43778462482bc122f3c0aabc361e8ccc

Observation c365b461-3df8-4e58-8af3-f790fd444351 · outbound

This paper cites Measuring Mathematical Problem Solving With the MATH Dataset.

AdaMTP: An Adaptive Training Paradigm for Multi-Token Prediction Measuring Mathematical Problem Solving With the MATH Dataset

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-15T15:25:35.683168Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T15:25:35.683168Z digest=sha256:c673b92ba297c42132aeaec86e9326d8e858eb0e20caba5bd6a05cd05ffc1f31

Observation 96ab64fb-ae7a-442a-ab22-c86884bbc001 · outbound

This paper cites WizardCoder: Empowering Code Large Language Models with Evol-Instruct.

AdaMTP: An Adaptive Training Paradigm for Multi-Token Prediction WizardCoder: Empowering Code Large Language Models with Evol-Instruct

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-15T15:25:35.687143Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T15:25:35.687143Z digest=sha256:aff713c080647100cc73e9c4ebf09a0578f250e1621c90ceeae817ef6a8c3e11

Observation dab79e4b-f641-400a-8fc0-638e4099e2a3 · outbound

This paper cites an unresolved cited work.

AdaMTP: An Adaptive Training Paradigm for Multi-Token Prediction Unresolved cited work

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-15T15:25:35.691476Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T15:25:35.691476Z digest=sha256:9dbc4f05f4e706073eb3991f32eac01b5e4cde27831b71eed7cbff5f76984fb2

Observation eabc3651-6d29-4a61-abc5-e0370d1727db · outbound

This paper cites The twelfth international conference on learning representations , year=.

AdaMTP: An Adaptive Training Paradigm for Multi-Token Prediction The twelfth international conference on learning representations , year=

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-15T15:25:35.696223Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T15:25:35.696223Z digest=sha256:8943f54574bfb756dcfb32d35f1f0242747b8568d8be5745573269a9edc494d8

Observation b11c1ef7-37a2-4809-bf71-58caf832463d · outbound

This paper cites Training Verifiers to Solve Math Word Problems.

AdaMTP: An Adaptive Training Paradigm for Multi-Token Prediction Training Verifiers to Solve Math Word Problems

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-15T15:25:35.700300Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T15:25:35.700300Z digest=sha256:64b38d9d36aea2e8a17a81f3ed0decd2b386c43c5bdb172694c043f89d39b104

Observation 8c450dd1-3485-47e4-9749-6d5e12d3efb1 · outbound

This paper cites Program Synthesis with Large Language Models.

AdaMTP: An Adaptive Training Paradigm for Multi-Token Prediction Program Synthesis with Large Language Models

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-15T15:25:35.704273Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T15:25:35.704273Z digest=sha256:52b1d40bbf91b920709bac6329f66b6a6d5a7487d4763737b18debd726827cba

Observation 56d7718e-3358-42a4-ad22-98e782a646cf · outbound

This paper cites Advances in neural information processing systems , volume=.

AdaMTP: An Adaptive Training Paradigm for Multi-Token Prediction Advances in neural information processing systems , volume=

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-15T15:25:35.708837Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T15:25:35.708837Z digest=sha256:7fc421103c4e20a4d08760d2961f502c2f22f165d321cec78236e1df006cb9ec

Observation 479bba2b-2309-4c94-acf7-449e546e8d59 · outbound

This paper cites Evaluating Large Language Models Trained on Code.

AdaMTP: An Adaptive Training Paradigm for Multi-Token Prediction Evaluating Large Language Models Trained on Code

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-15T15:25:35.713859Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T15:25:35.713859Z digest=sha256:48c3b0279cab02a5c1e3e45c1b440b639ca621112d41e1ec57ea2c887f81c400

Observation b2c7d324-7caa-48f1-aa34-ad9b4fa3b362 · outbound

This paper cites Measuring Massive Multitask Language Understanding.

AdaMTP: An Adaptive Training Paradigm for Multi-Token Prediction Measuring Massive Multitask Language Understanding

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-15T15:25:35.718483Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T15:25:35.718483Z digest=sha256:28dc9d82152646f07672dd3509ba804030fa7968a229dbb40b422806b6369225

Observation 66c737ae-6767-43ff-986b-b459f46ab715 · outbound

This paper cites Instruction-Following Evaluation for Large Language Models.

AdaMTP: An Adaptive Training Paradigm for Multi-Token Prediction Instruction-Following Evaluation for Large Language Models

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-15T15:25:35.722864Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T15:25:35.722864Z digest=sha256:2b5c2e91b20a1fd39c0bee5c9dfcec35c50e9aa44892b16b4d9efb2098b19e27

Observation 172356cb-0b9d-4c88-af09-8d1cb59fc6b5 · outbound

This paper cites , author=.

AdaMTP: An Adaptive Training Paradigm for Multi-Token Prediction , author=

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-15T15:25:35.727421Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T15:25:35.727421Z digest=sha256:93c77daa3067708a91101bc6b51f56cb85ef5d206cf01130072505f8939afcae

Pith citing papers

No inbound Pith citation observations are available.