Pith. sign in

Paper Citation Record · LEDGER

Next Token Prediction Towards Multimodal Intelligence: A Comprehensive Survey

As of 20 August 2026, this Paper Citation Record lists 100 of 292 outbound references and 19 inbound Pith citation observations for arXiv:2412.18619.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2412.18619 v2

Coverage vector

measured 100 of 292 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-11T14:59:01.581000Z

measured 119 of 119 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-20T06:33:59.587034+00:00

measured 19 of 19 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-15T22:21:30.258422Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T00:39:16.234771Z

Reference resolution

100 of 292 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved100
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation fe62bfa2-1b26-40e0-abd8-3c868a84bda5 · outbound

This paper cites Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone.

Next Token Prediction Towards Multimodal Intelligence: A Comprehensive Survey Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-11T14:59:01.080247Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T14:59:01.080247Z digest=sha256:380702d89b0f04d36012e95a1e0f4870055a9b3d5c7eee4a2adb6cde6c3a14b2

Observation b4cc7ec8-41c0-4af7-9a63-c219aa932e09 · outbound

This paper cites Scaling Laws for Generative Mixed-Modal Language Models.

Next Token Prediction Towards Multimodal Intelligence: A Comprehensive Survey Scaling Laws for Generative Mixed-Modal Language Models

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-11T14:59:01.086427Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T14:59:01.086427Z digest=sha256:f28f9652b46f61305f5dc6d23a2924afbafd903df63ee633bb85542cf3d5ad0c

Observation 1a3ce445-1301-4bf8-9187-1b1097c61834 · outbound

This paper cites an unresolved cited work.

Next Token Prediction Towards Multimodal Intelligence: A Comprehensive Survey Unresolved cited work

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-11T14:59:01.091260Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T14:59:01.091260Z digest=sha256:50eca3852d607c28c13b5cd62638658b818b3fd5ee0b4a6001fa3ae9393ee537

Observation 0c1fa7b2-a107-4df9-9f86-44bf8b70492c · outbound

This paper cites Seed-TTS: A Family of High-Quality Versatile Speech Generation Models.

Next Token Prediction Towards Multimodal Intelligence: A Comprehensive Survey Seed-TTS: A Family of High-Quality Versatile Speech Generation Models

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-11T14:59:01.095616Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T14:59:01.095616Z digest=sha256:1d3bcc404a75edec98746a59c8fc076ae6657931d5ccc99a58dfad9e9d68e088

Observation 03472a17-98a7-49ce-a5fd-cf900033a055 · outbound

This paper cites an unresolved cited work.

Next Token Prediction Towards Multimodal Intelligence: A Comprehensive Survey Unresolved cited work

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-11T14:59:01.101268Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T14:59:01.101268Z digest=sha256:f44a3c2174613178c67af2226c8843b665dff9495e0f6b828c4182260002f8e2

Observation 9dcf01d1-605b-4013-9a10-6db96d4dc130 · outbound

This paper cites ViViT: A Video Vision Transformer.

Next Token Prediction Towards Multimodal Intelligence: A Comprehensive Survey ViViT: A Video Vision Transformer

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-11T14:59:01.107233Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T14:59:01.107233Z digest=sha256:08e15779660a00560601431a05afc126604d8fd5744ab0dbaecdd3e5cf828ac0

Observation 263cd2fc-43bb-45aa-b9ec-e86db136fdcc · outbound

This paper cites OpenFlamingo: An Open-Source Framework for Training Large Autoregressive Vision-Language Models.

Next Token Prediction Towards Multimodal Intelligence: A Comprehensive Survey OpenFlamingo: An Open-Source Framework for Training Large Autoregressive Vision-Language Models

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-11T14:59:01.114006Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T14:59:01.114006Z digest=sha256:2d325f768c91415d942eccfead24219e7c92d4ae6e5f905f68b0ac6df06168d5

Observation 6459b60f-8c96-4e45-85d3-284d2d50662c · outbound

This paper cites Foundational Models Defining a New Era in Vision: A Survey and Outlook.

Next Token Prediction Towards Multimodal Intelligence: A Comprehensive Survey Foundational Models Defining a New Era in Vision: A Survey and Outlook

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-11T14:59:01.119847Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T14:59:01.119847Z digest=sha256:cf3f4d319f3c8a66099968965e50cd5ea9435b69544b87f8ac4eaf173ba15f8c

Observation c330543f-40bb-4fc7-9601-a6ce42b4e49f · outbound

This paper cites data2vec: A General Framework for Self-supervised Learning in Speech, Vision and Language.

Next Token Prediction Towards Multimodal Intelligence: A Comprehensive Survey data2vec: A General Framework for Self-supervised Learning in Speech, Vision and Language

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-11T14:59:01.125368Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T14:59:01.125368Z digest=sha256:37d62bf59a0eb823fd2fed72234453bb286037f28108fca81f4b4a494e509c7b

Observation 3eafa8b5-a3f6-48ba-b8ff-647d49e61366 · outbound

This paper cites vq-wav2vec: Self-Supervised Learning of Discrete Speech Representations.

Next Token Prediction Towards Multimodal Intelligence: A Comprehensive Survey vq-wav2vec: Self-Supervised Learning of Discrete Speech Representations

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-11T14:59:01.130923Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T14:59:01.130923Z digest=sha256:dd39a3d4c85dc3e9d09a74b6ca9d4931e50b0008adf999745a6674950b163083

Observation 1e764e5a-fd2e-46fe-926f-63b6e8289d0f · outbound

This paper cites an unresolved cited work.

Next Token Prediction Towards Multimodal Intelligence: A Comprehensive Survey Unresolved cited work

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-11T14:59:01.136222Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T14:59:01.136222Z digest=sha256:4316ce2e07a577675e8e420265df7000e7039d04131813e4bb05679d0fcfa1c1

Observation 373672a8-a805-4518-af54-30981f20cb28 · outbound

This paper cites an unresolved cited work.

Next Token Prediction Towards Multimodal Intelligence: A Comprehensive Survey Unresolved cited work

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-11T14:59:01.141710Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T14:59:01.141710Z digest=sha256:ed552ac41b727a5b49e34eebd46d34387f42ce9cc6996bcbfe874ac1d6987c58

Observation 7fb01719-74f9-4f02-84c7-772b48d6c044 · outbound

This paper cites Sequential Modeling Enables Scalable Learning for Large Vision Models.

Next Token Prediction Towards Multimodal Intelligence: A Comprehensive Survey Sequential Modeling Enables Scalable Learning for Large Vision Models

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-11T14:59:01.152569Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T14:59:01.152569Z digest=sha256:3dd525416c1a07d778fbdaf58c063e8fc603abbd94c621fcaeceea7ba0db24ff

Observation d138caa8-b876-4ec4-aba4-6e6dd832f8c4 · outbound

This paper cites an unresolved cited work.

Next Token Prediction Towards Multimodal Intelligence: A Comprehensive Survey Unresolved cited work

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-11T14:59:01.157928Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T14:59:01.157928Z digest=sha256:8dab85fa661eaa52369db2297a955e384f664f833c6916088e891da337d6d935

Observation 2bb294c4-e5f9-40d2-aeaa-d7633b9eb53b · outbound

This paper cites an unresolved cited work.

Next Token Prediction Towards Multimodal Intelligence: A Comprehensive Survey Unresolved cited work

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-11T14:59:01.162518Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T14:59:01.162518Z digest=sha256:2e4f03b94c55e7b3cf8b67980c55b7c696892263212713e4abf73327069bd629

Observation a102b79e-dabd-4371-b89e-226e17890273 · outbound

This paper cites an unresolved cited work.

Next Token Prediction Towards Multimodal Intelligence: A Comprehensive Survey Unresolved cited work

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-11T14:59:01.166219Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T14:59:01.166219Z digest=sha256:aeca129809c3e3847259bb82923cf5d4e42a9a54a6768df5bb53095ebec1a148

Observation 943587c3-60b6-42c5-a1ba-6568e0f3d3ef · outbound

This paper cites BEiT: BERT Pre-Training of Image Transformers.

Next Token Prediction Towards Multimodal Intelligence: A Comprehensive Survey BEiT: BERT Pre-Training of Image Transformers

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-11T14:59:01.170069Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T14:59:01.170069Z digest=sha256:3952da9159350014244badbab66b84e902b1b026f74acc7d2be381e1763bc9c6

Observation b1630485-7c8d-4c5f-ad1a-a693d67b0385 · outbound

This paper cites an unresolved cited work.

Next Token Prediction Towards Multimodal Intelligence: A Comprehensive Survey Unresolved cited work

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-11T14:59:01.174096Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T14:59:01.174096Z digest=sha256:089e56900709ffb02ec3e045c63589ee7e2dde037ec30d4679f1c5d054245143

Observation fcf02090-06d2-4758-a4da-98d8d0ecf100 · outbound

This paper cites EdVAE: Mitigating Codebook Collapse with Evidential Discrete Variational Autoencoders.

Next Token Prediction Towards Multimodal Intelligence: A Comprehensive Survey EdVAE: Mitigating Codebook Collapse with Evidential Discrete Variational Autoencoders

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-11T14:59:01.182588Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T14:59:01.182588Z digest=sha256:c0bbd54435c7a17c6f73fde284fab8e37a565135257a085396793da29ab9c865

Observation 4220cdff-e448-4ffb-bfbc-54a8176f4319 · outbound

This paper cites https://www.adept.ai/blog/fuyu-8b.

Next Token Prediction Towards Multimodal Intelligence: A Comprehensive Survey https://www.adept.ai/blog/fuyu-8b

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-11T14:59:01.178191Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T14:59:01.178191Z digest=sha256:45241d6b95e2598245abf41b80a359d1695a89966abd40733a96364b6b870608

Observation 994f7e8e-e4be-401e-ac5e-107a141be585 · outbound

This paper cites Estimating or Propagating Gradients Through Stochastic Neurons for Conditional Computation.

Next Token Prediction Towards Multimodal Intelligence: A Comprehensive Survey Estimating or Propagating Gradients Through Stochastic Neurons for Conditional Computation

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-11T14:59:01.192288Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T14:59:01.192288Z digest=sha256:22f779f2fb1d13829195882fb73036ca265a3122d6a3d7f48c71fd0082233c91

Observation 695b1b87-6d61-4298-b09f-98076e0bb3b5 · outbound

This paper cites Genomic Language Models: Opportunities and Challenges.

Next Token Prediction Towards Multimodal Intelligence: A Comprehensive Survey Genomic Language Models: Opportunities and Challenges

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-11T14:59:01.186593Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T14:59:01.186593Z digest=sha256:400ac8bb68c3bd29a42b1550c0cc53eee7d02a4d020fadf99da9d7addd589d09

Observation e2c93187-8fd5-4da0-9876-046d72576dae · outbound

This paper cites an unresolved cited work.

Next Token Prediction Towards Multimodal Intelligence: A Comprehensive Survey Unresolved cited work

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-11T14:59:01.202000Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T14:59:01.202000Z digest=sha256:f37aee92538ab3a50fb386de35358243693e78f860fbc606421e4da4ef238f19

Observation d399fe51-903e-4251-9345-031d01894d7e · outbound

This paper cites an unresolved cited work.

Next Token Prediction Towards Multimodal Intelligence: A Comprehensive Survey Unresolved cited work

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-11T14:59:01.197269Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T14:59:01.197269Z digest=sha256:66ad690ad78163c162481ce1f73bbc33d8c52ff83caeb4bd5e578afe2d8e2263

Observation 1c498f43-f05c-4559-93bd-fb15597138c6 · outbound

This paper cites an unresolved cited work.

Next Token Prediction Towards Multimodal Intelligence: A Comprehensive Survey Unresolved cited work

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-11T14:59:01.211479Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T14:59:01.211479Z digest=sha256:f84e085eaebd5eaa864b12464da799c7c3ad588d48ae9e22837b3bd5ec14ef44

Observation 5bb286b1-8954-429f-94cf-c25a58142809 · outbound

This paper cites FlexiViT: One Model for All Patch Sizes.

Next Token Prediction Towards Multimodal Intelligence: A Comprehensive Survey FlexiViT: One Model for All Patch Sizes

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-11T14:59:01.206564Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T14:59:01.206564Z digest=sha256:e8182d261f69dd21ba22522c5a668db9311724c35a7a9b59a8ee70dcb1c2b736

Observation 22e5f17f-4e8f-4e64-ac9d-45df9075c80d · outbound

This paper cites An Introduction to Vision-Language Modeling.

Next Token Prediction Towards Multimodal Intelligence: A Comprehensive Survey An Introduction to Vision-Language Modeling

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-11T14:59:01.220651Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T14:59:01.220651Z digest=sha256:30c02c5c8501014b9735bc9356831214e63e15c799b5c13f3de7f86d742329f3

Observation 92bf95ea-9dfa-4a06-9907-c1f8e28c1746 · outbound

This paper cites an unresolved cited work.

Next Token Prediction Towards Multimodal Intelligence: A Comprehensive Survey Unresolved cited work

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-11T14:59:01.215920Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T14:59:01.215920Z digest=sha256:5734b10b2350b13b81a35855c03daa220a7663711b3c100638f6578245971698

Observation 83b81543-e664-4409-bf0e-23b8fc776b8b · outbound

This paper cites an unresolved cited work.

Next Token Prediction Towards Multimodal Intelligence: A Comprehensive Survey Unresolved cited work

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-11T14:59:01.231103Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T14:59:01.231103Z digest=sha256:e4941075dbac2a1415bfb732671ca732fd8940738bbf9c9bcab51a359c1747ab

Observation 5253f172-e446-446c-97ef-5d69979731d1 · outbound

This paper cites an unresolved cited work.

Next Token Prediction Towards Multimodal Intelligence: A Comprehensive Survey Unresolved cited work

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-11T14:59:01.225979Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T14:59:01.225979Z digest=sha256:01b530a58c4f47d1cb5e53404588c46ec6c313ee51bed26511633654d329ab31

Observation efc3ee0d-2ed7-46d2-96d5-26f980be3a43 · outbound

This paper cites High-Performance Large-Scale Image Recognition Without Normalization.

Next Token Prediction Towards Multimodal Intelligence: A Comprehensive Survey High-Performance Large-Scale Image Recognition Without Normalization

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-11T14:59:01.240944Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T14:59:01.240944Z digest=sha256:4cce6fabda24282c7f1a7cd8f29afd44ffa1c8281d00570a3f9207d7f8aa87c4

Observation baa19661-bcab-4874-bfba-0f709cdfda9e · outbound

This paper cites Smith, and K.

Next Token Prediction Towards Multimodal Intelligence: A Comprehensive Survey Smith, and K

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-11T14:59:01.236113Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T14:59:01.236113Z digest=sha256:df653ddfa75c323950de7f3ea20caa767838d7868f4f99cfbfa2b8c6d73262a6

Observation bf7f4559-310e-4756-8883-351bd9fd997e · outbound

This paper cites an unresolved cited work.

Next Token Prediction Towards Multimodal Intelligence: A Comprehensive Survey Unresolved cited work

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-11T14:59:01.250696Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T14:59:01.250696Z digest=sha256:081992db68e35bf78e7242db1b834c81b77d91730fca6baa368c02c7c2ae4514

Observation ce96daae-8f69-4ae3-b931-b06f74a29a5f · outbound

This paper cites RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control.

Next Token Prediction Towards Multimodal Intelligence: A Comprehensive Survey RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-11T14:59:01.245779Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T14:59:01.245779Z digest=sha256:ab99023b5b206aa180015c6bb803841888f01407f7fe6f3f89a5d44e567dde5f

Observation 4d870116-35a3-429c-bc1f-8198808588bb · outbound

This paper cites an unresolved cited work.

Next Token Prediction Towards Multimodal Intelligence: A Comprehensive Survey Unresolved cited work

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-11T14:59:01.260501Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T14:59:01.260501Z digest=sha256:4206d38a37a629b5927aa1b8b89490d84cd061d3b2f88d30bf90f4eb17a83d42

Observation 3b539719-cc6b-4936-a75f-c2b827ad5d1f · outbound

This paper cites an unresolved cited work.

Next Token Prediction Towards Multimodal Intelligence: A Comprehensive Survey Unresolved cited work

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-11T14:59:01.255735Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T14:59:01.255735Z digest=sha256:85788793783542f32039e1b6f1272e92ed290ab2c981b22863276113062b20ad

Observation bf20845d-9085-4825-937d-cee5b6560c3c · outbound

This paper cites BenchLMM: Benchmarking Cross-style Visual Capability of Large Multimodal Models.

Next Token Prediction Towards Multimodal Intelligence: A Comprehensive Survey BenchLMM: Benchmarking Cross-style Visual Capability of Large Multimodal Models

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-11T14:59:01.270436Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T14:59:01.270436Z digest=sha256:55c97e22274419e4a15a74f713607cd95a4caa40b1f539060bc039f070e9e56f

Observation 42ea5248-9941-47dd-b393-ab422385372a · outbound

This paper cites an unresolved cited work.

Next Token Prediction Towards Multimodal Intelligence: A Comprehensive Survey Unresolved cited work

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-11T14:59:01.265306Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T14:59:01.265306Z digest=sha256:c8714476fa7b5dc193eb6d72c0d4ca742b852d98d3a1334d5477c18ee60fa981

Observation de3d030a-a87c-4f1b-a65e-2bf53078ce43 · outbound

This paper cites Efficient-VQGAN: Towards High-Resolution Image Generation with Efficient Vision Transformers.

Next Token Prediction Towards Multimodal Intelligence: A Comprehensive Survey Efficient-VQGAN: Towards High-Resolution Image Generation with Efficient Vision Transformers

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-11T14:59:01.280625Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T14:59:01.280625Z digest=sha256:7dfa3c4a83b1a4bdf04d05f1fc01e5e8efc7096718ffc894bec8c7c4f1745836

Observation a0b8e8c7-7211-42d6-a8fd-4e467016079d · outbound

This paper cites Swin-Unet: Unet-like Pure Transformer for Medical Image Segmentation.

Next Token Prediction Towards Multimodal Intelligence: A Comprehensive Survey Swin-Unet: Unet-like Pure Transformer for Medical Image Segmentation

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-11T14:59:01.275746Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T14:59:01.275746Z digest=sha256:917b063175eccdd6c268b8ec97d67c31b863fd3cefcc60b51d36380c3e90f95d

Observation f9279ae2-6997-4be2-97b8-0108a3bb73af · outbound

This paper cites HiP: Hierarchical Perceiver.

Next Token Prediction Towards Multimodal Intelligence: A Comprehensive Survey HiP: Hierarchical Perceiver

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-11T14:59:01.288757Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T14:59:01.288757Z digest=sha256:4826ce5ef15bae02cd378c26d103f95dabeec30a203914eb878a6c951c034e97

Observation c197ba0d-30aa-483c-8317-a26c04dff6b4 · outbound

This paper cites Emerging Properties in Self-Supervised Vision Transformers.

Next Token Prediction Towards Multimodal Intelligence: A Comprehensive Survey Emerging Properties in Self-Supervised Vision Transformers

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-11T14:59:01.284652Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T14:59:01.284652Z digest=sha256:cb56134069efa3fcad5ca7ae83b416e4450573189faa2ae7cedb2e046ee2e5fb

Observation 4e624e75-ff56-41b5-a1f9-609262452040 · outbound

This paper cites Muse: Text-To-Image Generation via Masked Generative Transformers.

Next Token Prediction Towards Multimodal Intelligence: A Comprehensive Survey Muse: Text-To-Image Generation via Masked Generative Transformers

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-11T14:59:01.298092Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T14:59:01.298092Z digest=sha256:024b6116392d5c923eebaeb230805d3cfca3b876a28f840ec2ab4242b954bf9b

Observation 9b2593d1-89a4-43a1-9aeb-8102f872e5da · outbound

This paper cites Visually Dehallucinative Instruction Generation: Know What You Don't Know.

Next Token Prediction Towards Multimodal Intelligence: A Comprehensive Survey Visually Dehallucinative Instruction Generation: Know What You Don't Know

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-11T14:59:01.292865Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T14:59:01.292865Z digest=sha256:eec38d332b2fc6dff34ee9fd255ba25b9990a2066db2d21840ab3e59eace0458

Observation cfd7f0a2-aeae-41e3-b9a7-dcbcdce6b204 · outbound

This paper cites an unresolved cited work.

Next Token Prediction Towards Multimodal Intelligence: A Comprehensive Survey Unresolved cited work

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-11T14:59:01.307983Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T14:59:01.307983Z digest=sha256:47dee6027d28bed328f041f02e80b40aaee4652c366a86aaa7777a9b09783132

Observation 00c77c29-e841-4869-aaf9-5b2ca1facf7c · outbound

This paper cites an unresolved cited work.

Next Token Prediction Towards Multimodal Intelligence: A Comprehensive Survey Unresolved cited work

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-11T14:59:01.302942Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T14:59:01.302942Z digest=sha256:110e4ce1e91a52d871eb53baa0c76ae09ad53204ffe96c49344ad1820735cabe

Observation 3571d774-5637-4e09-9f45-b0f0e67b5213 · outbound

This paper cites X-LLM: Bootstrapping Advanced Large Language Models by Treating Multi-Modalities as Foreign Languages.

Next Token Prediction Towards Multimodal Intelligence: A Comprehensive Survey X-LLM: Bootstrapping Advanced Large Language Models by Treating Multi-Modalities as Foreign Languages

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-11T14:59:01.317416Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T14:59:01.317416Z digest=sha256:c25f918df410ed6f1e48d5f964f191ce5c1f478625f5812dfa85741b873a78ad

Observation d0e2a89c-0f8e-4332-b173-843d15a89ea6 · outbound

This paper cites Visual Instruction Tuning with Polite Flamingo.

Next Token Prediction Towards Multimodal Intelligence: A Comprehensive Survey Visual Instruction Tuning with Polite Flamingo

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-11T14:59:01.312675Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T14:59:01.312675Z digest=sha256:8163f7eb4b196444dd97d407367c51fea6a2d9541bb067ff083d23be0ad577cc

Observation bd089ad7-cb8e-41bc-844b-29e6d0413fef · outbound

This paper cites an unresolved cited work.

Next Token Prediction Towards Multimodal Intelligence: A Comprehensive Survey Unresolved cited work

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-11T14:59:01.327041Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T14:59:01.327041Z digest=sha256:b1521f4c1e5ed8bf1bb4ef2fe16f7e9b39f9796644fc28874917b87df7c41872

Observation c846a0c2-7942-4035-9215-f616b19fe0a9 · outbound

This paper cites GigaSpeech: An Evolving, Multi-domain ASR Corpus with 10,000 Hours of Transcribed Audio.

Next Token Prediction Towards Multimodal Intelligence: A Comprehensive Survey GigaSpeech: An Evolving, Multi-domain ASR Corpus with 10,000 Hours of Transcribed Audio

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-11T14:59:01.322175Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T14:59:01.322175Z digest=sha256:89676170fd983c069465bc840ffc70c5c322e5e519483cadf44c69b7dbfd7511

Observation ab33ce5c-f8f6-4b53-a49d-96c561f8e8be · outbound

This paper cites A Spark of Vision-Language Intelligence: 2-Dimensional Autoregressive Transformer for Efficient Finegrained Image Generation.

Next Token Prediction Towards Multimodal Intelligence: A Comprehensive Survey A Spark of Vision-Language Intelligence: 2-Dimensional Autoregressive Transformer for Efficient Finegrained Image Generation

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-11T14:59:01.336510Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T14:59:01.336510Z digest=sha256:ae9470dcd929073ddb16b70641f7d8a716b77cefbd9fe2bd765ed7b3daa16a4f

Observation b8276a3e-45ec-4053-afcf-cc6c492fa4e0 · outbound

This paper cites ShareGPT4V: Improving Large Multi-Modal Models with Better Captions.

Next Token Prediction Towards Multimodal Intelligence: A Comprehensive Survey ShareGPT4V: Improving Large Multi-Modal Models with Better Captions

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-11T14:59:01.331481Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T14:59:01.331481Z digest=sha256:9432ad4b19ecfa36e34c2cb064817324f3b60507f4038623f8e59a910757ea3e

Observation fc3de712-dac5-4d35-9584-7f13d4cbd7dc · outbound

This paper cites PCA-Bench: Evaluating Multimodal Large Language Models in Perception-Cognition-Action Chain.

Next Token Prediction Towards Multimodal Intelligence: A Comprehensive Survey PCA-Bench: Evaluating Multimodal Large Language Models in Perception-Cognition-Action Chain

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-11T14:59:01.346015Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T14:59:01.346015Z digest=sha256:f5dea57d4a9d255eda8f2d327785ed1648f58dcd8cdd36659a715794173b93ff

Observation c38a2aaa-97f6-430c-9b2e-740de6224856 · outbound

This paper cites an unresolved cited work.

Next Token Prediction Towards Multimodal Intelligence: A Comprehensive Survey Unresolved cited work

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-11T14:59:01.341773Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T14:59:01.341773Z digest=sha256:e0e53f477e8e0b8be8e714eb7fac30d5488ecafd46f199acdc20ff34e286266b

Observation c902bd1b-f8df-4cc8-b0c7-e350b849d8d4 · outbound

This paper cites an unresolved cited work.

Next Token Prediction Towards Multimodal Intelligence: A Comprehensive Survey Unresolved cited work

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-11T14:59:01.354807Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T14:59:01.354807Z digest=sha256:f5523f2d9b0b2d408c39d5052845a8bea6195936dc4a3b3debed0e7e53442a3b

Observation f29064d1-35f5-4560-81fa-fbeecb340d41 · outbound

This paper cites An Image is Worth 1/2 Tokens After Layer 2: Plug-and-Play Inference Acceleration for Large Vision-Language Models.

Next Token Prediction Towards Multimodal Intelligence: A Comprehensive Survey An Image is Worth 1/2 Tokens After Layer 2: Plug-and-Play Inference Acceleration for Large Vision-Language Models

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-11T14:59:01.350120Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T14:59:01.350120Z digest=sha256:57b7a575999012b2af926c0edeb2e45849eb0a38279de6738519f17b537a0ec8

Observation 07018219-5482-4790-8ca2-1dce466a6770 · outbound

This paper cites BEATs: Audio Pre-Training with Acoustic Tokenizers.

Next Token Prediction Towards Multimodal Intelligence: A Comprehensive Survey BEATs: Audio Pre-Training with Acoustic Tokenizers

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-11T14:59:01.369162Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T14:59:01.369162Z digest=sha256:e0613c711a239323eb4d97de6df03ce51a986e5fa409c3aa8446a75d2d003b7d

Observation 13f12c5e-a11f-4e44-a4cd-0edf5c0fda77 · outbound

This paper cites an unresolved cited work.

Next Token Prediction Towards Multimodal Intelligence: A Comprehensive Survey Unresolved cited work

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-11T14:59:01.374638Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T14:59:01.374638Z digest=sha256:bb538a2154dbebc736f487cf1cb33dd268c6c73ebcfce11bc28b97fc4ac55bc0

Observation 509b8f1b-44d0-4033-88a3-3f49e6176625 · outbound

This paper cites an unresolved cited work.

Next Token Prediction Towards Multimodal Intelligence: A Comprehensive Survey Unresolved cited work

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-11T14:59:01.364278Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T14:59:01.364278Z digest=sha256:ef6961b6f2fc3cbc35605776bf3388a94d97c7cc834b5a6e4ed820f8fe4927b9

Observation e7257bb8-252f-4140-81b7-0b8aad065ed7 · outbound

This paper cites InternVL: Scaling up Vision Foundation Models and Aligning for Generic Visual-Linguistic Tasks.

Next Token Prediction Towards Multimodal Intelligence: A Comprehensive Survey InternVL: Scaling up Vision Foundation Models and Aligning for Generic Visual-Linguistic Tasks

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-11T14:59:01.389504Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T14:59:01.389504Z digest=sha256:a35bbc8490ab9c5596d0f95d3e5fd470026939ed7021033da493cbd1bd6d029e

Observation 75fee6e8-3154-41aa-9617-1ab5e12fcb06 · outbound

This paper cites SOLO: A Single Transformer for Scalable Vision-Language Modeling.

Next Token Prediction Towards Multimodal Intelligence: A Comprehensive Survey SOLO: A Single Transformer for Scalable Vision-Language Modeling

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-11T14:59:01.379180Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T14:59:01.379180Z digest=sha256:358f3f589bcd34830dd055fa2c9b246f4e50187caf9a309fb623f134b971afd3

Observation cc0af5e5-0672-4f6d-9f7c-4119c6531bf5 · outbound

This paper cites Gonzalez, Ion Stoica, and Eric P.

Next Token Prediction Towards Multimodal Intelligence: A Comprehensive Survey Gonzalez, Ion Stoica, and Eric P

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-11T14:59:01.398122Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T14:59:01.398122Z digest=sha256:dd43e768c3760d73cdc5ac5dbcbbc0ebfb8bfc092b0ae8e0c4ec4a6fc58a1e13

Observation cd8c3b4d-786f-47fb-a49a-916c31af0e49 · outbound

This paper cites an unresolved cited work.

Next Token Prediction Towards Multimodal Intelligence: A Comprehensive Survey Unresolved cited work

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-11T14:59:01.401875Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T14:59:01.401875Z digest=sha256:9b3e2b34c4c82c23b57500ca17e610d6aca95455c2a53664a547317952932d40

Observation badb4c3c-3c3e-4faa-a453-53460556e332 · outbound

This paper cites Diffusion Policy: Visuomotor Policy Learning via Action Diffusion.

Next Token Prediction Towards Multimodal Intelligence: A Comprehensive Survey Diffusion Policy: Visuomotor Policy Learning via Action Diffusion

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-11T14:59:01.393986Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T14:59:01.393986Z digest=sha256:4d5d4f04f1611592df60de0ae648c04b39cd4d45b42a48d453a580d116e0727d

Observation 72a67477-9312-4982-9d53-88482032ab57 · outbound

This paper cites Chi, Jeff Dean, Jacob Devlin, Adam Roberts, Denny Zhou, Quoc V.

Next Token Prediction Towards Multimodal Intelligence: A Comprehensive Survey Chi, Jeff Dean, Jacob Devlin, Adam Roberts, Denny Zhou, Quoc V

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-11T14:59:01.409466Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T14:59:01.409466Z digest=sha256:bdf3c0c8c95f85b93553680a8a69807ba3c7996e1b50c445642aab2bbce1565c

Observation afc1ca31-d0c4-4bb0-a47d-e4234d407cd6 · outbound

This paper cites an unresolved cited work.

Next Token Prediction Towards Multimodal Intelligence: A Comprehensive Survey Unresolved cited work

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-11T14:59:01.413324Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T14:59:01.413324Z digest=sha256:73029f68e540a310cda0efd88fe1285eba94ef90581f5426d8f6c9c14c339079

Observation 82f325f2-ca7f-45ae-a15e-fc1f267be666 · outbound

This paper cites Qwen-Audio: Advancing Universal Audio Understanding via Unified Large-Scale Audio-Language Models.

Next Token Prediction Towards Multimodal Intelligence: A Comprehensive Survey Qwen-Audio: Advancing Universal Audio Understanding via Unified Large-Scale Audio-Language Models

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-11T14:59:01.405469Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T14:59:01.405469Z digest=sha256:eb5bd8ebf6a6857fa3aa52003eba11db0df26ead9a029b796206082b0577ca50

Observation 53164fe2-376f-4cf1-bf58-52ad0b1476d4 · outbound

This paper cites an unresolved cited work.

Next Token Prediction Towards Multimodal Intelligence: A Comprehensive Survey Unresolved cited work

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-11T14:59:01.422391Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T14:59:01.422391Z digest=sha256:2d8784c50134a2de0ddb8de2b0c28a66160c2fca5374fda630616368e92a2ba3

Observation deaa9bcf-89b0-4249-96e5-2573fc335e86 · outbound

This paper cites an unresolved cited work.

Next Token Prediction Towards Multimodal Intelligence: A Comprehensive Survey Unresolved cited work

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-11T14:59:01.427039Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T14:59:01.427039Z digest=sha256:36315f7d13940e89ac4d2b1b115a4bb7d05ddef558155f7b79bb1f48c115c740

Observation 68ce523b-cb59-4d82-8f46-03f9ff65477d · outbound

This paper cites Training Verifiers to Solve Math Word Problems.

Next Token Prediction Towards Multimodal Intelligence: A Comprehensive Survey Training Verifiers to Solve Math Word Problems

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-11T14:59:01.417462Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T14:59:01.417462Z digest=sha256:e205612674a11ba0de08f461307273eb87c7c133688557a047cab885d9257966

Observation 3c62bd7f-0bd4-493b-86e1-a4b6e785a9fb · outbound

This paper cites A Survey on Multimodal Large Language Models for Autonomous Driving.

Next Token Prediction Towards Multimodal Intelligence: A Comprehensive Survey A Survey on Multimodal Large Language Models for Autonomous Driving

Reference 72

Resolution
unresolved
no resolver link, observed 2026-08-11T14:59:01.436221Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T14:59:01.436221Z digest=sha256:f368039d600f7171b08b9130a25898f0848688c9a0a0d2081ffc21df271f3799

Observation 2e138070-41d0-4e3f-bdc8-b36ac86a8f5a · outbound

This paper cites Holistic Analysis of Hallucination in GPT-4V(ision): Bias and Interference Challenges.

Next Token Prediction Towards Multimodal Intelligence: A Comprehensive Survey Holistic Analysis of Hallucination in GPT-4V(ision): Bias and Interference Challenges

Reference 73

Resolution
unresolved
no resolver link, observed 2026-08-11T14:59:01.441016Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T14:59:01.441016Z digest=sha256:7e4eb190fd6a976acb0c9f1aac1681368c54abdd96149e98378a5eccdd6dbe45

Observation 53ade0a4-d6bb-4f8e-a636-6bac7e340233 · outbound

This paper cites Multi-Task Learning with Deep Neural Networks: A Survey.

Next Token Prediction Towards Multimodal Intelligence: A Comprehensive Survey Multi-Task Learning with Deep Neural Networks: A Survey

Reference 74

Resolution
unresolved
no resolver link, observed 2026-08-11T14:59:01.431584Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T14:59:01.431584Z digest=sha256:dd57eb3f9ce776b1cad3ad7f8734c86292ff1d80040714959ddcbb56e4279466

Observation 8a2514d9-6259-4e23-8457-c78eed1d483a · outbound

This paper cites FlashAttention: Fast and Memory-Efficient Exact Attention with IO-Awareness.

Next Token Prediction Towards Multimodal Intelligence: A Comprehensive Survey FlashAttention: Fast and Memory-Efficient Exact Attention with IO-Awareness

Reference 75

Resolution
unresolved
no resolver link, observed 2026-08-11T14:59:01.450626Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T14:59:01.450626Z digest=sha256:5cdbe96fcc986055c226732b3cfd5dad85492bd917d235c81e7d84548156b073

Observation 0137b1e5-542f-4dbf-a1c7-5b910742c6b6 · outbound

This paper cites SpeechVerse: A Large-scale Generalizable Audio Language Model.

Next Token Prediction Towards Multimodal Intelligence: A Comprehensive Survey SpeechVerse: A Large-scale Generalizable Audio Language Model

Reference 76

Resolution
unresolved
no resolver link, observed 2026-08-11T14:59:01.455567Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T14:59:01.455567Z digest=sha256:4fbd791a469050ffded11f4612d7e0563cdcfc0f1bb44b7d8b88445cddb5cb28

Observation 00e51334-ec3f-4805-9bce-1ce34438fdb4 · outbound

This paper cites InstructBLIP: Towards General-purpose Vision-Language Models with Instruction Tuning.

Next Token Prediction Towards Multimodal Intelligence: A Comprehensive Survey InstructBLIP: Towards General-purpose Vision-Language Models with Instruction Tuning

Reference 77

Resolution
unresolved
no resolver link, observed 2026-08-11T14:59:01.445953Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T14:59:01.445953Z digest=sha256:8bf5b6ae2b8cf6be0fb04d8c33ea41f2a4cd4848f9d486d4485781087e5af241

Observation cf1f6e1d-f5f4-45a2-a2ad-888079543368 · outbound

This paper cites Heek, Matthias Minderer, Mathilde Caron, A.

Next Token Prediction Towards Multimodal Intelligence: A Comprehensive Survey Heek, Matthias Minderer, Mathilde Caron, A

Reference 79

Resolution
unresolved
no resolver link, observed 2026-08-11T14:59:01.470063Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T14:59:01.470063Z digest=sha256:159c5c44df9435344f7734adb4f569beb61eec4a05e12e557a2013bd0e845f79

Observation 7f1a0eae-bc95-49e0-94b6-db766ed7b782 · outbound

This paper cites FMA: A Dataset For Music Analysis.

Next Token Prediction Towards Multimodal Intelligence: A Comprehensive Survey FMA: A Dataset For Music Analysis

Reference 80

Resolution
unresolved
no resolver link, observed 2026-08-11T14:59:01.460484Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T14:59:01.460484Z digest=sha256:a526cd49ccac3118910a37effdf85e3c7b1b6fa5ae41916ab0fd7bf6d0e5c4a5

Observation fa457102-9ce3-46ca-8299-d1935f06b996 · outbound

This paper cites RedCaps: web-curated image-text data created by the people, for the people.

Next Token Prediction Towards Multimodal Intelligence: A Comprehensive Survey RedCaps: web-curated image-text data created by the people, for the people

Reference 81

Resolution
unresolved
no resolver link, observed 2026-08-11T14:59:01.484638Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T14:59:01.484638Z digest=sha256:ed0e550958dfaeb73230f3829f62cc84869d1a81fa94c4cc26ea9f1a80888f0a

Observation a55c6631-90df-4e81-a2a6-0f00f8b09eba · outbound

This paper cites an unresolved cited work.

Next Token Prediction Towards Multimodal Intelligence: A Comprehensive Survey Unresolved cited work

Reference 82

Resolution
unresolved
no resolver link, observed 2026-08-11T14:59:01.488476Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T14:59:01.488476Z digest=sha256:3b39731cef0e41b03ceb5187a68d2e2c45eea2a9a765bfdfecd4d54323918ed4

Observation 765fec69-3ccf-437c-8bda-76cccfd07904 · outbound

This paper cites Patch n' Pack: NaViT, a Vision Transformer for any Aspect Ratio and Resolution.

Next Token Prediction Towards Multimodal Intelligence: A Comprehensive Survey Patch n' Pack: NaViT, a Vision Transformer for any Aspect Ratio and Resolution

Reference 83

Resolution
unresolved
no resolver link, observed 2026-08-11T14:59:01.475206Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T14:59:01.475206Z digest=sha256:1b09ef29afcef89239fa4be2715d627a612de70149708141aa84f89da100ee38

Observation 1fc8a169-2507-4ff8-bdd8-d04f010a0bef · outbound

This paper cites Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models.

Next Token Prediction Towards Multimodal Intelligence: A Comprehensive Survey Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models

Reference 84

Resolution
unresolved
no resolver link, observed 2026-08-11T14:59:01.480020Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T14:59:01.480020Z digest=sha256:cc5a606775e12fad1fb57b1de48e26b9b7fe3931b70bffcb6d1fc89e3ebefbea

Observation 633b3996-342d-4b91-a586-7ff45ee4dde0 · outbound

This paper cites Jukebox: A Generative Model for Music.

Next Token Prediction Towards Multimodal Intelligence: A Comprehensive Survey Jukebox: A Generative Model for Music

Reference 85

Resolution
unresolved
no resolver link, observed 2026-08-11T14:59:01.500382Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T14:59:01.500382Z digest=sha256:d98bd69ab8f84aa6236de47ff1161a99cf2fad56da1b19d77e73b2d024ebe47b

Observation 8b092bff-42bc-475a-83b8-3cde6722c081 · outbound

This paper cites Diffusion Models Beat GANs on Image Synthesis.

Next Token Prediction Towards Multimodal Intelligence: A Comprehensive Survey Diffusion Models Beat GANs on Image Synthesis

Reference 86

Resolution
unresolved
no resolver link, observed 2026-08-11T14:59:01.504454Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T14:59:01.504454Z digest=sha256:28a15c667c9e022f09da28ecc9bf3b2e8e6a674ebf13b508136346c7af7e62a3

Observation 9f9ae898-7e3e-4421-b8b5-47f009d353fc · outbound

This paper cites ECAPA-TDNN: Emphasized Channel Attention, Propagation and Aggregation in TDNN Based Speaker Verification.

Next Token Prediction Towards Multimodal Intelligence: A Comprehensive Survey ECAPA-TDNN: Emphasized Channel Attention, Propagation and Aggregation in TDNN Based Speaker Verification

Reference 87

Resolution
unresolved
no resolver link, observed 2026-08-11T14:59:01.492409Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T14:59:01.492409Z digest=sha256:dbef484546d835288060107187b1016e4cd3f1b7f09c13d739ef6754327a720c

Observation 0db26017-e73b-4b5a-bbd4-f65a0efcc41d · outbound

This paper cites an unresolved cited work.

Next Token Prediction Towards Multimodal Intelligence: A Comprehensive Survey Unresolved cited work

Reference 88

Resolution
unresolved
no resolver link, observed 2026-08-11T14:59:01.496489Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T14:59:01.496489Z digest=sha256:83a7bda9d3f16683a4e9aa5e429c6ef119b894f1a947db3feaa311fdb29be852

Observation 30957332-f18c-47c2-8655-5cd706c375cb · outbound

This paper cites Temporal 3D ConvNets: New Architecture and Transfer Learning for Video Classification.

Next Token Prediction Towards Multimodal Intelligence: A Comprehensive Survey Temporal 3D ConvNets: New Architecture and Transfer Learning for Video Classification

Reference 89

Resolution
unresolved
no resolver link, observed 2026-08-11T14:59:01.518521Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T14:59:01.518521Z digest=sha256:c891f90e8e3f9153ae95d446fdd5cf883ded9dd0d3a7a4d86f25b43f1c942330

Observation 9348cce5-7787-4b82-a51d-f789edf73607 · outbound

This paper cites an unresolved cited work.

Next Token Prediction Towards Multimodal Intelligence: A Comprehensive Survey Unresolved cited work

Reference 90

Resolution
unresolved
no resolver link, observed 2026-08-11T14:59:01.523561Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T14:59:01.523561Z digest=sha256:d47cbda6353674979e786312ca5db42f17c74d38daea636a5178806a5e6b2fa2

Observation 45336d13-7523-4ecb-a115-66b85f1dd4c4 · outbound

This paper cites Unveiling Encoder-Free Vision-Language Models.

Next Token Prediction Towards Multimodal Intelligence: A Comprehensive Survey Unveiling Encoder-Free Vision-Language Models

Reference 91

Resolution
unresolved
no resolver link, observed 2026-08-11T14:59:01.509004Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T14:59:01.509004Z digest=sha256:88d6cf7508f933e88529bcd96188542fc5bfdb9084fa1aeedb7cc9d2e4a57dec

Observation 363f135a-4496-4ee9-8fe1-47b4341fa4e6 · outbound

This paper cites an unresolved cited work.

Next Token Prediction Towards Multimodal Intelligence: A Comprehensive Survey Unresolved cited work

Reference 92

Resolution
unresolved
no resolver link, observed 2026-08-11T14:59:01.513974Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T14:59:01.513974Z digest=sha256:61f782f3c4460802cc0835493bb1f90e15f75a443bf954b8290b002767e2dedc

Observation 3426ff5a-082f-44e3-a0a9-3a0b7ff33c0e · outbound

This paper cites LP-MusicCaps: LLM-Based Pseudo Music Captioning.

Next Token Prediction Towards Multimodal Intelligence: A Comprehensive Survey LP-MusicCaps: LLM-Based Pseudo Music Captioning

Reference 93

Resolution
unresolved
no resolver link, observed 2026-08-11T14:59:01.537265Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T14:59:01.537265Z digest=sha256:5b5ff08482133939b58b29602adc695b4514dd0f4569b867f49fa922f8b072c8

Observation 5d98913d-cb37-4594-b764-ac87293184d3 · outbound

This paper cites A Survey on In-context Learning.

Next Token Prediction Towards Multimodal Intelligence: A Comprehensive Survey A Survey on In-context Learning

Reference 94

Resolution
unresolved
no resolver link, observed 2026-08-11T14:59:01.542017Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T14:59:01.542017Z digest=sha256:80eb09777dada9f2b8238992e6892c65eb245fbb16897b4e40eb0b4538555880

Observation 7077c327-7bbb-428a-af1d-e3db90907591 · outbound

This paper cites an unresolved cited work.

Next Token Prediction Towards Multimodal Intelligence: A Comprehensive Survey Unresolved cited work

Reference 95

Resolution
unresolved
no resolver link, observed 2026-08-11T14:59:01.528111Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T14:59:01.528111Z digest=sha256:b3a32f47a15aae923fa6d5920442eaebe6ee201f9efc1b5128677764ad981f4e

Observation 13641ca5-614b-440a-8b41-d66e6fd25f4b · outbound

This paper cites an unresolved cited work.

Next Token Prediction Towards Multimodal Intelligence: A Comprehensive Survey Unresolved cited work

Reference 96

Resolution
unresolved
no resolver link, observed 2026-08-11T14:59:01.532673Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T14:59:01.532673Z digest=sha256:bba5b7d4000755c524db3fbf4d383fc83a642390df0b11756a00b5f6b9c724ec

Observation f9d7d0b9-eec7-4b4c-93b5-4edc22129e86 · outbound

This paper cites PaLM-E: An Embodied Multimodal Language Model.

Next Token Prediction Towards Multimodal Intelligence: A Comprehensive Survey PaLM-E: An Embodied Multimodal Language Model

Reference 97

Resolution
unresolved
no resolver link, observed 2026-08-11T14:59:01.556420Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T14:59:01.556420Z digest=sha256:05aa63524791923fc443334bc7200368a8fb20a864d4be586e734215278fb1ce

Observation 469ec685-3bfb-4f72-8275-502b27fef095 · outbound

This paper cites an unresolved cited work.

Next Token Prediction Towards Multimodal Intelligence: A Comprehensive Survey Unresolved cited work

Reference 98

Resolution
unresolved
no resolver link, observed 2026-08-11T14:59:01.561434Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T14:59:01.561434Z digest=sha256:997b4c7c8c8b6ab8d5c0bd16d46d0f4e71d275ee4d8e98b2095a7caf5c816394

Observation b820cdd7-daab-4380-885b-ae0376eb4fff · outbound

This paper cites DreamLLM: Synergistic Multimodal Comprehension and Creation.

Next Token Prediction Towards Multimodal Intelligence: A Comprehensive Survey DreamLLM: Synergistic Multimodal Comprehension and Creation

Reference 99

Resolution
unresolved
no resolver link, observed 2026-08-11T14:59:01.546678Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T14:59:01.546678Z digest=sha256:1bf5273ecd167d522ac39228dcdbccb465dc1ccafb58755fc66defee1810647e

Observation 0e19db84-16ba-4bf2-81e9-d598746136d2 · outbound

This paper cites An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale.

Next Token Prediction Towards Multimodal Intelligence: A Comprehensive Survey An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale

Reference 100

Resolution
unresolved
no resolver link, observed 2026-08-11T14:59:01.551495Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T14:59:01.551495Z digest=sha256:aaa07a6ccaddf37392ea1f76773175b22fb03d6da34b46c3a1d98d15afe00ed1

Observation d945b4da-fc3e-4a39-b2fb-33fc5d52be69 · outbound

This paper cites CosyVoice: A Scalable Multilingual Zero-shot Text-to-speech Synthesizer based on Supervised Semantic Tokens.

Next Token Prediction Towards Multimodal Intelligence: A Comprehensive Survey CosyVoice: A Scalable Multilingual Zero-shot Text-to-speech Synthesizer based on Supervised Semantic Tokens

Reference 101

Resolution
unresolved
no resolver link, observed 2026-08-11T14:59:01.576272Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T14:59:01.576272Z digest=sha256:ee482f4b25d490aefa32b03e95c3449da248f3f7baa47ddc2058cd80111e9b35

Observation 5a9409a4-187f-48ff-9a6d-3cffe028e505 · outbound

This paper cites LauraGPT: Listen, Attend, Understand, and Regenerate Audio with GPT.

Next Token Prediction Towards Multimodal Intelligence: A Comprehensive Survey LauraGPT: Listen, Attend, Understand, and Regenerate Audio with GPT

Reference 102

Resolution
unresolved
no resolver link, observed 2026-08-11T14:59:01.581000Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T14:59:01.581000Z digest=sha256:2128d0462d53a8a1bef9bef910ead156b4c75377c023129d6c02a81fd9d763cb

Pith citing papers

Observation c3c21c6f-1c09-4d2c-b84d-216470d0d4f6 · inbound

Parallelized Autoregressive Visual Generation cites this paper.

Parallelized Autoregressive Visual Generation Next Token Prediction Towards Multimodal Intelligence: A Comprehensive Survey

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-11T11:42:31.574249Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:42:31.574249Z digest=sha256:5bbff94cd3a43eb2f5dbef2a547605782654671793f203fffee1ed89d3aa35e6

Observation f71d18ae-288c-4b3e-b5ec-688d1f8a70a4 · inbound

Visual Autoregressive Modeling for Image Super-Resolution cites this paper.

Visual Autoregressive Modeling for Image Super-Resolution Next Token Prediction Towards Multimodal Intelligence: A Comprehensive Survey

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-09T21:47:59.515837Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T21:47:59.515837Z digest=sha256:4bf75a5fc3c21672af031cf7ee0f9726cc960fb1d1c10e44ac7ecb5e1b8933bf

Observation 22a69b2b-05a7-40c0-8e36-ad046e9930c2 · inbound

Next Block Prediction: Video Generation via Semi-Autoregressive Modeling cites this paper.

Next Block Prediction: Video Generation via Semi-Autoregressive Modeling Next Token Prediction Towards Multimodal Intelligence: A Comprehensive Survey

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-08T11:51:13.184615Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T11:51:13.184615Z digest=sha256:33f37277d115ac5f3ae0f15f34654221ab6560ac702a81475c3a8412abdd4cca

Observation 585b2bba-ca37-4592-acf3-41a4270dc873 · inbound

WISE: A World Knowledge-Informed Semantic Evaluation for Text-to-Image Generation cites this paper.

WISE: A World Knowledge-Informed Semantic Evaluation for Text-to-Image Generation Next Token Prediction Towards Multimodal Intelligence: A Comprehensive Survey

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-05-15T16:24:27.465541Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-15T16:24:27.407376Z digest=sha256:d46cd242637239c4e48b1f3616ecba29d5280fa7f9ea3a6ab804d9775376919b

Observation 2688273f-f0cc-40ae-ba27-e79b85be1831 · inbound

On The Landscape of Spoken Language Models: A Comprehensive Survey cites this paper.

On The Landscape of Spoken Language Models: A Comprehensive Survey Next Token Prediction Towards Multimodal Intelligence: A Comprehensive Survey

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-05-22T20:45:08.044990Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-22T20:44:57.476464Z digest=sha256:bd8c1d3d2a67c87f074c07f231256a11b083f233653b858c1b84355b2fcf3981

Observation 6e112dd2-b128-4dd6-8559-be9a6765b12e · inbound

Selftok: Discrete Visual Tokens of Autoregression, by Diffusion, and for Reasoning cites this paper.

Selftok: Discrete Visual Tokens of Autoregression, by Diffusion, and for Reasoning Next Token Prediction Towards Multimodal Intelligence: A Comprehensive Survey

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-15T22:21:30.258422Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:21:30.258422Z digest=sha256:a1b34e2728d5c1f076c923b6ab6e4bbf22f487f926e24924b83c59c8551a285f

Observation 63f2893d-8961-4a1b-8a61-e8cfd3cbb1e0 · inbound

Adapting Pretrained Language Models for Citation Classification via Self-Supervised Contrastive Learning cites this paper.

Adapting Pretrained Language Models for Citation Classification via Self-Supervised Contrastive Learning Next Token Prediction Towards Multimodal Intelligence: A Comprehensive Survey

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T15:38:27.366464Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:38:27.366464Z digest=sha256:7039bd83bbf589cb3842bcbfd5913ab31593d4cc7510ff3172020ba20d1804f1

Observation ca3c074a-5b51-4347-af22-492564f5c003 · inbound

Image Tokens Matter: Mitigating Hallucination in Discrete Tokenizer-based Large Vision-Language Models via Latent Editing cites this paper.

Image Tokens Matter: Mitigating Hallucination in Discrete Tokenizer-based Large Vision-Language Models via Latent Editing Next Token Prediction Towards Multimodal Intelligence: A Comprehensive Survey

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T14:27:58.571076Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:27:58.571076Z digest=sha256:769123c7787e3575faacc64cb8edf86db96efa5217e92760b3b7a270e7efc714

Observation 18e7c4a7-cc8a-4549-9bf3-db27cfb420f7 · inbound

MedErr-CT: A Visual Question Answering Benchmark for Identifying and Correcting Errors in CT Reports cites this paper.

MedErr-CT: A Visual Question Answering Benchmark for Identifying and Correcting Errors in CT Reports Next Token Prediction Towards Multimodal Intelligence: A Comprehensive Survey

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-06T23:10:44.228630Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:10:44.228630Z digest=sha256:d71137af2c4a66dfb7cd001c425d4fb6f1461b6cf16ebdc46b8076568297f31e

Observation 33a4a4e7-6b4a-4391-8c41-85d13f924c68 · inbound

Region-Aware Multimodal Large Language Model via SlowFast Tokenization and Pseudo-Mask Guidance for 3D CT Report Generation cites this paper.

Region-Aware Multimodal Large Language Model via SlowFast Tokenization and Pseudo-Mask Guidance for 3D CT Report Generation Next Token Prediction Towards Multimodal Intelligence: A Comprehensive Survey

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-06T21:55:45.970164Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:55:45.970164Z digest=sha256:fe5d7b5a9fe306a089191465f79c097edae8300a0c05ba14c11652b4e53a5362

Observation be4665de-631c-4724-af0e-127ef806586b · inbound

ViLaD: A Large Vision Language Diffusion Framework for End-to-End Autonomous Driving cites this paper.

ViLaD: A Large Vision Language Diffusion Framework for End-to-End Autonomous Driving Next Token Prediction Towards Multimodal Intelligence: A Comprehensive Survey

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-15T17:25:14.935617Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T17:25:14.935617Z digest=sha256:969a95ef22cd0e8f1f7518eda4f94485c7a14f0727aea50574df81bf7b30ae99

Observation 8eabab2d-0baa-405b-8d9a-8a88ffcd9422 · inbound

A Unified Low-level Foundation Model for Enhancing Pathology Image Quality cites this paper.

A Unified Low-level Foundation Model for Enhancing Pathology Image Quality Next Token Prediction Towards Multimodal Intelligence: A Comprehensive Survey

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-05T13:00:48.147977Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T13:00:48.147977Z digest=sha256:cece2ae3aa683081b06ca417bc041bc5fa73d187d49fcea31003831710b84e99

Observation b0e0be72-5092-499a-9868-616165665746 · inbound

SimReg: Achieving Higher Performance in the Pretraining via Embedding Similarity Regularization cites this paper.

SimReg: Achieving Higher Performance in the Pretraining via Embedding Similarity Regularization Next Token Prediction Towards Multimodal Intelligence: A Comprehensive Survey

Reference 1

Resolution
verified exact
arxiv_id, observed 2026-05-12T07:16:28.852066Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-12T03:33:45.192139Z digest=sha256:21f3c975c146e6de7296a033e81d8cbf13c881f13d68746d998d44df1e3e85b8

Observation 9e0d91d4-6eb3-4809-be35-12fa1b4a0948 · inbound

NITP: Next Implicit Token Prediction for LLM Pre-training cites this paper.

NITP: Next Implicit Token Prediction for LLM Pre-training Next Token Prediction Towards Multimodal Intelligence: A Comprehensive Survey

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-06-30T12:24:39.261088Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-06-30T12:23:42.587689Z digest=sha256:95f27d1efc8d12b85209937141c3b16d4e3b7a49cadf846346eeace4517302b7

Observation a72629b2-ab5b-4326-9104-a826e9231761 · inbound

NITP: Next Implicit Token Prediction for LLM Pre-training cites this paper.

NITP: Next Implicit Token Prediction for LLM Pre-training Next Token Prediction Towards Multimodal Intelligence: A Comprehensive Survey

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-07-04T00:39:16.236947Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-07-04T00:38:33.708836Z digest=sha256:aca09a4008cd51c416fa381e51be1e7d12b3fa384c39023736886c765506e40f

Observation 1c3ab298-d7ac-4ac3-b52e-158f88d19a75 · inbound

NITP: Next Implicit Token Prediction for LLM Pre-training cites this paper.

NITP: Next Implicit Token Prediction for LLM Pre-training Next Token Prediction Towards Multimodal Intelligence: A Comprehensive Survey

Reference 2

Resolution
unresolved
no resolver link, observed 2026-07-14T18:45:28.635910Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T18:45:28.635910Z digest=sha256:6a72222f203f87b4344deef936b0e05c377b54d403a215a40b865eee2692ab94

Observation 9765716f-0b0e-46d2-85c9-50c56f27fb6f · inbound

Toward Native Multimodal Modeling: A Roadmap cites this paper.

Toward Native Multimodal Modeling: A Roadmap Next Token Prediction Towards Multimodal Intelligence: A Comprehensive Survey

Reference 224

Resolution
verified exact
arxiv_id, observed 2026-06-29T23:04:02.082867Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-06-29T22:58:38.610609Z digest=sha256:777985fa490b674392104b710d8db494e0e8ce9cb42db5c76459298cdb97d16f

Observation 7d2e77ea-d6c9-45b9-921d-c920dc26abb9 · inbound

PathAR: Structure-First Autoregressive Synthesis of Multimodal Pathology Images cites this paper.

PathAR: Structure-First Autoregressive Synthesis of Multimodal Pathology Images Next Token Prediction Towards Multimodal Intelligence: A Comprehensive Survey

Reference 28

Resolution
verified exact
arxiv_id, observed 2026-07-01T22:06:15.887435Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-06-28T15:50:54.950899Z digest=sha256:a0e636bc1587e402911b16b9533548a902dd78477d7ec22ab9e7faac8722cea4

Observation 91cf0468-9629-4d60-b85b-94d377b9ac20 · inbound

PseudoBench: Measuring How Agentic Auto-Research Fuels Pseudoscience cites this paper.

PseudoBench: Measuring How Agentic Auto-Research Fuels Pseudoscience Next Token Prediction Towards Multimodal Intelligence: A Comprehensive Survey

Reference 87

Resolution
verified exact
arxiv_id, observed 2026-07-03T20:18:56.449179Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-06-27T01:29:12.725865Z digest=sha256:33a823cefd91ccaf37ec3fd0dec197bd754c67e0b8d273e8a800be23af7de21a