Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-15T22:21:30.601457Z
Paper Citation Record · LEDGER
As of 23 August 2026, this Paper Citation Record lists 100 of 136 outbound references and 10 inbound Pith citation observations for arXiv:2505.07538.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-15T22:21:30.601457Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-23T06:30:58.430688+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-07T12:34:11.795076Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-07-03T14:28:32.067048Z
100 of 136 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 853fae54-055d-4765-9ddd-2781eec2210b · outbound
Selftok: Discrete Visual Tokens of Autoregression, by Diffusion, and for Reasoning Cosmos World Foundation Model Platform for Physical AI
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation eabafcd6-b845-4f1c-a6ac-6b2586ac66b6 · outbound
Selftok: Discrete Visual Tokens of Autoregression, by Diffusion, and for Reasoning Llama: Open and efficient foundation language models
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3917c776-d4a0-42a9-a912-8f28d7a203ab · outbound
Selftok: Discrete Visual Tokens of Autoregression, by Diffusion, and for Reasoning Cogview4: Next-generation image creation
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0757000c-989f-48a4-92ce-05ab452a7c3c · outbound
Selftok: Discrete Visual Tokens of Autoregression, by Diffusion, and for Reasoning FlexTok: Resampling Images into 1D Token Sequences of Flexible Length
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a7a9f234-dc8a-4e09-bfe3-d6380bdfb96d · outbound
Selftok: Discrete Visual Tokens of Autoregression, by Diffusion, and for Reasoning Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 02dc6327-1738-45aa-8000-18c07da86843 · outbound
Selftok: Discrete Visual Tokens of Autoregression, by Diffusion, and for Reasoning Instructpix2pix: Learning to follow image editing instructions
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d662ea7e-8b52-448a-b85b-7dce66cfc2ce · outbound
Selftok: Discrete Visual Tokens of Autoregression, by Diffusion, and for Reasoning Coyo-700m: Image-text pair dataset
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3a6dd38f-3567-469d-a2a0-9842c5cbae80 · outbound
Selftok: Discrete Visual Tokens of Autoregression, by Diffusion, and for Reasoning Masactrl: Tuning-free mutual self-attention control for consistent image synthesis and editing
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation db27d10d-1d20-424a-bbac-e9512bcdb96c · outbound
Selftok: Discrete Visual Tokens of Autoregression, by Diffusion, and for Reasoning Muse: Text-To-Image Generation via Masked Generative Transformers
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e4e1cf9f-d723-45d4-9d4b-0ded1375c65e · outbound
Selftok: Discrete Visual Tokens of Autoregression, by Diffusion, and for Reasoning PixArt-$\alpha$: Fast Training of Diffusion Transformer for Photorealistic Text-to-Image Synthesis
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ff94986d-04fa-41b5-8edc-3dc52c7149f4 · outbound
Selftok: Discrete Visual Tokens of Autoregression, by Diffusion, and for Reasoning MMDetection: Open MMLab Detection Toolbox and Benchmark
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6e112dd2-b128-4dd6-8559-be9a6765b12e · outbound
Selftok: Discrete Visual Tokens of Autoregression, by Diffusion, and for Reasoning Next Token Prediction Towards Multimodal Intelligence: A Comprehensive Survey
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0c411dfa-871c-416a-bd1c-e7ceac585b17 · outbound
Selftok: Discrete Visual Tokens of Autoregression, by Diffusion, and for Reasoning Janus-Pro: Unified Multimodal Understanding and Generation with Data and Model Scaling
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d1b170c0-a401-4549-80ca-c1a75d75dc7f · outbound
Selftok: Discrete Visual Tokens of Autoregression, by Diffusion, and for Reasoning Expanding performance boundaries of open-source multimodal models with model, data, and test-time scaling, 2024
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a2fc0355-1853-434a-9b94-3571188d362d · outbound
Selftok: Discrete Visual Tokens of Autoregression, by Diffusion, and for Reasoning Internvl: Scaling up vision foundation models and aligning for generic visual-linguistic tasks
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 991e4d41-d404-46ac-88af-909d179c1477 · outbound
Selftok: Discrete Visual Tokens of Autoregression, by Diffusion, and for Reasoning Davidsonian Scene Graph: Improving Reliability in Fine-grained Evaluation for Text-to-Image Generation
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 29020a1a-e7cb-47ec-9512-81f5351c1ebd · outbound
Selftok: Discrete Visual Tokens of Autoregression, by Diffusion, and for Reasoning Instructblip: Towards general-purpose vision-language models with instruction tuning
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b07c87a6-8e2d-4c32-9d26-f3a4a23ca481 · outbound
Selftok: Discrete Visual Tokens of Autoregression, by Diffusion, and for Reasoning Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9c650bda-b151-4f03-91cc-27e349ad988a · outbound
Selftok: Discrete Visual Tokens of Autoregression, by Diffusion, and for Reasoning Imagenet: A large-scale hierarchical image database
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e33bdd08-ad5a-430d-9ce4-8e014621962d · outbound
Selftok: Discrete Visual Tokens of Autoregression, by Diffusion, and for Reasoning Diffusion models beat gans on image synthesis
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a994d005-c83a-4597-b7e5-2cf6bbafc688 · outbound
Selftok: Discrete Visual Tokens of Autoregression, by Diffusion, and for Reasoning InternLM-XComposer2: Mastering Free-form Text-Image Composition and Comprehension in Vision-Language Large Model
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ccec577d-8cd8-4e0d-a320-db7caa8647fa · outbound
Selftok: Discrete Visual Tokens of Autoregression, by Diffusion, and for Reasoning An image is worth 16x16 words: Transformers for image recognition at scale
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 62a9d35e-df87-4935-9485-09363b4645a4 · outbound
Selftok: Discrete Visual Tokens of Autoregression, by Diffusion, and for Reasoning Adaptive length image tokenization via recurrent allocation
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 80b3b58f-f02d-4e1b-9591-9e0af6ebb812 · outbound
Selftok: Discrete Visual Tokens of Autoregression, by Diffusion, and for Reasoning Scaling rectified flow transformers for high-resolution image synthesis
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 76fee686-1e78-43b2-8562-e6b77d14173c · outbound
Selftok: Discrete Visual Tokens of Autoregression, by Diffusion, and for Reasoning Taming transformers for high-resolution image synthesis
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8427b3ff-1693-4cd0-9532-d975ee6f87e1 · outbound
Selftok: Discrete Visual Tokens of Autoregression, by Diffusion, and for Reasoning Guiding Instruction-based Image Editing via Multimodal Large Language Models
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 13386499-bd3b-45c4-a6cd-e137f993a7f3 · outbound
Selftok: Discrete Visual Tokens of Autoregression, by Diffusion, and for Reasoning DataComp: In search of the next generation of multimodal datasets
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 18889c96-6f47-42c4-a1d7-b7a226c56420 · outbound
Selftok: Discrete Visual Tokens of Autoregression, by Diffusion, and for Reasoning Seedream 3.0 technical report, 2025
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6de8d79b-1cb1-42e6-9823-7387d02c7d23 · outbound
Selftok: Discrete Visual Tokens of Autoregression, by Diffusion, and for Reasoning SEED-X: Multimodal Models with Unified Multi-granularity Comprehension and Generation
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 27d72071-fbb4-48e4-b8a8-924177433665 · outbound
Selftok: Discrete Visual Tokens of Autoregression, by Diffusion, and for Reasoning Instructdiffusion: A generalist modeling interface for vision tasks
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 600b3329-c4b2-4fbc-ba4c-2ca452c5d1e7 · outbound
Selftok: Discrete Visual Tokens of Autoregression, by Diffusion, and for Reasoning Geneval: An object-focused framework for evaluating text-to- image alignment
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3e243bff-ef66-456c-bd6b-da1ab8b5c2ce · outbound
Selftok: Discrete Visual Tokens of Autoregression, by Diffusion, and for Reasoning Better & Faster Large Language Models via Multi-token Prediction
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 779507a9-4a73-46bd-8717-fdedb66aee14 · outbound
Selftok: Discrete Visual Tokens of Autoregression, by Diffusion, and for Reasoning Generative adversarial networks
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e74b460b-b72f-4871-8416-8e9308c7f2c5 · outbound
Selftok: Discrete Visual Tokens of Autoregression, by Diffusion, and for Reasoning Making the v in vqa matter: Elevating the role of image understanding in visual question answering
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 22104682-3486-4b52-be63-5681df19ef9c · outbound
Selftok: Discrete Visual Tokens of Autoregression, by Diffusion, and for Reasoning Multi-reward as condition for instruction- based image editing, 2024
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d0d1cd56-4605-432d-be82-cf9ccd8d4fb6 · outbound
Selftok: Discrete Visual Tokens of Autoregression, by Diffusion, and for Reasoning DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c375f887-d5d7-4c3d-aec4-040d63f3450c · outbound
Selftok: Discrete Visual Tokens of Autoregression, by Diffusion, and for Reasoning Vizwiz grand challenge: Answering visual questions from blind people
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9008339b-3c99-4b71-b876-bd54ceb307cc · outbound
Selftok: Discrete Visual Tokens of Autoregression, by Diffusion, and for Reasoning Infinity: Scaling bitwise autoregressive modeling for high-resolution image synthesis, 2024
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ea04f245-2ca7-43ce-86c9-de5e3eb8ccfd · outbound
Selftok: Discrete Visual Tokens of Autoregression, by Diffusion, and for Reasoning Prompt-to-Prompt Image Editing with Cross Attention Control
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation be4a4dcf-ca40-44ab-a1cc-4ae519122d56 · outbound
Selftok: Discrete Visual Tokens of Autoregression, by Diffusion, and for Reasoning Gans trained by a two time-scale update rule converge to a local nash equilibrium
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f6dfcb11-85d6-4a44-ab43-ed323029bb19 · outbound
Selftok: Discrete Visual Tokens of Autoregression, by Diffusion, and for Reasoning Hidream-i1
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f0981535-08e6-4647-857a-c8cc885f2a74 · outbound
Selftok: Discrete Visual Tokens of Autoregression, by Diffusion, and for Reasoning Towards a Definition of Disentangled Representations
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 75fc73ab-7688-47e6-ae52-f7588a43c1de · outbound
Selftok: Discrete Visual Tokens of Autoregression, by Diffusion, and for Reasoning Classifier-Free Diffusion Guidance
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 52adf81e-6273-4639-95ac-4b5d69bccd39 · outbound
Selftok: Discrete Visual Tokens of Autoregression, by Diffusion, and for Reasoning Disentanglement via latent quantization
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d12a027a-dfe5-4133-9a00-33d89758b782 · outbound
Selftok: Discrete Visual Tokens of Autoregression, by Diffusion, and for Reasoning MiniCPM: Unveiling the Potential of Small Language Models with Scalable Training Strategies
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c963e1d4-ab50-4d81-9524-bb2df27de48c · outbound
Selftok: Discrete Visual Tokens of Autoregression, by Diffusion, and for Reasoning Understanding square loss in training overparametrized neural network classifiers
Reference 46
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 83e7ce75-2c5e-4f4f-9023-2dc9c52a0acb · outbound
Selftok: Discrete Visual Tokens of Autoregression, by Diffusion, and for Reasoning Ella: Equip diffusion models with llm for enhanced semantic alignment, 2024
Reference 47
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6a26f0c4-bff6-4557-b824-6f5a2d25b334 · outbound
Selftok: Discrete Visual Tokens of Autoregression, by Diffusion, and for Reasoning ELLA: Equip Diffusion Models with LLM for Enhanced Semantic Alignment
Reference 48
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 51c076d4-e036-45df-8d6e-4ffa10990a75 · outbound
Selftok: Discrete Visual Tokens of Autoregression, by Diffusion, and for Reasoning ILLUME+: Illuminating Unified MLLM with Dual Visual Tokenization and Diffusion Refinement
Reference 49
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 14de0ede-ca8c-47b0-be8d-1d1c7705329f · outbound
Selftok: Discrete Visual Tokens of Autoregression, by Diffusion, and for Reasoning Vision-R1: Incentivizing Reasoning Capability in Multimodal Large Language Models
Reference 50
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cc83c1e5-7927-4d0a-9849-b6b5e2cc9f5f · outbound
Selftok: Discrete Visual Tokens of Autoregression, by Diffusion, and for Reasoning Gqa: A new dataset for real-world visual reasoning and compositional question answering
Reference 51
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 31569132-050a-4f63-9b7a-11fb77f66b6e · outbound
Selftok: Discrete Visual Tokens of Autoregression, by Diffusion, and for Reasoning GPT-4o System Card
Reference 52
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4fa7caed-f4d6-4004-b5e1-ed066c383b44 · outbound
Selftok: Discrete Visual Tokens of Autoregression, by Diffusion, and for Reasoning Direct Inversion: Boosting Diffusion-based Editing with 3 Lines of Code
Reference 53
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bf054325-6558-464b-bb1f-89483c73c190 · outbound
Selftok: Discrete Visual Tokens of Autoregression, by Diffusion, and for Reasoning How Far is Video Generation from World Model: A Physical Law Perspective
Reference 54
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ceab6c74-a369-4d70-a599-4527d22ebb47 · outbound
Selftok: Discrete Visual Tokens of Autoregression, by Diffusion, and for Reasoning A style-based generator architecture for generative adversarial networks
Reference 55
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4d0085df-d7ba-4c1c-8a1d-41f8b1f22403 · outbound
Selftok: Discrete Visual Tokens of Autoregression, by Diffusion, and for Reasoning VideoPoet: A Large Language Model for Zero-Shot Video Generation
Reference 56
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f702821b-c99b-41a7-b5a3-92b935fa8518 · outbound
Selftok: Discrete Visual Tokens of Autoregression, by Diffusion, and for Reasoning Unresolved cited work
Reference 57
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d8c11688-a2cf-40fa-8f0d-19d79cc1e8bc · outbound
Selftok: Discrete Visual Tokens of Autoregression, by Diffusion, and for Reasoning LLaVA-OneVision: Easy Visual Task Transfer
Reference 58
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 638fe5f2-d579-49c6-b51b-a6015cc8c741 · outbound
Selftok: Discrete Visual Tokens of Autoregression, by Diffusion, and for Reasoning mPLUG: Effective and Efficient Vision-Language Learning by Cross-modal Skip-connections
Reference 59
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 669ea675-efc2-4d0c-a58f-774f01abd0b6 · outbound
Selftok: Discrete Visual Tokens of Autoregression, by Diffusion, and for Reasoning Playground v2.5: Three Insights towards Enhancing Aesthetic Quality in Text-to-Image Generation
Reference 60
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4ff91796-5c1c-47e5-b54a-d5910d5e363a · outbound
Selftok: Discrete Visual Tokens of Autoregression, by Diffusion, and for Reasoning Treble Counterfactual VLMs: A Causal Approach to Hallucination
Reference 61
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 782a4c0a-4cfa-4dc9-a351-52f26267352f · outbound
Selftok: Discrete Visual Tokens of Autoregression, by Diffusion, and for Reasoning Autoregressive image generation without vector quantization
Reference 62
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2c8f42cf-5769-45ae-b5ad-a831411f96a8 · outbound
Selftok: Discrete Visual Tokens of Autoregression, by Diffusion, and for Reasoning ImageFolder: Autoregressive Image Generation with Folded Tokens
Reference 63
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f1d1ac06-1cbe-4fe1-b592-baa21ad21ed5 · outbound
Selftok: Discrete Visual Tokens of Autoregression, by Diffusion, and for Reasoning Dual Diffusion for Unified Image Generation and Understanding
Reference 64
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c3331cd3-151f-432c-8772-f813f68b1b87 · outbound
Selftok: Discrete Visual Tokens of Autoregression, by Diffusion, and for Reasoning Reasoning physical video generation with diffusion timestep tokens via reinforcement learning
Reference 65
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 650a9e5d-0bb0-4bc0-91cc-02f0d8a87581 · outbound
Selftok: Discrete Visual Tokens of Autoregression, by Diffusion, and for Reasoning Improved baselines with visual instruction tuning
Reference 66
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 38e8b85e-d58d-45d3-974e-de7f82eeba7d · outbound
Selftok: Discrete Visual Tokens of Autoregression, by Diffusion, and for Reasoning World Model on Million-Length Video And Language With Blockwise RingAttention
Reference 67
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d0a3c5b7-16bf-496b-9f39-89daf0ce0226 · outbound
Selftok: Discrete Visual Tokens of Autoregression, by Diffusion, and for Reasoning Flow straight and fast: Learning to generate and transfer data with rectified flow
Reference 68
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a8bb7a61-f1f1-4cb4-98b1-8dbc6d12f36e · outbound
Selftok: Discrete Visual Tokens of Autoregression, by Diffusion, and for Reasoning Challenging common assumptions in the unsupervised learning of disentangled representations
Reference 69
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 130c5350-07d0-4c92-861f-bccfcfe3b785 · outbound
Selftok: Discrete Visual Tokens of Autoregression, by Diffusion, and for Reasoning IconQA: A New Benchmark for Abstract Diagram Understanding and Visual Language Reasoning
Reference 70
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation acedce9a-1579-48bb-9fe1-06eb4d750e88 · outbound
Selftok: Discrete Visual Tokens of Autoregression, by Diffusion, and for Reasoning Towards Understanding Text Hallucination of Diffusion Models via Local Generation Bias
Reference 71
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4b433d4f-0686-4bbb-b350-b52ac039d73f · outbound
Selftok: Discrete Visual Tokens of Autoregression, by Diffusion, and for Reasoning Open-MAGVIT2: An Open-Source Project Toward Democratizing Auto-regressive Visual Generation
Reference 72
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8d68e5e7-dcc4-4377-9907-c7a83d748fe4 · outbound
Selftok: Discrete Visual Tokens of Autoregression, by Diffusion, and for Reasoning JanusFlow: Harmonizing Autoregression and Rectified Flow for Unified Multimodal Understanding and Generation
Reference 73
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3c96c7b6-7fbe-49f9-b0dd-369b6adf1f90 · outbound
Selftok: Discrete Visual Tokens of Autoregression, by Diffusion, and for Reasoning Ok-vqa: A visual question answering benchmark requiring external knowledge
Reference 74
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a91bb8fd-1208-4e09-bbe6-003023b3a0f2 · outbound
Selftok: Discrete Visual Tokens of Autoregression, by Diffusion, and for Reasoning Harnessing Discrete Representations For Continual Reinforcement Learning
Reference 75
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 295121ef-bf11-45ed-baad-4c4c7b2cd27b · outbound
Selftok: Discrete Visual Tokens of Autoregression, by Diffusion, and for Reasoning Ocr-vqa: Visual question answering by reading text in images
Reference 76
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 73b3651f-b484-44e8-95f0-3b96b46efae4 · outbound
Selftok: Discrete Visual Tokens of Autoregression, by Diffusion, and for Reasoning One-D-Piece: Image Tokenizer Meets Quality-Controllable Compression
Reference 77
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e4a2fde6-0611-4cbf-840f-c6c847e03f39 · outbound
Selftok: Discrete Visual Tokens of Autoregression, by Diffusion, and for Reasoning Null-text inversion for editing real images using guided diffusion models
Reference 78
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 14eed918-d081-43cb-9f97-983475a75d0f · outbound
Selftok: Discrete Visual Tokens of Autoregression, by Diffusion, and for Reasoning EditAR: Unified Conditional Generation with Autoregressive Models
Reference 79
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b9802df9-2822-45ff-bc41-d9f2cc0bf94d · outbound
Selftok: Discrete Visual Tokens of Autoregression, by Diffusion, and for Reasoning Introspective distillation for robust question answering
Reference 80
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 78feb04b-55de-456d-84a0-974af26d6921 · outbound
Selftok: Discrete Visual Tokens of Autoregression, by Diffusion, and for Reasoning Introducing 4o image generation
Reference 81
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 38130b36-fe3e-456e-af8f-8028d8c4f217 · outbound
Selftok: Discrete Visual Tokens of Autoregression, by Diffusion, and for Reasoning Vr-sampling: Accelerating flow generative model training with variance reduction sampling
Reference 82
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 187a22fe-1b94-4757-8902-7af3cbf38d28 · outbound
Selftok: Discrete Visual Tokens of Autoregression, by Diffusion, and for Reasoning Generative mutimodal pretraining with discrete diffusion timestep tokens
Reference 83
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 89ae4d76-8d74-4db0-9dfb-083496ba9fc5 · outbound
Selftok: Discrete Visual Tokens of Autoregression, by Diffusion, and for Reasoning Auto-encoding morph-tokens for multimodal llm
Reference 84
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4e8b4275-9989-4de7-9c77-75cbc84b8e59 · outbound
Selftok: Discrete Visual Tokens of Autoregression, by Diffusion, and for Reasoning Zero-shot image-to-image translation
Reference 85
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 77b23f67-bd63-4b4d-a0ee-b9dcbd2e5b57 · outbound
Selftok: Discrete Visual Tokens of Autoregression, by Diffusion, and for Reasoning Causality: Models, Reasoning, and Inference
Reference 86
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bb65c5c8-00e3-4078-a0d2-df32dae0e28e · outbound
Selftok: Discrete Visual Tokens of Autoregression, by Diffusion, and for Reasoning Skywork R1V: Pioneering Multimodal Reasoning with Chain-of-Thought
Reference 87
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 02149c5b-f559-4752-ac53-ef88093378a9 · outbound
Selftok: Discrete Visual Tokens of Autoregression, by Diffusion, and for Reasoning SDXL: Improving Latent Diffusion Models for High-Resolution Image Synthesis
Reference 88
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2356c3ea-3ece-449c-b702-cc2e7d541be6 · outbound
Selftok: Discrete Visual Tokens of Autoregression, by Diffusion, and for Reasoning TokenFlow: Unified Image Tokenizer for Multimodal Understanding and Generation
Reference 89
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 802ec8bc-4b74-460d-ae0c-f15b9c2b12b9 · outbound
Selftok: Discrete Visual Tokens of Autoregression, by Diffusion, and for Reasoning Learning transferable visual models from natural language supervision
Reference 90
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 75424c96-1a2c-4633-8e32-dafe9d1d6321 · outbound
Selftok: Discrete Visual Tokens of Autoregression, by Diffusion, and for Reasoning Improving language understanding by generative pre-training
Reference 91
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f2497048-f8a6-4268-89ee-4bcd31d7fed8 · outbound
Selftok: Discrete Visual Tokens of Autoregression, by Diffusion, and for Reasoning Direct preference optimization: Your language model is secretly a reward model
Reference 92
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 83a59e16-9733-4722-905a-dd575b9d595b · outbound
Selftok: Discrete Visual Tokens of Autoregression, by Diffusion, and for Reasoning Flow to the mode: Mode-seeking diffusion autoencoders for state-of-the-art image tokenization
Reference 93
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cb2ef13d-a3dd-4eca-83a2-c05ece23b561 · outbound
Selftok: Discrete Visual Tokens of Autoregression, by Diffusion, and for Reasoning Laion-5b: An open large-scale dataset for training next generation image-text models
Reference 94
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e6e89895-4e7d-4fca-baac-551e9fb179e6 · outbound
Selftok: Discrete Visual Tokens of Autoregression, by Diffusion, and for Reasoning A-okvqa: A benchmark for visual question answering using world knowledge
Reference 95
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e41c0708-7ad9-43ed-ba2d-1d4e640e7c82 · outbound
Selftok: Discrete Visual Tokens of Autoregression, by Diffusion, and for Reasoning Neural machine translation of rare words with subword units
Reference 96
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0cbe20a1-c749-408a-b1f0-27a512be7887 · outbound
Selftok: Discrete Visual Tokens of Autoregression, by Diffusion, and for Reasoning DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
Reference 97
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2955a866-7621-422d-b9f1-339861e7bc32 · outbound
Selftok: Discrete Visual Tokens of Autoregression, by Diffusion, and for Reasoning VLM-R1: A Stable and Generalizable R1-style Large Vision-Language Model
Reference 98
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 907a28fd-2965-4100-888f-d76998e92674 · outbound
Selftok: Discrete Visual Tokens of Autoregression, by Diffusion, and for Reasoning SeedEdit: Align Image Re-Generation to Image Editing
Reference 99
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 547786c1-c51d-4328-a325-36cbdf87e25f · outbound
Selftok: Discrete Visual Tokens of Autoregression, by Diffusion, and for Reasoning Improving Image Captioning with Better Use of Captions
Reference 100
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 03ed1463-9d7e-49f5-bd31-d818ef5b29e3 · inbound
IRBridge: Solving Image Restoration Bridge with Pre-trained Generative Diffusion Models Selftok: Discrete Visual Tokens of Autoregression, by Diffusion, and for Reasoning
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5a97a6a8-26bc-4f56-9a80-f7f0ead67f4b · inbound
UniCode$^2$: Cascaded Large-scale Codebooks for Unified Multimodal Understanding and Generation Selftok: Discrete Visual Tokens of Autoregression, by Diffusion, and for Reasoning
Reference 59
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 37013941-5f88-47bf-bfa3-6ab407143bea · inbound
TAP: Parameter-efficient Task-Aware Prompting for Adverse Weather Removal Selftok: Discrete Visual Tokens of Autoregression, by Diffusion, and for Reasoning
Reference 63
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 264d98ea-86c3-49f9-a435-05e7ea5f30e3 · inbound
(1D) Ordered Tokens Enable Efficient Test-Time Search Selftok: Discrete Visual Tokens of Autoregression, by Diffusion, and for Reasoning
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation bea56d38-c42d-462f-8a86-8972860f0c02 · inbound
Autoregressive Visual Generation Needs a Prologue Selftok: Discrete Visual Tokens of Autoregression, by Diffusion, and for Reasoning
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation a17bca25-1c3b-4e72-a6e5-dc9db4f3ed48 · inbound
Autoregressive Visual Generation Needs a Prologue Selftok: Discrete Visual Tokens of Autoregression, by Diffusion, and for Reasoning
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 4cb9248f-9406-4295-ad51-3e7438e6479f · inbound
HYDRA-X: Native Unified Multimodal Models with Holistic Visual Tokenizers Selftok: Discrete Visual Tokens of Autoregression, by Diffusion, and for Reasoning
Reference 83
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 0e76332c-b599-451e-b719-87cdbef12fb5 · inbound
Your Data Manifold is Secretly a Reward Model: Shell-LCC for Text-to-Video Generation Selftok: Discrete Visual Tokens of Autoregression, by Diffusion, and for Reasoning
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 7d076275-a792-43b0-b30f-23238cb1ce66 · inbound
Your Data Manifold is Secretly a Reward Model: Shell-LCC for Text-to-Video Generation Selftok: Discrete Visual Tokens of Autoregression, by Diffusion, and for Reasoning
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2f0020f5-58c2-4cf9-9e0e-596ba65d3cd1 · inbound
Twins: Learn to Predict Unified Representations with Focal Loss Selftok: Discrete Visual Tokens of Autoregression, by Diffusion, and for Reasoning
Reference 89
Source-reported events for the cited work
Unavailable: canonical work link unavailable.