Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-06-27T01:09:02.214590Z
Paper Citation Record · LEDGER
As of 5 August 2026, this Paper Citation Record lists 68 of 68 outbound references and 3 inbound Pith citation observations for arXiv:2606.17800.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-06-27T01:09:02.214590Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-05T06:32:48.257954+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-07-31T23:18:25.056705Z
A source-named dated measurement, never combined with another source.
Source: cited_works
68 of 68 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 75c27e2e-1861-4d1f-9fdb-4907178bbdad · outbound
MaineCoon: Pursuing A Real-Time Audio-Visual Social World Model On-policy distillation of language models: Learning from self-generated mistakes
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cb1b7e1f-1170-42a3-8547-6cbfcaad9416 · outbound
MaineCoon: Pursuing A Real-Time Audio-Visual Social World Model V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation bc1d0363-c7a1-407f-bb60-0e4ba5766704 · outbound
MaineCoon: Pursuing A Real-Time Audio-Visual Social World Model Optimizing few-step generation with adaptive matching distillation
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a1091932-944a-459f-bf77-f2fbfb17fd93 · outbound
MaineCoon: Pursuing A Real-Time Audio-Visual Social World Model Z-Image: An Efficient Image Generation Foundation Model with Single-Stream Diffusion Transformer
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation f04f1341-9fa0-4264-9662-34c9a5c9b0e2 · outbound
MaineCoon: Pursuing A Real-Time Audio-Visual Social World Model Q-dit: Accurate post-training quantization for diffusion transformers
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0e88e013-b541-46d2-af32-c39862a7bdfd · outbound
MaineCoon: Pursuing A Real-Time Audio-Visual Social World Model Out of time: automated lip sync in the wild
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation de2ffeb6-9f00-4801-9fb5-7bf50e4d538e · outbound
MaineCoon: Pursuing A Real-Time Audio-Visual Social World Model Deepseek-v4: Towards highly efficient million-token context intelligence, 2026
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c9033706-c690-4373-8671-5c78a9042a14 · outbound
MaineCoon: Pursuing A Real-Time Audio-Visual Social World Model Music Source Separation in the Waveform Domain
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation a8f6dc79-f9db-49f2-b0a0-efad0d2b543e · outbound
MaineCoon: Pursuing A Real-Time Audio-Visual Social World Model 8-bit Optimizers via Block-wise Quantization
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 1c95f4e5-5b87-4a9f-b90e-da4141c7b2fa · outbound
MaineCoon: Pursuing A Real-Time Audio-Visual Social World Model Scaling rectified flow transformers for high-resolution image synthesis
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1d529a82-51c7-4862-a7de-4ce607f1a83e · outbound
MaineCoon: Pursuing A Real-Time Audio-Visual Social World Model Revisiting On-Policy Distillation: Empirical Failure Modes and Simple Fixes
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 2e512e04-de91-41ed-a648-8ed6d77fee6d · outbound
MaineCoon: Pursuing A Real-Time Audio-Visual Social World Model Gemma 4 technical report
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0333484b-48bd-4a34-bc88-4fc2f1d6c2cd · outbound
MaineCoon: Pursuing A Real-Time Audio-Visual Social World Model Sample and computation redistribution for efficient face detection
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 48bc9efe-aa3b-4cd4-9666-f731eca95f54 · outbound
MaineCoon: Pursuing A Real-Time Audio-Visual Social World Model End-to-End Training for Autoregressive Video Diffusion via Self-Resampling
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 26e2f9da-c686-4f28-9605-e9788da2feac · outbound
MaineCoon: Pursuing A Real-Time Audio-Visual Social World Model LTX-2: Efficient Joint Audio-Visual Foundation Model
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation a05d971b-789d-4742-8f82-390a37df3c99 · outbound
MaineCoon: Pursuing A Real-Time Audio-Visual Social World Model CogVideo: Large-scale Pretraining for Text-to-Video Generation via Transformers
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 009317f0-85ea-47ae-acd8-d7a4a9cd9a34 · outbound
MaineCoon: Pursuing A Real-Time Audio-Visual Social World Model Self forcing: Bridging the train-test gap in autoregressive video diffusion.Advancesin Neural Information Processing Systems, 38:167283–167308, 2026
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d644cd62-1d19-40a8-b50e-07e65f48a20f · outbound
MaineCoon: Pursuing A Real-Time Audio-Visual Social World Model Live Avatar: Streaming Real-time Audio-Driven Avatar Generation with Infinite Length
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation e4a9914f-049c-428d-892d-d20e9821ebef · outbound
MaineCoon: Pursuing A Real-Time Audio-Visual Social World Model Vbench: Comprehensive benchmark suite for video generative models
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 091d319d-7938-4772-86fb-a18d882f950c · outbound
MaineCoon: Pursuing A Real-Time Audio-Visual Social World Model EasyOCR: Ready-to-use OCR with 80+ supported languages, 2020
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 878d4231-038b-4749-8217-d0e2e158bf9d · outbound
MaineCoon: Pursuing A Real-Time Audio-Visual Social World Model YOLOv11: An Overview of the Key Architectural Enhancements
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 491f773d-d92d-43d7-9962-6c495bff7273 · outbound
MaineCoon: Pursuing A Real-Time Audio-Visual Social World Model Adam: A Method for Stochastic Optimization
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 0fefe8df-5672-49e5-90b2-3b5add8a1019 · outbound
MaineCoon: Pursuing A Real-Time Audio-Visual Social World Model Pisa: Piecewise sparse attention is wiser for efficient diffusion transformers
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 6a98142c-d505-444b-9493-adab73c25454 · outbound
MaineCoon: Pursuing A Real-Time Audio-Visual Social World Model Joyai-echo: Pushing the frontier of long audio-visual generation
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c5880588-4c6a-4a29-8c18-babfafbf871d · outbound
MaineCoon: Pursuing A Real-Time Audio-Visual Social World Model DiffusionOPD: A Unified Perspective of On-Policy Distillation in Diffusion Models
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation aebee43d-4aa6-4ddc-986f-f15f8ce2d036 · outbound
MaineCoon: Pursuing A Real-Time Audio-Visual Social World Model Alignment of diffusion models: Fundamentals, challenges, and future.ACM Computing Surveys, 58(9):1–37, 2026
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9e8b5f29-9729-47aa-b534-0ad718dd6260 · outbound
MaineCoon: Pursuing A Real-Time Audio-Visual Social World Model Timestep embedding tells: It’s time to cache for video diffusion model
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d0b0c9e9-1d3d-445c-9102-a984054956e4 · outbound
MaineCoon: Pursuing A Real-Time Audio-Visual Social World Model Flow-grpo: Training flow matching models via online rl.Advances in neural information processing systems, 38:40783–40818, 2026
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f1610008-edf4-404e-b4b4-bc6cea6525ec · outbound
MaineCoon: Pursuing A Real-Time Audio-Visual Social World Model JavisDiT: Joint audio-video diffusion transformer with hierarchical spatio-temporal prior synchronization
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 60f84432-e8de-45cf-aa1a-e99349ae27f1 · outbound
MaineCoon: Pursuing A Real-Time Audio-Visual Social World Model Ilya Loshchilov and Frank Hutter
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 035dbc88-a657-4a14-9e78-4c3926a71000 · outbound
MaineCoon: Pursuing A Real-Time Audio-Visual Social World Model Rolling Forcing: Autoregressive Long Video Diffusion in Real Time
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 70ed78ee-6b3f-40d3-a518-924dd840dc03 · outbound
MaineCoon: Pursuing A Real-Time Audio-Visual Social World Model Decoupled weight decay regularization
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1ec2a4f5-740c-462b-a6d6-27fdab4168d8 · outbound
MaineCoon: Pursuing A Real-Time Audio-Visual Social World Model Ovi: Twin Backbone Cross-Modal Fusion for Audio-Video Generation
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 4125ae60-ca8a-4f86-aa00-8cdb46688c59 · outbound
MaineCoon: Pursuing A Real-Time Audio-Visual Social World Model Krea realtime 14b: Real-time video generation, 2025
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1a6e5ca9-1d58-45b9-8a21-7bbae1ebf267 · outbound
MaineCoon: Pursuing A Real-Time Audio-Visual Social World Model Robust speech recognition via large-scale weak supervision
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 20e6b02d-a125-4907-bf4b-56e732136e98 · outbound
MaineCoon: Pursuing A Real-Time Audio-Visual Social World Model MagicDistillation: Weak-to-Strong Video Distillation for Large-Scale Few-Step Synthesis
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation e670727a-1949-46f3-b617-4c777b83d4e4 · outbound
MaineCoon: Pursuing A Real-Time Audio-Visual Social World Model Efficient Video Diffusion Models: Advancements and Challenges
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 4dd845c9-dc2b-411f-9d80-ee4e96e18c7d · outbound
MaineCoon: Pursuing A Real-Time Audio-Visual Social World Model Fastlightgen: Fast and light video generation with fewer steps and parameters
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 382c8f9c-95c6-4f64-bced-79e63b815a04 · outbound
MaineCoon: Pursuing A Real-Time Audio-Visual Social World Model Liveditor-14b: Lightning unified video editing via in-context sparse attention
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation abd2b0b6-0f17-4010-a688-107281f5edd4 · outbound
MaineCoon: Pursuing A Real-Time Audio-Visual Social World Model Silero VAD: pre-trained enterprise-grade voice activity detector, 2024
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dc4d459e-c848-4167-99c6-bccd753b2a3f · outbound
MaineCoon: Pursuing A Real-Time Audio-Visual Social World Model A Survey of On-Policy Distillation for Large Language Models
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 73c0f837-14a9-40fd-995e-b5d72a304157 · outbound
MaineCoon: Pursuing A Real-Time Audio-Visual Social World Model TransNet V2: An effective deep network architecture for fast shot transition detection
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 088ea716-3d70-4593-8868-ba82ab46e464 · outbound
MaineCoon: Pursuing A Real-Time Audio-Visual Social World Model Omniforcing: Unleashing real-time joint audio-visual generation
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation bfc2bc9a-e6a9-4949-a7ce-da972c6b29fd · outbound
MaineCoon: Pursuing A Real-Time Audio-Visual Social World Model Hunyuan-gamecraft-2: Instruction-following interactive game world model
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation bde9766a-a6fe-4e29-a38f-d7faac39b441 · outbound
MaineCoon: Pursuing A Real-Time Audio-Visual Social World Model Gemini: A Family of Highly Capable Multimodal Models
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation b58486d8-8e1c-4d07-b2d8-70f5cb60de2f · outbound
MaineCoon: Pursuing A Real-Time Audio-Visual Social World Model Mova: Towards scalable and synchronized video-audio generation
Reference 47
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 1f136815-40fe-4e67-89da-ddc3b2d5acf7 · outbound
MaineCoon: Pursuing A Real-Time Audio-Visual Social World Model Diffusion model alignment using direct preference optimization
Reference 48
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6f371491-9900-4d09-91f9-f309d81ce86c · outbound
MaineCoon: Pursuing A Real-Time Audio-Visual Social World Model Wan: Open and Advanced Large-Scale Video Generative Models
Reference 49
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation a1900dbd-2862-4953-90ca-7f747d6346cb · outbound
MaineCoon: Pursuing A Real-Time Audio-Visual Social World Model HunyuanVideo 1.5 Technical Report
Reference 50
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 9fa2a114-21ab-4175-b7f8-480f72ab12cb · outbound
MaineCoon: Pursuing A Real-Time Audio-Visual Social World Model Qwen-Image Technical Report
Reference 51
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation cfd8da29-7b24-4233-9b19-7a127a3b6f40 · outbound
MaineCoon: Pursuing A Real-Time Audio-Visual Social World Model Longvideobench: A benchmark for long-context interleaved video-language understanding
Reference 52
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ccfe6001-d7b6-4452-a446-b7504c6998a8 · outbound
MaineCoon: Pursuing A Real-Time Audio-Visual Social World Model LongLive: Real-time Interactive Long Video Generation
Reference 53
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 9b55799b-a8e6-4398-9b97-70f9563e9d45 · outbound
MaineCoon: Pursuing A Real-Time Audio-Visual Social World Model Sparse videogen2: Accelerate video generation with sparse attention via semantic-aware permutation
Reference 54
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 87466778-c3a0-4d69-93cf-dda94dcc9c0d · outbound
MaineCoon: Pursuing A Real-Time Audio-Visual Social World Model Improved distribution matching distillation for fast image synthesis.Advancesin neural information processing systems, 37:47455–47487, 2024
Reference 55
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d3e7385c-6734-4025-b70e-ab9d31c3cce4 · outbound
MaineCoon: Pursuing A Real-Time Audio-Visual Social World Model One-step diffusion with distribution matching distillation
Reference 56
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d325a953-9702-456d-95e9-581c5f88c07b · outbound
MaineCoon: Pursuing A Real-Time Audio-Visual Social World Model From slow bidirectional to fast autoregressive video diffusion models
Reference 57
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 47a33905-e682-4d50-8ee9-3217b0333256 · outbound
MaineCoon: Pursuing A Real-Time Audio-Visual Social World Model Representation Alignment for Generation: Training Diffusion Transformers Is Easier Than You Think
Reference 58
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 11538768-8a02-43ff-b5c1-8d920cbdc22d · outbound
MaineCoon: Pursuing A Real-Time Audio-Visual Social World Model Soulx-flashtalk: Real-time infinite streaming of audio-driven avatars via self-correcting bidirectional distillation
Reference 59
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 12d65a8f-9490-460a-b98a-c693b6f49064 · outbound
MaineCoon: Pursuing A Real-Time Audio-Visual Social World Model Improved distribution matching distillation for fast image synthesis.Advances in neural information processing systems, 37:47455–47487, 2024a
Reference 60
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation a0500ec4-14d0-456f-b477-7315172c247a · outbound
MaineCoon: Pursuing A Real-Time Audio-Visual Social World Model Fast-WAM: Do World Action Models Need Test-time Future Imagination?
Reference 61
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation b4f63ff8-0c60-4d4c-897c-7c591a3a569c · outbound
MaineCoon: Pursuing A Real-Time Audio-Visual Social World Model Sigmoid loss for language image pre-training
Reference 62
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b0aea86a-5ca9-4b46-b8a0-b25e4fdda32c · outbound
MaineCoon: Pursuing A Real-Time Audio-Visual Social World Model Spargeattention: Accurate and training-free sparse attention accelerating any model inference
Reference 63
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 057113b3-6a28-4199-be9a-8b43b4710ac3 · outbound
MaineCoon: Pursuing A Real-Time Audio-Visual Social World Model Faster video diffusion with trainable sparse attention.Advances in Neural Information Processing Systems, 38: 152509–152534, 2026
Reference 64
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fa4e91e3-0bc4-46b1-9ce4-87de4c77cc3e · outbound
MaineCoon: Pursuing A Real-Time Audio-Visual Social World Model VideoREPA: Learning Physics for Video Generation through Relational Alignment with Foundation Models
Reference 65
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation ac6aa09c-3e61-43e6-a0df-5c65c2dffeb4 · outbound
MaineCoon: Pursuing A Real-Time Audio-Visual Social World Model ViDiT-Q: Efficient and Accurate Quantization of Diffusion Transformers for Image and Video Generation
Reference 66
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 0b5ccfd4-ed3d-4732-9302-09a5573ec089 · outbound
MaineCoon: Pursuing A Real-Time Audio-Visual Social World Model DiffusionNFT: Online Diffusion Reinforcement with Forward Process
Reference 67
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 946ddd3c-d0e1-4fb5-bb4b-c4b666eede11 · outbound
MaineCoon: Pursuing A Real-Time Audio-Visual Social World Model SANA-WM: Efficient Minute-Scale World Modeling with Hybrid Linear Diffusion Transformer
Reference 68
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation ee396425-4902-478c-a771-e8d737436aaa · outbound
MaineCoon: Pursuing A Real-Time Audio-Visual Social World Model Causal Forcing: Autoregressive Diffusion Distillation Done Right for High-Quality Real-Time Interactive Video Generation
Reference 69
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation e7c80bb1-33a2-44df-80f1-2c273baae2ec · inbound
Perceptual Flow Matching for Few-Step Generative Modeling MaineCoon: Pursuing A Real-Time Audio-Visual Social World Model
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation de100b1b-32d7-4f31-953b-d86a1aa7ae46 · inbound
Sol-Attn: Accelerating Video Generation Inference via On-the-Fly Attention Sparsification MaineCoon: Pursuing A Real-Time Audio-Visual Social World Model
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f3234e59-a1b2-4445-9564-286c13ffc4a2 · inbound
TaoMate: Anchor-Guided Memory Bridging Evolving and Reference States for Real-Time Audio-Video Digital Human Generation MaineCoon: Pursuing A Real-Time Audio-Visual Social World Model
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.