Pith. sign in

Paper Citation Record · LEDGER

MaineCoon: Pursuing A Real-Time Audio-Visual Social World Model

As of 5 August 2026, this Paper Citation Record lists 68 of 68 outbound references and 3 inbound Pith citation observations for arXiv:2606.17800.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2606.17800 v1

Coverage vector

measured 68 of 68 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-06-27T01:09:02.214590Z

measured 71 of 71 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-05T06:32:48.257954+00:00

measured 3 of 3 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-07-31T23:18:25.056705Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

68 of 68 outbound references displayed

  • verified exact33
  • verified fuzzy0
  • unresolved31
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch4

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 75c27e2e-1861-4d1f-9fdb-4907178bbdad · outbound

This paper cites On-policy distillation of language models: Learning from self-generated mistakes.

MaineCoon: Pursuing A Real-Time Audio-Visual Social World Model On-policy distillation of language models: Learning from self-generated mistakes

Reference 1

Resolution
unresolved
no resolver link, observed 2026-06-27T01:09:02.214590Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T01:09:02.214590Z digest=sha256:8f908a91207a9e2503547d02fabceb4c6281bdde7e6cb70102713e325390bff8

Observation cb1b7e1f-1170-42a3-8547-6cbfcaad9416 · outbound

This paper cites V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning.

MaineCoon: Pursuing A Real-Time Audio-Visual Social World Model V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning

Reference 3

Resolution
verified exact
local_arxiv, observed 2026-07-03T20:48:55.726546Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-27T01:09:02.214590Z digest=sha256:a195b4e4ead9bffbfe0bef58ad888181e20d38f4ab8e435a1af09c5ffe3575e8

Observation bc1d0363-c7a1-407f-bb60-0e4ba5766704 · outbound

This paper cites Optimizing few-step generation with adaptive matching distillation.

MaineCoon: Pursuing A Real-Time Audio-Visual Social World Model Optimizing few-step generation with adaptive matching distillation

Reference 4

Resolution
unresolved
no resolver link, observed 2026-06-27T01:09:02.214590Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T01:09:02.214590Z digest=sha256:962a1009d4973fc8954f1daf1ec286a321ee79bc3956a9f8d42bf6609fed8ffe

Observation a1091932-944a-459f-bf77-f2fbfb17fd93 · outbound

This paper cites Z-Image: An Efficient Image Generation Foundation Model with Single-Stream Diffusion Transformer.

MaineCoon: Pursuing A Real-Time Audio-Visual Social World Model Z-Image: An Efficient Image Generation Foundation Model with Single-Stream Diffusion Transformer

Reference 5

Resolution
verified exact
local_arxiv, observed 2026-07-03T20:48:55.730050Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-27T01:09:02.214590Z digest=sha256:f567db34661f70924fbd2f7245f51e12875c4319678a3331dc5629db697910a9

Observation f04f1341-9fa0-4264-9662-34c9a5c9b0e2 · outbound

This paper cites Q-dit: Accurate post-training quantization for diffusion transformers.

MaineCoon: Pursuing A Real-Time Audio-Visual Social World Model Q-dit: Accurate post-training quantization for diffusion transformers

Reference 6

Resolution
unresolved
no resolver link, observed 2026-06-27T01:09:02.214590Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T01:09:02.214590Z digest=sha256:078111c6fb5804db975ee0cc85273b71dc6e6571550cfb9236b2bcebec435158

Observation 0e88e013-b541-46d2-af32-c39862a7bdfd · outbound

This paper cites Out of time: automated lip sync in the wild.

MaineCoon: Pursuing A Real-Time Audio-Visual Social World Model Out of time: automated lip sync in the wild

Reference 7

Resolution
unresolved
no resolver link, observed 2026-06-27T01:09:02.214590Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T01:09:02.214590Z digest=sha256:1dcb521e40588221879d7f4aceb5f62f0d7fbbd03c4a006b23c8a94e05a27e5a

Observation de2ffeb6-9f00-4801-9fb5-7bf50e4d538e · outbound

This paper cites Deepseek-v4: Towards highly efficient million-token context intelligence, 2026.

MaineCoon: Pursuing A Real-Time Audio-Visual Social World Model Deepseek-v4: Towards highly efficient million-token context intelligence, 2026

Reference 8

Resolution
unresolved
no resolver link, observed 2026-06-27T01:09:02.214590Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T01:09:02.214590Z digest=sha256:ba0d26eaa2c5c28c2d6983a85f31893e89bba1aaf0be54a82319b262003152c6

Observation c9033706-c690-4373-8671-5c78a9042a14 · outbound

This paper cites Music Source Separation in the Waveform Domain.

MaineCoon: Pursuing A Real-Time Audio-Visual Social World Model Music Source Separation in the Waveform Domain

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-07-03T20:48:55.750717Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-27T01:09:02.214590Z digest=sha256:496c2462b4fde37b064c35de1ee017f8e0c0f97cc9583dc454e2f5b18d3dccf1

Observation a8f6dc79-f9db-49f2-b0a0-efad0d2b543e · outbound

This paper cites 8-bit Optimizers via Block-wise Quantization.

MaineCoon: Pursuing A Real-Time Audio-Visual Social World Model 8-bit Optimizers via Block-wise Quantization

Reference 10

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T20:48:55.723997Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-27T01:09:02.214590Z digest=sha256:a57e441a1ee3821c9e14700b5523f292afc1a8611c53c78299f492af5da06e5b

Observation 1c95f4e5-5b87-4a9f-b90e-da4141c7b2fa · outbound

This paper cites Scaling rectified flow transformers for high-resolution image synthesis.

MaineCoon: Pursuing A Real-Time Audio-Visual Social World Model Scaling rectified flow transformers for high-resolution image synthesis

Reference 11

Resolution
unresolved
no resolver link, observed 2026-06-27T01:09:02.214590Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T01:09:02.214590Z digest=sha256:f462ef9940537f52d7cec19d42cfac4e2758276d41b7f75fc8c73575e2749545

Observation 1d529a82-51c7-4862-a7de-4ce607f1a83e · outbound

This paper cites Revisiting On-Policy Distillation: Empirical Failure Modes and Simple Fixes.

MaineCoon: Pursuing A Real-Time Audio-Visual Social World Model Revisiting On-Policy Distillation: Empirical Failure Modes and Simple Fixes

Reference 12

Resolution
verified exact
local_arxiv, observed 2026-07-03T20:48:55.716162Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-27T01:09:02.214590Z digest=sha256:05b7468260b086a03c2247d723bc7990db9fe2f1ab7a980b192e94b49e3402a1

Observation 2e512e04-de91-41ed-a648-8ed6d77fee6d · outbound

This paper cites Gemma 4 technical report.

MaineCoon: Pursuing A Real-Time Audio-Visual Social World Model Gemma 4 technical report

Reference 13

Resolution
unresolved
no resolver link, observed 2026-06-27T01:09:02.214590Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T01:09:02.214590Z digest=sha256:d65d815c76cc70297883f3866cf08609052fa2c8107e319b734129815d1b9b53

Observation 0333484b-48bd-4a34-bc88-4fc2f1d6c2cd · outbound

This paper cites Sample and computation redistribution for efficient face detection.

MaineCoon: Pursuing A Real-Time Audio-Visual Social World Model Sample and computation redistribution for efficient face detection

Reference 14

Resolution
unresolved
no resolver link, observed 2026-06-27T01:09:02.214590Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T01:09:02.214590Z digest=sha256:d6d6fac52fa01cda5e6f217f036328ec007a74e408bfd48fc86ee50bc0b5adf5

Observation 48bc9efe-aa3b-4cd4-9666-f731eca95f54 · outbound

This paper cites End-to-End Training for Autoregressive Video Diffusion via Self-Resampling.

MaineCoon: Pursuing A Real-Time Audio-Visual Social World Model End-to-End Training for Autoregressive Video Diffusion via Self-Resampling

Reference 15

Resolution
verified exact
local_arxiv, observed 2026-07-03T20:48:55.713615Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-27T01:09:02.214590Z digest=sha256:77e76cb02475ec4f0a37335c22105d3bf3ce7a26ebfb14d30b837f7a3c992fb1

Observation 26e2f9da-c686-4f28-9605-e9788da2feac · outbound

This paper cites LTX-2: Efficient Joint Audio-Visual Foundation Model.

MaineCoon: Pursuing A Real-Time Audio-Visual Social World Model LTX-2: Efficient Joint Audio-Visual Foundation Model

Reference 16

Resolution
verified exact
local_arxiv, observed 2026-07-03T20:48:55.732603Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-27T01:09:02.214590Z digest=sha256:c9dd9e3350ffa338fb6cfe0e8b4d26f918fb683954113b8715c5e56cfc6527b1

Observation a05d971b-789d-4742-8f82-390a37df3c99 · outbound

This paper cites CogVideo: Large-scale Pretraining for Text-to-Video Generation via Transformers.

MaineCoon: Pursuing A Real-Time Audio-Visual Social World Model CogVideo: Large-scale Pretraining for Text-to-Video Generation via Transformers

Reference 17

Resolution
verified exact
local_arxiv, observed 2026-07-03T20:48:55.721396Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-27T01:09:02.214590Z digest=sha256:6f460c3e61a77297ac99f29c3949f0012b7d7402bafec372071f11048829b7d6

Observation 009317f0-85ea-47ae-acd8-d7a4a9cd9a34 · outbound

This paper cites Self forcing: Bridging the train-test gap in autoregressive video diffusion.Advancesin Neural Information Processing Systems, 38:167283–167308, 2026.

MaineCoon: Pursuing A Real-Time Audio-Visual Social World Model Self forcing: Bridging the train-test gap in autoregressive video diffusion.Advancesin Neural Information Processing Systems, 38:167283–167308, 2026

Reference 18

Resolution
unresolved
no resolver link, observed 2026-06-27T01:09:02.214590Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T01:09:02.214590Z digest=sha256:7089a77818c94a82f15a61bdd830d3082002ccd96b5a9f6b453530346ba837db

Observation d644cd62-1d19-40a8-b50e-07e65f48a20f · outbound

This paper cites Live Avatar: Streaming Real-time Audio-Driven Avatar Generation with Infinite Length.

MaineCoon: Pursuing A Real-Time Audio-Visual Social World Model Live Avatar: Streaming Real-time Audio-Driven Avatar Generation with Infinite Length

Reference 19

Resolution
verified exact
local_arxiv, observed 2026-07-03T20:48:55.732426Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-27T01:09:02.214590Z digest=sha256:00fa1508a5a35b17f87ae9629203eb2637403ae794669375fa7b69867df4a003

Observation e4a9914f-049c-428d-892d-d20e9821ebef · outbound

This paper cites Vbench: Comprehensive benchmark suite for video generative models.

MaineCoon: Pursuing A Real-Time Audio-Visual Social World Model Vbench: Comprehensive benchmark suite for video generative models

Reference 20

Resolution
unresolved
no resolver link, observed 2026-06-27T01:09:02.214590Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T01:09:02.214590Z digest=sha256:5c9f16e49fdca523783d8a68e0f270054fd1bd284a0d20aae5172a813ffd56d1

Observation 091d319d-7938-4772-86fb-a18d882f950c · outbound

This paper cites EasyOCR: Ready-to-use OCR with 80+ supported languages, 2020.

MaineCoon: Pursuing A Real-Time Audio-Visual Social World Model EasyOCR: Ready-to-use OCR with 80+ supported languages, 2020

Reference 21

Resolution
unresolved
no resolver link, observed 2026-06-27T01:09:02.214590Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T01:09:02.214590Z digest=sha256:76da5ea081ad5e61c271f2cb75baac22849e649a63cab295a35c500d9fe99094

Observation 878d4231-038b-4749-8217-d0e2e158bf9d · outbound

This paper cites YOLOv11: An Overview of the Key Architectural Enhancements.

MaineCoon: Pursuing A Real-Time Audio-Visual Social World Model YOLOv11: An Overview of the Key Architectural Enhancements

Reference 22

Resolution
verified exact
local_arxiv, observed 2026-07-03T20:48:55.734947Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-27T01:09:02.214590Z digest=sha256:c1eee3b637f808ca9ab5bb85e67bff6f4aaa8e80dde63a6936a03055e1d33664

Observation 491f773d-d92d-43d7-9962-6c495bff7273 · outbound

This paper cites Adam: A Method for Stochastic Optimization.

MaineCoon: Pursuing A Real-Time Audio-Visual Social World Model Adam: A Method for Stochastic Optimization

Reference 23

Resolution
verified exact
local_arxiv, observed 2026-07-03T20:48:55.752888Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-27T01:09:02.214590Z digest=sha256:529d8c1d0b1f12709c5d0b360baf97f89f98d99203f3b55efc820d14a4fa7732

Observation 0fefe8df-5672-49e5-90b2-3b5add8a1019 · outbound

This paper cites Pisa: Piecewise sparse attention is wiser for efficient diffusion transformers.

MaineCoon: Pursuing A Real-Time Audio-Visual Social World Model Pisa: Piecewise sparse attention is wiser for efficient diffusion transformers

Reference 24

Resolution
verified exact
arxiv_id, observed 2026-07-03T20:48:55.714071Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-27T01:09:02.214590Z digest=sha256:9afbaf07ff9c2e5929bb322f5bb3e58bb8d8b39cdb4d35415da4e62a1898e131

Observation 6a98142c-d505-444b-9493-adab73c25454 · outbound

This paper cites Joyai-echo: Pushing the frontier of long audio-visual generation.

MaineCoon: Pursuing A Real-Time Audio-Visual Social World Model Joyai-echo: Pushing the frontier of long audio-visual generation

Reference 25

Resolution
unresolved
no resolver link, observed 2026-06-27T01:09:02.214590Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T01:09:02.214590Z digest=sha256:59cd72dd8d2c8b3c76454c25e7b3e61f386adc6bda47d34c189aa784a8130c49

Observation c5880588-4c6a-4a29-8c18-babfafbf871d · outbound

This paper cites DiffusionOPD: A Unified Perspective of On-Policy Distillation in Diffusion Models.

MaineCoon: Pursuing A Real-Time Audio-Visual Social World Model DiffusionOPD: A Unified Perspective of On-Policy Distillation in Diffusion Models

Reference 26

Resolution
verified exact
local_arxiv, observed 2026-07-03T20:48:55.694810Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-27T01:09:02.214590Z digest=sha256:86b6ce8995c7a6135c615893d2bdb473478dfecd99328e41af4e02961f04396e

Observation aebee43d-4aa6-4ddc-986f-f15f8ce2d036 · outbound

This paper cites Alignment of diffusion models: Fundamentals, challenges, and future.ACM Computing Surveys, 58(9):1–37, 2026.

MaineCoon: Pursuing A Real-Time Audio-Visual Social World Model Alignment of diffusion models: Fundamentals, challenges, and future.ACM Computing Surveys, 58(9):1–37, 2026

Reference 27

Resolution
unresolved
no resolver link, observed 2026-06-27T01:09:02.214590Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T01:09:02.214590Z digest=sha256:231990ee122624a7d20c48a8289b82c2932b5af3f2755592f296571167a3757a

Observation 9e8b5f29-9729-47aa-b534-0ad718dd6260 · outbound

This paper cites Timestep embedding tells: It’s time to cache for video diffusion model.

MaineCoon: Pursuing A Real-Time Audio-Visual Social World Model Timestep embedding tells: It’s time to cache for video diffusion model

Reference 28

Resolution
unresolved
no resolver link, observed 2026-06-27T01:09:02.214590Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T01:09:02.214590Z digest=sha256:aad26d299c16b01543061d441541329c064abc97f2ac455cb19b98ccc35fe9d4

Observation d0b0c9e9-1d3d-445c-9102-a984054956e4 · outbound

This paper cites Flow-grpo: Training flow matching models via online rl.Advances in neural information processing systems, 38:40783–40818, 2026.

MaineCoon: Pursuing A Real-Time Audio-Visual Social World Model Flow-grpo: Training flow matching models via online rl.Advances in neural information processing systems, 38:40783–40818, 2026

Reference 29

Resolution
unresolved
no resolver link, observed 2026-06-27T01:09:02.214590Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T01:09:02.214590Z digest=sha256:c6eba8bbecf915a3322422320423f2a6ac7c6fe35b0f0214dc273bdb1349f7f0

Observation f1610008-edf4-404e-b4b4-bc6cea6525ec · outbound

This paper cites JavisDiT: Joint audio-video diffusion transformer with hierarchical spatio-temporal prior synchronization.

MaineCoon: Pursuing A Real-Time Audio-Visual Social World Model JavisDiT: Joint audio-video diffusion transformer with hierarchical spatio-temporal prior synchronization

Reference 30

Resolution
verified exact
arxiv_id, observed 2026-07-03T20:48:55.697492Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-27T01:09:02.214590Z digest=sha256:a4a6440ae267c410103be5cff0d4e0447c6e2bd9327a84d2024a94917935bcc5

Observation 60f84432-e8de-45cf-aa1a-e99349ae27f1 · outbound

This paper cites Ilya Loshchilov and Frank Hutter.

MaineCoon: Pursuing A Real-Time Audio-Visual Social World Model Ilya Loshchilov and Frank Hutter

Reference 31

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T20:48:55.711049Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-27T01:09:02.214590Z digest=sha256:82051b9cf0aede8beb6aba59fb082f3153ba1999649d92160c38b35489a39aea

Observation 035dbc88-a657-4a14-9e78-4c3926a71000 · outbound

This paper cites Rolling Forcing: Autoregressive Long Video Diffusion in Real Time.

MaineCoon: Pursuing A Real-Time Audio-Visual Social World Model Rolling Forcing: Autoregressive Long Video Diffusion in Real Time

Reference 32

Resolution
verified exact
local_arxiv, observed 2026-07-03T20:48:55.692199Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-27T01:09:02.214590Z digest=sha256:c8e4a3be4123bc3f3a0f297f3335b188da823e5104b4d815634fd38e52116c2c

Observation 70ed78ee-6b3f-40d3-a518-924dd840dc03 · outbound

This paper cites Decoupled weight decay regularization.

MaineCoon: Pursuing A Real-Time Audio-Visual Social World Model Decoupled weight decay regularization

Reference 33

Resolution
unresolved
no resolver link, observed 2026-06-27T01:09:02.214590Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T01:09:02.214590Z digest=sha256:e0a2c385236cc0f1f6f771aa0bd2078763383f70062e9bcb8ac19ba992f48674

Observation 1ec2a4f5-740c-462b-a6d6-27fdab4168d8 · outbound

This paper cites Ovi: Twin Backbone Cross-Modal Fusion for Audio-Video Generation.

MaineCoon: Pursuing A Real-Time Audio-Visual Social World Model Ovi: Twin Backbone Cross-Modal Fusion for Audio-Video Generation

Reference 34

Resolution
verified exact
local_arxiv, observed 2026-07-03T20:48:55.730338Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-27T01:09:02.214590Z digest=sha256:18f3d6655c146042a9d1cedf46661cd2c7a74bee86d9fd503875164fe5815f5e

Observation 4125ae60-ca8a-4f86-aa00-8cdb46688c59 · outbound

This paper cites Krea realtime 14b: Real-time video generation, 2025.

MaineCoon: Pursuing A Real-Time Audio-Visual Social World Model Krea realtime 14b: Real-time video generation, 2025

Reference 35

Resolution
unresolved
no resolver link, observed 2026-06-27T01:09:02.214590Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T01:09:02.214590Z digest=sha256:0660f3ba36a2aba6ab53c4f21427c436f4dfdda8ff5b0eb113f7f46f27278854

Observation 1a6e5ca9-1d58-45b9-8a21-7bbae1ebf267 · outbound

This paper cites Robust speech recognition via large-scale weak supervision.

MaineCoon: Pursuing A Real-Time Audio-Visual Social World Model Robust speech recognition via large-scale weak supervision

Reference 36

Resolution
unresolved
no resolver link, observed 2026-06-27T01:09:02.214590Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T01:09:02.214590Z digest=sha256:df606f88f92f1eabe8b5aba3393981a8750ab7d616b2e9db77ef905a3fc42518

Observation 20e6b02d-a125-4907-bf4b-56e732136e98 · outbound

This paper cites MagicDistillation: Weak-to-Strong Video Distillation for Large-Scale Few-Step Synthesis.

MaineCoon: Pursuing A Real-Time Audio-Visual Social World Model MagicDistillation: Weak-to-Strong Video Distillation for Large-Scale Few-Step Synthesis

Reference 37

Resolution
verified exact
arxiv_id, observed 2026-07-03T20:48:55.722024Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-27T01:09:02.214590Z digest=sha256:f0d71e1feda5ca98952678d65c9e4bb8af5da5b8e70a7eb3ae528f5f26cd588f

Observation e670727a-1949-46f3-b617-4c777b83d4e4 · outbound

This paper cites Efficient Video Diffusion Models: Advancements and Challenges.

MaineCoon: Pursuing A Real-Time Audio-Visual Social World Model Efficient Video Diffusion Models: Advancements and Challenges

Reference 38

Resolution
verified exact
local_arxiv, observed 2026-07-03T20:48:55.745991Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-27T01:09:02.214590Z digest=sha256:e9ee3cb7d4216f4da4f400e156c519804e37871028a03ca47faaaa380001e8b3

Observation 4dd845c9-dc2b-411f-9d80-ee4e96e18c7d · outbound

This paper cites Fastlightgen: Fast and light video generation with fewer steps and parameters.

MaineCoon: Pursuing A Real-Time Audio-Visual Social World Model Fastlightgen: Fast and light video generation with fewer steps and parameters

Reference 39

Resolution
unresolved
no resolver link, observed 2026-06-27T01:09:02.214590Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T01:09:02.214590Z digest=sha256:0e8ae05db5b45547ebc6a5401787dbf5c55397c8eeb7a28fc0e527672cb6639c

Observation 382c8f9c-95c6-4f64-bced-79e63b815a04 · outbound

This paper cites Liveditor-14b: Lightning unified video editing via in-context sparse attention.

MaineCoon: Pursuing A Real-Time Audio-Visual Social World Model Liveditor-14b: Lightning unified video editing via in-context sparse attention

Reference 40

Resolution
unresolved
no resolver link, observed 2026-06-27T01:09:02.214590Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T01:09:02.214590Z digest=sha256:960cfd840c22452badc95f4ab0ef0a3f2f34e731b64176069cc680c6eba921d4

Observation abd2b0b6-0f17-4010-a688-107281f5edd4 · outbound

This paper cites Silero VAD: pre-trained enterprise-grade voice activity detector, 2024.

MaineCoon: Pursuing A Real-Time Audio-Visual Social World Model Silero VAD: pre-trained enterprise-grade voice activity detector, 2024

Reference 41

Resolution
unresolved
no resolver link, observed 2026-06-27T01:09:02.214590Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T01:09:02.214590Z digest=sha256:75d85fec6404e02fff6aa67b7e78b73c5530433da2b2f38a8f2547b28fe2a5c0

Observation dc4d459e-c848-4167-99c6-bccd753b2a3f · outbound

This paper cites A Survey of On-Policy Distillation for Large Language Models.

MaineCoon: Pursuing A Real-Time Audio-Visual Social World Model A Survey of On-Policy Distillation for Large Language Models

Reference 42

Resolution
verified exact
local_arxiv, observed 2026-07-03T20:48:55.681098Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-27T01:09:02.214590Z digest=sha256:d5abe5371119e44dec20010147efeed54abda3d768191c4871c0ec340831e74d

Observation 73c0f837-14a9-40fd-995e-b5d72a304157 · outbound

This paper cites TransNet V2: An effective deep network architecture for fast shot transition detection.

MaineCoon: Pursuing A Real-Time Audio-Visual Social World Model TransNet V2: An effective deep network architecture for fast shot transition detection

Reference 43

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T20:48:55.750550Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-27T01:09:02.214590Z digest=sha256:6d323f09bb30fd4d35f95d94032f9375d1dd4815a9dde5430723f851ff4b6706

Observation 088ea716-3d70-4593-8868-ba82ab46e464 · outbound

This paper cites Omniforcing: Unleashing real-time joint audio-visual generation.

MaineCoon: Pursuing A Real-Time Audio-Visual Social World Model Omniforcing: Unleashing real-time joint audio-visual generation

Reference 44

Resolution
verified exact
arxiv_id, observed 2026-07-03T20:48:55.755568Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-27T01:09:02.214590Z digest=sha256:1dcc2e4bf0facdfadd51ef1bb4285c29c91cb0fef5ed1bf7148d4093199664c7

Observation bfc2bc9a-e6a9-4949-a7ce-da972c6b29fd · outbound

This paper cites Hunyuan-gamecraft-2: Instruction-following interactive game world model.

MaineCoon: Pursuing A Real-Time Audio-Visual Social World Model Hunyuan-gamecraft-2: Instruction-following interactive game world model

Reference 45

Resolution
verified exact
arxiv_id, observed 2026-07-03T20:48:55.740819Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-27T01:09:02.214590Z digest=sha256:474c57e9684d25186d786097a330ef3a58586c7433acc386294a184ea9b5d16f

Observation bde9766a-a6fe-4e29-a38f-d7faac39b441 · outbound

This paper cites Gemini: A Family of Highly Capable Multimodal Models.

MaineCoon: Pursuing A Real-Time Audio-Visual Social World Model Gemini: A Family of Highly Capable Multimodal Models

Reference 46

Resolution
verified exact
local_arxiv, observed 2026-07-03T20:48:55.667113Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-27T01:09:02.214590Z digest=sha256:79524edaa56472ac85a26b3dc35eff689920bb5256ec1b4d7869722aab170ae5

Observation b58486d8-8e1c-4d07-b2d8-70f5cb60de2f · outbound

This paper cites Mova: Towards scalable and synchronized video-audio generation.

MaineCoon: Pursuing A Real-Time Audio-Visual Social World Model Mova: Towards scalable and synchronized video-audio generation

Reference 47

Resolution
verified exact
arxiv_id, observed 2026-07-03T20:48:55.743399Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-27T01:09:02.214590Z digest=sha256:d2c3d596df60a32b81aeef36e35a9132e12f491ee57c286ea479b5e41caa0b70

Observation 1f136815-40fe-4e67-89da-ddc3b2d5acf7 · outbound

This paper cites Diffusion model alignment using direct preference optimization.

MaineCoon: Pursuing A Real-Time Audio-Visual Social World Model Diffusion model alignment using direct preference optimization

Reference 48

Resolution
unresolved
no resolver link, observed 2026-06-27T01:09:02.214590Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T01:09:02.214590Z digest=sha256:d7c88dfc044fc6096125deacb75fc442e7855549f7a7534a3420973d75b042bd

Observation 6f371491-9900-4d09-91f9-f309d81ce86c · outbound

This paper cites Wan: Open and Advanced Large-Scale Video Generative Models.

MaineCoon: Pursuing A Real-Time Audio-Visual Social World Model Wan: Open and Advanced Large-Scale Video Generative Models

Reference 49

Resolution
verified exact
local_arxiv, observed 2026-07-03T20:48:55.737672Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-27T01:09:02.214590Z digest=sha256:25d71194351c4fd06a577cb168df55482111e22b95f54076cbb40907ee9b0b67

Observation a1900dbd-2862-4953-90ca-7f747d6346cb · outbound

This paper cites HunyuanVideo 1.5 Technical Report.

MaineCoon: Pursuing A Real-Time Audio-Visual Social World Model HunyuanVideo 1.5 Technical Report

Reference 50

Resolution
verified exact
local_arxiv, observed 2026-07-03T20:48:55.745718Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-27T01:09:02.214590Z digest=sha256:6b7e62265db3fd5bf6d12e4a6debbec140bc5c52419e356603c74595ac32e273

Observation 9fa2a114-21ab-4175-b7f8-480f72ab12cb · outbound

This paper cites Qwen-Image Technical Report.

MaineCoon: Pursuing A Real-Time Audio-Visual Social World Model Qwen-Image Technical Report

Reference 51

Resolution
verified exact
local_arxiv, observed 2026-07-03T20:48:55.748232Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-27T01:09:02.214590Z digest=sha256:d6d41921e3ddc1d1e559f8b71cadc13dc2d9544c656c6284d39a6407c69774d6

Observation cfd8da29-7b24-4233-9b19-7a127a3b6f40 · outbound

This paper cites Longvideobench: A benchmark for long-context interleaved video-language understanding.

MaineCoon: Pursuing A Real-Time Audio-Visual Social World Model Longvideobench: A benchmark for long-context interleaved video-language understanding

Reference 52

Resolution
unresolved
no resolver link, observed 2026-06-27T01:09:02.214590Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T01:09:02.214590Z digest=sha256:38ec2dca7638523e0143b9d735cebe89dbfcbcd3fd5ed296c47ac1a6cdc0a6cc

Observation ccfe6001-d7b6-4452-a446-b7504c6998a8 · outbound

This paper cites LongLive: Real-time Interactive Long Video Generation.

MaineCoon: Pursuing A Real-Time Audio-Visual Social World Model LongLive: Real-time Interactive Long Video Generation

Reference 53

Resolution
verified exact
local_arxiv, observed 2026-07-03T20:48:55.683621Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-27T01:09:02.214590Z digest=sha256:4d0cdefdfbd07b047d87ff31b24236082a6fba4189d33754c987fa904bfc8336

Observation 9b55799b-a8e6-4398-9b97-70f9563e9d45 · outbound

This paper cites Sparse videogen2: Accelerate video generation with sparse attention via semantic-aware permutation.

MaineCoon: Pursuing A Real-Time Audio-Visual Social World Model Sparse videogen2: Accelerate video generation with sparse attention via semantic-aware permutation

Reference 54

Resolution
unresolved
no resolver link, observed 2026-06-27T01:09:02.214590Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T01:09:02.214590Z digest=sha256:95790879c7d1cd9a0fa5824a81d3325f10bb9bde1779bf3ae01a191a99b21932

Observation 87466778-c3a0-4d69-93cf-dda94dcc9c0d · outbound

This paper cites Improved distribution matching distillation for fast image synthesis.Advancesin neural information processing systems, 37:47455–47487, 2024.

MaineCoon: Pursuing A Real-Time Audio-Visual Social World Model Improved distribution matching distillation for fast image synthesis.Advancesin neural information processing systems, 37:47455–47487, 2024

Reference 55

Resolution
unresolved
no resolver link, observed 2026-06-27T01:09:02.214590Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T01:09:02.214590Z digest=sha256:15bcb206c3e6e381d0591e6825821a2414f4205cc60230cafa7b1db2612e8c60

Observation d3e7385c-6734-4025-b70e-ab9d31c3cce4 · outbound

This paper cites One-step diffusion with distribution matching distillation.

MaineCoon: Pursuing A Real-Time Audio-Visual Social World Model One-step diffusion with distribution matching distillation

Reference 56

Resolution
unresolved
no resolver link, observed 2026-06-27T01:09:02.214590Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T01:09:02.214590Z digest=sha256:599b2d0207495741790c5b53990fda968001a4fe601f4707a59dc2e4b01797bc

Observation d325a953-9702-456d-95e9-581c5f88c07b · outbound

This paper cites From slow bidirectional to fast autoregressive video diffusion models.

MaineCoon: Pursuing A Real-Time Audio-Visual Social World Model From slow bidirectional to fast autoregressive video diffusion models

Reference 57

Resolution
unresolved
no resolver link, observed 2026-06-27T01:09:02.214590Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T01:09:02.214590Z digest=sha256:f3a642b2f318afffdca07f9fe3aea7eb51974a38687d55faf58b57738ea9c23a

Observation 47a33905-e682-4d50-8ee9-3217b0333256 · outbound

This paper cites Representation Alignment for Generation: Training Diffusion Transformers Is Easier Than You Think.

MaineCoon: Pursuing A Real-Time Audio-Visual Social World Model Representation Alignment for Generation: Training Diffusion Transformers Is Easier Than You Think

Reference 58

Resolution
verified exact
local_arxiv, observed 2026-07-03T20:48:55.698823Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-27T01:09:02.214590Z digest=sha256:f4aed8d175b5ee090176193aa4ff23f892a6c517563aab40556db96868db0c81

Observation 11538768-8a02-43ff-b5c1-8d920cbdc22d · outbound

This paper cites Soulx-flashtalk: Real-time infinite streaming of audio-driven avatars via self-correcting bidirectional distillation.

MaineCoon: Pursuing A Real-Time Audio-Visual Social World Model Soulx-flashtalk: Real-time infinite streaming of audio-driven avatars via self-correcting bidirectional distillation

Reference 59

Resolution
unresolved
no resolver link, observed 2026-06-27T01:09:02.214590Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T01:09:02.214590Z digest=sha256:816e6cffced78a50c7070408b7bda5e3d0ed43368462abced79109963bc39018

Observation 12d65a8f-9490-460a-b98a-c693b6f49064 · outbound

This paper cites Improved distribution matching distillation for fast image synthesis.Advances in neural information processing systems, 37:47455–47487, 2024a.

MaineCoon: Pursuing A Real-Time Audio-Visual Social World Model Improved distribution matching distillation for fast image synthesis.Advances in neural information processing systems, 37:47455–47487, 2024a

Reference 60

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T20:48:55.724760Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-27T01:09:02.214590Z digest=sha256:5ab90bab4cb5b97a25bd5b144dd8898b922d56046532e331b56bf7b5eeb83d84

Observation a0500ec4-14d0-456f-b477-7315172c247a · outbound

This paper cites Fast-WAM: Do World Action Models Need Test-time Future Imagination?.

MaineCoon: Pursuing A Real-Time Audio-Visual Social World Model Fast-WAM: Do World Action Models Need Test-time Future Imagination?

Reference 61

Resolution
verified exact
local_arxiv, observed 2026-07-03T20:48:55.716418Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-27T01:09:02.214590Z digest=sha256:5c73d7ec4197d37d46d3cb59ca8db6fee6ef9bf02217e9d5d2489fde573143b3

Observation b4f63ff8-0c60-4d4c-897c-7c591a3a569c · outbound

This paper cites Sigmoid loss for language image pre-training.

MaineCoon: Pursuing A Real-Time Audio-Visual Social World Model Sigmoid loss for language image pre-training

Reference 62

Resolution
unresolved
no resolver link, observed 2026-06-27T01:09:02.214590Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T01:09:02.214590Z digest=sha256:da1c19d3c9b7132fc73abd213117b15ade45e158184af16443d0a7df53a59ad6

Observation b0aea86a-5ca9-4b46-b8a0-b25e4fdda32c · outbound

This paper cites Spargeattention: Accurate and training-free sparse attention accelerating any model inference.

MaineCoon: Pursuing A Real-Time Audio-Visual Social World Model Spargeattention: Accurate and training-free sparse attention accelerating any model inference

Reference 63

Resolution
unresolved
no resolver link, observed 2026-06-27T01:09:02.214590Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T01:09:02.214590Z digest=sha256:b763c7b369067b9faed3e362212c377e37a365ce969e1bc1d6c23420d16ed9c7

Observation 057113b3-6a28-4199-be9a-8b43b4710ac3 · outbound

This paper cites Faster video diffusion with trainable sparse attention.Advances in Neural Information Processing Systems, 38: 152509–152534, 2026.

MaineCoon: Pursuing A Real-Time Audio-Visual Social World Model Faster video diffusion with trainable sparse attention.Advances in Neural Information Processing Systems, 38: 152509–152534, 2026

Reference 64

Resolution
unresolved
no resolver link, observed 2026-06-27T01:09:02.214590Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T01:09:02.214590Z digest=sha256:c7c80344dd8f0ee1aa1f41d7ac3dbb35f0352f82fecba09583fc3ef04fecf5b8

Observation fa4e91e3-0bc4-46b1-9ce4-87de4c77cc3e · outbound

This paper cites VideoREPA: Learning Physics for Video Generation through Relational Alignment with Foundation Models.

MaineCoon: Pursuing A Real-Time Audio-Visual Social World Model VideoREPA: Learning Physics for Video Generation through Relational Alignment with Foundation Models

Reference 65

Resolution
verified exact
arxiv_id, observed 2026-07-03T20:48:55.711321Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-27T01:09:02.214590Z digest=sha256:46d001a7b6fb7429391926f31bc7e740c09e38dd4eee340c592381fddf2671ec

Observation ac6aa09c-3e61-43e6-a0df-5c65c2dffeb4 · outbound

This paper cites ViDiT-Q: Efficient and Accurate Quantization of Diffusion Transformers for Image and Video Generation.

MaineCoon: Pursuing A Real-Time Audio-Visual Social World Model ViDiT-Q: Efficient and Accurate Quantization of Diffusion Transformers for Image and Video Generation

Reference 66

Resolution
verified exact
arxiv_id, observed 2026-07-03T20:48:55.688078Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-27T01:09:02.214590Z digest=sha256:185d599566753552f6b7dde25aa12e0d9ac919d2e8aa3b0058e60ff4bd166d73

Observation 0b5ccfd4-ed3d-4732-9302-09a5573ec089 · outbound

This paper cites DiffusionNFT: Online Diffusion Reinforcement with Forward Process.

MaineCoon: Pursuing A Real-Time Audio-Visual Social World Model DiffusionNFT: Online Diffusion Reinforcement with Forward Process

Reference 67

Resolution
verified exact
local_arxiv, observed 2026-07-03T20:48:55.696397Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-27T01:09:02.214590Z digest=sha256:6c33e46040de2bfa4c448d4a55b6786a218adc013d374c4765e5017c24ba4b6c

Observation 946ddd3c-d0e1-4fb5-bb4b-c4b666eede11 · outbound

This paper cites SANA-WM: Efficient Minute-Scale World Modeling with Hybrid Linear Diffusion Transformer.

MaineCoon: Pursuing A Real-Time Audio-Visual Social World Model SANA-WM: Efficient Minute-Scale World Modeling with Hybrid Linear Diffusion Transformer

Reference 68

Resolution
verified exact
local_arxiv, observed 2026-07-03T20:48:55.735119Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-27T01:09:02.214590Z digest=sha256:309053bc2175bff60da6d8d8f37ec777ae02d6b15ba4b044e870bf1ed1e6188d

Observation ee396425-4902-478c-a771-e8d737436aaa · outbound

This paper cites Causal Forcing: Autoregressive Diffusion Distillation Done Right for High-Quality Real-Time Interactive Video Generation.

MaineCoon: Pursuing A Real-Time Audio-Visual Social World Model Causal Forcing: Autoregressive Diffusion Distillation Done Right for High-Quality Real-Time Interactive Video Generation

Reference 69

Resolution
verified exact
local_arxiv, observed 2026-07-03T20:48:55.748077Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-27T01:09:02.214590Z digest=sha256:f2462090b2f30675722576f5ab015aba7ae43067275fed29d2bc01a1d6f6a3c0

Pith citing papers

Observation e7c80bb1-33a2-44df-80f1-2c273baae2ec · inbound

Perceptual Flow Matching for Few-Step Generative Modeling cites this paper.

Perceptual Flow Matching for Few-Step Generative Modeling MaineCoon: Pursuing A Real-Time Audio-Visual Social World Model

Reference 1

Resolution
unresolved
no resolver link, observed 2026-07-12T01:53:57.900599Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T01:53:57.900599Z digest=sha256:f96ab73731d67dc4bb27e305d868501da29e834cccdcea38deed420fc646d585

Observation de100b1b-32d7-4f31-953b-d86a1aa7ae46 · inbound

Sol-Attn: Accelerating Video Generation Inference via On-the-Fly Attention Sparsification cites this paper.

Sol-Attn: Accelerating Video Generation Inference via On-the-Fly Attention Sparsification MaineCoon: Pursuing A Real-Time Audio-Visual Social World Model

Reference 6

Resolution
unresolved
no resolver link, observed 2026-07-31T23:18:25.056705Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T23:18:25.056705Z digest=sha256:5b750e260902231b1b9b4595c0d562cecd171eaeb3c2672d774b8acb8fde3513

Observation f3234e59-a1b2-4445-9564-286c13ffc4a2 · inbound

TaoMate: Anchor-Guided Memory Bridging Evolving and Reference States for Real-Time Audio-Video Digital Human Generation cites this paper.

TaoMate: Anchor-Guided Memory Bridging Evolving and Reference States for Real-Time Audio-Video Digital Human Generation MaineCoon: Pursuing A Real-Time Audio-Visual Social World Model

Reference 8

Resolution
unresolved
no resolver link, observed 2026-07-31T17:08:11.387051Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-31T17:08:11.387051Z digest=sha256:e8073abf5ec657cd2834bd2052d7421d330f6dc0c5ac19c7dd169d2e183827b0