Pith. sign in

Paper Citation Record · LEDGER

KVAE: Family of Tokenizers for Multimodal Generative Models

As of 8 August 2026, this Paper Citation Record lists 100 of 122 outbound references and 0 inbound Pith citation observations for arXiv:2608.05798.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2608.05798 v1

Coverage vector

measured 100 of 122 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T23:24:49.346598Z

measured 100 of 100 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

100 of 122 outbound references displayed

  • verified exact2
  • verified fuzzy41
  • unresolved57
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 738a38b0-545f-4185-a25e-5273ff6c1ee1 · outbound

This paper cites Bitdance: Scaling autoregressive generative models with binary tokens, 2026.

KVAE: Family of Tokenizers for Multimodal Generative Models Bitdance: Scaling autoregressive generative models with binary tokens, 2026

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T23:24:48.884725Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:24:48.884725Z digest=sha256:0fa4e5b934493854c5a551bca59f57e0a1f7428d8ea2c6d8c32ce465e6100d36

Observation 2cf5ec29-295d-4960-8f0d-451f2c896b44 · outbound

This paper cites OmniDoc-TokenBench.

KVAE: Family of Tokenizers for Multimodal Generative Models OmniDoc-TokenBench

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T23:24:48.890347Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:24:48.890347Z digest=sha256:d508d5e1ae0c6dde79e91c2c7491b812188be06a72e563f3e2c34bd46354bef5

Observation 58e03af8-44a4-4176-9f42-fb71009e891d · outbound

This paper cites AOM Common Test Conditions v5.0.

KVAE: Family of Tokenizers for Multimodal Generative Models AOM Common Test Conditions v5.0

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T23:24:48.895514Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:24:48.895514Z digest=sha256:d1b525f80c69726cb01fcc05697625413905373c3b10c4c0ac5a06553643f9e3

Observation e0c328e6-e129-4c3b-94e2-3bc18372faf5 · outbound

This paper cites Kandinsky 5.0: A family of foundation models for image and video generation,.

KVAE: Family of Tokenizers for Multimodal Generative Models Kandinsky 5.0: A family of foundation models for image and video generation,

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T23:24:48.900423Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:24:48.900423Z digest=sha256:1a5ab3a5bf8d116686f46a1e6b2df2d8d39e877b4cc745072bf2c73ba02de409

Observation 1c9e3033-ed5a-4ebb-b5c2-9c9bc343525b · outbound

This paper cites Qwen2.5-vl technical report, 2025.

KVAE: Family of Tokenizers for Multimodal Generative Models Qwen2.5-vl technical report, 2025

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T23:24:48.905419Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:24:48.905419Z digest=sha256:1b7204e710261f210d7708d1bfa0d9157b2ea6b5450bb78baec58a038d74e846

Observation e7aec70c-cff9-4657-95a8-14e5ec270626 · outbound

This paper cites Stable video diffusion: Scaling latent video diffusion models to large datasets,.

KVAE: Family of Tokenizers for Multimodal Generative Models Stable video diffusion: Scaling latent video diffusion models to large datasets,

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T23:24:48.915306Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:24:48.915306Z digest=sha256:5a72e383469443d4458f95929c1d272e7ecb5e8fb6c40c504848bf5db0c87d8a

Observation f4e01222-a462-47f1-8769-f251b1d97829 · outbound

This paper cites Align your latents: High-resolution video synthesis with latent diffusion models, 2023.

KVAE: Family of Tokenizers for Multimodal Generative Models Align your latents: High-resolution video synthesis with latent diffusion models, 2023

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T23:24:48.919954Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:24:48.919954Z digest=sha256:d28ac8499a9f6f281e7bbc665ec5231b8a945bdc34b3a564750493f6870e0866

Observation 08a77cb1-62c7-4212-be39-0d505c00db6c · outbound

This paper cites Bruinsma, Ana Lucic, Megan Stanley, Anna Vaughan, Johannes Brandstetter, Patrick Garvan, Maik Riechert, Jonathan A.

KVAE: Family of Tokenizers for Multimodal Generative Models Bruinsma, Ana Lucic, Megan Stanley, Anna Vaughan, Johannes Brandstetter, Patrick Garvan, Maik Riechert, Jonathan A

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T23:24:48.924532Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:24:48.924532Z digest=sha256:c25a3a3791efeb21b8b0ad9cb5ada3100501663d5c1ff5858c2e86007f427178

Observation 2d1a6791-9ac9-4d11-a4a2-c6189f7f9025 · outbound

This paper cites Bradley and Milton E.

KVAE: Family of Tokenizers for Multimodal Generative Models Bradley and Milton E

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T23:24:48.929922Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:24:48.929922Z digest=sha256:7d4c30f01edcee2599e0053b759a848e6c6fd0e14771c5c65bdb9118e3a8876d

Observation 1dafd9e2-d524-463c-8765-c667503a3beb · outbound

This paper cites Efros, and Tero Karras.

KVAE: Family of Tokenizers for Multimodal Generative Models Efros, and Tero Karras

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T23:24:48.934742Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:24:48.934742Z digest=sha256:1a72571aea5267a5fd4e8de33847565c2c712aea83f0ce86c15c4608060d09e7

Observation 8f99be9f-2c46-4c89-a8e8-634720a099f3 · outbound

This paper cites Video generation models as world simulators.

KVAE: Family of Tokenizers for Multimodal Generative Models Video generation models as world simulators

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T23:24:48.939667Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:24:48.939667Z digest=sha256:52059032b5b65cf71fc058b6de9f5c3bd1934f6661e28c4d8df1f8d7e6c7ea6d

Observation ac82de2f-508f-42f8-94f2-b3442585c127 · outbound

This paper cites Deep compression autoencoder for efficient high-resolution diffusion models, 2025.

KVAE: Family of Tokenizers for Multimodal Generative Models Deep compression autoencoder for efficient high-resolution diffusion models, 2025

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T23:24:48.944198Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:24:48.944198Z digest=sha256:5bf4db871b5c53646673d4870ff6d20ed762effa4c147950695b077b8d1ac0c0

Observation a01b3ac9-6c80-4101-a447-d258f2d05261 · outbound

This paper cites Dc-videogen: Efficient video generation with deep compression video autoencoder, 2025.

KVAE: Family of Tokenizers for Multimodal Generative Models Dc-videogen: Efficient video generation with deep compression video autoencoder, 2025

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T23:24:48.948798Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:24:48.948798Z digest=sha256:53789e6b265e4cbdd5f42ba6acd3b65e0c211e8b36bd0f569e1d1875e77cbcc6

Observation 66f14d0a-fbab-440d-9a23-60e4be4f9113 · outbound

This paper cites Dc-ae 1.5: Accelerating diffusion model convergence with structured latent space, 2025.

KVAE: Family of Tokenizers for Multimodal Generative Models Dc-ae 1.5: Accelerating diffusion model convergence with structured latent space, 2025

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T23:24:48.953175Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:24:48.953175Z digest=sha256:297a23b4b87702a9a77aab4bfaf22625c32102fabede9ad8e0ff3fefdeb9893c

Observation 8cfa0b25-621d-4ebf-a247-cdecb6287abc · outbound

This paper cites MMAudio: Taming Multimodal Joint Training for High-Quality Video-to-Audio Synthesis.

KVAE: Family of Tokenizers for Multimodal Generative Models MMAudio: Taming Multimodal Joint Training for High-Quality Video-to-Audio Synthesis

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T23:24:48.957682Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:24:48.957682Z digest=sha256:27cf622efb96974e9dcf4a08857e42da6ca6d30a5464c436cffca44627f78bde

Observation 39300b56-8bb8-49ca-a936-b46c8e8fdf26 · outbound

This paper cites Chien, Liuzixuan Lin, Hai Nguyen, Varsha Rao, Tristan Sharma, and Rajini Wijayawardana.

KVAE: Family of Tokenizers for Multimodal Generative Models Chien, Liuzixuan Lin, Hai Nguyen, Varsha Rao, Tristan Sharma, and Rajini Wijayawardana

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T23:24:48.962483Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:24:48.962483Z digest=sha256:68a965be0330cf1fd831ca3f4ab84175fd232ab1c3912f6f9cf619def0dd0618

Observation 3208e4f5-b9a4-4f4e-9113-e86893969d43 · outbound

This paper cites Chi, Jeff Dean, Jacob Devlin, Adam Roberts, Denny Zhou, Quoc V.

KVAE: Family of Tokenizers for Multimodal Generative Models Chi, Jeff Dean, Jacob Devlin, Adam Roberts, Denny Zhou, Quoc V

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T23:24:48.966861Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:24:48.966861Z digest=sha256:8ba43cc9215c4734a93f0dd90f58acebe94aaf83619bacb55880fbc8f11e66c1

Observation 68383886-6f37-4b0b-8661-df2658afe5cf · outbound

This paper cites Adversarial video generation on complex datasets, 2019.

KVAE: Family of Tokenizers for Multimodal Generative Models Adversarial video generation on complex datasets, 2019

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T23:24:48.971343Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:24:48.971343Z digest=sha256:1fc586b1b9424381efb4cfff0e298d3b22b94f4f71daa28b453e5338026819b7

Observation 81dcc4f3-3506-4684-ad39-9ab84632b47e · outbound

This paper cites High Fidelity Neural Audio Compression.

KVAE: Family of Tokenizers for Multimodal Generative Models High Fidelity Neural Audio Compression

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T23:24:48.975675Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:24:48.975675Z digest=sha256:44a298029226f330ed64c424f7f4de3bddf6a1b2419aeb4750bf9ec73c86d3a2

Observation 79456876-3f1f-4d7b-b238-e098f9074619 · outbound

This paper cites Irc-gan: Introspective recurrent convolutional gan for text-to-video generation.

KVAE: Family of Tokenizers for Multimodal Generative Models Irc-gan: Introspective recurrent convolutional gan for text-to-video generation

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T23:24:48.980726Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:24:48.980726Z digest=sha256:5a14fb722144e3534330ba42410888bc912fa6a4de942029e60fdc948398107a

Observation 90a9e832-bdce-496d-afdd-d01c2ade3351 · outbound

This paper cites Taming transformers for high-resolution image synthesis, 2021.

KVAE: Family of Tokenizers for Multimodal Generative Models Taming transformers for high-resolution image synthesis, 2021

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T23:24:48.985253Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:24:48.985253Z digest=sha256:3f3d029d75649784d4f282fa429fd350f62f93d8490f0e04c7d960754e8b121c

Observation 4502afe0-0465-4f0f-9d01-890af4e52d42 · outbound

This paper cites Stable Audio Open.

KVAE: Family of Tokenizers for Multimodal Generative Models Stable Audio Open

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T23:24:48.989694Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:24:48.989694Z digest=sha256:fc24f892bcf526d2813db27349cc34f903c6c04d0c6449cbf74a9f36b670395e

Observation 5b1c7234-92ae-487b-b756-b6f47c997c35 · outbound

This paper cites Parker, Matthew Rice, CJ Carr, Zack Zukowski, Josiah Taylor, and Jordi Pons.

KVAE: Family of Tokenizers for Multimodal Generative Models Parker, Matthew Rice, CJ Carr, Zack Zukowski, Josiah Taylor, and Jordi Pons

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T23:24:48.994502Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:24:48.994502Z digest=sha256:3b6b809bfee3286096996073efe353f9da432ed3183c6c9fe49aac9b0c7ff31c

Observation 42f007fe-896c-4768-a598-b7200889aeab · outbound

This paper cites The prism hypothesis: Harmonizing semantic and pixel representations via unified autoencoding, 2026.

KVAE: Family of Tokenizers for Multimodal Generative Models The prism hypothesis: Harmonizing semantic and pixel representations via unified autoencoding, 2026

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T23:24:48.998841Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:24:48.998841Z digest=sha256:b6bc5867b263b7c58e5f4a2a6c5d89159808864e0995271c7e91b4b823068de2

Observation 9afbd4dd-4573-449e-b4c3-10ea5566898c · outbound

This paper cites Gemmeke, Daniel P.

KVAE: Family of Tokenizers for Multimodal Generative Models Gemmeke, Daniel P

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T23:24:49.003333Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:24:49.003333Z digest=sha256:bd0ffa79191d77bf0ffa4b854ffcc61b8a9ab6e4bdcbf3061a7e41a6e278171e

Observation 6fe1651f-184c-4137-a0e1-f14a18e6255a · outbound

This paper cites BigVGAN: A universal neural vocoder with large-scale training.

KVAE: Family of Tokenizers for Multimodal Generative Models BigVGAN: A universal neural vocoder with large-scale training

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T23:24:49.007973Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:24:49.007973Z digest=sha256:21c4712ee193ec7e9f0324f8acb961cb870f893da6435f7dc59ea3b8e2bd868d

Observation cc499c7a-24e3-4967-9a5e-f6aaf1ad7721 · outbound

This paper cites Goodfellow, Jean Pouget-Abadie, Mehdi Mirza, Bing Xu, David Warde-Farley, Sherjil Ozair, Aaron Courville, and Yoshua Bengio.

KVAE: Family of Tokenizers for Multimodal Generative Models Goodfellow, Jean Pouget-Abadie, Mehdi Mirza, Bing Xu, David Warde-Farley, Sherjil Ozair, Aaron Courville, and Yoshua Bengio

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T23:24:49.012406Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:24:49.012406Z digest=sha256:74c49e0e567d2e339cec15a82057908edd4d3b1adee07efa909d74c9e680daf1

Observation aee2644b-0f0f-44a3-ac52-1f0c0a3d1f37 · outbound

This paper cites Veo 3.1: Our leading video generation model.

KVAE: Family of Tokenizers for Multimodal Generative Models Veo 3.1: Our leading video generation model

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T23:24:49.016907Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:24:49.016907Z digest=sha256:7164797276469b2afe9dbe11adf418fd96a2f93266b178d6955b7b04c7b5444d

Observation b3bd5b17-09ff-45ff-9cea-0b4fa494d047 · outbound

This paper cites Ltx-2: Efficient joint audio-visual foundation model, 2026.

KVAE: Family of Tokenizers for Multimodal Generative Models Ltx-2: Efficient joint audio-visual foundation model, 2026

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T23:24:49.021526Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:24:49.021526Z digest=sha256:fb013b2aa00c8989854fd8c93ddb0b2758aefd4499a1e81ca3ba9e6d1331eb29

Observation d25e578e-3c43-4d82-b3b6-a4d102f65412 · outbound

This paper cites an unresolved cited work.

KVAE: Family of Tokenizers for Multimodal Generative Models Unresolved cited work

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T23:24:49.026296Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:24:49.026296Z digest=sha256:55e691e115f18e99c6dab0210942dfe2020b8e57d1ed0a4f99c6a9f9dc09eda3

Observation 51815ec0-39d4-4233-b438-7d728e498fc2 · outbound

This paper cites Gans trained by a two time-scale update rule converge to a local nash equilibrium.Advances in neural information processing systems, 30, 2017.

KVAE: Family of Tokenizers for Multimodal Generative Models Gans trained by a two time-scale update rule converge to a local nash equilibrium.Advances in neural information processing systems, 30, 2017

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T23:24:49.031152Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:24:49.031152Z digest=sha256:3eac632b9612717e80ffd0d25b8ac8a0f253660a025c987b1a1d84ed64577270

Observation acfa54a4-7194-4056-9435-82f1d3aada79 · outbound

This paper cites Kingma, Ben Poole, Mohammad Norouzi, David J.

KVAE: Family of Tokenizers for Multimodal Generative Models Kingma, Ben Poole, Mohammad Norouzi, David J

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T23:24:49.035612Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:24:49.035612Z digest=sha256:6c185dbf64f8dfec391ea607437522beb44b086c047c2f2408b759306e30eba3

Observation ac1880fa-d3b4-45ba-b19f-e5d33fbb0059 · outbound

This paper cites an unresolved cited work.

KVAE: Family of Tokenizers for Multimodal Generative Models Unresolved cited work

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T23:24:49.040278Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:24:49.040278Z digest=sha256:3dd995e0159e97769f714795a80ce754f24a9f27e00404793327d75d4adc46ee

Observation 74b0e03a-b622-43c5-9b45-0d0faf6572eb · outbound

This paper cites Cogvideo: Large-scale pretraining for text-to-video generation via transformers, 2022.

KVAE: Family of Tokenizers for Multimodal Generative Models Cogvideo: Large-scale pretraining for text-to-video generation via transformers, 2022

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T23:24:49.045380Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:24:49.045380Z digest=sha256:e2e8199397b7b373e5478400ce0bdf6697f83ba5deee6606614d2cdb19627e95

Observation 388dabe9-9483-4020-a64f-bf40609068c2 · outbound

This paper cites Tangoflux: Super fast and faithful text to audio generation with flow matching and clap-ranked preference optimization,.

KVAE: Family of Tokenizers for Multimodal Generative Models Tangoflux: Super fast and faithful text to audio generation with flow matching and clap-ranked preference optimization,

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T23:24:49.049876Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:24:49.049876Z digest=sha256:f0a760a00b29f6c13e844c8f3d5fa82dc3cb54658bf91f392b03ee094b2a56c4

Observation 62f3ccf0-45f6-46a3-bae0-92d1aad6f2da · outbound

This paper cites Perceptual evaluation of speech quality (PESQ).International Telecommunication Union, 2001.

KVAE: Family of Tokenizers for Multimodal Generative Models Perceptual evaluation of speech quality (PESQ).International Telecommunication Union, 2001

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-07T23:24:49.055309Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:24:49.055309Z digest=sha256:19cd6d4624d268c2d86a65a468407d809dd7d1e44e6370b55565fb8f910dc6e7

Observation 2f138c9e-758c-405e-b102-950da176a488 · outbound

This paper cites Video pixel networks, 2016.

KVAE: Family of Tokenizers for Multimodal Generative Models Video pixel networks, 2016

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-07T23:24:49.059886Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:24:49.059886Z digest=sha256:88c68d16170b93b56540fe999d242cf8a573881ac73734c16baa281b03a2ac1f

Observation 2b8624e0-78ca-499b-bff4-74489982f7a5 · outbound

This paper cites Fréchet audio distance: A reference-free metric for evaluating music enhancement algorithms.Interspeech,.

KVAE: Family of Tokenizers for Multimodal Generative Models Fréchet audio distance: A reference-free metric for evaluating music enhancement algorithms.Interspeech,

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-07T23:24:49.064527Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:24:49.064527Z digest=sha256:dd59e85678d6967011b2ff3fa80f6100ff9b352d37924cf76c6b54e3e9348750

Observation 232dab4f-1757-46a5-87c7-363c5b97a577 · outbound

This paper cites AudioCaps: Generating captions for audios in the wild.

KVAE: Family of Tokenizers for Multimodal Generative Models AudioCaps: Generating captions for audios in the wild

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-07T23:24:49.069235Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:24:49.069235Z digest=sha256:ac08d190f95aaef7a3cf921da030b59c60257ad0e731175b2b6fdc975a807ab8

Observation 9cddac63-115e-4f59-b7dc-a7ff81cc025e · outbound

This paper cites Kingma and Jimmy Ba.

KVAE: Family of Tokenizers for Multimodal Generative Models Kingma and Jimmy Ba

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-07T23:24:49.073665Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:24:49.073665Z digest=sha256:33a7a0f66e971e140f60e21e6eeac7411359331243c85fc000e0751d26f6c14c

Observation ac8ebf5b-0045-4547-8dd7-3bdecc1b4e8b · outbound

This paper cites Auto-encoding variational bayes, 2022.

KVAE: Family of Tokenizers for Multimodal Generative Models Auto-encoding variational bayes, 2022

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-07T23:24:49.078139Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:24:49.078139Z digest=sha256:2e7f484579fd77b243ef0736c928338f0e2fbb178b983be219b1352c113a89ca

Observation eed9165a-132b-46a2-9a61-0628ff3086bd · outbound

This paper cites Klingai enters the 3.0 era: All in one, one for all! kling 3.0 model now fully rolled out.

KVAE: Family of Tokenizers for Multimodal Generative Models Klingai enters the 3.0 era: All in one, one for all! kling 3.0 model now fully rolled out

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-07T23:24:49.082831Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:24:49.082831Z digest=sha256:e8faab32b083cc4b92fcf11e73b00eb90fafaa4004aa0c4d78016660bd4a59ac

Observation c3bfd2f4-fbfe-402a-bb65-80e0edfc4a3e · outbound

This paper cites Carbon Emissions in the Tailpipe of Gen- erative AI.Harvard Data Science Review, 15(Special Issue 5), aug 20 2024.

KVAE: Family of Tokenizers for Multimodal Generative Models Carbon Emissions in the Tailpipe of Gen- erative AI.Harvard Data Science Review, 15(Special Issue 5), aug 20 2024

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-07T23:24:49.087019Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:24:49.087019Z digest=sha256:092e4c08736b7e5e98d5376af4d0cbf309c81a7f9f7340b92ad63179a3901cd0

Observation 5c12fc1a-6a0c-4df5-9492-f8bdbf88ce97 · outbound

This paper cites Ross, Bryan Seybold, and Lu Jiang.

KVAE: Family of Tokenizers for Multimodal Generative Models Ross, Bryan Seybold, and Lu Jiang

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-07T23:24:49.091331Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:24:49.091331Z digest=sha256:d58dabbecaed4dc68bf6ff27cda7610ec9dc6f675f11d950c2083c1e3388091c

Observation 11284e29-edbb-49d5-a292-19aed98a3d65 · outbound

This paper cites HiFi-GAN: Generative adversarial networks for efficient and high fidelity speech synthesis.

KVAE: Family of Tokenizers for Multimodal Generative Models HiFi-GAN: Generative adversarial networks for efficient and high fidelity speech synthesis

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-07T23:24:49.095870Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:24:49.095870Z digest=sha256:99abb1b4980c3ebeb4b76c7021c7836d497f18fd5cd677170bec82d63d41c5e4

Observation 6df06d21-2e6f-4c84-9e10-0f5f86539d0c · outbound

This paper cites Plumb- ley.

KVAE: Family of Tokenizers for Multimodal Generative Models Plumb- ley

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T23:24:51.103658Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T23:24:49.100167Z digest=sha256:5be4c9d8e1b648adda510cb5780a145fa99e91a356e717ed86770353d4fe2883

Observation 0aaa8769-e8f5-470d-a322-8fcd01577bca · outbound

This paper cites Hunyuanvideo: A systematic framework for large video generative models,.

KVAE: Family of Tokenizers for Multimodal Generative Models Hunyuanvideo: A systematic framework for large video generative models,

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-07T23:24:49.105028Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:24:49.105028Z digest=sha256:23b36ed31c3c32146080b7ca3a7521749cc7596367ef08153939a499590869d1

Observation 36e12cc6-4dc9-4036-9ac1-ee3691524038 · outbound

This paper cites Kandinsky video tools.

KVAE: Family of Tokenizers for Multimodal Generative Models Kandinsky video tools

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T23:24:51.078140Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T23:24:49.109643Z digest=sha256:1f740a1f2d209474da44174b161ec2c6fca36d20e0920d95ad9153b3dc21391d

Observation 31a7a776-c7af-4f4f-8bc8-1312dd7dcb93 · outbound

This paper cites Efficient training of audio transformers with patchout.

KVAE: Family of Tokenizers for Multimodal Generative Models Efficient training of audio transformers with patchout

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T23:24:51.062027Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T23:24:49.114113Z digest=sha256:77b56c1eeee314152327e26c145149279812c5ea8115257567badce6c8b39bb3

Observation 127290f6-9b74-49d9-b344-b37a41322fd7 · outbound

This paper cites Eq-vae: Equivariance regularized latent space for improved generative image modeling, 2025.

KVAE: Family of Tokenizers for Multimodal Generative Models Eq-vae: Equivariance regularized latent space for improved generative image modeling, 2025

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T23:24:51.046580Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T23:24:49.118762Z digest=sha256:0e90818bad9af5c5b7b3ac5c426b99dc8d6a5e237f08b03b59feb5f3910bafd6

Observation 998362ab-e4c9-4bf6-a90c-0ca4b683a950 · outbound

This paper cites High-fidelity audio compression with improved RVQGAN.

KVAE: Family of Tokenizers for Multimodal Generative Models High-fidelity audio compression with improved RVQGAN

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T23:24:51.031097Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T23:24:49.123265Z digest=sha256:a29a980140484adc6c6c931a072f71e51a22cf0d8c1d81a994fbbaeb2fe63d37

Observation 5848a0cc-c2f6-482d-954d-dfa79c488168 · outbound

This paper cites REPA-E: Unlocking vae for end-to-end tuning of latent diffusion transformers.

KVAE: Family of Tokenizers for Multimodal Generative Models REPA-E: Unlocking vae for end-to-end tuning of latent diffusion transformers

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T23:24:51.015069Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T23:24:49.127893Z digest=sha256:8a0e7e893b4d7af7a791e5d1a258dc2bbee90909e256075499f7f59ee69863d5

Observation cc608454-847f-40a0-8459-3c9ccee14f2b · outbound

This paper cites DiffusionBench: On Holistic Evaluation of Diffusion Transformers.

KVAE: Family of Tokenizers for Multimodal Generative Models DiffusionBench: On Holistic Evaluation of Diffusion Transformers

Reference 54

Resolution
verified exact
local_arxiv, observed 2026-08-07T23:24:50.096776Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T23:24:49.132533Z digest=sha256:f0584c088645ba2945db5d5e3f12b4c1159b5be6db03204a5d53e7c42d2e82f6

Observation 83adf514-2156-4f71-a374-ffa46d3c452f · outbound

This paper cites Wf-vae: Enhancing video vae by wavelet-driven energy flow for latent video diffusion model,.

KVAE: Family of Tokenizers for Multimodal Generative Models Wf-vae: Enhancing video vae by wavelet-driven energy flow for latent video diffusion model,

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T23:24:50.999575Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T23:24:49.137542Z digest=sha256:c519b3ee3abf6f3700ac123d9a3c0f97bf16d4ced29583c30c99d88f6e4bedc7

Observation e5363bd6-7328-439c-add9-f9aedcb225ab · outbound

This paper cites Generating novel, designable, and diverse protein structures by equivariantly diffusing oriented residue clouds, 2023.

KVAE: Family of Tokenizers for Multimodal Generative Models Generating novel, designable, and diverse protein structures by equivariantly diffusing oriented residue clouds, 2023

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T23:24:50.984146Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T23:24:49.142269Z digest=sha256:39671f1b7e569478f4cf360e8cb71d7d60baac1ae98e5f1eb5c9c975eea565c5

Observation 267aeaba-7451-409e-b2f4-67858bde79ba · outbound

This paper cites AudioLDM 2: Learning Holistic Audio Generation with Self-supervised Pretraining.

KVAE: Family of Tokenizers for Multimodal Generative Models AudioLDM 2: Learning Holistic Audio Generation with Self-supervised Pretraining

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-07T23:24:49.146653Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:24:49.146653Z digest=sha256:89d491467c4095b3aa690191fc425ff61b8ab62cb3450db069f51308f75f56d7

Observation 66ca06a4-e630-4e22-b35d-8cac8283534a · outbound

This paper cites Delving into Latent Spectral Biasing of Video VAEs for Superior Diffusability.

KVAE: Family of Tokenizers for Multimodal Generative Models Delving into Latent Spectral Biasing of Video VAEs for Superior Diffusability

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-07T23:24:49.151448Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:24:49.151448Z digest=sha256:50c7620e448fa5b22dc335c26cc318c81375eece2e53d4af276ddc8837f1ed41

Observation acc391df-336e-415e-be70-460d0f658a11 · outbound

This paper cites an unresolved cited work.

KVAE: Family of Tokenizers for Multimodal Generative Models Unresolved cited work

Reference 59

Resolution
unresolved
raw_fallback, observed 2026-08-07T23:24:50.969040Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T23:24:49.156389Z digest=sha256:0626833ea74a27c07183df4460080bd02eb5a348dc115b4dd4866b64bff4539f

Observation b4353548-027b-4ca2-a114-f6b0f9e70e21 · outbound

This paper cites The song describer dataset: A corpus of audio captions for music-and- language evaluation.

KVAE: Family of Tokenizers for Multimodal Generative Models The song describer dataset: A corpus of audio captions for music-and- language evaluation

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T23:24:50.954476Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T23:24:49.160855Z digest=sha256:d521c1c873d6613956825c8ce7c14f7e8e949b41691b1471f273886221bda8bc

Observation c0d76d52-e54f-48d8-a3be-b4b6bbb0005b · outbound

This paper cites Balasubramanian.

KVAE: Family of Tokenizers for Multimodal Generative Models Balasubramanian

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T23:24:50.939257Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T23:24:49.165459Z digest=sha256:65f5b01378bab07e3b4f106d86acfa35408ee6bf0b970906ad9cde6925b4ed4b

Observation b51194c7-d6ad-4f0a-8fe1-844428a199cb · outbound

This paper cites Transition matching distillation for fast video generation, 2026.

KVAE: Family of Tokenizers for Multimodal Generative Models Transition matching distillation for fast video generation, 2026

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T23:24:50.924914Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T23:24:49.170003Z digest=sha256:3070c4a9d25fd2ed38f8472308044024a9ffe74234daa08da135da929a179f0a

Observation d1a2b48e-5f39-4930-9d4b-2590079c194c · outbound

This paper cites Blaschko, Albert Ali Salah, and Itir Onal Ertugrul.

KVAE: Family of Tokenizers for Multimodal Generative Models Blaschko, Albert Ali Salah, and Itir Onal Ertugrul

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-07T23:24:49.174272Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:24:49.174272Z digest=sha256:a4cc5194297d45fef082ae16f9b69e418f82a3c32de12e759a83fc3042468eb6

Observation 81b25390-cc78-405e-975f-b87ef68d0aea · outbound

This paper cites Alpamayo-r1: Bridging reasoning and action prediction for generalizable autonomous driving in the long tail, 2026.

KVAE: Family of Tokenizers for Multimodal Generative Models Alpamayo-r1: Bridging reasoning and action prediction for generalizable autonomous driving in the long tail, 2026

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T23:24:50.910624Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T23:24:49.178680Z digest=sha256:83342a4df25605c135a21e7a8b44e5aa7909aa8aec478f235048909c558b7bcf

Observation 6377f506-b873-48fb-9081-51f4cfa0c397 · outbound

This paper cites Cosmos World Foundation Model Platform for Physical AI.

KVAE: Family of Tokenizers for Multimodal Generative Models Cosmos World Foundation Model Platform for Physical AI

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-07T23:24:49.183042Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:24:49.183042Z digest=sha256:3802e793130dc5431e7c133cd91c80bc2f15c89a73f45917bea8c81bbde4fa29

Observation 800b8ce9-b089-4322-971e-87cb2867cc2f · outbound

This paper cites To create what you tell: Generating videos from captions, 2018.

KVAE: Family of Tokenizers for Multimodal Generative Models To create what you tell: Generating videos from captions, 2018

Reference 66

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T23:24:50.895129Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T23:24:49.187278Z digest=sha256:999e017b7eb24f094dc283dcc2e85c67455eed4214712e370cec39e6fbc9fcc6

Observation bbe6bb1f-7dfe-4235-8b32-0602637df5de · outbound

This paper cites Librispeech: An ASR corpus based on public domain audio books.

KVAE: Family of Tokenizers for Multimodal Generative Models Librispeech: An ASR corpus based on public domain audio books

Reference 67

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T23:24:50.880432Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T23:24:49.192154Z digest=sha256:176bf80e01f81b4601a5369caec1cb4b243288a851d5bf5eac314affb304c9e3

Observation be565eb9-cfe7-46b6-bb9b-8be96549aa1e · outbound

This paper cites Parker, Zach Evans, CJ Carr, Zack Zukowski, Josiah Taylor, Matthew Rice, and Jordi Pons.

KVAE: Family of Tokenizers for Multimodal Generative Models Parker, Zach Evans, CJ Carr, Zack Zukowski, Josiah Taylor, Matthew Rice, and Jordi Pons

Reference 68

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T23:24:50.864864Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T23:24:49.196758Z digest=sha256:4a533605a976c66d239eb078b753249450c17d6d0776623733e3c62a234f1dba

Observation 8b31fb12-eda1-4f4a-bf0f-82d1d1e71768 · outbound

This paper cites Sampson, Shikai Li, Simone Parmeggiani, Steve Fine, Tara Fowler, Vladan Petrovic, and Yuming Du.

KVAE: Family of Tokenizers for Multimodal Generative Models Sampson, Shikai Li, Simone Parmeggiani, Steve Fine, Tara Fowler, Vladan Petrovic, and Yuming Du

Reference 69

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T23:24:50.849477Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T23:24:49.201249Z digest=sha256:5dec4bf00992e7bffe1d0fa210c1263eba071e117146389e570234ae370e73f4

Observation ca168e10-08e7-4eab-87da-b326e642041e · outbound

This paper cites Andersson, Andrew El-Kadi, Dominic Masters, Timo Ewalds, Jacklynn Stott, Shakir Mohamed, Peter Battaglia, Remi Lam, and Matthew Willson.

KVAE: Family of Tokenizers for Multimodal Generative Models Andersson, Andrew El-Kadi, Dominic Masters, Timo Ewalds, Jacklynn Stott, Shakir Mohamed, Peter Battaglia, Remi Lam, and Matthew Willson

Reference 70

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T23:24:50.833535Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T23:24:49.205992Z digest=sha256:3ef2a6b854fb0c53e2d3981c3ae750c63a4af704ebcf33912d9e27ec06d9dfb0

Observation 0ff90aab-604b-4284-bacd-ad48092a822b · outbound

This paper cites Qwen-Audio-VAE Technical Report.

KVAE: Family of Tokenizers for Multimodal Generative Models Qwen-Audio-VAE Technical Report

Reference 71

Resolution
verified exact
local_arxiv, observed 2026-08-07T23:24:49.823936Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T23:24:49.210558Z digest=sha256:a6281e184abd84e7740abf2ad4ede7487d0fbd2ccbdb88bdd5eef28ff628db55

Observation 396ca9dd-f10f-4d1b-9a33-54a7aca03ca7 · outbound

This paper cites Qwen-Image-VAE-2.0 Technical Report.

KVAE: Family of Tokenizers for Multimodal Generative Models Qwen-Image-VAE-2.0 Technical Report

Reference 72

Resolution
unresolved
no resolver link, observed 2026-08-07T23:24:49.215146Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:24:49.215146Z digest=sha256:a9eb08779dc851eb3ac75de1b04339394f18f59428b3fc6be528ab083afcf52f

Observation a1a0de29-3159-47e1-bae3-88f9a00de7e8 · outbound

This paper cites Learning transferable visual models from natural language supervision.

KVAE: Family of Tokenizers for Multimodal Generative Models Learning transferable visual models from natural language supervision

Reference 73

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T23:24:50.817684Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T23:24:49.219882Z digest=sha256:c677388b1676ad3fe1682f94a85088e4d87bcf80079467ecc6ce2171f2925e85

Observation 56dab5a3-9387-4627-a2a5-79b447587466 · outbound

This paper cites MUSDB18-HQ — an uncompressed version of MUSDB18, 2019.

KVAE: Family of Tokenizers for Multimodal Generative Models MUSDB18-HQ — an uncompressed version of MUSDB18, 2019

Reference 74

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T23:24:50.802189Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T23:24:49.224296Z digest=sha256:a31ca389a103ab0500e0f1c94a01862166a7024048589aacdde66a96436c3627

Observation 403980d2-0bb7-4e58-8a4e-b51a022855e0 · outbound

This paper cites EARS: An anechoic fullband speech dataset benchmarked for speech enhancement and dereverberation.

KVAE: Family of Tokenizers for Multimodal Generative Models EARS: An anechoic fullband speech dataset benchmarked for speech enhancement and dereverberation

Reference 75

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T23:24:50.787304Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T23:24:49.228947Z digest=sha256:26990c608e485fb59076b4bdc246daa570107bbeb8c8d25e3d785829bca5847a

Observation 908a6963-34df-4061-9ae6-a918534a033b · outbound

This paper cites High-resolution image synthesis with latent diffusion models, 2022.

KVAE: Family of Tokenizers for Multimodal Generative Models High-resolution image synthesis with latent diffusion models, 2022

Reference 76

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T23:24:50.772663Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T23:24:49.233612Z digest=sha256:59c78d9e72477251852cdc61d8006330f5b3680b4a4aaa5f3126cffef80d1a4c

Observation df019351-9721-4388-8474-f420028c64c2 · outbound

This paper cites High-resolution image synthesis with latent diffusion models, 2022.

KVAE: Family of Tokenizers for Multimodal Generative Models High-resolution image synthesis with latent diffusion models, 2022

Reference 77

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T23:24:50.757578Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T23:24:49.238165Z digest=sha256:93da813ce6ac80747da241c88bd3f49765c371eedc5c604e4ce24eacef2bae01

Observation d2ee0a73-a87d-4d66-8420-375818288256 · outbound

This paper cites Runway gen-4: Ai video generation with world consistency.

KVAE: Family of Tokenizers for Multimodal Generative Models Runway gen-4: Ai video generation with world consistency

Reference 78

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T23:24:50.742254Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T23:24:49.242655Z digest=sha256:e260afa8995fb2e8d075fc28b77e998d345680bdc677d13bc2c6265264bb777e

Observation 1381ca67-e28f-4a49-b640-88ed3f1cd99e · outbound

This paper cites Flow to the mode: Mode-seeking diffusion autoencoders for state-of-the-art image tokenization, 2025.

KVAE: Family of Tokenizers for Multimodal Generative Models Flow to the mode: Mode-seeking diffusion autoencoders for state-of-the-art image tokenization, 2025

Reference 79

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T23:24:50.725842Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T23:24:49.246801Z digest=sha256:c21e465dc5587ef5c849922b9d2fd774d8587099eac2730a5c1fcf81ad3ecb9d

Observation f28f6329-029f-4897-a6f0-4573d0b56e2d · outbound

This paper cites Make- a-video: Text-to-video generation without text-video data, 2022.

KVAE: Family of Tokenizers for Multimodal Generative Models Make- a-video: Text-to-video generation without text-video data, 2022

Reference 80

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T23:24:50.710306Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T23:24:49.251220Z digest=sha256:5b116d0d96ccbcc0e44a6f55d9cea06cd00c94c8e21fe706059c8f032dba8c42

Observation a188e79e-5a70-4a33-b517-84db242a8fa8 · outbound

This paper cites What matters for representation alignment: Global information or spatial structure?arXiv preprint arXiv:2512.10794, 2025.

KVAE: Family of Tokenizers for Multimodal Generative Models What matters for representation alignment: Global information or spatial structure?arXiv preprint arXiv:2512.10794, 2025

Reference 81

Resolution
unresolved
no resolver link, observed 2026-08-07T23:24:49.255842Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:24:49.255842Z digest=sha256:ce813a0d6e9229761f567cf22be5003f485bc098abfbaf1484a7fb3f2fa77bb0

Observation 6cc40914-4d56-4056-b59a-1fdd5dd4ac21 · outbound

This paper cites Improving the diffusability of autoencoders, 2025.

KVAE: Family of Tokenizers for Multimodal Generative Models Improving the diffusability of autoencoders, 2025

Reference 82

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T23:24:50.693904Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T23:24:49.260320Z digest=sha256:9841efe0b28d0a8ea4e1327f42a27261ec5804da252cba7746fe89f6df93a82a

Observation 76dfbbe9-c306-4d70-bad0-2e7965de56f9 · outbound

This paper cites Stylegan-v: A continuous video generator with the price, image quality and perks of stylegan2, 2022.

KVAE: Family of Tokenizers for Multimodal Generative Models Stylegan-v: A continuous video generator with the price, image quality and perks of stylegan2, 2022

Reference 83

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T23:24:50.678729Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T23:24:49.264895Z digest=sha256:6fbe89c57f461e86fca39c8cf9bccd47c7b15d21052231661263a20cd2506518

Observation 8570bb5d-e8b1-4fe9-8313-b7bdf698ea11 · outbound

This paper cites Ucf101: A dataset of 101 human actions classes from videos in the wild, 2012.

KVAE: Family of Tokenizers for Multimodal Generative Models Ucf101: A dataset of 101 human actions classes from videos in the wild, 2012

Reference 84

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T23:24:50.663478Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T23:24:49.269466Z digest=sha256:f6236f8b05d439d39d3de983b9f210aea0a2883ccc724ea517c3501d251f243a

Observation 19383de8-b3f2-4913-8671-bc2eb2dd16c3 · outbound

This paper cites Scenediffuser++: City-scale traffic simulation via a generative world model, 2025.

KVAE: Family of Tokenizers for Multimodal Generative Models Scenediffuser++: City-scale traffic simulation via a generative world model, 2025

Reference 85

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T23:24:50.647889Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T23:24:49.273993Z digest=sha256:128cc8fceb2493111b27bad08be373210b18a37088c287cb628eb3209088ae0e

Observation 95573368-f86d-4172-9e9e-868bfa40d269 · outbound

This paper cites Z-image: An efficient image generation foundation model with single-stream diffusion transformer,.

KVAE: Family of Tokenizers for Multimodal Generative Models Z-image: An efficient image generation foundation model with single-stream diffusion transformer,

Reference 86

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T23:24:50.632963Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T23:24:49.278379Z digest=sha256:b1145ad5a186b45f3eb59a67dca934c2663b4d7af770250935f3c41e3c99c96e

Observation 5405c1df-101f-4983-925d-d8fbf8ec6146 · outbound

This paper cites an unresolved cited work.

KVAE: Family of Tokenizers for Multimodal Generative Models Unresolved cited work

Reference 87

Resolution
unresolved
raw_fallback, observed 2026-08-07T23:24:50.617480Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T23:24:49.282806Z digest=sha256:064d05cdfa93ab5e26e4cb992effb24062d9958ae840d3123d516b377f96f89c

Observation 49e126f2-5c89-408a-bd6a-c64e1f68d805 · outbound

This paper cites Nextstep-1: Toward autoregressive image generation with continuous tokens at scale, 2025.

KVAE: Family of Tokenizers for Multimodal Generative Models Nextstep-1: Toward autoregressive image generation with continuous tokens at scale, 2025

Reference 88

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T23:24:50.602974Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T23:24:49.287263Z digest=sha256:07558672487b082b9aa7f477b2f2a7113a9aa992e302e07328e5e9c725fe8f18

Observation 54f8ed22-adc4-477d-aa77-894691e0c336 · outbound

This paper cites HunyuanVideo-Foley: Multimodal Diffusion with Representation Alignment for High-Fidelity Foley Audio Generation.

KVAE: Family of Tokenizers for Multimodal Generative Models HunyuanVideo-Foley: Multimodal Diffusion with Representation Alignment for High-Fidelity Foley Audio Generation

Reference 89

Resolution
unresolved
no resolver link, observed 2026-08-07T23:24:49.291826Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:24:49.291826Z digest=sha256:6912cdc3731051b3b1529ecf219599e2097c93f933cd3e43bf030be9a543ad4e

Observation fa0cbf27-e6fc-431d-af41-a2baebeba92c · outbound

This paper cites Reducio! generating 1k video within 16 seconds using extremely compressed motion latents, 2025.

KVAE: Family of Tokenizers for Multimodal Generative Models Reducio! generating 1k video within 16 seconds using extremely compressed motion latents, 2025

Reference 90

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T23:24:50.587939Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T23:24:49.296716Z digest=sha256:c65c31ebf534ca1b7b4bb4f11a03a86cb427a0780e61d980d0527581985c9ed6

Observation efee10b7-6d58-44ed-b07b-e9ab1e95acaf · outbound

This paper cites Metaxas, and Sergey Tulyakov.

KVAE: Family of Tokenizers for Multimodal Generative Models Metaxas, and Sergey Tulyakov

Reference 91

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T23:24:50.573136Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T23:24:49.301476Z digest=sha256:b7157f6aae22fadefd3e543c889a792ce93ceed52f69fcf69ec7502e79148a15

Observation 23c4e403-bc4d-43c7-8fee-7ada83688ad1 · outbound

This paper cites Meta Audiobox Aesthetics: Unified Automatic Quality Assessment for Speech, Music, and Sound.

KVAE: Family of Tokenizers for Multimodal Generative Models Meta Audiobox Aesthetics: Unified Automatic Quality Assessment for Speech, Music, and Sound

Reference 92

Resolution
unresolved
no resolver link, observed 2026-08-07T23:24:49.306053Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:24:49.306053Z digest=sha256:4add20e860cf283b1ea9caad95648498c3c15ee720c74691bbbe779994a7a54d

Observation efeb5b68-bd37-45d3-a383-dc5adebe7538 · outbound

This paper cites Ssdd: Single-step diffusion decoder for efficient image tokenization, 2026.

KVAE: Family of Tokenizers for Multimodal Generative Models Ssdd: Single-step diffusion decoder for efficient image tokenization, 2026

Reference 93

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T23:24:50.558721Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T23:24:49.310789Z digest=sha256:0c4dfc40f70546f6a4e099f61e729274c9575d659a30f95f794fd249fe9defa9

Observation f8046108-5b8b-4da1-b9d0-9f8cde99f1d4 · outbound

This paper cites Conditional image generation with pixelcnn decoders, 2016.

KVAE: Family of Tokenizers for Multimodal Generative Models Conditional image generation with pixelcnn decoders, 2016

Reference 94

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T23:24:50.544106Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T23:24:49.315228Z digest=sha256:237ed23e96d92b58e76ac054ea6f40b83e5c46ffd04b347939cae7e18d36b058

Observation 235a4e58-ae68-4a14-8a35-8677b551f010 · outbound

This paper cites Neural discrete representation learning, 2018.

KVAE: Family of Tokenizers for Multimodal Generative Models Neural discrete representation learning, 2018

Reference 95

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T23:24:50.529444Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T23:24:49.319830Z digest=sha256:d985f6177255ac33551f35567361d85fa2996238f285820ca6b1d4fa7280364c

Observation fcb04103-a5e0-40db-82c6-a454a4cc5187 · outbound

This paper cites Gomez, Lukasz Kaiser, and Illia Polosukhin.

KVAE: Family of Tokenizers for Multimodal Generative Models Gomez, Lukasz Kaiser, and Illia Polosukhin

Reference 96

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T23:24:50.514926Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T23:24:49.324257Z digest=sha256:6401454005fb3fa84a8d36a3a89ec94de5f5cd42879dac0c98619f1b32bf2a07

Observation 2283e7e1-6c9b-45bc-a3ee-a12990febe3d · outbound

This paper cites Generating videos with scene dynamics, 2016.

KVAE: Family of Tokenizers for Multimodal Generative Models Generating videos with scene dynamics, 2016

Reference 97

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T23:24:50.499693Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T23:24:49.328663Z digest=sha256:f3e46e08fdaa6a65c8e3c0be4beb7983df5ddefa1c8832196fdf99432db0cc11

Observation 3d01af25-d7ad-4d26-9e55-82118e8e8c08 · outbound

This paper cites Wan: Open and advanced large-scale video generative models, 2025.

KVAE: Family of Tokenizers for Multimodal Generative Models Wan: Open and advanced large-scale video generative models, 2025

Reference 98

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T23:24:50.484897Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T23:24:49.333015Z digest=sha256:fdf69a7ef003cff6e7e400371c7525ccc72c28e713dd7d33dd4c2d682ae36195

Observation 69657ab6-1c35-43ea-8ef2-6a4476db2986 · outbound

This paper cites Wan-2.2 anouncement.

KVAE: Family of Tokenizers for Multimodal Generative Models Wan-2.2 anouncement

Reference 99

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T23:24:50.470707Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T23:24:49.337918Z digest=sha256:1d33e45dd1938875474e221a81d3c393cc3b73ed4efea60376e1d22830908908

Observation 76571650-0562-4483-9884-118e27a89523 · outbound

This paper cites an unresolved cited work.

KVAE: Family of Tokenizers for Multimodal Generative Models Unresolved cited work

Reference 100

Resolution
unresolved
raw_fallback, observed 2026-08-07T23:24:50.456314Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T23:24:49.342282Z digest=sha256:b9fadd3ab758da8602ab2ae24215b959411318f4ceec9bd8cc98c84e4764bc20

Observation 56c741e4-ba26-43b8-a506-dd18820a4bb4 · outbound

This paper cites Videomae v2: Scaling video masked autoencoders with dual masking.

KVAE: Family of Tokenizers for Multimodal Generative Models Videomae v2: Scaling video masked autoencoders with dual masking

Reference 101

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T23:24:50.442381Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T23:24:49.346598Z digest=sha256:6b94e55de3a20e724927f293a00f998ccc832ca1b60659e566710536681adc69

Pith citing papers

No inbound Pith citation observations are available.