Pith. sign in

Paper Citation Record · LEDGER

KVAE: Family of Tokenizers for Multimodal Generative Models

As of 8 August 2026, this Paper Citation Record lists 100 of 122 outbound references and 0 inbound Pith citation observations for arXiv:2608.05798.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2608.05798 v1

Coverage vector

measured 100 of 122 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T23:24:49.346598Z

measured 100 of 100 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

100 of 122 outbound references displayed

  • verified exact2
  • verified fuzzy41
  • unresolved57
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 738a38b0-545f-4185-a25e-5273ff6c1ee1 · outbound

This paper cites Bitdance: Scaling autoregressive generative models with binary tokens, 2026.

KVAE: Family of Tokenizers for Multimodal Generative Models Bitdance: Scaling autoregressive generative models with binary tokens, 2026

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T23:24:48.884725Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:24:48.884725Z digest=sha256:60241458f4f1a177832d0e595112373a3f2a2aa6779eaed63f3f47ac2b32259b

Observation 2cf5ec29-295d-4960-8f0d-451f2c896b44 · outbound

This paper cites OmniDoc-TokenBench.

KVAE: Family of Tokenizers for Multimodal Generative Models OmniDoc-TokenBench

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T23:24:48.890347Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:24:48.890347Z digest=sha256:bbcb690e7ffa5c5e03d7926a8915a976a4c9b1f59cc37731df309bf993635851

Observation 58e03af8-44a4-4176-9f42-fb71009e891d · outbound

This paper cites AOM Common Test Conditions v5.0.

KVAE: Family of Tokenizers for Multimodal Generative Models AOM Common Test Conditions v5.0

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T23:24:48.895514Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:24:48.895514Z digest=sha256:83a344b40a07f495ce1dc1b0226d3f454cfeb9b27e258e353fb8c7752227c4a9

Observation e0c328e6-e129-4c3b-94e2-3bc18372faf5 · outbound

This paper cites Kandinsky 5.0: A family of foundation models for image and video generation,.

KVAE: Family of Tokenizers for Multimodal Generative Models Kandinsky 5.0: A family of foundation models for image and video generation,

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T23:24:48.900423Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:24:48.900423Z digest=sha256:9a0b8308063eac5ef8825c98ff3cd75d66c5ec282c002be8dd3d6a73337fbd10

Observation 1c9e3033-ed5a-4ebb-b5c2-9c9bc343525b · outbound

This paper cites Qwen2.5-vl technical report, 2025.

KVAE: Family of Tokenizers for Multimodal Generative Models Qwen2.5-vl technical report, 2025

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T23:24:48.905419Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:24:48.905419Z digest=sha256:7bd054f49253c6397c0fc43a8d0bff672d09b0e944aa8b05564c54be55fd888e

Observation e7aec70c-cff9-4657-95a8-14e5ec270626 · outbound

This paper cites Stable video diffusion: Scaling latent video diffusion models to large datasets,.

KVAE: Family of Tokenizers for Multimodal Generative Models Stable video diffusion: Scaling latent video diffusion models to large datasets,

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T23:24:48.915306Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:24:48.915306Z digest=sha256:3e89229f56faca087f174c77adea476fa77a645de5d36c4db975071119aebcfc

Observation f4e01222-a462-47f1-8769-f251b1d97829 · outbound

This paper cites Align your latents: High-resolution video synthesis with latent diffusion models, 2023.

KVAE: Family of Tokenizers for Multimodal Generative Models Align your latents: High-resolution video synthesis with latent diffusion models, 2023

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T23:24:48.919954Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:24:48.919954Z digest=sha256:5e13e9ebce1fb6b2c6c24ad0f32a44cdb4128b014550e71def985e3a415b4063

Observation 08a77cb1-62c7-4212-be39-0d505c00db6c · outbound

This paper cites Bruinsma, Ana Lucic, Megan Stanley, Anna Vaughan, Johannes Brandstetter, Patrick Garvan, Maik Riechert, Jonathan A.

KVAE: Family of Tokenizers for Multimodal Generative Models Bruinsma, Ana Lucic, Megan Stanley, Anna Vaughan, Johannes Brandstetter, Patrick Garvan, Maik Riechert, Jonathan A

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T23:24:48.924532Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:24:48.924532Z digest=sha256:0d4367c0819539b4f7c512eb865864f86770acbcf96753fc1f0c1c4c6d4f17b4

Observation 2d1a6791-9ac9-4d11-a4a2-c6189f7f9025 · outbound

This paper cites Bradley and Milton E.

KVAE: Family of Tokenizers for Multimodal Generative Models Bradley and Milton E

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T23:24:48.929922Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:24:48.929922Z digest=sha256:f7d32ad493f3a606ef38e001065bc90902cb7232912fdfed5e3947e226cb4da6

Observation 1dafd9e2-d524-463c-8765-c667503a3beb · outbound

This paper cites Efros, and Tero Karras.

KVAE: Family of Tokenizers for Multimodal Generative Models Efros, and Tero Karras

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T23:24:48.934742Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:24:48.934742Z digest=sha256:a8a886966698b9090225e3576b621e0b589d5c59af1cb10b59e3472b486997fe

Observation 8f99be9f-2c46-4c89-a8e8-634720a099f3 · outbound

This paper cites Video generation models as world simulators.

KVAE: Family of Tokenizers for Multimodal Generative Models Video generation models as world simulators

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T23:24:48.939667Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:24:48.939667Z digest=sha256:cc28a888f5488b8fde3f4c619d906d577f8a22d5164e977815dc83dc76142cff

Observation ac82de2f-508f-42f8-94f2-b3442585c127 · outbound

This paper cites Deep compression autoencoder for efficient high-resolution diffusion models, 2025.

KVAE: Family of Tokenizers for Multimodal Generative Models Deep compression autoencoder for efficient high-resolution diffusion models, 2025

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T23:24:48.944198Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:24:48.944198Z digest=sha256:0d16ab7dbe6ffc90e82e3e01278aab2aa32b6b1ffe544326df8cb332578d35fa

Observation a01b3ac9-6c80-4101-a447-d258f2d05261 · outbound

This paper cites Dc-videogen: Efficient video generation with deep compression video autoencoder, 2025.

KVAE: Family of Tokenizers for Multimodal Generative Models Dc-videogen: Efficient video generation with deep compression video autoencoder, 2025

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T23:24:48.948798Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:24:48.948798Z digest=sha256:1192d840181082439108b53ad9d246c9cb30487470eb27b039f37e38c5a5d09b

Observation 66f14d0a-fbab-440d-9a23-60e4be4f9113 · outbound

This paper cites Dc-ae 1.5: Accelerating diffusion model convergence with structured latent space, 2025.

KVAE: Family of Tokenizers for Multimodal Generative Models Dc-ae 1.5: Accelerating diffusion model convergence with structured latent space, 2025

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T23:24:48.953175Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:24:48.953175Z digest=sha256:8421ab349f6d8618357ca20435d88efb633ef01e0be3cbdc600b5f28a439860d

Observation 8cfa0b25-621d-4ebf-a247-cdecb6287abc · outbound

This paper cites MMAudio: Taming Multimodal Joint Training for High-Quality Video-to-Audio Synthesis.

KVAE: Family of Tokenizers for Multimodal Generative Models MMAudio: Taming Multimodal Joint Training for High-Quality Video-to-Audio Synthesis

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T23:24:48.957682Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:24:48.957682Z digest=sha256:95346f30b8d54214a77ef282385d02086cf973a5720d69ced80dba29bd32f0bb

Observation 39300b56-8bb8-49ca-a936-b46c8e8fdf26 · outbound

This paper cites Chien, Liuzixuan Lin, Hai Nguyen, Varsha Rao, Tristan Sharma, and Rajini Wijayawardana.

KVAE: Family of Tokenizers for Multimodal Generative Models Chien, Liuzixuan Lin, Hai Nguyen, Varsha Rao, Tristan Sharma, and Rajini Wijayawardana

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T23:24:48.962483Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:24:48.962483Z digest=sha256:56260ec394936f0c4bb09909b4cef9d8dd170ba188d89f729d21f15c50bfef58

Observation 3208e4f5-b9a4-4f4e-9113-e86893969d43 · outbound

This paper cites Chi, Jeff Dean, Jacob Devlin, Adam Roberts, Denny Zhou, Quoc V.

KVAE: Family of Tokenizers for Multimodal Generative Models Chi, Jeff Dean, Jacob Devlin, Adam Roberts, Denny Zhou, Quoc V

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T23:24:48.966861Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:24:48.966861Z digest=sha256:8dc3b5efa7dfcc40f30a016a3301df9835835e14e446758913cace39d06337fd

Observation 68383886-6f37-4b0b-8661-df2658afe5cf · outbound

This paper cites Adversarial video generation on complex datasets, 2019.

KVAE: Family of Tokenizers for Multimodal Generative Models Adversarial video generation on complex datasets, 2019

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T23:24:48.971343Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:24:48.971343Z digest=sha256:00dadb2aa542a79d52459ab78e80e0f58cd92a84c3fee15437c9dcb0ff33a7dc

Observation 81dcc4f3-3506-4684-ad39-9ab84632b47e · outbound

This paper cites High Fidelity Neural Audio Compression.

KVAE: Family of Tokenizers for Multimodal Generative Models High Fidelity Neural Audio Compression

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T23:24:48.975675Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:24:48.975675Z digest=sha256:11661e17d6b23b3af955e3c820dbd4ba35918784ab6bef098cac35d0c0933a2e

Observation 79456876-3f1f-4d7b-b238-e098f9074619 · outbound

This paper cites Irc-gan: Introspective recurrent convolutional gan for text-to-video generation.

KVAE: Family of Tokenizers for Multimodal Generative Models Irc-gan: Introspective recurrent convolutional gan for text-to-video generation

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T23:24:48.980726Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:24:48.980726Z digest=sha256:4196c0eaff1e60a8d2f7493dac075f2a2be4ed0fa601b759d5c78821d0efe2c5

Observation 90a9e832-bdce-496d-afdd-d01c2ade3351 · outbound

This paper cites Taming transformers for high-resolution image synthesis, 2021.

KVAE: Family of Tokenizers for Multimodal Generative Models Taming transformers for high-resolution image synthesis, 2021

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T23:24:48.985253Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:24:48.985253Z digest=sha256:1bb1b7a39bc2a2cf926e46b6be4dbee688ddbc8d66bf68f2ae3c258fd0ca160e

Observation 4502afe0-0465-4f0f-9d01-890af4e52d42 · outbound

This paper cites Stable Audio Open.

KVAE: Family of Tokenizers for Multimodal Generative Models Stable Audio Open

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T23:24:48.989694Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:24:48.989694Z digest=sha256:2ff7af3187f3f9b9a3abc0a23e98b0a4c2186115bbd3260a8ded0021fd2c39bf

Observation 5b1c7234-92ae-487b-b756-b6f47c997c35 · outbound

This paper cites Parker, Matthew Rice, CJ Carr, Zack Zukowski, Josiah Taylor, and Jordi Pons.

KVAE: Family of Tokenizers for Multimodal Generative Models Parker, Matthew Rice, CJ Carr, Zack Zukowski, Josiah Taylor, and Jordi Pons

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T23:24:48.994502Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:24:48.994502Z digest=sha256:1175abe2170b779ff34c802a2ebdda1f9630eaec02b86df2ff09aed26cf4b2c2

Observation 42f007fe-896c-4768-a598-b7200889aeab · outbound

This paper cites The prism hypothesis: Harmonizing semantic and pixel representations via unified autoencoding, 2026.

KVAE: Family of Tokenizers for Multimodal Generative Models The prism hypothesis: Harmonizing semantic and pixel representations via unified autoencoding, 2026

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T23:24:48.998841Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:24:48.998841Z digest=sha256:9fb21a9c24802b738eabf21b95ebbd34b31bb45b7044eb78f25a19f157abbfe3

Observation 9afbd4dd-4573-449e-b4c3-10ea5566898c · outbound

This paper cites Gemmeke, Daniel P.

KVAE: Family of Tokenizers for Multimodal Generative Models Gemmeke, Daniel P

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T23:24:49.003333Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:24:49.003333Z digest=sha256:7b96e1fdf27c83010d8e702617b4ddf65705e9d0de054301c8769fc35ff2dddd

Observation 6fe1651f-184c-4137-a0e1-f14a18e6255a · outbound

This paper cites BigVGAN: A universal neural vocoder with large-scale training.

KVAE: Family of Tokenizers for Multimodal Generative Models BigVGAN: A universal neural vocoder with large-scale training

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T23:24:49.007973Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:24:49.007973Z digest=sha256:17ec3c9d95377d020d87a2324b4c8a471a763c2c6dc06c9fb8ebb594b16fa5b9

Observation cc499c7a-24e3-4967-9a5e-f6aaf1ad7721 · outbound

This paper cites Goodfellow, Jean Pouget-Abadie, Mehdi Mirza, Bing Xu, David Warde-Farley, Sherjil Ozair, Aaron Courville, and Yoshua Bengio.

KVAE: Family of Tokenizers for Multimodal Generative Models Goodfellow, Jean Pouget-Abadie, Mehdi Mirza, Bing Xu, David Warde-Farley, Sherjil Ozair, Aaron Courville, and Yoshua Bengio

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T23:24:49.012406Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:24:49.012406Z digest=sha256:bce8c1f6e17a982e7df65d97387897dabc2e284d1d406025da243f842d9d461b

Observation aee2644b-0f0f-44a3-ac52-1f0c0a3d1f37 · outbound

This paper cites Veo 3.1: Our leading video generation model.

KVAE: Family of Tokenizers for Multimodal Generative Models Veo 3.1: Our leading video generation model

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T23:24:49.016907Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:24:49.016907Z digest=sha256:ee856f204ac7a13f81ae3bcf8d20b7149926ba1901dbeba10e2de829bf73464c

Observation b3bd5b17-09ff-45ff-9cea-0b4fa494d047 · outbound

This paper cites Ltx-2: Efficient joint audio-visual foundation model, 2026.

KVAE: Family of Tokenizers for Multimodal Generative Models Ltx-2: Efficient joint audio-visual foundation model, 2026

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T23:24:49.021526Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:24:49.021526Z digest=sha256:de9400050dea670aa5657673529ee741bcaa26072c6aa360bcb2e81251efdef9

Observation d25e578e-3c43-4d82-b3b6-a4d102f65412 · outbound

This paper cites an unresolved cited work.

KVAE: Family of Tokenizers for Multimodal Generative Models Unresolved cited work

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T23:24:49.026296Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:24:49.026296Z digest=sha256:378636663ae36fdf4d0e757c647c2f8cab4bd6d87f3efea3e1938387d990b3ba

Observation 51815ec0-39d4-4233-b438-7d728e498fc2 · outbound

This paper cites Gans trained by a two time-scale update rule converge to a local nash equilibrium.Advances in neural information processing systems, 30, 2017.

KVAE: Family of Tokenizers for Multimodal Generative Models Gans trained by a two time-scale update rule converge to a local nash equilibrium.Advances in neural information processing systems, 30, 2017

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T23:24:49.031152Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:24:49.031152Z digest=sha256:f0e2d15071b8ad5e8dd02c68744399bca631a509edb8c08b090dad7302594554

Observation acfa54a4-7194-4056-9435-82f1d3aada79 · outbound

This paper cites Kingma, Ben Poole, Mohammad Norouzi, David J.

KVAE: Family of Tokenizers for Multimodal Generative Models Kingma, Ben Poole, Mohammad Norouzi, David J

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T23:24:49.035612Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:24:49.035612Z digest=sha256:9727c0e76254caba8668b36e327220b32f48a43ec29b9f5f6a5ee451f82b7834

Observation ac1880fa-d3b4-45ba-b19f-e5d33fbb0059 · outbound

This paper cites an unresolved cited work.

KVAE: Family of Tokenizers for Multimodal Generative Models Unresolved cited work

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T23:24:49.040278Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:24:49.040278Z digest=sha256:d632a9b038ed71c7172060d71960c0ade236aaae36bca5427e9ee337ebde4d2f

Observation 74b0e03a-b622-43c5-9b45-0d0faf6572eb · outbound

This paper cites Cogvideo: Large-scale pretraining for text-to-video generation via transformers, 2022.

KVAE: Family of Tokenizers for Multimodal Generative Models Cogvideo: Large-scale pretraining for text-to-video generation via transformers, 2022

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T23:24:49.045380Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:24:49.045380Z digest=sha256:9afad11de5ef206df4c4041a43d35b8aa48f0d6a333ebf3aca341ad4996b090a

Observation 388dabe9-9483-4020-a64f-bf40609068c2 · outbound

This paper cites Tangoflux: Super fast and faithful text to audio generation with flow matching and clap-ranked preference optimization,.

KVAE: Family of Tokenizers for Multimodal Generative Models Tangoflux: Super fast and faithful text to audio generation with flow matching and clap-ranked preference optimization,

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T23:24:49.049876Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:24:49.049876Z digest=sha256:332356119622ce3772c53029c0f69c01228ac8f9cfa33e4d3bb35e64ac86a17d

Observation 62f3ccf0-45f6-46a3-bae0-92d1aad6f2da · outbound

This paper cites Perceptual evaluation of speech quality (PESQ).International Telecommunication Union, 2001.

KVAE: Family of Tokenizers for Multimodal Generative Models Perceptual evaluation of speech quality (PESQ).International Telecommunication Union, 2001

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-07T23:24:49.055309Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:24:49.055309Z digest=sha256:325f1d9775ea65e4fc7a5f5fb8abc410e97fcaecb1a72ecd013936cad80fc338

Observation 2f138c9e-758c-405e-b102-950da176a488 · outbound

This paper cites Video pixel networks, 2016.

KVAE: Family of Tokenizers for Multimodal Generative Models Video pixel networks, 2016

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-07T23:24:49.059886Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:24:49.059886Z digest=sha256:cd3f062ffa0739b734f834d05c61b956935cd653e055fa88c2b3fc31ffdfc317

Observation 2b8624e0-78ca-499b-bff4-74489982f7a5 · outbound

This paper cites Fréchet audio distance: A reference-free metric for evaluating music enhancement algorithms.Interspeech,.

KVAE: Family of Tokenizers for Multimodal Generative Models Fréchet audio distance: A reference-free metric for evaluating music enhancement algorithms.Interspeech,

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-07T23:24:49.064527Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:24:49.064527Z digest=sha256:f28e774bab69d4a2ac9c30fe73a5084a14a0c71e9acd5eca2f7ba09fe5b9aec3

Observation 232dab4f-1757-46a5-87c7-363c5b97a577 · outbound

This paper cites AudioCaps: Generating captions for audios in the wild.

KVAE: Family of Tokenizers for Multimodal Generative Models AudioCaps: Generating captions for audios in the wild

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-07T23:24:49.069235Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:24:49.069235Z digest=sha256:0f3435f6ffbec8d5b5285295d3e9eacc0204114ce2dd5875b85dd879f33e5655

Observation 9cddac63-115e-4f59-b7dc-a7ff81cc025e · outbound

This paper cites Kingma and Jimmy Ba.

KVAE: Family of Tokenizers for Multimodal Generative Models Kingma and Jimmy Ba

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-07T23:24:49.073665Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:24:49.073665Z digest=sha256:01f79bc23132ac694b228d718ea8a39920a6e2a5a519f5febfcbf9fc051108da

Observation ac8ebf5b-0045-4547-8dd7-3bdecc1b4e8b · outbound

This paper cites Auto-encoding variational bayes, 2022.

KVAE: Family of Tokenizers for Multimodal Generative Models Auto-encoding variational bayes, 2022

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-07T23:24:49.078139Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:24:49.078139Z digest=sha256:3de9c4d692875047944df4f4f0f0e4ddd9729dadd686624bc4d8581f146d0c9f

Observation eed9165a-132b-46a2-9a61-0628ff3086bd · outbound

This paper cites Klingai enters the 3.0 era: All in one, one for all! kling 3.0 model now fully rolled out.

KVAE: Family of Tokenizers for Multimodal Generative Models Klingai enters the 3.0 era: All in one, one for all! kling 3.0 model now fully rolled out

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-07T23:24:49.082831Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:24:49.082831Z digest=sha256:0ba2fc7f2052c4150788a5b98aad56bd60ac7f726094d93c874259da02dbae2c

Observation c3bfd2f4-fbfe-402a-bb65-80e0edfc4a3e · outbound

This paper cites Carbon Emissions in the Tailpipe of Gen- erative AI.Harvard Data Science Review, 15(Special Issue 5), aug 20 2024.

KVAE: Family of Tokenizers for Multimodal Generative Models Carbon Emissions in the Tailpipe of Gen- erative AI.Harvard Data Science Review, 15(Special Issue 5), aug 20 2024

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-07T23:24:49.087019Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:24:49.087019Z digest=sha256:8213299ac282d0b7c222cd5bb2c7cdb4e9f4b208dd3a04ae4a265e82f5effd76

Observation 5c12fc1a-6a0c-4df5-9492-f8bdbf88ce97 · outbound

This paper cites Ross, Bryan Seybold, and Lu Jiang.

KVAE: Family of Tokenizers for Multimodal Generative Models Ross, Bryan Seybold, and Lu Jiang

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-07T23:24:49.091331Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:24:49.091331Z digest=sha256:9a9a8924e744adb479f715e9d062b249270df44141676c74c241a18907c8f2dd

Observation 11284e29-edbb-49d5-a292-19aed98a3d65 · outbound

This paper cites HiFi-GAN: Generative adversarial networks for efficient and high fidelity speech synthesis.

KVAE: Family of Tokenizers for Multimodal Generative Models HiFi-GAN: Generative adversarial networks for efficient and high fidelity speech synthesis

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-07T23:24:49.095870Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:24:49.095870Z digest=sha256:c6fea87948853e990304ecdaa88311a193d8ff12d8b62fc4d8647e32b773f16e

Observation 6df06d21-2e6f-4c84-9e10-0f5f86539d0c · outbound

This paper cites Plumb- ley.

KVAE: Family of Tokenizers for Multimodal Generative Models Plumb- ley

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T23:24:51.103658Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T23:24:49.100167Z digest=sha256:f73fbcb1ae016be689229d1d10b1eeaf1a6f1ef02ac7cd5ec07d5c5c6070ac8d

Observation 0aaa8769-e8f5-470d-a322-8fcd01577bca · outbound

This paper cites Hunyuanvideo: A systematic framework for large video generative models,.

KVAE: Family of Tokenizers for Multimodal Generative Models Hunyuanvideo: A systematic framework for large video generative models,

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-07T23:24:49.105028Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:24:49.105028Z digest=sha256:ed696d06cd5d2be567fff5f1befc4b5cbf4e612f94ce6e8d0d75b68b5d38875e

Observation 36e12cc6-4dc9-4036-9ac1-ee3691524038 · outbound

This paper cites Kandinsky video tools.

KVAE: Family of Tokenizers for Multimodal Generative Models Kandinsky video tools

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T23:24:51.078140Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T23:24:49.109643Z digest=sha256:245aaf85c25c5a4d827e9f195cfe853e65b58b98b1da52b6f5205dde24cadb94

Observation 31a7a776-c7af-4f4f-8bc8-1312dd7dcb93 · outbound

This paper cites Efficient training of audio transformers with patchout.

KVAE: Family of Tokenizers for Multimodal Generative Models Efficient training of audio transformers with patchout

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T23:24:51.062027Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T23:24:49.114113Z digest=sha256:e1df759f2dbfbaa6bc1edc25f3f851497f27d0a95e82414fc360c4b1f60907a4

Observation 127290f6-9b74-49d9-b344-b37a41322fd7 · outbound

This paper cites Eq-vae: Equivariance regularized latent space for improved generative image modeling, 2025.

KVAE: Family of Tokenizers for Multimodal Generative Models Eq-vae: Equivariance regularized latent space for improved generative image modeling, 2025

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T23:24:51.046580Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T23:24:49.118762Z digest=sha256:07fc3aaf309479afd370f1d1db6d2a02427f7b30bd7d429cb49ee75cd54c7f00

Observation 998362ab-e4c9-4bf6-a90c-0ca4b683a950 · outbound

This paper cites High-fidelity audio compression with improved RVQGAN.

KVAE: Family of Tokenizers for Multimodal Generative Models High-fidelity audio compression with improved RVQGAN

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T23:24:51.031097Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T23:24:49.123265Z digest=sha256:2b600f3dca4b5e2d99a447311bf3865765a04955c1ccc590191ddc7e59a8a74f

Observation 5848a0cc-c2f6-482d-954d-dfa79c488168 · outbound

This paper cites REPA-E: Unlocking vae for end-to-end tuning of latent diffusion transformers.

KVAE: Family of Tokenizers for Multimodal Generative Models REPA-E: Unlocking vae for end-to-end tuning of latent diffusion transformers

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T23:24:51.015069Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T23:24:49.127893Z digest=sha256:bd7ddf6d31f30b90aa49508796e33388967811a62b6ee3a9546de8a5bab33fca

Observation cc608454-847f-40a0-8459-3c9ccee14f2b · outbound

This paper cites DiffusionBench: On Holistic Evaluation of Diffusion Transformers.

KVAE: Family of Tokenizers for Multimodal Generative Models DiffusionBench: On Holistic Evaluation of Diffusion Transformers

Reference 54

Resolution
verified exact
local_arxiv, observed 2026-08-07T23:24:50.096776Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T23:24:49.132533Z digest=sha256:1b3c40ac002a958c0fe5fef5bb85851720fb0f13d00fdd088367fbc7c96897c8

Observation 83adf514-2156-4f71-a374-ffa46d3c452f · outbound

This paper cites Wf-vae: Enhancing video vae by wavelet-driven energy flow for latent video diffusion model,.

KVAE: Family of Tokenizers for Multimodal Generative Models Wf-vae: Enhancing video vae by wavelet-driven energy flow for latent video diffusion model,

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T23:24:50.999575Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T23:24:49.137542Z digest=sha256:c8931222dab404d89e7a368e5d149ef7c47e37273776b090588758e4e3d948bd

Observation e5363bd6-7328-439c-add9-f9aedcb225ab · outbound

This paper cites Generating novel, designable, and diverse protein structures by equivariantly diffusing oriented residue clouds, 2023.

KVAE: Family of Tokenizers for Multimodal Generative Models Generating novel, designable, and diverse protein structures by equivariantly diffusing oriented residue clouds, 2023

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T23:24:50.984146Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T23:24:49.142269Z digest=sha256:27a1abd3ad0a10a3928ac3aca1d77dfd34679b77c74bde0f61c83d1ff2c02d4d

Observation 267aeaba-7451-409e-b2f4-67858bde79ba · outbound

This paper cites AudioLDM 2: Learning Holistic Audio Generation with Self-supervised Pretraining.

KVAE: Family of Tokenizers for Multimodal Generative Models AudioLDM 2: Learning Holistic Audio Generation with Self-supervised Pretraining

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-07T23:24:49.146653Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:24:49.146653Z digest=sha256:6789d2036049428a73a8587bdae2aeaca0d39ea1a1cc5581f84db1d118950306

Observation 66ca06a4-e630-4e22-b35d-8cac8283534a · outbound

This paper cites Delving into Latent Spectral Biasing of Video VAEs for Superior Diffusability.

KVAE: Family of Tokenizers for Multimodal Generative Models Delving into Latent Spectral Biasing of Video VAEs for Superior Diffusability

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-07T23:24:49.151448Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:24:49.151448Z digest=sha256:fafa552fa1819ec970a91442e5dc3dbd88fe123f8b5fd5dbb5283cecfb6c89f3

Observation acc391df-336e-415e-be70-460d0f658a11 · outbound

This paper cites an unresolved cited work.

KVAE: Family of Tokenizers for Multimodal Generative Models Unresolved cited work

Reference 59

Resolution
unresolved
raw_fallback, observed 2026-08-07T23:24:50.969040Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T23:24:49.156389Z digest=sha256:534250d9d713346556078bfae207c4a15c4d331497690e0bc4370ce848eba47e

Observation b4353548-027b-4ca2-a114-f6b0f9e70e21 · outbound

This paper cites The song describer dataset: A corpus of audio captions for music-and- language evaluation.

KVAE: Family of Tokenizers for Multimodal Generative Models The song describer dataset: A corpus of audio captions for music-and- language evaluation

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T23:24:50.954476Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T23:24:49.160855Z digest=sha256:e0ea887d910291cbf5131a6e3adbf02d00070fb2c1b8b4a3a388fc900074f9b9

Observation c0d76d52-e54f-48d8-a3be-b4b6bbb0005b · outbound

This paper cites Balasubramanian.

KVAE: Family of Tokenizers for Multimodal Generative Models Balasubramanian

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T23:24:50.939257Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T23:24:49.165459Z digest=sha256:e93ffc5257b51cb40b7a15016fc2725c22a9f3cf4951ad8621163a429ab85ed5

Observation b51194c7-d6ad-4f0a-8fe1-844428a199cb · outbound

This paper cites Transition matching distillation for fast video generation, 2026.

KVAE: Family of Tokenizers for Multimodal Generative Models Transition matching distillation for fast video generation, 2026

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T23:24:50.924914Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T23:24:49.170003Z digest=sha256:c4ff7cad40b88bf0ebd9b3f5bbd64ff47bae7b66f087b7c4db9cf30a827948bf

Observation d1a2b48e-5f39-4930-9d4b-2590079c194c · outbound

This paper cites Blaschko, Albert Ali Salah, and Itir Onal Ertugrul.

KVAE: Family of Tokenizers for Multimodal Generative Models Blaschko, Albert Ali Salah, and Itir Onal Ertugrul

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-07T23:24:49.174272Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:24:49.174272Z digest=sha256:780a71fdee7bb6ae45068df92890c11f1c5c1c7a465f74f7a727d688fc195f79

Observation 81b25390-cc78-405e-975f-b87ef68d0aea · outbound

This paper cites Alpamayo-r1: Bridging reasoning and action prediction for generalizable autonomous driving in the long tail, 2026.

KVAE: Family of Tokenizers for Multimodal Generative Models Alpamayo-r1: Bridging reasoning and action prediction for generalizable autonomous driving in the long tail, 2026

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T23:24:50.910624Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T23:24:49.178680Z digest=sha256:7ede3420b2bd673f25bcafa29f7f58d15fd187234eb1474ca5418b232a705251

Observation 6377f506-b873-48fb-9081-51f4cfa0c397 · outbound

This paper cites Cosmos World Foundation Model Platform for Physical AI.

KVAE: Family of Tokenizers for Multimodal Generative Models Cosmos World Foundation Model Platform for Physical AI

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-07T23:24:49.183042Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:24:49.183042Z digest=sha256:5ea86b1910491d75ddcf994daf94ea38bfdc7c4eea44286d3962ab005039b71e

Observation 800b8ce9-b089-4322-971e-87cb2867cc2f · outbound

This paper cites To create what you tell: Generating videos from captions, 2018.

KVAE: Family of Tokenizers for Multimodal Generative Models To create what you tell: Generating videos from captions, 2018

Reference 66

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T23:24:50.895129Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T23:24:49.187278Z digest=sha256:6c394591486066631502c894d29e0935a608bff4e9302fdebd5345348d1f77de

Observation bbe6bb1f-7dfe-4235-8b32-0602637df5de · outbound

This paper cites Librispeech: An ASR corpus based on public domain audio books.

KVAE: Family of Tokenizers for Multimodal Generative Models Librispeech: An ASR corpus based on public domain audio books

Reference 67

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T23:24:50.880432Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T23:24:49.192154Z digest=sha256:f1943fbb0cd2fd0ed53a702e947ba740b7289fce1e92cddd4185390f79b732b1

Observation be565eb9-cfe7-46b6-bb9b-8be96549aa1e · outbound

This paper cites Parker, Zach Evans, CJ Carr, Zack Zukowski, Josiah Taylor, Matthew Rice, and Jordi Pons.

KVAE: Family of Tokenizers for Multimodal Generative Models Parker, Zach Evans, CJ Carr, Zack Zukowski, Josiah Taylor, Matthew Rice, and Jordi Pons

Reference 68

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T23:24:50.864864Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T23:24:49.196758Z digest=sha256:85beb6a9bce515a3857e771ff839313418dfaf8d166de53ebc2dfee2235c8d92

Observation 8b31fb12-eda1-4f4a-bf0f-82d1d1e71768 · outbound

This paper cites Sampson, Shikai Li, Simone Parmeggiani, Steve Fine, Tara Fowler, Vladan Petrovic, and Yuming Du.

KVAE: Family of Tokenizers for Multimodal Generative Models Sampson, Shikai Li, Simone Parmeggiani, Steve Fine, Tara Fowler, Vladan Petrovic, and Yuming Du

Reference 69

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T23:24:50.849477Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T23:24:49.201249Z digest=sha256:57ffe063f849b105ac75ebac8a4996d663c23d026c9029dafc4407e12d933f60

Observation ca168e10-08e7-4eab-87da-b326e642041e · outbound

This paper cites Andersson, Andrew El-Kadi, Dominic Masters, Timo Ewalds, Jacklynn Stott, Shakir Mohamed, Peter Battaglia, Remi Lam, and Matthew Willson.

KVAE: Family of Tokenizers for Multimodal Generative Models Andersson, Andrew El-Kadi, Dominic Masters, Timo Ewalds, Jacklynn Stott, Shakir Mohamed, Peter Battaglia, Remi Lam, and Matthew Willson

Reference 70

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T23:24:50.833535Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T23:24:49.205992Z digest=sha256:90c074d5e68d93cd9e81930057f2cdbbb446982bba629067a2577b08474eaf10

Observation 0ff90aab-604b-4284-bacd-ad48092a822b · outbound

This paper cites Qwen-Audio-VAE Technical Report.

KVAE: Family of Tokenizers for Multimodal Generative Models Qwen-Audio-VAE Technical Report

Reference 71

Resolution
verified exact
local_arxiv, observed 2026-08-07T23:24:49.823936Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T23:24:49.210558Z digest=sha256:76a915a652b75e9330d26d2cbc47328be36e53d2155c1676006e698c1214699a

Observation 396ca9dd-f10f-4d1b-9a33-54a7aca03ca7 · outbound

This paper cites Qwen-Image-VAE-2.0 Technical Report.

KVAE: Family of Tokenizers for Multimodal Generative Models Qwen-Image-VAE-2.0 Technical Report

Reference 72

Resolution
unresolved
no resolver link, observed 2026-08-07T23:24:49.215146Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:24:49.215146Z digest=sha256:1c5ba31f28fb37b9123cfc0c45a0a75c1e3a5fbdfc4a1c92e5d38d774d088b01

Observation a1a0de29-3159-47e1-bae3-88f9a00de7e8 · outbound

This paper cites Learning transferable visual models from natural language supervision.

KVAE: Family of Tokenizers for Multimodal Generative Models Learning transferable visual models from natural language supervision

Reference 73

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T23:24:50.817684Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T23:24:49.219882Z digest=sha256:a0d5f9eb08097c4cb5f0f024e643954dde756e75a6c1cee0b39d985dff8d80e8

Observation 56dab5a3-9387-4627-a2a5-79b447587466 · outbound

This paper cites MUSDB18-HQ — an uncompressed version of MUSDB18, 2019.

KVAE: Family of Tokenizers for Multimodal Generative Models MUSDB18-HQ — an uncompressed version of MUSDB18, 2019

Reference 74

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T23:24:50.802189Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T23:24:49.224296Z digest=sha256:f656ef7cca9d7a30113bc100c2f9a0a400734fc1cd436948d50b326e4af6f9f7

Observation 403980d2-0bb7-4e58-8a4e-b51a022855e0 · outbound

This paper cites EARS: An anechoic fullband speech dataset benchmarked for speech enhancement and dereverberation.

KVAE: Family of Tokenizers for Multimodal Generative Models EARS: An anechoic fullband speech dataset benchmarked for speech enhancement and dereverberation

Reference 75

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T23:24:50.787304Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T23:24:49.228947Z digest=sha256:9624d6672c24cebe0076af5781f1cbb49b5382e508c44a83e13c4d14b9a5b7e3

Observation 908a6963-34df-4061-9ae6-a918534a033b · outbound

This paper cites High-resolution image synthesis with latent diffusion models, 2022.

KVAE: Family of Tokenizers for Multimodal Generative Models High-resolution image synthesis with latent diffusion models, 2022

Reference 76

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T23:24:50.772663Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T23:24:49.233612Z digest=sha256:7c50cc988b04cd338019930854e95a2050b25930101a64c04fdd641831f1eb31

Observation df019351-9721-4388-8474-f420028c64c2 · outbound

This paper cites High-resolution image synthesis with latent diffusion models, 2022.

KVAE: Family of Tokenizers for Multimodal Generative Models High-resolution image synthesis with latent diffusion models, 2022

Reference 77

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T23:24:50.757578Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T23:24:49.238165Z digest=sha256:b073234859cd9614e8167ce0e05de4906fef8d3722ce0a319556477f6722dcc7

Observation d2ee0a73-a87d-4d66-8420-375818288256 · outbound

This paper cites Runway gen-4: Ai video generation with world consistency.

KVAE: Family of Tokenizers for Multimodal Generative Models Runway gen-4: Ai video generation with world consistency

Reference 78

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T23:24:50.742254Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T23:24:49.242655Z digest=sha256:2d9e6c14dbeddeb662b50214c2cac0c70b9c4fad1e7f4e9db7dda56b13c2b633

Observation 1381ca67-e28f-4a49-b640-88ed3f1cd99e · outbound

This paper cites Flow to the mode: Mode-seeking diffusion autoencoders for state-of-the-art image tokenization, 2025.

KVAE: Family of Tokenizers for Multimodal Generative Models Flow to the mode: Mode-seeking diffusion autoencoders for state-of-the-art image tokenization, 2025

Reference 79

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T23:24:50.725842Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T23:24:49.246801Z digest=sha256:35654e484df98ecc50c9ce36a9c4395436c945808ceecca8374164971393cbe5

Observation f28f6329-029f-4897-a6f0-4573d0b56e2d · outbound

This paper cites Make- a-video: Text-to-video generation without text-video data, 2022.

KVAE: Family of Tokenizers for Multimodal Generative Models Make- a-video: Text-to-video generation without text-video data, 2022

Reference 80

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T23:24:50.710306Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T23:24:49.251220Z digest=sha256:9e0f0771bcba67abe83806bc3ac6ffb8c41fef0f09d284b9fab89eb243a2ee02

Observation a188e79e-5a70-4a33-b517-84db242a8fa8 · outbound

This paper cites What matters for representation alignment: Global information or spatial structure?arXiv preprint arXiv:2512.10794, 2025.

KVAE: Family of Tokenizers for Multimodal Generative Models What matters for representation alignment: Global information or spatial structure?arXiv preprint arXiv:2512.10794, 2025

Reference 81

Resolution
unresolved
no resolver link, observed 2026-08-07T23:24:49.255842Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:24:49.255842Z digest=sha256:b3c49b605a87dc16680e6fdff8d04cc54b73df437def99586c59355a287df8ff

Observation 6cc40914-4d56-4056-b59a-1fdd5dd4ac21 · outbound

This paper cites Improving the diffusability of autoencoders, 2025.

KVAE: Family of Tokenizers for Multimodal Generative Models Improving the diffusability of autoencoders, 2025

Reference 82

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T23:24:50.693904Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T23:24:49.260320Z digest=sha256:d503f822fd68f2c7d84b92a19a932f84f65b5da47f1a4b4c61340084ef1e7b82

Observation 76dfbbe9-c306-4d70-bad0-2e7965de56f9 · outbound

This paper cites Stylegan-v: A continuous video generator with the price, image quality and perks of stylegan2, 2022.

KVAE: Family of Tokenizers for Multimodal Generative Models Stylegan-v: A continuous video generator with the price, image quality and perks of stylegan2, 2022

Reference 83

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T23:24:50.678729Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T23:24:49.264895Z digest=sha256:1ed108b0705e82a9525af8cc84b14c7843a05409e5f4f0e1670ca6e99e0215b0

Observation 8570bb5d-e8b1-4fe9-8313-b7bdf698ea11 · outbound

This paper cites Ucf101: A dataset of 101 human actions classes from videos in the wild, 2012.

KVAE: Family of Tokenizers for Multimodal Generative Models Ucf101: A dataset of 101 human actions classes from videos in the wild, 2012

Reference 84

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T23:24:50.663478Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T23:24:49.269466Z digest=sha256:8ecffae0df816b1edb95444adbc1b8a5a3cf5e09f8089ca3c0f23b619f6f0d8a

Observation 19383de8-b3f2-4913-8671-bc2eb2dd16c3 · outbound

This paper cites Scenediffuser++: City-scale traffic simulation via a generative world model, 2025.

KVAE: Family of Tokenizers for Multimodal Generative Models Scenediffuser++: City-scale traffic simulation via a generative world model, 2025

Reference 85

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T23:24:50.647889Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T23:24:49.273993Z digest=sha256:585211c5997fa7d7643d7a20bc17487878fb72813bcfccd8d4e3659bacfdf4d9

Observation 95573368-f86d-4172-9e9e-868bfa40d269 · outbound

This paper cites Z-image: An efficient image generation foundation model with single-stream diffusion transformer,.

KVAE: Family of Tokenizers for Multimodal Generative Models Z-image: An efficient image generation foundation model with single-stream diffusion transformer,

Reference 86

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T23:24:50.632963Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T23:24:49.278379Z digest=sha256:f98c086ccd5d345e83df348d5768c06386499d4570bbb9cadb2d1c9dea429068

Observation 5405c1df-101f-4983-925d-d8fbf8ec6146 · outbound

This paper cites an unresolved cited work.

KVAE: Family of Tokenizers for Multimodal Generative Models Unresolved cited work

Reference 87

Resolution
unresolved
raw_fallback, observed 2026-08-07T23:24:50.617480Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T23:24:49.282806Z digest=sha256:d715533592b44c0be27cf6251c621be046e8ce2a6ede6b41fbe0430195e67747

Observation 49e126f2-5c89-408a-bd6a-c64e1f68d805 · outbound

This paper cites Nextstep-1: Toward autoregressive image generation with continuous tokens at scale, 2025.

KVAE: Family of Tokenizers for Multimodal Generative Models Nextstep-1: Toward autoregressive image generation with continuous tokens at scale, 2025

Reference 88

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T23:24:50.602974Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T23:24:49.287263Z digest=sha256:8c3366e184c5c6bc24f17e5a3c5bf9ed38c430c2dd2e0b7fed770816ac22324c

Observation 54f8ed22-adc4-477d-aa77-894691e0c336 · outbound

This paper cites HunyuanVideo-Foley: Multimodal Diffusion with Representation Alignment for High-Fidelity Foley Audio Generation.

KVAE: Family of Tokenizers for Multimodal Generative Models HunyuanVideo-Foley: Multimodal Diffusion with Representation Alignment for High-Fidelity Foley Audio Generation

Reference 89

Resolution
unresolved
no resolver link, observed 2026-08-07T23:24:49.291826Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:24:49.291826Z digest=sha256:90cb880a84bc92af4496efafe3278599bb3a4f806b3b678af8560d28aaa12d25

Observation fa0cbf27-e6fc-431d-af41-a2baebeba92c · outbound

This paper cites Reducio! generating 1k video within 16 seconds using extremely compressed motion latents, 2025.

KVAE: Family of Tokenizers for Multimodal Generative Models Reducio! generating 1k video within 16 seconds using extremely compressed motion latents, 2025

Reference 90

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T23:24:50.587939Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T23:24:49.296716Z digest=sha256:b2e75c18299c35fd01a51f519d4d5696f9c70fc5ecc1b436f00ca38782fc6d90

Observation efee10b7-6d58-44ed-b07b-e9ab1e95acaf · outbound

This paper cites Metaxas, and Sergey Tulyakov.

KVAE: Family of Tokenizers for Multimodal Generative Models Metaxas, and Sergey Tulyakov

Reference 91

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T23:24:50.573136Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T23:24:49.301476Z digest=sha256:ee4236ea97a2496a67ad321b6c4b88535d8cbfefc897f8b575f9c834ef3e5f85

Observation 23c4e403-bc4d-43c7-8fee-7ada83688ad1 · outbound

This paper cites Meta Audiobox Aesthetics: Unified Automatic Quality Assessment for Speech, Music, and Sound.

KVAE: Family of Tokenizers for Multimodal Generative Models Meta Audiobox Aesthetics: Unified Automatic Quality Assessment for Speech, Music, and Sound

Reference 92

Resolution
unresolved
no resolver link, observed 2026-08-07T23:24:49.306053Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:24:49.306053Z digest=sha256:1bc60c277b010e5d4170803bef4789071c1bd980f6a20535ed856075efbcacd4

Observation efeb5b68-bd37-45d3-a383-dc5adebe7538 · outbound

This paper cites Ssdd: Single-step diffusion decoder for efficient image tokenization, 2026.

KVAE: Family of Tokenizers for Multimodal Generative Models Ssdd: Single-step diffusion decoder for efficient image tokenization, 2026

Reference 93

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T23:24:50.558721Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T23:24:49.310789Z digest=sha256:4e813472c4cfe2d8c5b4e9375f10f431f97c934e3d9a5b11e8b53e3368a4559d

Observation f8046108-5b8b-4da1-b9d0-9f8cde99f1d4 · outbound

This paper cites Conditional image generation with pixelcnn decoders, 2016.

KVAE: Family of Tokenizers for Multimodal Generative Models Conditional image generation with pixelcnn decoders, 2016

Reference 94

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T23:24:50.544106Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T23:24:49.315228Z digest=sha256:852b77bb5f55f5cc193ac2d89fed57e63a789fee8257635aaef957dc3427d6f0

Observation 235a4e58-ae68-4a14-8a35-8677b551f010 · outbound

This paper cites Neural discrete representation learning, 2018.

KVAE: Family of Tokenizers for Multimodal Generative Models Neural discrete representation learning, 2018

Reference 95

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T23:24:50.529444Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T23:24:49.319830Z digest=sha256:cc07387e6b9631e940f7a9347c07547ce3a5e851969cf82a3f1b888ccecd6bc5

Observation fcb04103-a5e0-40db-82c6-a454a4cc5187 · outbound

This paper cites Gomez, Lukasz Kaiser, and Illia Polosukhin.

KVAE: Family of Tokenizers for Multimodal Generative Models Gomez, Lukasz Kaiser, and Illia Polosukhin

Reference 96

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T23:24:50.514926Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T23:24:49.324257Z digest=sha256:583efe738714be21ecf7e17643b6a27690eed37cf9a4b61b6bce6b1bb05b177e

Observation 2283e7e1-6c9b-45bc-a3ee-a12990febe3d · outbound

This paper cites Generating videos with scene dynamics, 2016.

KVAE: Family of Tokenizers for Multimodal Generative Models Generating videos with scene dynamics, 2016

Reference 97

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T23:24:50.499693Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T23:24:49.328663Z digest=sha256:7196cc24d4a21fedb8d8be023b4a005e187348ad276ab5a72ec3635cf44766be

Observation 3d01af25-d7ad-4d26-9e55-82118e8e8c08 · outbound

This paper cites Wan: Open and advanced large-scale video generative models, 2025.

KVAE: Family of Tokenizers for Multimodal Generative Models Wan: Open and advanced large-scale video generative models, 2025

Reference 98

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T23:24:50.484897Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T23:24:49.333015Z digest=sha256:d4c3233a280d100976608a2372810da1ce54dca244bb42516ac2594493879740

Observation 69657ab6-1c35-43ea-8ef2-6a4476db2986 · outbound

This paper cites Wan-2.2 anouncement.

KVAE: Family of Tokenizers for Multimodal Generative Models Wan-2.2 anouncement

Reference 99

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T23:24:50.470707Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T23:24:49.337918Z digest=sha256:acfb4d0e7b9a1a8c6f8cecbf46bc0a626d3d6610574c9b2ad43819d32419628a

Observation 76571650-0562-4483-9884-118e27a89523 · outbound

This paper cites an unresolved cited work.

KVAE: Family of Tokenizers for Multimodal Generative Models Unresolved cited work

Reference 100

Resolution
unresolved
raw_fallback, observed 2026-08-07T23:24:50.456314Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T23:24:49.342282Z digest=sha256:29c8e54945681730eb3ae8c3b70e2caa329ed22768e4e947c91c84430fc8099b

Observation 56c741e4-ba26-43b8-a506-dd18820a4bb4 · outbound

This paper cites Videomae v2: Scaling video masked autoencoders with dual masking.

KVAE: Family of Tokenizers for Multimodal Generative Models Videomae v2: Scaling video masked autoencoders with dual masking

Reference 101

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T23:24:50.442381Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T23:24:49.346598Z digest=sha256:cce47b4c8940701a3616d40ae39e3a5b6eedce468569735dea533da1505fac8a

Pith citing papers

No inbound Pith citation observations are available.