Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-05-08T03:57:38.298693Z
Paper Citation Record · LEDGER
As of 6 August 2026, this Paper Citation Record lists 63 of 63 outbound references and 0 inbound Pith citation observations for arXiv:2605.06407.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-05-08T03:57:38.298693Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-06T06:34:29.942622+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
63 of 63 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 1a5b5db7-b8e1-4500-a3b8-e1e22eb9396b · outbound
WavCube: Unifying Speech Representation for Understanding and Generation via Semantic-Acoustic Joint Modeling wav2vec 2.0: A framework for self-supervised learning of speech representations.Proc.NIPS
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 0b4fa7bb-89d3-4046-8aa2-64ea35599fd8 · outbound
WavCube: Unifying Speech Representation for Understanding and Generation via Semantic-Acoustic Joint Modeling Semanticgen: Video generation in semantic space
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 5567082c-745b-434d-8021-3dec5a131289 · outbound
WavCube: Unifying Speech Representation for Understanding and Generation via Semantic-Acoustic Joint Modeling Dino-sae: Dino spherical autoencoder for high-fidelity image reconstruction and generation
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 5b4c785c-d610-4ca8-801a-7025db119237 · outbound
WavCube: Unifying Speech Representation for Understanding and Generation via Semantic-Acoustic Joint Modeling Aligning visual foundation encoders to tokenizers for diffusion models
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation a7981266-e00d-4c78-93b1-4050edcec904 · outbound
WavCube: Unifying Speech Representation for Understanding and Generation via Semantic-Acoustic Joint Modeling Wavlm: Large-scale self-supervised pre-training for full stack speech processing.Proc.JSTSP
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation cdc84f77-62c6-4362-8f66-e0bb374087ce · outbound
WavCube: Unifying Speech Representation for Understanding and Generation via Semantic-Acoustic Joint Modeling F5-tts: A fairytaler that fakes fluent and faithful speech with flow matching
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 1d4b5985-8f0d-49fe-9ba2-468c013fdcbc · outbound
WavCube: Unifying Speech Representation for Understanding and Generation via Semantic-Acoustic Joint Modeling Large-scale self-supervised speech representation learning for automatic speaker verification
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation f89d1a5d-f419-47b1-8823-3a8a06858779 · outbound
WavCube: Unifying Speech Representation for Understanding and Generation via Semantic-Acoustic Joint Modeling On the distillation loss functions of speech vae for unified reconstruction, understanding, and generation
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 72c49639-6ed0-4037-8d83-073618180500 · outbound
WavCube: Unifying Speech Representation for Understanding and Generation via Semantic-Acoustic Joint Modeling Emerging properties in unified multimodal pretraining
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 3acfc358-8e54-4a0f-baa5-9e983f9ab79b · outbound
WavCube: Unifying Speech Representation for Understanding and Generation via Semantic-Acoustic Joint Modeling Dashengtokenizer: One layer is enough for unified audio understanding and generation
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 47a5ce0d-6d48-441d-ad9e-c73dcd679740 · outbound
WavCube: Unifying Speech Representation for Understanding and Generation via Semantic-Acoustic Joint Modeling RePack: Representation packing of vision foundation model features enhances diffusion transformer
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 353da288-f00d-4c7c-82e4-05118952e1d6 · outbound
WavCube: Unifying Speech Representation for Understanding and Generation via Semantic-Acoustic Joint Modeling Vqrae: Representation quantization autoencoders for multimodal understanding, generation and reconstruction
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 0ef03bc7-bb72-4882-b7ca-15ef767d6c2f · outbound
WavCube: Unifying Speech Representation for Understanding and Generation via Semantic-Acoustic Joint Modeling Cosyvoice: A scalable multilingual zero-shot text-to-speech synthesizer based on supervised semantic tokens
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation ee5c5b21-5cb8-40f3-a5ba-329e68b0eb97 · outbound
WavCube: Unifying Speech Representation for Understanding and Generation via Semantic-Acoustic Joint Modeling E2 tts: Embarrassingly easy fully non-autoregressive zero-shot tts
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation b34e2d68-2dea-4b97-bb10-76e40495fc58 · outbound
WavCube: Unifying Speech Representation for Understanding and Generation via Semantic-Acoustic Joint Modeling Scaling rectified flow transform- ers for high-resolution image synthesis
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 3c21039e-63ef-429a-b36a-5b8deb1eb1e9 · outbound
WavCube: Unifying Speech Representation for Understanding and Generation via Semantic-Acoustic Joint Modeling Stable audio open
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation c876adfc-0a34-4e5d-9d48-ceffa4fb810a · outbound
WavCube: Unifying Speech Representation for Understanding and Generation via Semantic-Acoustic Joint Modeling Unified autoregressive visual generation and under- standing with continuous tokens
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 8ffc97e6-beab-49b3-8f52-0b89a3d59fff · outbound
WavCube: Unifying Speech Representation for Understanding and Generation via Semantic-Acoustic Joint Modeling The prism hypothesis: Harmonizing semantic and pixel representations via unified autoencoding
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation fc7ea41a-3ae4-4a53-9f58-df1cf0c49091 · outbound
WavCube: Unifying Speech Representation for Understanding and Generation via Semantic-Acoustic Joint Modeling One layer is enough: Adapting pretrained visual encoders for image generation
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 0dc84dfe-0978-4ed0-87f2-25d5d6048935 · outbound
WavCube: Unifying Speech Representation for Understanding and Generation via Semantic-Acoustic Joint Modeling Rpiae: A representation-pivoted autoencoder enhancing both image generation and editing
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 9341eaba-4edd-468d-9e4d-b5ee815f93cb · outbound
WavCube: Unifying Speech Representation for Understanding and Generation via Semantic-Acoustic Joint Modeling Fireredtts: A foundation text-to-speech framework for industry-level generative speech applications
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 076ee37e-d19c-455c-a442-e2621a9818a4 · outbound
WavCube: Unifying Speech Representation for Understanding and Generation via Semantic-Acoustic Joint Modeling Dera: Decoupled representation alignment for video tokenization
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 9c78bd01-f01b-45a4-94c3-6730cf6ed50e · outbound
WavCube: Unifying Speech Representation for Understanding and Generation via Semantic-Acoustic Joint Modeling Emilia: An extensive, multilingual, and diverse speech dataset for large-scale speech generation
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 03569692-1b96-46e4-83a5-b9f04cd56c34 · outbound
WavCube: Unifying Speech Representation for Understanding and Generation via Semantic-Acoustic Joint Modeling Unified latents (ul): How to train your latents
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation a4396a53-4b04-4535-b732-0b1d66927549 · outbound
WavCube: Unifying Speech Representation for Understanding and Generation via Semantic-Acoustic Joint Modeling Hubert: Self-supervised speech representation learning by masked prediction of hidden units.Proc.TASLP
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation f8fdc8e7-4e69-427a-9aa7-0f7e6f29dcee · outbound
WavCube: Unifying Speech Representation for Understanding and Generation via Semantic-Acoustic Joint Modeling Meanflow trans- formers with representation autoencoders
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 22d56740-0d77-42c4-833b-d8d335682d04 · outbound
WavCube: Unifying Speech Representation for Understanding and Generation via Semantic-Acoustic Joint Modeling Libriheavy: A 50,000 hours asr corpus with punctuation casing and context
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 11020210-052a-4376-839f-5138a8b20b59 · outbound
WavCube: Unifying Speech Representation for Understanding and Generation via Semantic-Acoustic Joint Modeling Planning in 8 tokens: A compact discrete tokenizer for latent world model
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation cfa8f662-60a1-4d3a-80b7-225f264b407d · outbound
WavCube: Unifying Speech Representation for Understanding and Generation via Semantic-Acoustic Joint Modeling Toward diffusible high-dimensional latent spaces: A frequency perspective
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation e3be5ef7-acee-4ad8-990a-b33b1c854a3a · outbound
WavCube: Unifying Speech Representation for Understanding and Generation via Semantic-Acoustic Joint Modeling Repa-e: Unlocking vae for end-to-end tuning of latent diffusion transformers
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 03ee103d-472b-4605-9486-19260405146d · outbound
WavCube: Unifying Speech Representation for Understanding and Generation via Semantic-Acoustic Joint Modeling Mogao: An omni foundation model for interleaved multi-modal generation
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 624348d6-389d-4c26-8b94-adefa348426e · outbound
WavCube: Unifying Speech Representation for Understanding and Generation via Semantic-Acoustic Joint Modeling Tuna: Taming unified visual representations for native unified multimodal models
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 8df0f97f-ea1e-4bc3-8070-0bd9d116822b · outbound
WavCube: Unifying Speech Representation for Understanding and Generation via Semantic-Acoustic Joint Modeling Pixelgen: Pixel diffusion beats latent diffusion with perceptual loss
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 8420f11e-0cd0-4a15-b670-087c054041f8 · outbound
WavCube: Unifying Speech Representation for Understanding and Generation via Semantic-Acoustic Joint Modeling Self-supervised speech representation learning: A review.Proc.JSTSP
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 945f4e9e-b6b8-41ec-b291-8ae955989e45 · outbound
WavCube: Unifying Speech Representation for Understanding and Generation via Semantic-Acoustic Joint Modeling Semantic-vae: Semantic-alignment latent representation for better speech synthesis
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 796b709a-fe72-40f3-9692-6147da8e07fc · outbound
WavCube: Unifying Speech Representation for Understanding and Generation via Semantic-Acoustic Joint Modeling Dinov2: Learning robust visual features without supervision
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 44bd08bb-c6d5-4ca5-86b7-2db5f2765ffc · outbound
WavCube: Unifying Speech Representation for Understanding and Generation via Semantic-Acoustic Joint Modeling Semantics lead the way: Harmonizing semantic and texture modeling with asynchronous latent diffusion
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation e38fe8ac-e9ae-4342-9b55-c9d5ac22e624 · outbound
WavCube: Unifying Speech Representation for Understanding and Generation via Semantic-Acoustic Joint Modeling Librispeech: an asr corpus based on public domain audio books
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation a39dfb3d-a2a5-4ee9-8988-d09a3d266a28 · outbound
WavCube: Unifying Speech Representation for Understanding and Generation via Semantic-Acoustic Joint Modeling Esc: Dataset for environmental sound classification
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 2229547e-7223-4e48-943c-24d7f0b42212 · outbound
WavCube: Unifying Speech Representation for Understanding and Generation via Semantic-Acoustic Joint Modeling Robust speech recognition via large-scale weak supervision
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation ea19573a-d43a-4006-9497-f4c4432b4e5b · outbound
WavCube: Unifying Speech Representation for Understanding and Generation via Semantic-Acoustic Joint Modeling Utmos: Utokyo-sarulab system for voicemos challenge 2022
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 23434823-e6c4-45d3-b204-74bc250ad0c9 · outbound
WavCube: Unifying Speech Representation for Understanding and Generation via Semantic-Acoustic Joint Modeling Latent diffusion model without variational autoencoder
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 49757e23-4e06-484c-afb0-bf0185213c8d · outbound
WavCube: Unifying Speech Representation for Understanding and Generation via Semantic-Acoustic Joint Modeling V ocos: Closing the gap between time-domain and fourier-based neural vocoders for high-quality audio synthesis
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation cbec838a-4299-4777-b1b0-59f35dad94b2 · outbound
WavCube: Unifying Speech Representation for Understanding and Generation via Semantic-Acoustic Joint Modeling Magicodec: Simple masked gaussian-injected codec for high-fidelity reconstruction and generation
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 19d63470-fa9c-4956-9cec-7f38ec747a8d · outbound
WavCube: Unifying Speech Representation for Understanding and Generation via Semantic-Acoustic Joint Modeling Multimodal latent language modeling with next-token diffusion
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation c4762025-a3f0-4acc-9922-744fd3f7c3c8 · outbound
WavCube: Unifying Speech Representation for Understanding and Generation via Semantic-Acoustic Joint Modeling Scaling text-to-image diffusion transformers with representation autoencoders
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation bfa511e5-6a30-4a24-8a69-7d9bdac8dd37 · outbound
WavCube: Unifying Speech Representation for Understanding and Generation via Semantic-Acoustic Joint Modeling Semanticvocoder: Bridging audio generation and audio understanding via semantic latents
Reference 47
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 09af182e-1063-4306-8987-4057aad00c6f · outbound
WavCube: Unifying Speech Representation for Understanding and Generation via Semantic-Acoustic Joint Modeling Ming-uniaudio: Speech llm for joint understanding, generation and editing with unified representation
Reference 48
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 97890df2-f0ed-4595-87fc-22fde36dd534 · outbound
WavCube: Unifying Speech Representation for Understanding and Generation via Semantic-Acoustic Joint Modeling Superb: Speech processing universal performance benchmark
Reference 49
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation b7188616-fdcd-436f-ac93-ee5710d1ad99 · outbound
WavCube: Unifying Speech Representation for Understanding and Generation via Semantic-Acoustic Joint Modeling A survey of unified multimodal understanding and generation: Advances and challenges.Authorea Preprints
Reference 50
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 1646acd4-4f89-4498-b244-ec4353f4435e · outbound
WavCube: Unifying Speech Representation for Understanding and Generation via Semantic-Acoustic Joint Modeling Towards scalable pre-training of visual tokenizers for generation
Reference 51
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 788c1e23-d045-45c1-a7c8-7b37dedef0c6 · outbound
WavCube: Unifying Speech Representation for Understanding and Generation via Semantic-Acoustic Joint Modeling Reconstruction vs
Reference 52
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation dcbd8e9b-0b9f-4b8e-997b-d1238c2b5cc1 · outbound
WavCube: Unifying Speech Representation for Understanding and Generation via Semantic-Acoustic Joint Modeling Distribution matching variational autoencoder
Reference 53
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation c001eb19-247e-4fe6-adf4-30b8c11ee4e0 · outbound
WavCube: Unifying Speech Representation for Understanding and Generation via Semantic-Acoustic Joint Modeling Representation alignment for generation: Training diffusion transformers is easier than you think
Reference 54
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 0ef2fefd-703a-48e8-8a01-8279a4538fc4 · outbound
WavCube: Unifying Speech Representation for Understanding and Generation via Semantic-Acoustic Joint Modeling Libritts: A corpus derived from librispeech for text-to-speech
Reference 55
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 2d172464-4a49-4daa-abf4-4ffbd514521d · outbound
WavCube: Unifying Speech Representation for Understanding and Generation via Semantic-Acoustic Joint Modeling Sigmoid loss for language image pre-training
Reference 56
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation a063cb65-f995-49b8-b9c2-ea39021f0339 · outbound
WavCube: Unifying Speech Representation for Understanding and Generation via Semantic-Acoustic Joint Modeling Mimo-audio: Audio language models are few-shot learners
Reference 57
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation fdf06353-1a0e-485f-826e-9dc74de59bac · outbound
WavCube: Unifying Speech Representation for Understanding and Generation via Semantic-Acoustic Joint Modeling Openvision 3: A family of unified visual encoder for both understanding and generation
Reference 58
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 2d5bee4b-a855-4cda-a0ff-051c04601585 · outbound
WavCube: Unifying Speech Representation for Understanding and Generation via Semantic-Acoustic Joint Modeling Rae- nwm: Navigation world model in dense visual representation space
Reference 59
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation f0a7d532-0664-4f0c-a804-7c531c4942ec · outbound
WavCube: Unifying Speech Representation for Understanding and Generation via Semantic-Acoustic Joint Modeling Both semantics and reconstruction matter: Making representation encoders ready for text-to-image generation and editing
Reference 60
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 1d5392e2-3d7e-49ff-a75b-ee31b52980b6 · outbound
WavCube: Unifying Speech Representation for Understanding and Generation via Semantic-Acoustic Joint Modeling Efficient image-goal navigation with representative latent world model
Reference 61
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 4a81d87f-a1b4-4d32-af54-fbc6f9e9a846 · outbound
WavCube: Unifying Speech Representation for Understanding and Generation via Semantic-Acoustic Joint Modeling Unified multimodal understanding and generation models: Advances, challenges, and opportunities
Reference 62
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 986a14f4-00de-4b34-b7bd-bb6594d5f021 · outbound
WavCube: Unifying Speech Representation for Understanding and Generation via Semantic-Acoustic Joint Modeling Diffusion transformers with representation autoencoders
Reference 63
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
No inbound Pith citation observations are available.