Pith. sign in

Paper Citation Record · LEDGER

WavCube: Unifying Speech Representation for Understanding and Generation via Semantic-Acoustic Joint Modeling

As of 6 August 2026, this Paper Citation Record lists 63 of 63 outbound references and 0 inbound Pith citation observations for arXiv:2605.06407.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2605.06407 v1

Coverage vector

measured 63 of 63 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-05-08T03:57:38.298693Z

measured 63 of 63 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-06T06:34:29.942622+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

63 of 63 outbound references displayed

  • verified exact0
  • verified fuzzy63
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 1a5b5db7-b8e1-4500-a3b8-e1e22eb9396b · outbound

This paper cites wav2vec 2.0: A framework for self-supervised learning of speech representations.Proc.NIPS.

WavCube: Unifying Speech Representation for Understanding and Generation via Semantic-Acoustic Joint Modeling wav2vec 2.0: A framework for self-supervised learning of speech representations.Proc.NIPS

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T21:28:22.811214Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-08T03:57:38.298693Z digest=sha256:9dfa827467898926a1ca9c84252183633d61b19f4ef30ac4153af115c865b2dc

Observation 0b4fa7bb-89d3-4046-8aa2-64ea35599fd8 · outbound

This paper cites Semanticgen: Video generation in semantic space.

WavCube: Unifying Speech Representation for Understanding and Generation via Semantic-Acoustic Joint Modeling Semanticgen: Video generation in semantic space

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T21:28:22.806449Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-08T03:57:38.298693Z digest=sha256:b2820d22a5a59b5b780401db07655b406344ebe00c747e4679ff8b8146743158

Observation 5567082c-745b-434d-8021-3dec5a131289 · outbound

This paper cites Dino-sae: Dino spherical autoencoder for high-fidelity image reconstruction and generation.

WavCube: Unifying Speech Representation for Understanding and Generation via Semantic-Acoustic Joint Modeling Dino-sae: Dino spherical autoencoder for high-fidelity image reconstruction and generation

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T21:28:22.793701Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-08T03:57:38.298693Z digest=sha256:eae0515bbe867e74f98ba9fb03fc95af51e938836548a6bdd7b3676a252a342a

Observation 5b4c785c-d610-4ca8-801a-7025db119237 · outbound

This paper cites Aligning visual foundation encoders to tokenizers for diffusion models.

WavCube: Unifying Speech Representation for Understanding and Generation via Semantic-Acoustic Joint Modeling Aligning visual foundation encoders to tokenizers for diffusion models

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T21:28:22.775224Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-08T03:57:38.298693Z digest=sha256:83b969cf52def46bb852637ac6fee03218eb786fe5150c903e93d2aada423648

Observation a7981266-e00d-4c78-93b1-4050edcec904 · outbound

This paper cites Wavlm: Large-scale self-supervised pre-training for full stack speech processing.Proc.JSTSP.

WavCube: Unifying Speech Representation for Understanding and Generation via Semantic-Acoustic Joint Modeling Wavlm: Large-scale self-supervised pre-training for full stack speech processing.Proc.JSTSP

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T21:28:22.771771Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-08T03:57:38.298693Z digest=sha256:2c8ba2329e4d94a00cea946bdc09cdad1b4c9de75c2342a28ea1c78d3cd0b351

Observation cdc84f77-62c6-4362-8f66-e0bb374087ce · outbound

This paper cites F5-tts: A fairytaler that fakes fluent and faithful speech with flow matching.

WavCube: Unifying Speech Representation for Understanding and Generation via Semantic-Acoustic Joint Modeling F5-tts: A fairytaler that fakes fluent and faithful speech with flow matching

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T21:28:22.778667Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-08T03:57:38.298693Z digest=sha256:8f8f2a4db46d12a506baa66a134f4f2297bc57bfb75dfb3b03fee2f4b87f4b25

Observation 1d4b5985-8f0d-49fe-9ba2-468c013fdcbc · outbound

This paper cites Large-scale self-supervised speech representation learning for automatic speaker verification.

WavCube: Unifying Speech Representation for Understanding and Generation via Semantic-Acoustic Joint Modeling Large-scale self-supervised speech representation learning for automatic speaker verification

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T21:28:22.764782Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-08T03:57:38.298693Z digest=sha256:567cc1a03e9647fe65f15b14dc849f50f18ccb829cc96c0ce7b0691aa447c137

Observation f89d1a5d-f419-47b1-8823-3a8a06858779 · outbound

This paper cites On the distillation loss functions of speech vae for unified reconstruction, understanding, and generation.

WavCube: Unifying Speech Representation for Understanding and Generation via Semantic-Acoustic Joint Modeling On the distillation loss functions of speech vae for unified reconstruction, understanding, and generation

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T21:28:22.710147Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-08T03:57:38.298693Z digest=sha256:78c23ad4acd74011b1a927fd8cd169a66af1c40433324408c9e230274223e5e8

Observation 72c49639-6ed0-4037-8d83-073618180500 · outbound

This paper cites Emerging properties in unified multimodal pretraining.

WavCube: Unifying Speech Representation for Understanding and Generation via Semantic-Acoustic Joint Modeling Emerging properties in unified multimodal pretraining

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T21:28:22.713185Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-08T03:57:38.298693Z digest=sha256:26cdaa466c0cb0ae2ea3723e11b6f2b5726360c50504e409392b4c0c5ef189c9

Observation 3acfc358-8e54-4a0f-baa5-9e983f9ab79b · outbound

This paper cites Dashengtokenizer: One layer is enough for unified audio understanding and generation.

WavCube: Unifying Speech Representation for Understanding and Generation via Semantic-Acoustic Joint Modeling Dashengtokenizer: One layer is enough for unified audio understanding and generation

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T21:28:22.761462Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-08T03:57:38.298693Z digest=sha256:8f327f93d813562ec139883724fbfccac863c9e305c1a11474e93020bee4bc64

Observation 47a5ce0d-6d48-441d-ad9e-c73dcd679740 · outbound

This paper cites RePack: Representation packing of vision foundation model features enhances diffusion transformer.

WavCube: Unifying Speech Representation for Understanding and Generation via Semantic-Acoustic Joint Modeling RePack: Representation packing of vision foundation model features enhances diffusion transformer

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T21:28:22.768413Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-08T03:57:38.298693Z digest=sha256:362578cc92563f2987d3d97f06502fb131d2cd770dd2492c0026f7bdce5857f4

Observation 353da288-f00d-4c7c-82e4-05118952e1d6 · outbound

This paper cites Vqrae: Representation quantization autoencoders for multimodal understanding, generation and reconstruction.

WavCube: Unifying Speech Representation for Understanding and Generation via Semantic-Acoustic Joint Modeling Vqrae: Representation quantization autoencoders for multimodal understanding, generation and reconstruction

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T21:28:22.782309Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-08T03:57:38.298693Z digest=sha256:b53502140a2fc60867f4326dbaddffdc2bc6cc8c050de0969d58d65267a608f1

Observation 0ef03bc7-bb72-4882-b7ca-15ef767d6c2f · outbound

This paper cites Cosyvoice: A scalable multilingual zero-shot text-to-speech synthesizer based on supervised semantic tokens.

WavCube: Unifying Speech Representation for Understanding and Generation via Semantic-Acoustic Joint Modeling Cosyvoice: A scalable multilingual zero-shot text-to-speech synthesizer based on supervised semantic tokens

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T21:28:22.785861Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-08T03:57:38.298693Z digest=sha256:4c89daaf1ae06e098860e39639520a647ae8170f87105d9300f4951a1e7826ee

Observation ee5c5b21-5cb8-40f3-a5ba-329e68b0eb97 · outbound

This paper cites E2 tts: Embarrassingly easy fully non-autoregressive zero-shot tts.

WavCube: Unifying Speech Representation for Understanding and Generation via Semantic-Acoustic Joint Modeling E2 tts: Embarrassingly easy fully non-autoregressive zero-shot tts

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T21:28:22.789668Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-08T03:57:38.298693Z digest=sha256:61fea9504033154b8808f2d26eaeae3410e2b5e0e7b5e4089149eb0d0b4a7842

Observation b34e2d68-2dea-4b97-bb10-76e40495fc58 · outbound

This paper cites Scaling rectified flow transform- ers for high-resolution image synthesis.

WavCube: Unifying Speech Representation for Understanding and Generation via Semantic-Acoustic Joint Modeling Scaling rectified flow transform- ers for high-resolution image synthesis

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T21:28:22.820045Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-08T03:57:38.298693Z digest=sha256:503a4d51d9789d9564e5cf1e5b5522d6bba2deccc14defe11398a0591111e38d

Observation 3c21039e-63ef-429a-b36a-5b8deb1eb1e9 · outbound

This paper cites Stable audio open.

WavCube: Unifying Speech Representation for Understanding and Generation via Semantic-Acoustic Joint Modeling Stable audio open

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T21:28:22.861520Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-08T03:57:38.298693Z digest=sha256:e447b240e1996d03f6f87cda1e4cf1509c1d7c5d5bcde5631e306534f27b1b82

Observation c876adfc-0a34-4e5d-9d48-ceffa4fb810a · outbound

This paper cites Unified autoregressive visual generation and under- standing with continuous tokens.

WavCube: Unifying Speech Representation for Understanding and Generation via Semantic-Acoustic Joint Modeling Unified autoregressive visual generation and under- standing with continuous tokens

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T21:28:22.743188Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-08T03:57:38.298693Z digest=sha256:ba2b711a43834a5c385efa8e81c1cd9078a6be3a2df6e4b1b3e7f8dd8c35e608

Observation 8ffc97e6-beab-49b3-8f52-0b89a3d59fff · outbound

This paper cites The prism hypothesis: Harmonizing semantic and pixel representations via unified autoencoding.

WavCube: Unifying Speech Representation for Understanding and Generation via Semantic-Acoustic Joint Modeling The prism hypothesis: Harmonizing semantic and pixel representations via unified autoencoding

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T21:28:22.852504Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-08T03:57:38.298693Z digest=sha256:c6ab95ac993ac2890c44a615b204cfe8a8043f58a889b9e82114f31fe745106e

Observation fc7ea41a-3ae4-4a53-9f58-df1cf0c49091 · outbound

This paper cites One layer is enough: Adapting pretrained visual encoders for image generation.

WavCube: Unifying Speech Representation for Understanding and Generation via Semantic-Acoustic Joint Modeling One layer is enough: Adapting pretrained visual encoders for image generation

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T21:28:22.823676Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-08T03:57:38.298693Z digest=sha256:8d5970671b4ba4bde1982ce9ed14a79ab92d57b5e1d6d2d7114f7e0cf001bef7

Observation 0dc84dfe-0978-4ed0-87f2-25d5d6048935 · outbound

This paper cites Rpiae: A representation-pivoted autoencoder enhancing both image generation and editing.

WavCube: Unifying Speech Representation for Understanding and Generation via Semantic-Acoustic Joint Modeling Rpiae: A representation-pivoted autoencoder enhancing both image generation and editing

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T21:28:22.753465Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-08T03:57:38.298693Z digest=sha256:dee274860d015f90c935554a35211bba6d2c49d746cadf09075ce0b183cfd16a

Observation 9341eaba-4edd-468d-9e4d-b5ee815f93cb · outbound

This paper cites Fireredtts: A foundation text-to-speech framework for industry-level generative speech applications.

WavCube: Unifying Speech Representation for Understanding and Generation via Semantic-Acoustic Joint Modeling Fireredtts: A foundation text-to-speech framework for industry-level generative speech applications

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T21:28:22.746642Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-08T03:57:38.298693Z digest=sha256:51998db8bd62c324e9849b0ff0f21e5e7800284beea0a563810a0113a3f38a77

Observation 076ee37e-d19c-455c-a442-e2621a9818a4 · outbound

This paper cites Dera: Decoupled representation alignment for video tokenization.

WavCube: Unifying Speech Representation for Understanding and Generation via Semantic-Acoustic Joint Modeling Dera: Decoupled representation alignment for video tokenization

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T21:28:22.750012Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-08T03:57:38.298693Z digest=sha256:f5ac9ba8c2842e8e5a79cd9356f1e2bcff79edc8749200b7d8e4dbc57b7dfab4

Observation 9c78bd01-f01b-45a4-94c3-6730cf6ed50e · outbound

This paper cites Emilia: An extensive, multilingual, and diverse speech dataset for large-scale speech generation.

WavCube: Unifying Speech Representation for Understanding and Generation via Semantic-Acoustic Joint Modeling Emilia: An extensive, multilingual, and diverse speech dataset for large-scale speech generation

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T21:28:22.723789Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-08T03:57:38.298693Z digest=sha256:efdcd010e4a8d19f1d892e0ed40295f881283ef553fcae3e6f80fca7e306d81b

Observation 03569692-1b96-46e4-83a5-b9f04cd56c34 · outbound

This paper cites Unified latents (ul): How to train your latents.

WavCube: Unifying Speech Representation for Understanding and Generation via Semantic-Acoustic Joint Modeling Unified latents (ul): How to train your latents

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T21:28:22.731260Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-08T03:57:38.298693Z digest=sha256:c6b925e681a1ed9153304b797e18b69fc9c5a357d9ebc3529af3b08618b9958f

Observation a4396a53-4b04-4535-b732-0b1d66927549 · outbound

This paper cites Hubert: Self-supervised speech representation learning by masked prediction of hidden units.Proc.TASLP.

WavCube: Unifying Speech Representation for Understanding and Generation via Semantic-Acoustic Joint Modeling Hubert: Self-supervised speech representation learning by masked prediction of hidden units.Proc.TASLP

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T21:28:22.830842Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-08T03:57:38.298693Z digest=sha256:723a8cf3e7e70e9c57de1464cef2e0b0d882fcd8fe268c2d05da42b5ced40806

Observation f8fdc8e7-4e69-427a-9aa7-0f7e6f29dcee · outbound

This paper cites Meanflow trans- formers with representation autoencoders.

WavCube: Unifying Speech Representation for Understanding and Generation via Semantic-Acoustic Joint Modeling Meanflow trans- formers with representation autoencoders

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T21:28:22.716037Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-08T03:57:38.298693Z digest=sha256:db4e1d8da003b71eccda3353bb32a935a8813e2e11577c43408edfdb9feb5555

Observation 22d56740-0d77-42c4-833b-d8d335682d04 · outbound

This paper cites Libriheavy: A 50,000 hours asr corpus with punctuation casing and context.

WavCube: Unifying Speech Representation for Understanding and Generation via Semantic-Acoustic Joint Modeling Libriheavy: A 50,000 hours asr corpus with punctuation casing and context

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T21:28:22.849191Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-08T03:57:38.298693Z digest=sha256:a0a6bc4cdecff5565ec1f3f83981b3398c98e356af19c1e353696e4b1b2a8f06

Observation 11020210-052a-4376-839f-5138a8b20b59 · outbound

This paper cites Planning in 8 tokens: A compact discrete tokenizer for latent world model.

WavCube: Unifying Speech Representation for Understanding and Generation via Semantic-Acoustic Joint Modeling Planning in 8 tokens: A compact discrete tokenizer for latent world model

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T21:28:22.712448Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-08T03:57:38.298693Z digest=sha256:2eba1f2b74d1b9a3409ce75c5c5138948680ff2cfc94391dc3995ded12b5cdd1

Observation cfa8f662-60a1-4d3a-80b7-225f264b407d · outbound

This paper cites Toward diffusible high-dimensional latent spaces: A frequency perspective.

WavCube: Unifying Speech Representation for Understanding and Generation via Semantic-Acoustic Joint Modeling Toward diffusible high-dimensional latent spaces: A frequency perspective

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T21:28:22.760944Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-08T03:57:38.298693Z digest=sha256:e744d0f48e8d21684e3ed8f57c293d24945523958461f296c6638faf6544cfd7

Observation e3be5ef7-acee-4ad8-990a-b33b1c854a3a · outbound

This paper cites Repa-e: Unlocking vae for end-to-end tuning of latent diffusion transformers.

WavCube: Unifying Speech Representation for Understanding and Generation via Semantic-Acoustic Joint Modeling Repa-e: Unlocking vae for end-to-end tuning of latent diffusion transformers

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T21:28:22.786638Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-08T03:57:38.298693Z digest=sha256:3fc95d36ba13a3936eb594498a33d7272324890784773fcc22085d78968baba1

Observation 03ee103d-472b-4605-9486-19260405146d · outbound

This paper cites Mogao: An omni foundation model for interleaved multi-modal generation.

WavCube: Unifying Speech Representation for Understanding and Generation via Semantic-Acoustic Joint Modeling Mogao: An omni foundation model for interleaved multi-modal generation

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T21:28:22.753367Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-08T03:57:38.298693Z digest=sha256:7ca8d0d7ec3742236265487635cf20c73e83103de4d7a52d416500c2114fe848

Observation 624348d6-389d-4c26-8b94-adefa348426e · outbound

This paper cites Tuna: Taming unified visual representations for native unified multimodal models.

WavCube: Unifying Speech Representation for Understanding and Generation via Semantic-Acoustic Joint Modeling Tuna: Taming unified visual representations for native unified multimodal models

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T21:28:22.708830Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-08T03:57:38.298693Z digest=sha256:bcd1dca6c979895e29b1c7ffa91aebfb5f66e8a8e75371c52f329277cc6a8d6f

Observation 8df0f97f-ea1e-4bc3-8070-0bd9d116822b · outbound

This paper cites Pixelgen: Pixel diffusion beats latent diffusion with perceptual loss.

WavCube: Unifying Speech Representation for Understanding and Generation via Semantic-Acoustic Joint Modeling Pixelgen: Pixel diffusion beats latent diffusion with perceptual loss

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T21:28:22.725000Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-08T03:57:38.298693Z digest=sha256:72cdd44b19a17ea8ca270b5fd20b2019d7eb7ada53ac4ee3c38e8909c08c61a7

Observation 8420f11e-0cd0-4a15-b670-087c054041f8 · outbound

This paper cites Self-supervised speech representation learning: A review.Proc.JSTSP.

WavCube: Unifying Speech Representation for Understanding and Generation via Semantic-Acoustic Joint Modeling Self-supervised speech representation learning: A review.Proc.JSTSP

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T21:28:22.767124Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-08T03:57:38.298693Z digest=sha256:4671ce19af70dd6b3ac68fe10cdee5a87064d087cd0b277cdd8c0ea206805552

Observation 945f4e9e-b6b8-41ec-b291-8ae955989e45 · outbound

This paper cites Semantic-vae: Semantic-alignment latent representation for better speech synthesis.

WavCube: Unifying Speech Representation for Understanding and Generation via Semantic-Acoustic Joint Modeling Semantic-vae: Semantic-alignment latent representation for better speech synthesis

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T21:28:22.827390Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-08T03:57:38.298693Z digest=sha256:d0d1676c192b21e5556f007a6e00470ec0160ca64f596b6e340b1b578f90cc14

Observation 796b709a-fe72-40f3-9692-6147da8e07fc · outbound

This paper cites Dinov2: Learning robust visual features without supervision.

WavCube: Unifying Speech Representation for Understanding and Generation via Semantic-Acoustic Joint Modeling Dinov2: Learning robust visual features without supervision

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T21:28:22.722298Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-08T03:57:38.298693Z digest=sha256:1c82d6ace99c01d1ce2a388fd91379cf6a95ab58c864ba54b5a3381baccbca7d

Observation 44bd08bb-c6d5-4ca5-86b7-2db5f2765ffc · outbound

This paper cites Semantics lead the way: Harmonizing semantic and texture modeling with asynchronous latent diffusion.

WavCube: Unifying Speech Representation for Understanding and Generation via Semantic-Acoustic Joint Modeling Semantics lead the way: Harmonizing semantic and texture modeling with asynchronous latent diffusion

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T21:28:22.834443Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-08T03:57:38.298693Z digest=sha256:3c782009e88ee1af3da7aa1e99db42c58c544e85aeb093504a586ad9646f748d

Observation e38fe8ac-e9ae-4342-9b55-c9d5ac22e624 · outbound

This paper cites Librispeech: an asr corpus based on public domain audio books.

WavCube: Unifying Speech Representation for Understanding and Generation via Semantic-Acoustic Joint Modeling Librispeech: an asr corpus based on public domain audio books

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T21:28:22.682833Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-08T03:57:38.298693Z digest=sha256:c9c3de77c145be839100b13353a2f1aee8eef35cd4f646cb4ff05cd4614a1305

Observation a39dfb3d-a2a5-4ee9-8988-d09a3d266a28 · outbound

This paper cites Esc: Dataset for environmental sound classification.

WavCube: Unifying Speech Representation for Understanding and Generation via Semantic-Acoustic Joint Modeling Esc: Dataset for environmental sound classification

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T21:28:22.686294Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-08T03:57:38.298693Z digest=sha256:0f391950362925c58cfe25152146c20ba36dfefc9e64524d13b77717dc2b36fe

Observation 2229547e-7223-4e48-943c-24d7f0b42212 · outbound

This paper cites Robust speech recognition via large-scale weak supervision.

WavCube: Unifying Speech Representation for Understanding and Generation via Semantic-Acoustic Joint Modeling Robust speech recognition via large-scale weak supervision

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T21:28:22.675581Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-08T03:57:38.298693Z digest=sha256:ea40d8726d53e8b7b41b8bf86e8a5653284bc2be4283b3969616270f45e8744a

Observation ea19573a-d43a-4006-9497-f4c4432b4e5b · outbound

This paper cites Utmos: Utokyo-sarulab system for voicemos challenge 2022.

WavCube: Unifying Speech Representation for Understanding and Generation via Semantic-Acoustic Joint Modeling Utmos: Utokyo-sarulab system for voicemos challenge 2022

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T21:28:22.841850Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-08T03:57:38.298693Z digest=sha256:688d8dc079f6faaaba5ca79eef0ddc1183ca7b74adea6068664d2cbae58b3691

Observation 23434823-e6c4-45d3-b204-74bc250ad0c9 · outbound

This paper cites Latent diffusion model without variational autoencoder.

WavCube: Unifying Speech Representation for Understanding and Generation via Semantic-Acoustic Joint Modeling Latent diffusion model without variational autoencoder

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T21:28:22.838072Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-08T03:57:38.298693Z digest=sha256:c8c61b871b5f99f3282f4c8085f77ad2bd7792d4fba9b6bfaffad68e53e0c185

Observation 49757e23-4e06-484c-afb0-bf0185213c8d · outbound

This paper cites V ocos: Closing the gap between time-domain and fourier-based neural vocoders for high-quality audio synthesis.

WavCube: Unifying Speech Representation for Understanding and Generation via Semantic-Acoustic Joint Modeling V ocos: Closing the gap between time-domain and fourier-based neural vocoders for high-quality audio synthesis

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T21:28:22.769903Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-08T03:57:38.298693Z digest=sha256:0ddf8a18492e414e1a29682ead4af805e8b839d1f1841b9797d7cf4912c75bdf

Observation cbec838a-4299-4777-b1b0-59f35dad94b2 · outbound

This paper cites Magicodec: Simple masked gaussian-injected codec for high-fidelity reconstruction and generation.

WavCube: Unifying Speech Representation for Understanding and Generation via Semantic-Acoustic Joint Modeling Magicodec: Simple masked gaussian-injected codec for high-fidelity reconstruction and generation

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T21:28:22.845697Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-08T03:57:38.298693Z digest=sha256:fb6cac5d69dfc12586bec903ca426de216fad2620af2f964fe806ffac0c12d10

Observation 19d63470-fa9c-4956-9cec-7f38ec747a8d · outbound

This paper cites Multimodal latent language modeling with next-token diffusion.

WavCube: Unifying Speech Representation for Understanding and Generation via Semantic-Acoustic Joint Modeling Multimodal latent language modeling with next-token diffusion

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T21:28:22.679137Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-08T03:57:38.298693Z digest=sha256:a784ecc395d257ff64b8fb69bb02b4083966f5e12cb212491f4d0177544077ef

Observation c4762025-a3f0-4acc-9922-744fd3f7c3c8 · outbound

This paper cites Scaling text-to-image diffusion transformers with representation autoencoders.

WavCube: Unifying Speech Representation for Understanding and Generation via Semantic-Acoustic Joint Modeling Scaling text-to-image diffusion transformers with representation autoencoders

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T21:28:22.778369Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-08T03:57:38.298693Z digest=sha256:a1d24a323bc309c16f11c03d83e070d055b91f01b2440c3dac6b47292ae27638

Observation bfa511e5-6a30-4a24-8a69-7d9bdac8dd37 · outbound

This paper cites Semanticvocoder: Bridging audio generation and audio understanding via semantic latents.

WavCube: Unifying Speech Representation for Understanding and Generation via Semantic-Acoustic Joint Modeling Semanticvocoder: Bridging audio generation and audio understanding via semantic latents

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T21:28:22.772913Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-08T03:57:38.298693Z digest=sha256:39f78d27f716fa1fb77a98c5829d6cef8d2978bfa8953094f672abc8ba7f17e4

Observation 09af182e-1063-4306-8987-4057aad00c6f · outbound

This paper cites Ming-uniaudio: Speech llm for joint understanding, generation and editing with unified representation.

WavCube: Unifying Speech Representation for Understanding and Generation via Semantic-Acoustic Joint Modeling Ming-uniaudio: Speech llm for joint understanding, generation and editing with unified representation

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T21:28:22.757823Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-08T03:57:38.298693Z digest=sha256:bccf353e4a71c92a1e9aa5f3bdf3e0677f7d54cd09aac1f2bff6b33d794fb047

Observation 97890df2-f0ed-4595-87fc-22fde36dd534 · outbound

This paper cites Superb: Speech processing universal performance benchmark.

WavCube: Unifying Speech Representation for Understanding and Generation via Semantic-Acoustic Joint Modeling Superb: Speech processing universal performance benchmark

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T21:28:22.775601Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-08T03:57:38.298693Z digest=sha256:7d36d5cf5deb33e05594092f91340bf55dfa7c2c17f337445b8eaa7136ab274d

Observation b7188616-fdcd-436f-ac93-ee5710d1ad99 · outbound

This paper cites A survey of unified multimodal understanding and generation: Advances and challenges.Authorea Preprints.

WavCube: Unifying Speech Representation for Understanding and Generation via Semantic-Acoustic Joint Modeling A survey of unified multimodal understanding and generation: Advances and challenges.Authorea Preprints

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T21:28:22.780948Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-08T03:57:38.298693Z digest=sha256:fe38d36fb3e9b72ef1ac71161d7e490fb30c598eadb6d466b539880d4a6d0de5

Observation 1646acd4-4f89-4498-b244-ec4353f4435e · outbound

This paper cites Towards scalable pre-training of visual tokenizers for generation.

WavCube: Unifying Speech Representation for Understanding and Generation via Semantic-Acoustic Joint Modeling Towards scalable pre-training of visual tokenizers for generation

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T21:28:22.668542Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-08T03:57:38.298693Z digest=sha256:a6d3608284196c661097a4bbb264f4ca5236766bc55e30ce71082465b087228f

Observation 788c1e23-d045-45c1-a7c8-7b37dedef0c6 · outbound

This paper cites Reconstruction vs.

WavCube: Unifying Speech Representation for Understanding and Generation via Semantic-Acoustic Joint Modeling Reconstruction vs

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T21:28:22.858438Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-08T03:57:38.298693Z digest=sha256:6ff06cd997770a4aa182b459d06fe2d6dc851a1972deac8a20d2e5031e37f7f8

Observation dcbd8e9b-0b9f-4b8e-997b-d1238c2b5cc1 · outbound

This paper cites Distribution matching variational autoencoder.

WavCube: Unifying Speech Representation for Understanding and Generation via Semantic-Acoustic Joint Modeling Distribution matching variational autoencoder

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T21:28:22.649604Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-08T03:57:38.298693Z digest=sha256:2a6f76a82f8bc5b7ca6537f3ca01fba6faa72e3d04db71eb44e2d7877b8a00d1

Observation c001eb19-247e-4fe6-adf4-30b8c11ee4e0 · outbound

This paper cites Representation alignment for generation: Training diffusion transformers is easier than you think.

WavCube: Unifying Speech Representation for Understanding and Generation via Semantic-Acoustic Joint Modeling Representation alignment for generation: Training diffusion transformers is easier than you think

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T21:28:22.660563Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-08T03:57:38.298693Z digest=sha256:6bc7cfd93e3a7a2bd235682d2e8a8d67ed68356147db330f2113f0baaa4761a5

Observation 0ef2fefd-703a-48e8-8a01-8279a4538fc4 · outbound

This paper cites Libritts: A corpus derived from librispeech for text-to-speech.

WavCube: Unifying Speech Representation for Understanding and Generation via Semantic-Acoustic Joint Modeling Libritts: A corpus derived from librispeech for text-to-speech

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T21:28:22.670159Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-08T03:57:38.298693Z digest=sha256:ebbafa65cf0aeac3b06002b2aec4c86f5ce44a3da158ab5257db26437e7464f1

Observation 2d172464-4a49-4daa-abf4-4ffbd514521d · outbound

This paper cites Sigmoid loss for language image pre-training.

WavCube: Unifying Speech Representation for Understanding and Generation via Semantic-Acoustic Joint Modeling Sigmoid loss for language image pre-training

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T21:28:22.719709Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-08T03:57:38.298693Z digest=sha256:4bd02334b9ca38e87f9b68ca45a01090773d3055c2cc5459dc651e6fd45bb03f

Observation a063cb65-f995-49b8-b9c2-ea39021f0339 · outbound

This paper cites Mimo-audio: Audio language models are few-shot learners.

WavCube: Unifying Speech Representation for Understanding and Generation via Semantic-Acoustic Joint Modeling Mimo-audio: Audio language models are few-shot learners

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T21:28:22.737874Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-08T03:57:38.298693Z digest=sha256:eb2db113bba122a6bbd69a39b83256a9896ec50367fcaacf40976676188e0330

Observation fdf06353-1a0e-485f-826e-9dc74de59bac · outbound

This paper cites Openvision 3: A family of unified visual encoder for both understanding and generation.

WavCube: Unifying Speech Representation for Understanding and Generation via Semantic-Acoustic Joint Modeling Openvision 3: A family of unified visual encoder for both understanding and generation

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T21:28:22.634948Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-08T03:57:38.298693Z digest=sha256:147f2e837180d727efa2a4c2c6432320a0a6e5b845b11714a19cfe81e21f0b8b

Observation 2d5bee4b-a855-4cda-a0ff-051c04601585 · outbound

This paper cites Rae- nwm: Navigation world model in dense visual representation space.

WavCube: Unifying Speech Representation for Understanding and Generation via Semantic-Acoustic Joint Modeling Rae- nwm: Navigation world model in dense visual representation space

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T21:28:22.638498Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-08T03:57:38.298693Z digest=sha256:68290dcf0ba2980f76432d11ab6334d4f7fc8aeadbcb0dea871636b08380d953

Observation f0a7d532-0664-4f0c-a804-7c531c4942ec · outbound

This paper cites Both semantics and reconstruction matter: Making representation encoders ready for text-to-image generation and editing.

WavCube: Unifying Speech Representation for Understanding and Generation via Semantic-Acoustic Joint Modeling Both semantics and reconstruction matter: Making representation encoders ready for text-to-image generation and editing

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T21:28:22.646338Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-08T03:57:38.298693Z digest=sha256:f50d2e480b0ef97e76eeeef18dbb39d2ef379a3d466808d4fd21e505dcd19fe8

Observation 1d5392e2-3d7e-49ff-a75b-ee31b52980b6 · outbound

This paper cites Efficient image-goal navigation with representative latent world model.

WavCube: Unifying Speech Representation for Understanding and Generation via Semantic-Acoustic Joint Modeling Efficient image-goal navigation with representative latent world model

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T21:28:22.628030Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-08T03:57:38.298693Z digest=sha256:d7e9be1c23a6373913df6450b018b8fb552f24f6fa37c9781bd077a310c90acd

Observation 4a81d87f-a1b4-4d32-af54-fbc6f9e9a846 · outbound

This paper cites Unified multimodal understanding and generation models: Advances, challenges, and opportunities.

WavCube: Unifying Speech Representation for Understanding and Generation via Semantic-Acoustic Joint Modeling Unified multimodal understanding and generation models: Advances, challenges, and opportunities

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T21:28:22.631689Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-08T03:57:38.298693Z digest=sha256:3d12c52d7242b59aed6e1b4f6d4a0c816ea13f753fbeb305ee0351ba1cad47f3

Observation 986a14f4-00de-4b34-b7bd-bb6594d5f021 · outbound

This paper cites Diffusion transformers with representation autoencoders.

WavCube: Unifying Speech Representation for Understanding and Generation via Semantic-Acoustic Joint Modeling Diffusion transformers with representation autoencoders

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T21:28:22.816029Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-08T03:57:38.298693Z digest=sha256:eb21af2d4e60beb176c50c6edeb23b8d0522d788f399b5b8d8e7bd442969b5a6

Pith citing papers

No inbound Pith citation observations are available.