Pith. sign in

Paper Citation Record · LEDGER

Breaking the Barriers of Text-Hungry and Audio-Deficient AI

As of 7 August 2026, this Paper Citation Record lists 100 of 165 outbound references and 3 inbound Pith citation observations for arXiv:2506.02443.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.02443 v1

Coverage vector

measured 100 of 165 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T11:28:54.976320Z

measured 103 of 103 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 3 of 3 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-06-26T09:46:15.701501Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T09:39:46.872455Z

Reference resolution

100 of 165 outbound references displayed

  • verified exact17
  • verified fuzzy0
  • unresolved83
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation c6c2304a-23ac-4914-ba28-4f55fc9e71c3 · outbound

This paper cites Simultron: On-device simultaneous speech to speech translation.

Breaking the Barriers of Text-Hungry and Audio-Deficient AI Simultron: On-device simultaneous speech to speech translation

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T11:28:45.767223Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:28:45.767223Z digest=sha256:dd63621502df4da106dad1ada67474190f5b5ad9051a73f8a7d9bf8d884a6d74

Observation 873835c3-1a3e-4f91-b08d-eda15d57db18 · outbound

This paper cites Qwen 2.5: A comprehensive review of the leading resource-efficient llm with potentioal to surpass all competitors.

Breaking the Barriers of Text-Hungry and Audio-Deficient AI Qwen 2.5: A comprehensive review of the leading resource-efficient llm with potentioal to surpass all competitors

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T11:28:45.821059Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:28:45.821059Z digest=sha256:d08d243663224a74aeba495fbfffc863da3371ec06b9ac0fe465be4bebe3a134

Observation 32a0ea4b-8888-4f91-83a8-5d35e1858e85 · outbound

This paper cites Analysis of layer- wise training in direct speech to speech translation using bi-lstm.

Breaking the Barriers of Text-Hungry and Audio-Deficient AI Analysis of layer- wise training in direct speech to speech translation using bi-lstm

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T11:28:45.924856Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:28:45.924856Z digest=sha256:48ace40d6df1f02d625ac3f730fdb5e315cc99a79faea551625c80f287309658

Observation 2a273406-4bde-489e-bdd9-71d54fcd955a · outbound

This paper cites Precipitation nowcasting with generative diffusion models.

Breaking the Barriers of Text-Hungry and Audio-Deficient AI Precipitation nowcasting with generative diffusion models

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T11:28:46.038879Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:28:46.038879Z digest=sha256:a0549ce1b70d254dd34f9e11496c99f7d458d994b4262cff14dcaf7b63c52d74

Observation 3b8fe074-38c4-4a16-aaab-be8e14324918 · outbound

This paper cites Mean-Field-Type Game Theory: Applica- tions, volume 2.

Breaking the Barriers of Text-Hungry and Audio-Deficient AI Mean-Field-Type Game Theory: Applica- tions, volume 2

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T11:28:46.191304Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:28:46.191304Z digest=sha256:4ad7c06b0266e1eb55a434f88d5fef392c0107e535ed98a5d7600be0f9f48223

Observation 50a12541-1011-418a-890b-61774aaf52b1 · outbound

This paper cites an unresolved cited work.

Breaking the Barriers of Text-Hungry and Audio-Deficient AI Unresolved cited work

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T11:28:46.297443Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:28:46.297443Z digest=sha256:e6c69cfeecd51795daac621b6d6a79a948a915d8a2d7ec8cd6992530f71fee3c

Observation f7990542-4be9-4a87-90b4-55da5b2b7ea4 · outbound

This paper cites Listen and Translate: A Proof of Concept for End-to-End Speech-to-Text Translation.

Breaking the Barriers of Text-Hungry and Audio-Deficient AI Listen and Translate: A Proof of Concept for End-to-End Speech-to-Text Translation

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T11:28:46.397405Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:28:46.397405Z digest=sha256:13c0b1ec5ecf63fd04fcc18d4aec7f2bbb1724c66c939f9a80a0c900e0c62c1a

Observation 1c201fa9-12a9-4fbd-933a-bcdce127e9ac · outbound

This paper cites an unresolved cited work.

Breaking the Barriers of Text-Hungry and Audio-Deficient AI Unresolved cited work

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T11:28:46.487794Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:28:46.487794Z digest=sha256:b953fb0f60417d2e8b1f97703489ae3d91fab269c01114b4f667fd4a8f66f8ab

Observation c72b7324-91de-4b01-9e1e-f8fae1dd1ebd · outbound

This paper cites Large language models are strong audio-visual speech recognition learners.

Breaking the Barriers of Text-Hungry and Audio-Deficient AI Large language models are strong audio-visual speech recognition learners

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T11:28:46.586257Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:28:46.586257Z digest=sha256:70f62e4364fa624270745701eed80b4ea5b88ab144b0437f3e515975163099e4

Observation 82501d3c-2e97-4271-9597-0295d15f6637 · outbound

This paper cites Low frame-rate speech codec: a codec designed for fast high-quality speech llm training and inference.

Breaking the Barriers of Text-Hungry and Audio-Deficient AI Low frame-rate speech codec: a codec designed for fast high-quality speech llm training and inference

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T11:28:46.705659Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:28:46.705659Z digest=sha256:08732ef1be3c4af2e0afea06ce676894696679238ba4e4ba71bc0cd873fb2458

Observation 8355f08a-cb85-47f3-8dae-a183f17c899a · outbound

This paper cites A speech-to-speech translation based interface for tourism.

Breaking the Barriers of Text-Hungry and Audio-Deficient AI A speech-to-speech translation based interface for tourism

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T11:28:46.827097Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:28:46.827097Z digest=sha256:1590d4c603d512abf31cf42dbf1718b0445236fd8b4ec036dc7de62cc902e605

Observation 257d59fa-52f2-480f-942e-00eab9818c01 · outbound

This paper cites Exploring in-context learning of textless speech language model for speech classification tasks.

Breaking the Barriers of Text-Hungry and Audio-Deficient AI Exploring in-context learning of textless speech language model for speech classification tasks

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T11:28:46.942799Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:28:46.942799Z digest=sha256:f44061a2f17966611cad468bdf9945fb41d51fd49edc8c590c918d27c06cb78e

Observation ffd1783b-d96b-45b3-875b-cd32e9e3f7cc · outbound

This paper cites Audio Large Language Models Can Be Descriptive Speech Quality Evaluators.

Breaking the Barriers of Text-Hungry and Audio-Deficient AI Audio Large Language Models Can Be Descriptive Speech Quality Evaluators

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T11:28:47.037457Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:28:47.037457Z digest=sha256:b9636c209135e9c5e32aca352112ebac188db9643e720c1a4e3c8ffc75c6602c

Observation 50a4e804-f797-4395-8160-ed86c96a8184 · outbound

This paper cites Multi-modal generative ai: Multi-modal llm, diffusion and beyond.

Breaking the Barriers of Text-Hungry and Audio-Deficient AI Multi-modal generative ai: Multi-modal llm, diffusion and beyond

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T11:28:47.131408Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:28:47.131408Z digest=sha256:ace1c15c0573e72f557c806a4f780d40d1a3a629a8a6570365fdae4fccc5b9a9

Observation cfb96779-2c01-4c02-8f85-b926d2500d09 · outbound

This paper cites BLASER: A Text-Free Speech-to-Speech Translation Evaluation Metric.

Breaking the Barriers of Text-Hungry and Audio-Deficient AI BLASER: A Text-Free Speech-to-Speech Translation Evaluation Metric

Reference 15

Resolution
verified exact
local_arxiv, observed 2026-08-07T11:29:09.300688Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T11:28:47.272554Z digest=sha256:b6c0c5f6fbf2e47d9f19097734e93f36d3a4312935423d2ac2b890f2d870b3cd

Observation e0a91960-f798-408d-9327-e419336192c6 · outbound

This paper cites Opportunities and challenges of diffusion models for generative ai.

Breaking the Barriers of Text-Hungry and Audio-Deficient AI Opportunities and challenges of diffusion models for generative ai

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T11:28:47.395714Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:28:47.395714Z digest=sha256:4177911a42586f719c5ca0d6024bc3a8a9186beb03493b2beb2aa94e495cd9b2

Observation d123d686-eb6c-44ea-b2e1-1c3d8cda3ec5 · outbound

This paper cites Speech-to-Speech Translation For A Real-world Unwritten Language.

Breaking the Barriers of Text-Hungry and Audio-Deficient AI Speech-to-Speech Translation For A Real-world Unwritten Language

Reference 17

Resolution
verified exact
local_arxiv, observed 2026-08-07T11:29:08.984427Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T11:28:47.506438Z digest=sha256:9a2176b3f3dc72e1f37fa0278fdedc310d931fa3f4593612dc02a2e37adab01b

Observation 45848d7e-2c1f-42fb-ad7b-5181aebd560e · outbound

This paper cites Beyond Single-Audio: Advancing Multi-Audio Processing in Audio Large Language Models.

Breaking the Barriers of Text-Hungry and Audio-Deficient AI Beyond Single-Audio: Advancing Multi-Audio Processing in Audio Large Language Models

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T11:28:47.620955Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:28:47.620955Z digest=sha256:8ed346869caf6c2b89828f16d3f9f106abc41b02a2dbf661005e924535569710

Observation f85aa4a1-dd53-4ecd-855f-9464586307fb · outbound

This paper cites MAVFlow: Preserving Paralinguistic Elements with Conditional Flow Matching for Zero-Shot AV2AV Multilingual Translation.

Breaking the Barriers of Text-Hungry and Audio-Deficient AI MAVFlow: Preserving Paralinguistic Elements with Conditional Flow Matching for Zero-Shot AV2AV Multilingual Translation

Reference 19

Resolution
verified exact
local_arxiv, observed 2026-08-07T11:29:08.710391Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T11:28:47.738849Z digest=sha256:6e88ecc3cf135b500cc8b0a0b42f1c8534993c610748d33a873e868f0027b87e

Observation 4cd6b87b-9d0a-4756-914e-c97a4f230078 · outbound

This paper cites V2sflow: Video-to-speech generation with speech decomposition and rectified flow.

Breaking the Barriers of Text-Hungry and Audio-Deficient AI V2sflow: Video-to-speech generation with speech decomposition and rectified flow

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T11:28:47.864422Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:28:47.864422Z digest=sha256:f8de28723884f1e6349ac2700b0479af3ce44d981ffbba5b35b03c34560967ae

Observation ed150bb4-fd36-4975-ba85-8a9f662e12e9 · outbound

This paper cites Qwen2-Audio Technical Report.

Breaking the Barriers of Text-Hungry and Audio-Deficient AI Qwen2-Audio Technical Report

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T11:28:47.996341Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:28:47.996341Z digest=sha256:13a500212e9bced5ca387ee6c38d0dd90e4f686c2e1c2424994449b1ffd3ebc7

Observation 70e5c870-ac9f-41f7-bfa6-0dd9bddde940 · outbound

This paper cites Qwen-Audio: Advancing Universal Audio Understanding via Unified Large-Scale Audio-Language Models.

Breaking the Barriers of Text-Hungry and Audio-Deficient AI Qwen-Audio: Advancing Universal Audio Understanding via Unified Large-Scale Audio-Language Models

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T11:28:48.113094Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:28:48.113094Z digest=sha256:aaa6453b3f2b651c3896673978d202c21bc4eb2a5d19d3636ab5ad30a97b9a86

Observation b4b26d0b-0a3d-4474-ab62-54688b122d44 · outbound

This paper cites Diffusion models in vision: A survey.

Breaking the Barriers of Text-Hungry and Audio-Deficient AI Diffusion models in vision: A survey

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T11:28:48.258509Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:28:48.258509Z digest=sha256:dbe4c48b842495bfa11d1d23f13a809a1923552175bbb97cda9806136982d021

Observation 499a2de5-a6d6-49b1-9bbe-b68e66bb5ecb · outbound

This paper cites Exploring the Benefits of Tokenization of Discrete Acoustic Units.

Breaking the Barriers of Text-Hungry and Audio-Deficient AI Exploring the Benefits of Tokenization of Discrete Acoustic Units

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T11:28:48.396469Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:28:48.396469Z digest=sha256:a3998b0658c102d33d6de2e7ef43c392bd45a551722c053ac8b85fee69b2e411

Observation aceb2984-3db4-4261-9616-02507bf01c93 · outbound

This paper cites ADIFF: Explaining audio difference using natural language.

Breaking the Barriers of Text-Hungry and Audio-Deficient AI ADIFF: Explaining audio difference using natural language

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T11:28:48.510861Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:28:48.510861Z digest=sha256:579f4d4f866baeef116847100417fadb2f7b97c291066c279b0962566d73a051

Observation ab55b59d-fcab-4bd1-88ca-d70844f0a013 · outbound

This paper cites French- fulfulde textless and cascading speech translation: Towards a dual architecture.

Breaking the Barriers of Text-Hungry and Audio-Deficient AI French- fulfulde textless and cascading speech translation: Towards a dual architecture

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T11:28:48.632625Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:28:48.632625Z digest=sha256:5a1758c68c8d23c442e3364b3bd1ecb7f4a55a9a8fb443f761b19e790194e31c

Observation d1013065-7116-412b-8efe-45c6e9a37755 · outbound

This paper cites Textless Speech-to-Speech Translation With Limited Parallel Data.

Breaking the Barriers of Text-Hungry and Audio-Deficient AI Textless Speech-to-Speech Translation With Limited Parallel Data

Reference 27

Resolution
verified exact
local_arxiv, observed 2026-08-07T11:29:08.377385Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T11:28:48.766848Z digest=sha256:b0566c45eb9f21a1f4c086c66d7e83271db7c1e5126f638845ffd6d4e306aa98

Observation a17fd686-12f2-4ec2-b5a1-979ad80e7a4b · outbound

This paper cites PolyVoice: Language Models for Speech to Speech Translation.

Breaking the Barriers of Text-Hungry and Audio-Deficient AI PolyVoice: Language Models for Speech to Speech Translation

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T11:28:48.876266Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:28:48.876266Z digest=sha256:5fd9e25d3b6f5f064d776b86c80c97e3e4c4339385416ce1a4abf409fc9cde7c

Observation 0ed51a5e-b3e9-4b28-9310-5c372fc528a5 · outbound

This paper cites LauraGPT: Listen, Attend, Understand, and Regenerate Audio with GPT.

Breaking the Barriers of Text-Hungry and Audio-Deficient AI LauraGPT: Listen, Attend, Understand, and Regenerate Audio with GPT

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T11:28:48.957781Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:28:48.957781Z digest=sha256:a812621455ea27ba252402ea17c5b24f68431b7ecad7cf21805d3f2139ee26e0

Observation 4826969a-6b2d-429a-8547-47185ea11b6d · outbound

This paper cites Analyzing Speech Unit Selection for Textless Speech-to-Speech Translation.

Breaking the Barriers of Text-Hungry and Audio-Deficient AI Analyzing Speech Unit Selection for Textless Speech-to-Speech Translation

Reference 30

Resolution
verified exact
local_arxiv, observed 2026-08-07T11:29:08.081160Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T11:28:49.074528Z digest=sha256:6524b3f9145da2e0753a32acf2b8a7359e52673815da898a3ee72e5a38df04ea

Observation 0eac4a58-04d9-427d-bad9-8e49c446503e · outbound

This paper cites Enhancing expressivity transfer in textless speech-to-speech translation.

Breaking the Barriers of Text-Hungry and Audio-Deficient AI Enhancing expressivity transfer in textless speech-to-speech translation

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T11:28:49.184864Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:28:49.184864Z digest=sha256:be465d6fd9f40d3b33bf51c0359fe9a6945a8342d402d9c0fac35a95fd699f02

Observation bb15fd2f-4b1b-47a7-8885-81dd702209da · outbound

This paper cites Towards massive parallel corpus creation for hausa-to-english machine translation.

Breaking the Barriers of Text-Hungry and Audio-Deficient AI Towards massive parallel corpus creation for hausa-to-english machine translation

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T11:28:49.309000Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:28:49.309000Z digest=sha256:7f7a1c5c0b84b5c0d66acd5b69cc28b9478dd495fc7891a90775b0c2d96752f1

Observation 43061353-529b-4bd1-b9e2-4e95189b84e7 · outbound

This paper cites Auditory-visual perception of speech.

Breaking the Barriers of Text-Hungry and Audio-Deficient AI Auditory-visual perception of speech

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T11:28:49.422521Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:28:49.422521Z digest=sha256:5491f94da9a635bd873a11395bf3a3214ff64a27d202f6d674ed023102117b1e

Observation 91bd5a83-bffa-449d-ab01-199c8b972f30 · outbound

This paper cites Cascade or direct speech translation? a case study.

Breaking the Barriers of Text-Hungry and Audio-Deficient AI Cascade or direct speech translation? a case study

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T11:28:49.526893Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:28:49.526893Z digest=sha256:e406cf092f1b1a39a276c8b5b3acbefd0e8b6bb5993f018a2ed7d1686ef98f16

Observation e5b44b4b-c15b-4aa9-b989-426f7342c4ab · outbound

This paper cites CTC-based Non-autoregressive Textless Speech-to-Speech Translation.

Breaking the Barriers of Text-Hungry and Audio-Deficient AI CTC-based Non-autoregressive Textless Speech-to-Speech Translation

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T11:28:49.651108Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:28:49.651108Z digest=sha256:2b20d7aa98ef6e746df61dc8f71a3a4fca15419e382d03d6b060f1c5e3dfd5fd

Observation b5a56457-cccd-405c-a2f8-3d656195576b · outbound

This paper cites Can We Achieve High-quality Direct Speech-to-Speech Translation without Parallel Speech Data?.

Breaking the Barriers of Text-Hungry and Audio-Deficient AI Can We Achieve High-quality Direct Speech-to-Speech Translation without Parallel Speech Data?

Reference 36

Resolution
verified exact
local_arxiv, observed 2026-08-07T11:29:07.737070Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T11:28:49.736451Z digest=sha256:521865733437fc3f8482de713466c29196c851911d4d4efdbd9f24124d7fdef7

Observation aef21d98-ff4e-4d1a-b6bc-c81f4bd7ba12 · outbound

This paper cites Generative learning of the solution of parametric partial differential equations using guided diffusion models and virtual observations.

Breaking the Barriers of Text-Hungry and Audio-Deficient AI Generative learning of the solution of parametric partial differential equations using guided diffusion models and virtual observations

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-07T11:28:49.823671Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:28:49.823671Z digest=sha256:c3dd2a6575a1676a589eac5434fb480b3b03083b715bb976c50880d3564367f2

Observation 2623a34d-a06c-4916-b297-75ecf9cc5921 · outbound

This paper cites Unsupervised speech technology for low-resource languages.

Breaking the Barriers of Text-Hungry and Audio-Deficient AI Unsupervised speech technology for low-resource languages

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-07T11:28:49.953797Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:28:49.953797Z digest=sha256:03a053772bbbfb0e99035f108b9c424485f50f05dc90006c6c37b643f6c3e242

Observation 0da42731-d665-45ed-a794-9f9946f11ffa · outbound

This paper cites Speech-to-speech translation.

Breaking the Barriers of Text-Hungry and Audio-Deficient AI Speech-to-speech translation

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-07T11:28:50.060496Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:28:50.060496Z digest=sha256:c02d8e4cc9ba919ac2c4d9b4539e711dccd9057fdadc0570f5587e7fed86da0b

Observation 6e31a667-aa47-4940-9947-2d4e8bbd8fba · outbound

This paper cites Audio Dialogues: Dialogues dataset for audio and music understanding.

Breaking the Barriers of Text-Hungry and Audio-Deficient AI Audio Dialogues: Dialogues dataset for audio and music understanding

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-07T11:28:50.157071Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:28:50.157071Z digest=sha256:a8d34f6ae600f617ed2d76c837f530aacfbe379338ddcb0e43e20a41da39876d

Observation 60b14585-6067-444f-a72e-8091a3e33557 · outbound

This paper cites Multilingual Speech-to-Speech Translation into Multiple Target Languages.

Breaking the Barriers of Text-Hungry and Audio-Deficient AI Multilingual Speech-to-Speech Translation into Multiple Target Languages

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-07T11:28:50.237516Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:28:50.237516Z digest=sha256:67ef08d46623f26dded2ce59f8fdaf5bff816dd42594a2ac6bd74849b7cdccb0

Observation f8cd0254-62a1-47a9-9fda-2f42e927c5f1 · outbound

This paper cites Joint audio and speech understanding.

Breaking the Barriers of Text-Hungry and Audio-Deficient AI Joint audio and speech understanding

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-07T11:28:50.328623Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:28:50.328623Z digest=sha256:af378c053aba2b657adb1e471c7612210710a12136ae8285b3b2591cf3214638

Observation 39fa5909-c21d-46eb-a46f-2ee384003423 · outbound

This paper cites Tibetan–chinese speech-to-speech translation based on dis- crete units.

Breaking the Barriers of Text-Hungry and Audio-Deficient AI Tibetan–chinese speech-to-speech translation based on dis- crete units

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-07T11:28:50.423111Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:28:50.423111Z digest=sha256:d4f9842364360b77701c253269cc61c5e776bba9e564580a1f0bb10c2bb9c84c

Observation 95e3ae8a-f5a0-4b8e-b1b8-f29255c55b25 · outbound

This paper cites Recent advances in discrete speech tokens: A review.

Breaking the Barriers of Text-Hungry and Audio-Deficient AI Recent advances in discrete speech tokens: A review

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-07T11:28:50.514977Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:28:50.514977Z digest=sha256:befab95df08a8b83ac72c662a4ad4308f718c3dddc1fcf9f1c234b0ca8901c0f

Observation 666aa822-1a2a-4348-bae2-6e5fe90247a1 · outbound

This paper cites Direct Speech-to-Speech Neural Machine Translation: A Survey.

Breaking the Barriers of Text-Hungry and Audio-Deficient AI Direct Speech-to-Speech Neural Machine Translation: A Survey

Reference 45

Resolution
verified exact
local_arxiv, observed 2026-08-07T11:29:07.434642Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T11:28:50.617509Z digest=sha256:17becb48417ed1ca12e50f23ac025e0ab8ab89748de8f46fafea23cbddb5d2c1

Observation 2022b7aa-ae5c-427c-a578-058301bac9b2 · outbound

This paper cites Onellm: One framework to align all modalities with language.

Breaking the Barriers of Text-Hungry and Audio-Deficient AI Onellm: One framework to align all modalities with language

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-07T11:28:50.713865Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:28:50.713865Z digest=sha256:b0c14f0c2e33f96cf6b6295e863a3296252c647000d5b10d83cbd4b564fcddb8

Observation fd866ec3-b451-408d-b0a0-8748559e606d · outbound

This paper cites Physics-inspired approaches in generative diffusion models.

Breaking the Barriers of Text-Hungry and Audio-Deficient AI Physics-inspired approaches in generative diffusion models

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-07T11:28:50.772569Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:28:50.772569Z digest=sha256:68f3dc48d57c96c8f6ec884567352c18ee2c307a3cdce92c933fe6cd09b1f683

Observation f355817d-0e17-4169-9cbd-8afe0eb53b49 · outbound

This paper cites Exploring In-Context Learning of Textless Speech Language Model for Speech Classification Tasks.

Breaking the Barriers of Text-Hungry and Audio-Deficient AI Exploring In-Context Learning of Textless Speech Language Model for Speech Classification Tasks

Reference 48

Resolution
verified exact
local_arxiv, observed 2026-08-07T11:29:07.249632Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T11:28:50.849055Z digest=sha256:6a0ccb6fa03ca53e813f5f8053819cd2b23d61c6c084aaec8fc5a1ff091af65a

Observation fbc096ac-803e-475f-9254-687f59be384f · outbound

This paper cites Chain-of-thought prompting for speech translation.

Breaking the Barriers of Text-Hungry and Audio-Deficient AI Chain-of-thought prompting for speech translation

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-07T11:28:50.906902Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:28:50.906902Z digest=sha256:b6719598b0771ae3143461b0fd27136c9097d251ff10639dca8d1eb292a012c5

Observation a16b2aec-5c45-419e-bdd6-f59e89cb8817 · outbound

This paper cites TranSpeech: Speech-to-Speech Translation With Bilateral Perturbation.

Breaking the Barriers of Text-Hungry and Audio-Deficient AI TranSpeech: Speech-to-Speech Translation With Bilateral Perturbation

Reference 50

Resolution
verified exact
local_arxiv, observed 2026-08-07T11:29:07.050233Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T11:28:50.968083Z digest=sha256:327ed901ffadb01d0e1e187d4d198e742abb4e7e0e34ee430874a0b91ff8bbfb

Observation 645479fc-6505-4d81-8dc5-c3e9f1e7c109 · outbound

This paper cites Textless Acoustic Model with Self-Supervised Distillation for Noise-Robust Expressive Speech-to-Speech Translation.

Breaking the Barriers of Text-Hungry and Audio-Deficient AI Textless Acoustic Model with Self-Supervised Distillation for Noise-Robust Expressive Speech-to-Speech Translation

Reference 51

Resolution
verified exact
local_arxiv, observed 2026-08-07T11:29:06.877355Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T11:28:51.039084Z digest=sha256:b05bd4f1d875476f96a613a8a9f9b60ce3db3e0389acee058fa972f9792532d3

Observation 4f7aa9c0-3d17-439d-b26a-4b34f3d7c969 · outbound

This paper cites Massively multi- lingual forced aligner leveraging self-supervised discrete units.

Breaking the Barriers of Text-Hungry and Audio-Deficient AI Massively multi- lingual forced aligner leveraging self-supervised discrete units

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-07T11:28:51.104629Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:28:51.104629Z digest=sha256:c01616a34c165056d666068cce5319de9a76605057f7d87803c1fcb0f77c4641

Observation 411f20dd-1332-4924-ac1f-3640152ea1a7 · outbound

This paper cites LibriS2S: A German-English Speech-to-Speech Translation Corpus.

Breaking the Barriers of Text-Hungry and Audio-Deficient AI LibriS2S: A German-English Speech-to-Speech Translation Corpus

Reference 53

Resolution
verified exact
local_arxiv, observed 2026-08-07T11:29:06.641946Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T11:28:51.219694Z digest=sha256:8e75a4f5853245622c79d4f3e28d4035235f320f0cc8a06bfb39c9dad8261588

Observation b3c742b7-aa8e-4ea6-b906-76cf35b79278 · outbound

This paper cites WavChat: A Survey of Spoken Dialogue Models.

Breaking the Barriers of Text-Hungry and Audio-Deficient AI WavChat: A Survey of Spoken Dialogue Models

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-07T11:28:51.286122Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:28:51.286122Z digest=sha256:1fcd28e83c0e2e22cb74809d925f66214cd63b995b12f44f543329ebe47b175b

Observation 73b92d89-70e8-41c1-936a-f94529374e76 · outbound

This paper cites Can Generative Geospatial Diffusion Models Excel as Discriminative Geospatial Foundation Models?.

Breaking the Barriers of Text-Hungry and Audio-Deficient AI Can Generative Geospatial Diffusion Models Excel as Discriminative Geospatial Foundation Models?

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-07T11:28:51.380557Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:28:51.380557Z digest=sha256:c33875d547d9e83d41e8eb86cb6d2de03296523ad793cf8ec76629d842f1bff7

Observation 2b4fde17-57c6-4b7f-b51f-693b03ae020e · outbound

This paper cites Listra automatic speech translation: English to lingala case study.

Breaking the Barriers of Text-Hungry and Audio-Deficient AI Listra automatic speech translation: English to lingala case study

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-07T11:28:51.433996Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:28:51.433996Z digest=sha256:413c0f3e355081dd1c0623a4d8a344059d0a782cbde28451b8af7119edcf1cdc

Observation 8164b2a9-f011-4b13-8074-1a8b4d7bb00d · outbound

This paper cites Gdplan: Generative network planning via graph diffusion model.

Breaking the Barriers of Text-Hungry and Audio-Deficient AI Gdplan: Generative network planning via graph diffusion model

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-07T11:28:51.582236Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:28:51.582236Z digest=sha256:b187a0f3134da6a74e4a33ffabb20d599b374dba2f9e8c934d51216b0636e02c

Observation ffa65f22-3545-4468-a606-d1383c7257be · outbound

This paper cites Direct Punjabi to English speech translation using discrete units.

Breaking the Barriers of Text-Hungry and Audio-Deficient AI Direct Punjabi to English speech translation using discrete units

Reference 58

Resolution
verified exact
local_arxiv, observed 2026-08-07T11:29:06.313102Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T11:28:51.726334Z digest=sha256:15db8414f7db8693f7b5a0f757754bdfec7ae075a7e42b2b018835ddf77210f8

Observation 47a8994c-a3da-440a-93fb-7ce5beed79f2 · outbound

This paper cites Textless unit-to-unit training for many- to-many multilingual speech-to-speech translation.

Breaking the Barriers of Text-Hungry and Audio-Deficient AI Textless unit-to-unit training for many- to-many multilingual speech-to-speech translation

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-07T11:28:51.872948Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:28:51.872948Z digest=sha256:db446ccee461b0e9cb528b8837568c8e2bca0214777514d460f97bfa6f35fd55

Observation f1df9749-84f9-43fa-a302-e888b5d5135b · outbound

This paper cites Phi dm-dialog: an experimental speech-to-speech dialog translation system.

Breaking the Barriers of Text-Hungry and Audio-Deficient AI Phi dm-dialog: an experimental speech-to-speech dialog translation system

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-07T11:28:52.075185Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:28:52.075185Z digest=sha256:cf2e92acadb4800009dd416dc12a444b41d93e338857a0f212a5a2396f43a8be

Observation e1bcde44-36e0-4eca-ae94-34ec2c0fdf4f · outbound

This paper cites High-Fidelity Simultaneous Speech-To-Speech Translation.

Breaking the Barriers of Text-Hungry and Audio-Deficient AI High-Fidelity Simultaneous Speech-To-Speech Translation

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-07T11:28:52.222146Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:28:52.222146Z digest=sha256:289d228431cfcb6b0ce00d9b503e1877ea3c893a9b33a5c5f6c53b69baebb291

Observation ee857ae4-18f4-4465-91a5-e7e86832f1ea · outbound

This paper cites Janus-iii: Speech-to-speech translation in multiple languages.

Breaking the Barriers of Text-Hungry and Audio-Deficient AI Janus-iii: Speech-to-speech translation in multiple languages

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-07T11:28:52.338122Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:28:52.338122Z digest=sha256:d028a03e8f03135e8380f26229051baf31d82f9df9c6fb449f5081986a037e84

Observation a544d151-2432-4379-afe7-49bd57a39dac · outbound

This paper cites Textless Speech-to-Speech Translation on Real Data.

Breaking the Barriers of Text-Hungry and Audio-Deficient AI Textless Speech-to-Speech Translation on Real Data

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-07T11:28:52.430942Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:28:52.430942Z digest=sha256:9f234dd5f40be585603a44c92cf41983758b6b2566bca37f2d39de7586960e3b

Observation 7bd8c795-7df6-4383-9fdd-0cfa0c5444aa · outbound

This paper cites Video diffusion models are strong video inpainter.

Breaking the Barriers of Text-Hungry and Audio-Deficient AI Video diffusion models are strong video inpainter

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-07T11:28:52.488633Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:28:52.488633Z digest=sha256:30c2213cb05e135d1df75339ffafad573070fc101e946c065a6840bd5586935a

Observation 6f212478-dc0f-4540-8470-5ec1cb2360fd · outbound

This paper cites Speech proportion and accuracy in simultaneous interpretation from english into korean.

Breaking the Barriers of Text-Hungry and Audio-Deficient AI Speech proportion and accuracy in simultaneous interpretation from english into korean

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-07T11:28:52.561720Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:28:52.561720Z digest=sha256:8f1f8ba6a3f50410ef789b3df844b51ec96cd115244339ffed777ef79ebd8a98

Observation c5462798-f80a-43c7-a9e3-edd0a298c7b6 · outbound

This paper cites Diffusion models for audio restoration: A review [special issue on model-based and data-driven audio signal processing].

Breaking the Barriers of Text-Hungry and Audio-Deficient AI Diffusion models for audio restoration: A review [special issue on model-based and data-driven audio signal processing]

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-07T11:28:52.627257Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:28:52.627257Z digest=sha256:3ddd611840fc1ddbda68f0aa6be34af1be532638bff0761458f3b4c996288a5e

Observation 2bfdd863-1ecb-41df-8902-eb01fd97d21a · outbound

This paper cites Conditional diffusion model for missing value imputation.

Breaking the Barriers of Text-Hungry and Audio-Deficient AI Conditional diffusion model for missing value imputation

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-07T11:28:52.698356Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:28:52.698356Z digest=sha256:7aec5c42d1291b1c47f1166674b64c4c100c87eff4aaf3ed86fc61aab5d773d1

Observation 69f88707-5b02-41bd-b2e2-f5ae64cbdcff · outbound

This paper cites BrainECHO: Semantic Brain Signal Decoding through Vector-Quantized Spectrogram Reconstruction for Whisper-Enhanced Text Generation.

Breaking the Barriers of Text-Hungry and Audio-Deficient AI BrainECHO: Semantic Brain Signal Decoding through Vector-Quantized Spectrogram Reconstruction for Whisper-Enhanced Text Generation

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-07T11:28:52.756373Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:28:52.756373Z digest=sha256:47a05a1e95b794e7d6de34266c6b7e4b72cb1cfe53cba12226640b80a6af957e

Observation a861450a-14e7-4ae6-b631-4ebdec81cd59 · outbound

This paper cites Baichuan-Audio: A Unified Framework for End-to-End Speech Interaction.

Breaking the Barriers of Text-Hungry and Audio-Deficient AI Baichuan-Audio: A Unified Framework for End-to-End Speech Interaction

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-07T11:28:52.822450Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:28:52.822450Z digest=sha256:06956f9660cb2a630106e7de5f92fb116d296588c98bb984a9106c4edd1d47da

Observation c7976c58-03ca-45cc-b81f-0ef8ea4a3323 · outbound

This paper cites Textless direct speech-to-speech translation with discrete speech representation.

Breaking the Barriers of Text-Hungry and Audio-Deficient AI Textless direct speech-to-speech translation with discrete speech representation

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-07T11:28:52.889256Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:28:52.889256Z digest=sha256:eb600594c578771a8c8a1c08d919c6a0e27aebf7319e1b53009c79e77ba1b397

Observation 396661f3-0e18-4ef1-a595-31a025022a83 · outbound

This paper cites Beyond Words: AuralLLM and SignMST-C for Sign Language Production and Bidirectional Accessibility.

Breaking the Barriers of Text-Hungry and Audio-Deficient AI Beyond Words: AuralLLM and SignMST-C for Sign Language Production and Bidirectional Accessibility

Reference 71

Resolution
verified exact
local_arxiv, observed 2026-08-07T11:29:06.009246Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T11:28:52.948623Z digest=sha256:c46eaa8ea8d613be08bbb9cc7cbc480ee0a9c112094d4ada259e5fb0ea260a65

Observation ccef7ffc-1061-4497-8b00-b6dc666de485 · outbound

This paper cites Align-SLM: Textless Spoken Language Models with Reinforcement Learning from AI Feedback.

Breaking the Barriers of Text-Hungry and Audio-Deficient AI Align-SLM: Textless Spoken Language Models with Reinforcement Learning from AI Feedback

Reference 72

Resolution
unresolved
no resolver link, observed 2026-08-07T11:28:53.047690Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:28:53.047690Z digest=sha256:44bd00a9ea9c3158bcb2271f7c3e98af8985d8885e692ef5d34d5db7ecdabd35

Observation 70b4be9b-ed03-4861-b1d3-1846e98fd97e · outbound

This paper cites an unresolved cited work.

Breaking the Barriers of Text-Hungry and Audio-Deficient AI Unresolved cited work

Reference 73

Resolution
unresolved
no resolver link, observed 2026-08-07T11:28:53.113989Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:28:53.113989Z digest=sha256:1bca64f68f4423bbc8e582de33b498cd7239fd7687ec02258d7304b7745e7300

Observation 79af8f6c-b061-4ac3-88c1-3a856ac09d70 · outbound

This paper cites Handdiffuse: generative controllers for two-hand interactions via diffusion models.

Breaking the Barriers of Text-Hungry and Audio-Deficient AI Handdiffuse: generative controllers for two-hand interactions via diffusion models

Reference 74

Resolution
unresolved
no resolver link, observed 2026-08-07T11:28:53.190421Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:28:53.190421Z digest=sha256:87497f7576b57f418bd973da56d81b22cd6c60b6a7e5842c2d20deb26f3cf9b2

Observation 0222dac1-65b2-4a42-b3f4-ba5a3b496eda · outbound

This paper cites A Preliminary Exploration with GPT-4o Voice Mode.

Breaking the Barriers of Text-Hungry and Audio-Deficient AI A Preliminary Exploration with GPT-4o Voice Mode

Reference 75

Resolution
unresolved
no resolver link, observed 2026-08-07T11:28:53.249841Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:28:53.249841Z digest=sha256:d7adc0e199d4b6c2ef0338cce0666634e23b9594702f27a67b2b2fee6b326510

Observation 550e152a-e5dc-42bc-8b1f-74d8a62b5a70 · outbound

This paper cites Recent highlights in multilingual and multimodal speech translation.

Breaking the Barriers of Text-Hungry and Audio-Deficient AI Recent highlights in multilingual and multimodal speech translation

Reference 76

Resolution
unresolved
no resolver link, observed 2026-08-07T11:28:53.332770Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:28:53.332770Z digest=sha256:93011fad3e542630ae3fc67e7b9d6ed1b304c71d7d8f7f667e838afa53dc3ed7

Observation a738617c-b7fd-458e-825c-4d7de2be9be0 · outbound

This paper cites Speech-to-speech low-resource translation.

Breaking the Barriers of Text-Hungry and Audio-Deficient AI Speech-to-speech low-resource translation

Reference 77

Resolution
unresolved
no resolver link, observed 2026-08-07T11:28:53.418689Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:28:53.418689Z digest=sha256:5bb1683d30bb60d7c70b23f06037795b96c5426cd1aa5d23d1ba87603ccec2c2

Observation 1b92ec96-557f-43f8-a985-c861c919bd5f · outbound

This paper cites Listening and seeing again: Generative error correction for audio-visual speech recognition.

Breaking the Barriers of Text-Hungry and Audio-Deficient AI Listening and seeing again: Generative error correction for audio-visual speech recognition

Reference 78

Resolution
unresolved
no resolver link, observed 2026-08-07T11:28:53.480148Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:28:53.480148Z digest=sha256:196727f9da469315596af15701b1beec6f29c562284ae59572561d2967c5f99e

Observation 497f7294-6426-4c88-8ee4-a722d4b124f6 · outbound

This paper cites SLIDE: Integrating Speech Language Model with LLM for Spontaneous Spoken Dialogue Generation.

Breaking the Barriers of Text-Hungry and Audio-Deficient AI SLIDE: Integrating Speech Language Model with LLM for Spontaneous Spoken Dialogue Generation

Reference 79

Resolution
unresolved
no resolver link, observed 2026-08-07T11:28:53.542723Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:28:53.542723Z digest=sha256:1ac87c0a550eef9d4c1ec7aa47be622da080dc49db2d0a855e7f91f1aa0aaef2

Observation 51da33f3-43af-4f58-b49d-4360c75d864c · outbound

This paper cites Llamapartialspoof: An llm-driven fake speech dataset simulating disinformation generation.

Breaking the Barriers of Text-Hungry and Audio-Deficient AI Llamapartialspoof: An llm-driven fake speech dataset simulating disinformation generation

Reference 80

Resolution
unresolved
no resolver link, observed 2026-08-07T11:28:53.615822Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:28:53.615822Z digest=sha256:057035add1714893b14352600fe3390eab103427f1693f10c6fdd579b8d1af73

Observation d8a008de-eb3f-46d5-a025-6441456e2a81 · outbound

This paper cites Build llm-based zero-shot streaming tts system with cosyvoice.

Breaking the Barriers of Text-Hungry and Audio-Deficient AI Build llm-based zero-shot streaming tts system with cosyvoice

Reference 81

Resolution
unresolved
no resolver link, observed 2026-08-07T11:28:53.691843Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:28:53.691843Z digest=sha256:fabc725a81cc619481ad8f18f7141723f0f2279c9be803296223869d7a2db8ce

Observation 7da2b580-0764-4db3-813a-999550cfc9a4 · outbound

This paper cites Auto-avsr: Audio-visual speech recognition with automatic labels.

Breaking the Barriers of Text-Hungry and Audio-Deficient AI Auto-avsr: Audio-visual speech recognition with automatic labels

Reference 82

Resolution
unresolved
no resolver link, observed 2026-08-07T11:28:53.754082Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:28:53.754082Z digest=sha256:0779889e561aa39d86ef10b7db269804bb20ebfbf5fe6250c1f94ea6b9f17ea9

Observation 7fa3d4a9-2460-499f-a34e-4f104d4c58c9 · outbound

This paper cites Real-Time Textless Dialogue Generation.

Breaking the Barriers of Text-Hungry and Audio-Deficient AI Real-Time Textless Dialogue Generation

Reference 83

Resolution
unresolved
no resolver link, observed 2026-08-07T11:28:53.813975Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:28:53.813975Z digest=sha256:4184043bac2d71a7ca3c8c02829b79c6c3935e1a4f2f3a8ae1f12a8d0a33e1e2

Observation 3cb01864-40b8-4d19-ae41-18a246831544 · outbound

This paper cites Slamming: Training a Speech Language Model on One GPU in a Day.

Breaking the Barriers of Text-Hungry and Audio-Deficient AI Slamming: Training a Speech Language Model on One GPU in a Day

Reference 84

Resolution
unresolved
no resolver link, observed 2026-08-07T11:28:53.898225Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:28:53.898225Z digest=sha256:8fb8a7a0e0020795279c57a3199ac63281b5c91ed86274749c3bb8cabb50fb53

Observation 809a6685-b873-4b14-9281-a9ece5f55ba6 · outbound

This paper cites Enhancing Low-Resource Language and Instruction Following Capabilities of Audio Language Models.

Breaking the Barriers of Text-Hungry and Audio-Deficient AI Enhancing Low-Resource Language and Instruction Following Capabilities of Audio Language Models

Reference 85

Resolution
unresolved
no resolver link, observed 2026-08-07T11:28:53.960789Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:28:53.960789Z digest=sha256:e10b2fee3fc6b3d7a4a09f776a9b283b5976bf955b76bca8ce650af38f3aeb4d

Observation f53f924c-240e-4287-b891-99000f3eec12 · outbound

This paper cites Make some noise: Towards llm audio reasoning and generation using sound tokens.

Breaking the Barriers of Text-Hungry and Audio-Deficient AI Make some noise: Towards llm audio reasoning and generation using sound tokens

Reference 86

Resolution
unresolved
no resolver link, observed 2026-08-07T11:28:54.025981Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:28:54.025981Z digest=sha256:fe4d36abdf93245961f98f2ba0dd3758aa481c96d7e04a012a9d9ccd9c2998c8

Observation 8e180aeb-c0ca-4e72-9646-087233593a67 · outbound

This paper cites Deep networks as denoising algorithms: Sample-efficient learning of diffusion models in high-dimensional graphical models.

Breaking the Barriers of Text-Hungry and Audio-Deficient AI Deep networks as denoising algorithms: Sample-efficient learning of diffusion models in high-dimensional graphical models

Reference 87

Resolution
unresolved
no resolver link, observed 2026-08-07T11:28:54.074608Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:28:54.074608Z digest=sha256:1257a1ba944b05032ac7a73aef297de4859d18bf38a3f80ec7f21eb96c716def

Observation e4e61795-1735-49f3-b201-3cbf5270002a · outbound

This paper cites Amharic speech recognition for speech translation.

Breaking the Barriers of Text-Hungry and Audio-Deficient AI Amharic speech recognition for speech translation

Reference 88

Resolution
unresolved
no resolver link, observed 2026-08-07T11:28:54.129617Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:28:54.129617Z digest=sha256:137c56d507abdccded2b4918e06a012be5444d357dc91e96d8d20c22a3391946

Observation 85a5e050-475a-4d8f-a090-e283bb7612b0 · outbound

This paper cites Parrot: Autoregressive spoken dialogue language modeling with decoder-only transform- ers.

Breaking the Barriers of Text-Hungry and Audio-Deficient AI Parrot: Autoregressive spoken dialogue language modeling with decoder-only transform- ers

Reference 89

Resolution
unresolved
no resolver link, observed 2026-08-07T11:28:54.215355Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:28:54.215355Z digest=sha256:f9b89f5ac5a6091e55a83a3a405a6092c9cc603ab90f2de3f04527e9f1d0fdff

Observation beb0db37-1d76-4902-a215-7357bd88f8eb · outbound

This paper cites Towards to a direct speech to speech for endangered languages in africa.

Breaking the Barriers of Text-Hungry and Audio-Deficient AI Towards to a direct speech to speech for endangered languages in africa

Reference 90

Resolution
unresolved
no resolver link, observed 2026-08-07T11:28:54.281049Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:28:54.281049Z digest=sha256:b171aec81cef24273ad7f27cd8badbbb30970a3a612c1c828e22b6db8d1c507c

Observation a87d303b-cef5-4093-a2d6-25f6d4b7597b · outbound

This paper cites A Unit-based System and Dataset for Expressive Direct Speech-to-Speech Translation.

Breaking the Barriers of Text-Hungry and Audio-Deficient AI A Unit-based System and Dataset for Expressive Direct Speech-to-Speech Translation

Reference 91

Resolution
verified exact
local_arxiv, observed 2026-08-07T11:29:05.757265Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T11:28:54.358626Z digest=sha256:ad0c85d3396fd0282bd1fb99a1f6c08fd7d39cf9f61ea077b8bc27a3f450f4fe

Observation 5d30e757-ba22-422a-b20a-df145a1b26f0 · outbound

This paper cites Spoken Question Answering and Speech Continuation Using Spectrogram-Powered LLM.

Breaking the Barriers of Text-Hungry and Audio-Deficient AI Spoken Question Answering and Speech Continuation Using Spectrogram-Powered LLM

Reference 92

Resolution
unresolved
no resolver link, observed 2026-08-07T11:28:54.440358Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:28:54.440358Z digest=sha256:fe5c8d716649519e18d635540f5d7b0ceb48310258b1bbba4492a8465ebdf29c

Observation 404457b9-4d9e-466b-89c6-ef2fde59496c · outbound

This paper cites Towards real-time multilingual multimodal speech-to-speech translation.

Breaking the Barriers of Text-Hungry and Audio-Deficient AI Towards real-time multilingual multimodal speech-to-speech translation

Reference 93

Resolution
unresolved
no resolver link, observed 2026-08-07T11:28:54.514589Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:28:54.514589Z digest=sha256:86b4153bebeca892c70fc5a717b0b557df684f4276ae508c5dfa4e53909e7c96

Observation 29d563db-e77d-4365-8701-2432871951b3 · outbound

This paper cites One Model, Many Languages: Meta-learning for Multilingual Text-to-Speech.

Breaking the Barriers of Text-Hungry and Audio-Deficient AI One Model, Many Languages: Meta-learning for Multilingual Text-to-Speech

Reference 94

Resolution
verified exact
local_arxiv, observed 2026-08-07T11:29:05.595083Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T11:28:54.554883Z digest=sha256:a8b897345095d692fab6b22a76a341106630bd0aca303743bc644941e64db2b9

Observation 6f82a242-5629-4085-b9f6-6be3c3184587 · outbound

This paper cites Spoken Language Modeling from Raw Audio.

Breaking the Barriers of Text-Hungry and Audio-Deficient AI Spoken Language Modeling from Raw Audio

Reference 95

Resolution
unresolved
no resolver link, observed 2026-08-07T11:28:54.618527Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:28:54.618527Z digest=sha256:47a55c4275ede7476289e3eaf08f840abb7153ea988da4ac8b56d7e8db1f8ec9

Observation a296a4c6-ff87-4be2-ac85-65a5381ba322 · outbound

This paper cites Visually Grounded Speech Models for Low-resource Languages and Cognitive Modelling.

Breaking the Barriers of Text-Hungry and Audio-Deficient AI Visually Grounded Speech Models for Low-resource Languages and Cognitive Modelling

Reference 96

Resolution
verified exact
local_arxiv, observed 2026-08-07T11:29:05.415512Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T11:28:54.661961Z digest=sha256:d8be85cec982104d13dd4232524aa1fafe2b0a19d7e020bc7cb2ce88702e0da7

Observation 04064344-4acb-46fc-8596-1a702b74407a · outbound

This paper cites Verbmobil: The use of prosody in the linguistic components of a speech understanding system.

Breaking the Barriers of Text-Hungry and Audio-Deficient AI Verbmobil: The use of prosody in the linguistic components of a speech understanding system

Reference 97

Resolution
unresolved
no resolver link, observed 2026-08-07T11:28:54.778940Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:28:54.778940Z digest=sha256:a7db7806862355a0cbb9bff8cbd94a646d10894633f4581d23e0bbba0ada981d

Observation 34d099fb-c96f-4cec-85e7-4e0ec50d7b47 · outbound

This paper cites Phonology-Guided Speech-to-Speech Translation for African Languages.

Breaking the Barriers of Text-Hungry and Audio-Deficient AI Phonology-Guided Speech-to-Speech Translation for African Languages

Reference 98

Resolution
verified exact
local_arxiv, observed 2026-08-07T11:29:05.278205Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T11:28:54.842017Z digest=sha256:12a1369dae783d826f37472fd8e122b6a21786a611460ace6a3f9a52719c831c

Observation 899aff8c-87f0-48a0-b348-c1531fdc2af2 · outbound

This paper cites Let's Go Real Talk: Spoken Dialogue Model for Face-to-Face Conversation.

Breaking the Barriers of Text-Hungry and Audio-Deficient AI Let's Go Real Talk: Spoken Dialogue Model for Face-to-Face Conversation

Reference 99

Resolution
unresolved
no resolver link, observed 2026-08-07T11:28:54.880977Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:28:54.880977Z digest=sha256:e0fd6508c9a645944372be75e38826a157bfbf77dc87041129c092404f6236ca

Observation 69f5a34b-30cf-4b75-8c5b-9544d342ba84 · outbound

This paper cites Long-Form Speech Generation with Spoken Language Models.

Breaking the Barriers of Text-Hungry and Audio-Deficient AI Long-Form Speech Generation with Spoken Language Models

Reference 100

Resolution
unresolved
no resolver link, observed 2026-08-07T11:28:54.976320Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:28:54.976320Z digest=sha256:39c5e28dd0d5eddf52dc9fe525e81b466b52959fe3a8da7125daa5f52744ce3c

Pith citing papers

Observation b97f05c3-f9ac-40d9-8c1e-c8c2a4cb0a72 · inbound

Achieving Generational Peace in Mali through Intergenerational Mean-Field-Type Game-based Incentives cites this paper.

Achieving Generational Peace in Mali through Intergenerational Mean-Field-Type Game-based Incentives Breaking the Barriers of Text-Hungry and Audio-Deficient AI

Reference 53

Resolution
verified exact
arxiv_id, observed 2026-05-21T00:59:19.261302Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-21T00:58:13.403170Z digest=sha256:240e31ee0689ce2cb2cc5efdbf0992a9f9402f87cf3c617005d2e4bee4c25052

Observation e83bdcdb-5e9f-4667-9220-e8c24470eaa5 · inbound

Mitigating Polycentric Conflict-Trap Risk in Mali via Intergenerational Volterra Mean-Field-Type Games cites this paper.

Mitigating Polycentric Conflict-Trap Risk in Mali via Intergenerational Volterra Mean-Field-Type Games Breaking the Barriers of Text-Hungry and Audio-Deficient AI

Reference 20

Resolution
verified exact
arxiv_id, observed 2026-07-04T09:39:46.381030Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-26T09:46:15.701501Z digest=sha256:6562abfa1cce7775804b4cba6530f10b9e5663bc1d9ec076fe278631852e2a2e

Observation f7b6c78e-a369-4e67-867a-2a4f67fa9319 · inbound

Risk-Aware Information Theory cites this paper.

Risk-Aware Information Theory Breaking the Barriers of Text-Hungry and Audio-Deficient AI

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-07-04T09:39:46.873728Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-26T09:41:41.049881Z digest=sha256:5de11e4866329d2c16f7790531e4c3bfde16102f4ba003e18cce00c98e80cca7