Pith. sign in

Paper Citation Record · LEDGER

Seamless: Multilingual Expressive and Streaming Speech Translation

As of 11 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 75 inbound Pith citation observations for arXiv:2312.05187.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2312.05187 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 75 of 75 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00

measured 75 of 75 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-11T04:45:04.446640Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-05T02:28:24.338817Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

41
pith, observed 2026-08-05T02:28:24.338817Z

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 8f5c02d6-c97e-4a85-9161-d221f8354af9 · inbound

How "Real" is Your Real-Time Simultaneous Speech-to-Text Translation System? cites this paper.

How "Real" is Your Real-Time Simultaneous Speech-to-Text Translation System? Seamless: Multilingual Expressive and Streaming Speech Translation

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-11T04:45:04.446640Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:45:04.446640Z digest=sha256:4bd8f089c9a31b6e5b26f9215887614d3f216a561a0f8c1e2561c02ae069f62a

Observation 15670c3a-ba8f-4be6-aee4-b1d914f7afee · inbound

EmoReg: Directional Latent Vector Modeling for Emotional Intensity Regularization in Diffusion-based Voice Conversion cites this paper.

EmoReg: Directional Latent Vector Modeling for Emotional Intensity Regularization in Diffusion-based Voice Conversion Seamless: Multilingual Expressive and Streaming Speech Translation

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-10T23:28:04.157974Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:28:04.157974Z digest=sha256:f29ed115dba3bf9370b1ee687614d961bd618bab42471965181f0340480f7e6f

Observation 9bae3ad8-329a-4835-a9ec-4a62ecd9315c · inbound

Addressing speaker gender bias in large scale speech translation systems cites this paper.

Addressing speaker gender bias in large scale speech translation systems Seamless: Multilingual Expressive and Streaming Speech Translation

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-10T21:10:24.764561Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:10:24.764561Z digest=sha256:34491f6dde43f3ea668da20ba4a4c1278fc9b42f79b38d4e0771ebf873a50d95

Observation 7b44b743-b89c-4310-a8b2-d64a3a4eee3f · inbound

WhiSPA: Semantically and Psychologically Aligned Whisper with Self-Supervised Contrastive and Student-Teacher Learning cites this paper.

WhiSPA: Semantically and Psychologically Aligned Whisper with Self-Supervised Contrastive and Student-Teacher Learning Seamless: Multilingual Expressive and Streaming Speech Translation

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-10T20:27:22.440781Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T20:27:22.440781Z digest=sha256:b5d6465fcbb3f1487493a08505ae7bb32e0452eaafce251fb607f55c8635033e

Observation 6438701b-a5ec-4a37-abf9-f6a4f3844d0a · inbound

A Unit-based System and Dataset for Expressive Direct Speech-to-Speech Translation cites this paper.

A Unit-based System and Dataset for Expressive Direct Speech-to-Speech Translation Seamless: Multilingual Expressive and Streaming Speech Translation

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-09T19:17:55.392128Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T19:17:55.392128Z digest=sha256:9ed2da1de7646df9f6dec709d79494a0462c561c4076cf63effa8c3fb8d7ea0d

Observation db04918e-1588-4b62-94e2-e2ab7ccfb8f7 · inbound

High-Fidelity Simultaneous Speech-To-Speech Translation cites this paper.

High-Fidelity Simultaneous Speech-To-Speech Translation Seamless: Multilingual Expressive and Streaming Speech Translation

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-09T05:01:56.004378Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T05:01:56.004378Z digest=sha256:8d6238a64507f6de23855b46144d19dc72a0bd042eef5e546a56cef7efca5f59

Observation fb740150-cfe6-42f7-abdd-edf326726cb4 · inbound

Llasa: Scaling Train-Time and Inference-Time Compute for Llama-based Speech Synthesis cites this paper.

Llasa: Scaling Train-Time and Inference-Time Compute for Llama-based Speech Synthesis Seamless: Multilingual Expressive and Streaming Speech Translation

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-08T23:29:25.829951Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T23:29:25.829951Z digest=sha256:708375aa1d59ebb37da7a671bd59771edd1cad7e3ea6064da9668816adde4583

Observation 616006b8-a807-4b2e-a31f-151df3c44c06 · inbound

OWLS: Scaling Laws for Multilingual Speech Recognition and Translation Models cites this paper.

OWLS: Scaling Laws for Multilingual Speech Recognition and Translation Models Seamless: Multilingual Expressive and Streaming Speech Translation

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T18:32:13.574777Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T18:32:13.574777Z digest=sha256:2b04ee9d894b58f2d6ed4c7aa293576d82ecd5ac37317d33f0c1c7d49ea00e5d

Observation 237bd786-7a1b-4f89-b000-8c1889add4aa · inbound

Bridging the Linguistic Divide: A Survey on Leveraging Large Language Models for Machine Translation cites this paper.

Bridging the Linguistic Divide: A Survey on Leveraging Large Language Models for Machine Translation Seamless: Multilingual Expressive and Streaming Speech Translation

Reference 105

Resolution
verified exact
arxiv_id, observed 2026-05-22T21:42:11.401276Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-22T21:38:29.497183Z digest=sha256:8ab15ad9e246c16134f5cd1df801b8231f0fc7e0c9b25ba7b29936dea84d0947

Observation 66071322-9757-4aa5-b7b7-669446609592 · inbound

ToxicTone: A Mandarin Audio Dataset Annotated for Toxicity and Toxic Utterance Tonality cites this paper.

ToxicTone: A Mandarin Audio Dataset Annotated for Toxicity and Toxic Utterance Tonality Seamless: Multilingual Expressive and Streaming Speech Translation

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T15:14:32.771868Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:14:32.771868Z digest=sha256:18e945b1f45a6dd415bc22a248da5dc37d7259e47b6c14232c961701782835c6

Observation 851d97c0-f3eb-4d2c-b8c6-bb67bb7072d4 · inbound

From Tens of Hours to Tens of Thousands: Scaling Back-Translation for Speech Recognition cites this paper.

From Tens of Hours to Tens of Thousands: Scaling Back-Translation for Speech Recognition Seamless: Multilingual Expressive and Streaming Speech Translation

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T14:55:54.912281Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:55:54.912281Z digest=sha256:8be98b22c80e615673e1967dfe4a24764a031894617e6b344136da50418ded2f

Observation f9c63014-3cb0-48d8-8af5-6f85d8fdbb59 · inbound

Analyzing Mitigation Strategies for Catastrophic Forgetting in End-to-End Training of Spoken Language Models cites this paper.

Analyzing Mitigation Strategies for Catastrophic Forgetting in End-to-End Training of Spoken Language Models Seamless: Multilingual Expressive and Streaming Speech Translation

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T14:49:28.537697Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:49:28.537697Z digest=sha256:b4c10dcf57460e1b764f33689ea8e5d6f47b790fdd7efc768322bfc99819d255

Observation 4c46a8f5-927f-4f61-8981-28f0fa111913 · inbound

ALAS: An Automatic Latent Alignment Score for Audio Language Models cites this paper.

ALAS: An Automatic Latent Alignment Score for Audio Language Models Seamless: Multilingual Expressive and Streaming Speech Translation

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-07T14:07:21.971099Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:07:21.971099Z digest=sha256:d399e789805f378aa4db7caa4a47825773805fe8fc45996da6f859bcf83ab907

Observation 6dbcbbbd-4370-4a97-9aba-c82c3a6a2db4 · inbound

The NaijaVoices Dataset: Cultivating Large-Scale, High-Quality, Culturally-Rich Speech Data for African Languages cites this paper.

The NaijaVoices Dataset: Cultivating Large-Scale, High-Quality, Culturally-Rich Speech Data for African Languages Seamless: Multilingual Expressive and Streaming Speech Translation

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-07T13:56:29.523320Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:56:29.523320Z digest=sha256:09a9301999f4d952429e574a876c9a8c8c56af94665a185370115f1f85269c8e

Observation b43460e0-e64a-4507-9949-d9ba69f1826a · inbound

SeqPO-SiMT: Sequential Policy Optimization for Simultaneous Machine Translation cites this paper.

SeqPO-SiMT: Sequential Policy Optimization for Simultaneous Machine Translation Seamless: Multilingual Expressive and Streaming Speech Translation

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T13:57:10.263287Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:57:10.263287Z digest=sha256:6a04d4794719a9a651b035a24f024ebdd5e05e97962a977798bf1093af536161

Observation ff5e8b65-dc66-4d44-b1e0-7d44e285d4cb · inbound

PSRB: A Comprehensive Benchmark for Evaluating Persian ASR Systems cites this paper.

PSRB: A Comprehensive Benchmark for Evaluating Persian ASR Systems Seamless: Multilingual Expressive and Streaming Speech Translation

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T13:42:54.912484Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:42:54.912484Z digest=sha256:71b512520929e518129d6d5c0b86550c3f5ce44a2212ca42e866897c3bad7b74

Observation 30ea2ba9-1172-490f-ae57-4e5662940919 · inbound

GMU Systems for the IWSLT 2025 Low-Resource Speech Translation Shared Task cites this paper.

GMU Systems for the IWSLT 2025 Low-Resource Speech Translation Shared Task Seamless: Multilingual Expressive and Streaming Speech Translation

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T13:30:21.788261Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:30:21.788261Z digest=sha256:0d16e2dc09c9a410ce79c78f1c10425af0d0abfc0e5f3d1759b163897840afaa

Observation 79a2e2f6-b622-4482-b707-70e9e8e0fc9c · inbound

Identifying Primary Stress Across Related Languages and Dialects with Transformer-based Speech Encoder Models cites this paper.

Identifying Primary Stress Across Related Languages and Dialects with Transformer-based Speech Encoder Models Seamless: Multilingual Expressive and Streaming Speech Translation

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-07T12:26:47.266093Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:26:47.266093Z digest=sha256:07f3546d11b71c5cff5a9508125026235306e0f9173cc8a82ad5be71c0020a09

Observation 50be0880-8e66-45fc-97f8-da8b100ac724 · inbound

SwitchLingua: The First Large-Scale Multilingual and Multi-Ethnic Code-Switching Dataset cites this paper.

SwitchLingua: The First Large-Scale Multilingual and Multi-Ethnic Code-Switching Dataset Seamless: Multilingual Expressive and Streaming Speech Translation

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T12:37:58.867229Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:37:58.867229Z digest=sha256:df73986b0bde8483b6ec4c76dba85e0798ccc8e7c73025e1d7b7e2d68030b4cc

Observation e3a921c1-e254-4f16-97f5-8c17d3bf56f9 · inbound

Translate With Care: Addressing Gender Bias, Neutrality, and Reasoning in Large Language Model Translations cites this paper.

Translate With Care: Addressing Gender Bias, Neutrality, and Reasoning in Large Language Model Translations Seamless: Multilingual Expressive and Streaming Speech Translation

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T12:04:11.190594Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:04:11.190594Z digest=sha256:9fc1051a0fbf70c1cb7de7786a65c7df47e287efac3cf6a3370b17d6afdc5709

Observation 6161839f-97d8-4820-b17f-79b608a8c284 · inbound

Prompt-Unseen-Emotion: Zero-shot Expressive Speech Synthesis with Prompt-LLM Contextual Knowledge for Mixed Emotions cites this paper.

Prompt-Unseen-Emotion: Zero-shot Expressive Speech Synthesis with Prompt-LLM Contextual Knowledge for Mixed Emotions Seamless: Multilingual Expressive and Streaming Speech Translation

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T11:21:08.605193Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:21:08.605193Z digest=sha256:761364a387157af4459c8bc58ced7ffe86ad31332835a4ead881f1c866447fbb

Observation d0036907-2891-45ad-b5d5-0b56035ca3c8 · inbound

It's Not a Walk in the Park! Challenges of Idiom Translation in Speech-to-text Systems cites this paper.

It's Not a Walk in the Park! Challenges of Idiom Translation in Speech-to-text Systems Seamless: Multilingual Expressive and Streaming Speech Translation

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T11:15:58.484341Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:15:58.484341Z digest=sha256:fa0db60ee01c9b1a46f28396681eb93bc29e911f9a830d6d681635d174a4694b

Observation c9f3b311-7eea-48d9-9bb0-b849850a81b0 · inbound

Towards Better Disentanglement in Non-Autoregressive Zero-Shot Expressive Voice Conversion cites this paper.

Towards Better Disentanglement in Non-Autoregressive Zero-Shot Expressive Voice Conversion Seamless: Multilingual Expressive and Streaming Speech Translation

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T10:55:20.148186Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:55:20.148186Z digest=sha256:4590101d46fd41bab7a108628418e4ac6d505543d4e5b91b5fb598d59d3d2272

Observation b40de9ad-a150-4fdc-8ab2-08f8dcfa8685 · inbound

Double Entendre: Robust Audio-Based AI-Generated Lyrics Detection via Multi-View Fusion cites this paper.

Double Entendre: Robust Audio-Based AI-Generated Lyrics Detection via Multi-View Fusion Seamless: Multilingual Expressive and Streaming Speech Translation

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-06T23:48:37.233627Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T23:48:37.233627Z digest=sha256:2335b116a681d44dc2d9374f2bd1e326fe9c49f8202a4b74062f795bcdac2f73

Observation 8117339d-05be-41c1-85d4-dc3e9f67baea · inbound

Instituto de Telecomunica\c{c}\~oes at IWSLT 2025: Aligning Small-Scale Speech and Language Models for Speech-to-Text Learning cites this paper.

Instituto de Telecomunica\c{c}\~oes at IWSLT 2025: Aligning Small-Scale Speech and Language Models for Speech-to-Text Learning Seamless: Multilingual Expressive and Streaming Speech Translation

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-06T23:34:06.852611Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:34:06.852611Z digest=sha256:1c9e47e00ba0d1a81a5f8b9c113523eb50ecf16408b874c6ba3066b76cb2eba5

Observation 39cf408f-2267-471a-813e-88f5f227bf17 · inbound

TTS-CtrlNet: Time varying emotion aligned text-to-speech generation with ControlNet cites this paper.

TTS-CtrlNet: Time varying emotion aligned text-to-speech generation with ControlNet Seamless: Multilingual Expressive and Streaming Speech Translation

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-06T19:55:03.608406Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:55:03.608406Z digest=sha256:8f6b46b30002e01eb9593d409ae233a8602fcaea9a3d2b27602b7f3716454cc2

Observation d3dfb43f-ea9c-4f1d-a106-8ed571e35bf2 · inbound

StreamUni: Achieving Streaming Speech Translation with a Unified Large Speech-Language Model cites this paper.

StreamUni: Achieving Streaming Speech Translation with a Unified Large Speech-Language Model Seamless: Multilingual Expressive and Streaming Speech Translation

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-06T18:37:34.111997Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:37:34.111997Z digest=sha256:ec0d7bc8e63a05f42a5bfe7e619b2ef94bee2159816203d9d42ea334966b2180

Observation ec7eb0f9-ab53-4ee6-bd44-073d934ca742 · inbound

Seed LiveInterpret 2.0: End-to-end Simultaneous Speech-to-speech Translation with Your Voice cites this paper.

Seed LiveInterpret 2.0: End-to-end Simultaneous Speech-to-speech Translation with Your Voice Seamless: Multilingual Expressive and Streaming Speech Translation

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-06T14:52:42.956099Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:52:42.956099Z digest=sha256:f47ba33627018d3376286ec4ba12bfadd877720a158328d16f7e561c237d9eb5

Observation 82e01ac4-351c-4ffe-8a4d-2934366cf7f1 · inbound

VN-MTEB: Vietnamese Massive Text Embedding Benchmark cites this paper.

VN-MTEB: Vietnamese Massive Text Embedding Benchmark Seamless: Multilingual Expressive and Streaming Speech Translation

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-06T12:46:44.500438Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T12:46:44.500438Z digest=sha256:af85339c161cd886dae1a5815c91c58b7e630dff0ff79598e3e88973c1d75bc1

Observation 63998669-d7b1-46d9-8169-8bd0665d063d · inbound

The Prosody of Emojis cites this paper.

The Prosody of Emojis Seamless: Multilingual Expressive and Streaming Speech Translation

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-06T10:09:15.914528Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T10:09:15.914528Z digest=sha256:144a7b4dccdbd85051983571ff8fab5379251b039702697ce4fb5ba4209114c4

Observation 629c8f73-b06a-4b22-b96e-9b67c73425ec · inbound

ProMode: A Speech Prosody Model Conditioned on Acoustic and Textual Inputs cites this paper.

ProMode: A Speech Prosody Model Conditioned on Acoustic and Textual Inputs Seamless: Multilingual Expressive and Streaming Speech Translation

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-05T21:07:04.621533Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T21:07:04.621533Z digest=sha256:3ba355b6e3823bb58c6309cf332189a465f81c18c8250ad5a3880bc4f086f3f7

Observation 09ff79f7-f56e-4818-b82a-14dd7a7c85e9 · inbound

Geolocation-Aware Robust Spoken Language Identification cites this paper.

Geolocation-Aware Robust Spoken Language Identification Seamless: Multilingual Expressive and Streaming Speech Translation

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-05T17:06:13.588262Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T17:06:13.588262Z digest=sha256:f49ed1c5dcb3aa942edfdd307229557538b1c02f8dc89bbb3c6c3faa8dc81f84

Observation 33b76601-05f0-4649-a81e-a99f6f0883b8 · inbound

CAM\~OES: A Comprehensive Automatic Speech Recognition Benchmark for European Portuguese cites this paper.

CAM\~OES: A Comprehensive Automatic Speech Recognition Benchmark for European Portuguese Seamless: Multilingual Expressive and Streaming Speech Translation

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-05T15:39:21.280859Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T15:39:21.280859Z digest=sha256:e42271c80552154ff963ef4e9ff464650ed79cb335d959cffb3a7b9165be304b

Observation e0ef36f3-4ac7-42e1-8eb6-cd379d1c8086 · inbound

Automatic Pronunciation Error Detection and Correction of the Holy Quran's Learners Using Deep Learning cites this paper.

Automatic Pronunciation Error Detection and Correction of the Holy Quran's Learners Using Deep Learning Seamless: Multilingual Expressive and Streaming Speech Translation

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-05T15:22:20.585051Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T15:22:20.585051Z digest=sha256:f1957fcc42e765f3eeee4410d58aeb0289235dae5b60673d6db8b61e5bf74162

Observation 9bc5e83a-d93f-40bd-909d-dd5cff15bdaf · inbound

NADI 2025: The First Multidialectal Arabic Speech Processing Shared Task cites this paper.

NADI 2025: The First Multidialectal Arabic Speech Processing Shared Task Seamless: Multilingual Expressive and Streaming Speech Translation

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-05T12:00:47.916052Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T12:00:47.916052Z digest=sha256:e6c928d4bb0f26e38c640fc62a51e8ca05898bb788988fd5f18b3655ad319c2f

Observation 5717e209-f013-4b6d-90b2-82114cc72974 · inbound

Denoising GER: A Noise-Robust Generative Error Correction with LLM for Speech Recognition cites this paper.

Denoising GER: A Noise-Robust Generative Error Correction with LLM for Speech Recognition Seamless: Multilingual Expressive and Streaming Speech Translation

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-05T10:16:04.849538Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T10:16:04.849538Z digest=sha256:43f769ea1fb7635b09c9cd1cb9f95bb07311478246e111c35605eb5581067ab3

Observation 3e685189-5188-4e6d-b25d-2e0eaeddabad · inbound

On the Contribution of Lexical Features to Speech Emotion Recognition cites this paper.

On the Contribution of Lexical Features to Speech Emotion Recognition Seamless: Multilingual Expressive and Streaming Speech Translation

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-05T05:21:20.819131Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T05:21:20.819131Z digest=sha256:777917bb4033c2488850b29a11dbafb692d0f882b9a06f5d79ddcff5c3788f62

Observation e4649a6e-8678-42f5-b023-fdafd2a84640 · inbound

Optimal Multi-Task Learning at Regularization Horizon for Speech Translation Task cites this paper.

Optimal Multi-Task Learning at Regularization Horizon for Speech Translation Task Seamless: Multilingual Expressive and Streaming Speech Translation

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-05T10:17:22.398113Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T10:17:22.398113Z digest=sha256:239aee18517943fed7aac033a513d1477e46b9c35b466eb98ef03718e039243c

Observation 418ccc03-7c3c-45e2-92f8-f9f8ea3d7ded · inbound

Simultaneous Speech-to-Speech Translation Without Aligned Data cites this paper.

Simultaneous Speech-to-Speech Translation Without Aligned Data Seamless: Multilingual Expressive and Streaming Speech Translation

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-03T00:57:00.966475Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T00:57:00.966475Z digest=sha256:1fb63a7fb5bfc33d24aba47bd1184b0113cbcbeb42ee0774423cc78f89439b38

Observation c6a820ec-8067-4536-aa61-c6865d1513a3 · inbound

Cross-Modal Robustness Transfer (CMRT): Training Robust Speech Translation Models Using Adversarial Text cites this paper.

Cross-Modal Robustness Transfer (CMRT): Training Robust Speech Translation Models Using Adversarial Text Seamless: Multilingual Expressive and Streaming Speech Translation

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-03T00:03:40.825775Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T00:03:40.825775Z digest=sha256:83a4606a3ebceaaa874634d38ba871d79c618729bb7880d3af6532882c0fbefd

Observation cadf4a1d-ff1e-4eff-8a47-19514088bcfa · inbound

Benchmarking Multilingual Speech Models on Pashto: Zero-Shot ASR, Script Failure, and Cross-Domain Evaluation cites this paper.

Benchmarking Multilingual Speech Models on Pashto: Zero-Shot ASR, Script Failure, and Cross-Domain Evaluation Seamless: Multilingual Expressive and Streaming Speech Translation

Reference 5

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T22:35:48.711956Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T19:44:30.762851Z digest=sha256:9ee9df54a098b3f2133e70a1758362d140d1451ea6079ebaadd1d32820cc29c6

Observation 59ea09d2-9d34-4f74-aa0a-62a6c3e0f26c · inbound

"OK Aura, Be Fair With Me": Demographics-Agnostic Training for Bias Mitigation in Wake-up Word Detection cites this paper.

"OK Aura, Be Fair With Me": Demographics-Agnostic Training for Bias Mitigation in Wake-up Word Detection Seamless: Multilingual Expressive and Streaming Speech Translation

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-05-10T23:40:53.652838Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T18:55:56.713141Z digest=sha256:c75cc8818d80cf55256e5a21343f4d46c206a7366e3082f523abe2f08ce482e8

Observation b66f182e-ac22-428d-97ac-65388974cd4b · inbound

DialogueSidon: Recovering Full-Duplex Dialogue Tracks from In-the-Wild Dialogue Audio cites this paper.

DialogueSidon: Recovering Full-Duplex Dialogue Tracks from In-the-Wild Dialogue Audio Seamless: Multilingual Expressive and Streaming Speech Translation

Reference 1

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T06:05:59.059056Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T17:47:58.217280Z digest=sha256:b67ee2c529f7c430e357f46debeba7de45645802d45b935b2d7cd00d2dfaf637

Observation cacd9948-47fc-48db-8242-a2f87096caa5 · inbound

NaijaS2ST: A Multi-Accent Benchmark for Speech-to-Speech Translation in Low-Resource Nigerian Languages cites this paper.

NaijaS2ST: A Multi-Accent Benchmark for Speech-to-Speech Translation in Low-Resource Nigerian Languages Seamless: Multilingual Expressive and Streaming Speech Translation

Reference 6

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T07:11:53.108467Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T07:11:32.500279Z digest=sha256:ae9a11efda16a14fa95a91c9100ca97273570e2cf1c4fb48c2a7b17fa48c677a

Observation 7f6d5090-705a-4231-aac5-8fb3eb72ef3d · inbound

MoVE: Translating Laughter and Tears via Mixture of Vocalization Experts in Speech-to-Speech Translation cites this paper.

MoVE: Translating Laughter and Tears via Mixture of Vocalization Experts in Speech-to-Speech Translation Seamless: Multilingual Expressive and Streaming Speech Translation

Reference 18

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T05:41:02.471102Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T05:36:07.090627Z digest=sha256:5a3f93dbcfc96e9a3f8825591b5c8a88faf00ef15c34cc0858083d07f4e46d45

Observation eed19b61-eade-457a-9af2-1887c7153432 · inbound

DM-ASR: Diarization-aware Multi-speaker ASR with Large Language Models cites this paper.

DM-ASR: Diarization-aware Multi-speaker ASR with Large Language Models Seamless: Multilingual Expressive and Streaming Speech Translation

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-05-11T20:21:12.553074Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-08T09:18:53.285951Z digest=sha256:af7191c53cf99f63f1262c68c8cbcc14c534c67819f7e410a2348508ecc1b12c

Observation 06d85084-87c1-42ae-b0b3-e2b7ba997c1e · inbound

The False Resonance: A Critical Examination of Emotion Embedding Similarity for Speech Generation Evaluation cites this paper.

The False Resonance: A Critical Examination of Emotion Embedding Similarity for Speech Generation Evaluation Seamless: Multilingual Expressive and Streaming Speech Translation

Reference 14

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T09:11:27.152485Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-07T12:34:40.888089Z digest=sha256:df9f73a62c602b244c425ac8cbe537813b60a34bdf32a9bdddbe100c21ef0fc2

Observation 756cc432-8472-4580-a337-dad97d50d64c · inbound

The False Resonance: A Critical Examination of Emotion Embedding Similarity for Speech Generation Evaluation cites this paper.

The False Resonance: A Critical Examination of Emotion Embedding Similarity for Speech Generation Evaluation Seamless: Multilingual Expressive and Streaming Speech Translation

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-02T15:21:28.794596Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T15:21:28.794596Z digest=sha256:6e53cd470d0eefb309dd546879aec48a74b62094dc8058616151fb96abb12a07

Observation e72b486e-93b3-4cf2-8330-0855f686fecf · inbound

PoDAR: Power-Disentangled Audio Representation for Generative Modeling cites this paper.

PoDAR: Power-Disentangled Audio Representation for Generative Modeling Seamless: Multilingual Expressive and Streaming Speech Translation

Reference 25

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T06:41:35.366926Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-12T04:04:08.420600Z digest=sha256:7100183be6711e3e82434841590912f04b1654affc44c649cca85f58728b8215

Observation 30ff1b09-6ce2-401d-8c9b-ba86950305c4 · inbound

AffectCodec: Emotion-Preserving Neural Speech Codec for Expressive Speech Modeling cites this paper.

AffectCodec: Emotion-Preserving Neural Speech Codec for Expressive Speech Modeling Seamless: Multilingual Expressive and Streaming Speech Translation

Reference 22

Resolution
verified exact
arxiv_id, observed 2026-05-13T01:07:00.315034Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-13T01:04:54.506749Z digest=sha256:7445b8193f0280f7d8744b3e9354dd4280a4f9fd3b39e400769ed8b9f0918e40

Observation 87ec1725-7956-4b75-b2a5-e25ff2d67f63 · inbound

MindVoice: Reconstructing Intelligible Speech from Non-invasive Neural Signals with Pretrained Priors cites this paper.

MindVoice: Reconstructing Intelligible Speech from Non-invasive Neural Signals with Pretrained Priors Seamless: Multilingual Expressive and Streaming Speech Translation

Reference 28

Resolution
metadata mismatch
arxiv_id, observed 2026-07-01T20:26:12.753138Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-28T21:08:39.636966Z digest=sha256:ad74f6117866e0e07848b2a2e1c12ce1e8452b663a50b60814691e81da04f770

Observation 7474c107-281c-44fb-9a4b-ce980bb1e5cb · inbound

Benchmarking Speech-to-Speech Translation Models cites this paper.

Benchmarking Speech-to-Speech Translation Models Seamless: Multilingual Expressive and Streaming Speech Translation

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-07-02T02:46:29.477885Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-06-28T10:34:10.510906Z digest=sha256:6c201c9a162defdfe0dcf084a50e35dad247833c897db35c444deb98426d3e3b

Observation a7005d31-ff81-4f4e-8b79-86aed1c59557 · inbound

Audio Interaction Model cites this paper.

Audio Interaction Model Seamless: Multilingual Expressive and Streaming Speech Translation

Reference 4

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T10:46:52.407400Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-28T04:57:05.062465Z digest=sha256:f167143596facceb72b38d5d342fc77a855c95a34506393642aa292a71a991ad

Observation 096a4347-caba-4d24-beac-f6c1c3973ea4 · inbound

Multilingual Multi-Speaker Unit Vocoders: A Systematic Analysis of Discrete Speech Representations cites this paper.

Multilingual Multi-Speaker Unit Vocoders: A Systematic Analysis of Discrete Speech Representations Seamless: Multilingual Expressive and Streaming Speech Translation

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-07-02T15:47:06.285476Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-27T23:29:48.769390Z digest=sha256:82c91189de0a8c640d881123f07ba991e3f5465bb9371c962101191281c30820

Observation df00d1c9-6a16-4eab-b16d-22244251e401 · inbound

HybridCodec: Fast Dual-Stream, Semantically Enhanced Neural Audio Codec cites this paper.

HybridCodec: Fast Dual-Stream, Semantically Enhanced Neural Audio Codec Seamless: Multilingual Expressive and Streaming Speech Translation

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-07-02T15:47:06.437959Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-27T23:27:29.942344Z digest=sha256:13e0fb501bb72397060bed802cc2e7d8b377532250ae73040930e0e688952db4

Observation 94f4db8b-a6c4-497d-8893-3890574eabc1 · inbound

Anchoring the Unknown: Open-Set Model Attribution via Proxy-Anchor Learning cites this paper.

Anchoring the Unknown: Open-Set Model Attribution via Proxy-Anchor Learning Seamless: Multilingual Expressive and Streaming Speech Translation

Reference 18

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T07:47:44.712442Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-27T11:48:10.283990Z digest=sha256:c40a424d499ac2a71845bf4c43d57f3d7cd61a195976313b1161497c170be59c

Observation 86ac0d2e-d37f-417e-848d-eea1f8ef4494 · inbound

Gumbel-BEARD: Automatic Layer Selection for Self-Supervised Adaptation of Whisper in Low-Resource Domains cites this paper.

Gumbel-BEARD: Automatic Layer Selection for Self-Supervised Adaptation of Whisper in Low-Resource Domains Seamless: Multilingual Expressive and Streaming Speech Translation

Reference 9

Resolution
metadata mismatch
arxiv_id, observed 2026-06-28T03:11:31.188666Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-27T11:21:16.723967Z digest=sha256:c4689c5d9fecc9eb71450787a16d7f601e365fd16cbb445027f3028d6cb8bc27

Observation 505cbd3d-906a-4bfd-a491-1c88ca960e0a · inbound

Pretrained self-supervised speech models can recognize unseen consonants cites this paper.

Pretrained self-supervised speech models can recognize unseen consonants Seamless: Multilingual Expressive and Streaming Speech Translation

Reference 30

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T10:07:56.072798Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-27T10:14:47.932613Z digest=sha256:75a9afb59077e6fd90554a47940f0d1875af8676be8ee6cb46dc820849a7001c

Observation 46b8da44-a37f-4631-accb-50f80bbbd1b5 · inbound

Self-Guidance: Enhancing Neural Codecs via Decoder Manifold Alignment cites this paper.

Self-Guidance: Enhancing Neural Codecs via Decoder Manifold Alignment Seamless: Multilingual Expressive and Streaming Speech Translation

Reference 93

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T16:08:37.520098Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-06-27T06:05:26.735340Z digest=sha256:558e6ca97fb0a56f5487807b83795995a52d82cd248b488a0ef0290bb12f0696

Observation 2a138061-4672-4382-ba64-ee1c02dd87ad · inbound

BrainWorld: A Structural-Prior-Conditioned Generative Model for Whole-Brain 4D fMRI Dynamics cites this paper.

BrainWorld: A Structural-Prior-Conditioned Generative Model for Whole-Brain 4D fMRI Dynamics Seamless: Multilingual Expressive and Streaming Speech Translation

Reference 5

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T20:28:55.750398Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-27T01:18:46.958501Z digest=sha256:9852ca3743af106b40f0a01aaa3d7157eab91c31793de55dc0a8976ff7fcd64f

Observation f2924d50-7e66-4653-8e36-8368c1e3c2bb · inbound

Next-Turn: Duration-Aware Streaming Endpoint Detection via Time-to-Next-Speech-Onset Prediction cites this paper.

Next-Turn: Duration-Aware Streaming Endpoint Detection via Time-to-Next-Speech-Onset Prediction Seamless: Multilingual Expressive and Streaming Speech Translation

Reference 11

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T23:19:03.188350Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-26T22:43:07.952982Z digest=sha256:e940ab7770c527bfe5fe3005c276cc9ae82d2260bdefe5148be745001e930286

Observation 84bea949-1a59-4de3-8a61-7b53ccd35fa1 · inbound

Zero-VC: Zero-Lookahead Streaming Voice Conversion via Speaker Anonymization cites this paper.

Zero-VC: Zero-Lookahead Streaming Voice Conversion via Speaker Anonymization Seamless: Multilingual Expressive and Streaming Speech Translation

Reference 21

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T05:39:39.687965Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-26T15:49:21.097255Z digest=sha256:4653857e3f7534b3ac476bd461d8816773ffdd0ddfd7e618d3cc3179b9f6589b

Observation d39fe1b6-f753-4493-a467-39386c86c7af · inbound

OpenWER: Improving Cross-Lingual ASR Evaluation and Enabling Token-Based Accuracy Metrics cites this paper.

OpenWER: Improving Cross-Lingual ASR Evaluation and Enabling Token-Based Accuracy Metrics Seamless: Multilingual Expressive and Streaming Speech Translation

Reference 14

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T06:19:37.722796Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-26T14:39:16.717868Z digest=sha256:8eef1e9bcf5b1dad7caa1957f06dfa6f1f075cbaff46476b2cfc0b1737dcae07

Observation 0a89363e-c645-41f9-aa85-98655d7b56f7 · inbound

Adaptive Oscillatory Inductive Bias for Modeling Sharp Prosodic Dynamics in Diffusion-Based TTS cites this paper.

Adaptive Oscillatory Inductive Bias for Modeling Sharp Prosodic Dynamics in Diffusion-Based TTS Seamless: Multilingual Expressive and Streaming Speech Translation

Reference 31

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T20:20:07.945691Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-25T20:11:18.045777Z digest=sha256:88e07f07a61b4a521d4c08e46c1a453e72d3cfe069763ea12661f32a008616b1

Observation 0a31f934-31ab-432d-9ee1-1ef2a8c5f835 · inbound

KM-Speaker: Keypoint-Based Style Control for High-Quality Speech-Driven 3D Facial Animation and Dialogue Localization cites this paper.

KM-Speaker: Keypoint-Based Style Control for High-Quality Speech-Driven 3D Facial Animation and Dialogue Localization Seamless: Multilingual Expressive and Streaming Speech Translation

Reference 2

Resolution
metadata mismatch
arxiv_id, observed 2026-07-01T15:45:48.558494Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-30T01:07:49.064263Z digest=sha256:ba105698384b9e16046809d12de695179a313df4d9fc83162ee7d4d84d6b717f

Observation 53facf2d-0be8-4702-a28a-68034568f1c2 · inbound

FlexiSLM: A Dynamic and Controllable Frame Rate Spoken Language Model cites this paper.

FlexiSLM: A Dynamic and Controllable Frame Rate Spoken Language Model Seamless: Multilingual Expressive and Streaming Speech Translation

Reference 240

Resolution
verified exact
arxiv_id, observed 2026-07-01T11:55:42.189673Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-07-01T03:50:26.873406Z digest=sha256:c4efa87de8c377c5a2626d5c86ab8f39510d0d897723d760538b8287eed0bb62

Observation 0b7d441c-1ce7-4721-a804-9212530cdf22 · inbound

Is Natural Always Appropriate? Investigating Naturalness and Appropriateness Across Different Domains for TTS Evaluation cites this paper.

Is Natural Always Appropriate? Investigating Naturalness and Appropriateness Across Different Domains for TTS Evaluation Seamless: Multilingual Expressive and Streaming Speech Translation

Reference 41

Resolution
metadata mismatch
arxiv_id, observed 2026-07-01T11:55:42.960525Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-07-01T03:32:23.838961Z digest=sha256:0c25e714b31378c9254701b016942f21c94828a4d01e468432e21f8094c288e5

Observation 4f908b46-439e-4ac9-8e7e-65d3fa8bd2da · inbound

Unified Audio Intelligence Without Regressing on Text Intelligence cites this paper.

Unified Audio Intelligence Without Regressing on Text Intelligence Seamless: Multilingual Expressive and Streaming Speech Translation

Reference 233

Resolution
verified exact
local_arxiv, observed 2026-07-08T00:04:22.694856Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-07-07T23:59:38.702609Z digest=sha256:ecd1de293f3c667e5b177956589c2297e30017d0f05ce7015411385d29df87c7

Observation 4d47e0f6-4983-4343-b674-ac7667bed826 · inbound

Unified Audio Intelligence Without Regressing on Text Intelligence cites this paper.

Unified Audio Intelligence Without Regressing on Text Intelligence Seamless: Multilingual Expressive and Streaming Speech Translation

Reference 233

Resolution
unresolved
no resolver link, observed 2026-07-11T07:46:49.059192Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-11T07:46:49.059192Z digest=sha256:537eb4ddabf53716e43196fbcdb604692da0484eee46aa8e55151b4327370955

Observation d9c7ef2d-ae9c-4011-a625-d203dc28a578 · inbound

GigaAM Multilingual: Foundation Model for Underrepresented Languages cites this paper.

GigaAM Multilingual: Foundation Model for Underrepresented Languages Seamless: Multilingual Expressive and Streaming Speech Translation

Reference 42

Resolution
unresolved
no resolver link, observed 2026-07-14T12:15:06.470439Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T12:15:06.470439Z digest=sha256:fe12e42f5f2ce061624836ba0750f497f1faa0b8ce76228b8ed07642dfcceae0

Observation 8c533b46-5142-4199-a6da-7df8b69f18f0 · inbound

AMECxSV: Adaptive Metadata-Driven Embedding-Fusion Calibration for X-Lingual Speaker Verification cites this paper.

AMECxSV: Adaptive Metadata-Driven Embedding-Fusion Calibration for X-Lingual Speaker Verification Seamless: Multilingual Expressive and Streaming Speech Translation

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-01T20:47:02.397442Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T20:47:02.397442Z digest=sha256:e3869fc11cff2ec46ec5684cd65451b43cb5c32c48453cf76179859fa366bef7

Observation c8ff292f-e501-46d2-b0ea-5a210388216d · inbound

X-Translator: A Real-Time Multilingual Speaker-Aware Speech-to-Speech Translation System cites this paper.

X-Translator: A Real-Time Multilingual Speaker-Aware Speech-to-Speech Translation System Seamless: Multilingual Expressive and Streaming Speech Translation

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-01T17:45:42.261982Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T17:45:42.261982Z digest=sha256:8c1bc41b75cc24509c427e8202b3e26f181b866446cf0d1d6e82eac738d6bb73

Observation be2d0bd3-ed51-44f4-9f59-792e38fcd503 · inbound

FlashRT: Agent Harness for Guiding Agents to Deploy Real-Time Multimodal Applications cites this paper.

FlashRT: Agent Harness for Guiding Agents to Deploy Real-Time Multimodal Applications Seamless: Multilingual Expressive and Streaming Speech Translation

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-01T15:52:51.239863Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T15:52:51.239863Z digest=sha256:b4b6d17137f8cc17465d72de4b115dcc8229914fddc803db75068c2944cd24d2

Observation b2df21b3-0f26-4327-b575-024c73baea70 · inbound

SimulS2ST-Omni: Data-Efficient Streaming Speech-to-Speech Translation via Explicit Trajectory Supervision cites this paper.

SimulS2ST-Omni: Data-Efficient Streaming Speech-to-Speech Translation via Explicit Trajectory Supervision Seamless: Multilingual Expressive and Streaming Speech Translation

Reference 192

Resolution
unresolved
no resolver link, observed 2026-08-01T11:43:06.630799Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T11:43:06.630799Z digest=sha256:9a2eb5dc38dc0afa49a055cbc4fd28344d1b87d3286b1460fdc749bbdf7c6435

Observation 5a0d7f71-e7bf-4804-8b48-f7515130997f · inbound

Teffic-Audio: Tell Fact from Fiction cites this paper.

Teffic-Audio: Tell Fact from Fiction Seamless: Multilingual Expressive and Streaming Speech Translation

Reference 2018

Resolution
unresolved
no resolver link, observed 2026-07-31T10:28:22.922154Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T10:28:22.922154Z digest=sha256:4e34848ea4cd1d23cf65610554e0cbdadae4af0b5ebc24c8a58dd49c4444a810