Pith. sign in

Paper Citation Record · LEDGER

High-Fidelity Simultaneous Speech-To-Speech Translation

As of 22 August 2026, this Paper Citation Record lists 45 of 45 outbound references and 5 inbound Pith citation observations for arXiv:2502.03382.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2502.03382 v2

Coverage vector

measured 45 of 45 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-09T05:01:56.219187Z

measured 50 of 50 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-22T06:32:14.747728+00:00

measured 5 of 5 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T11:28:52.222146Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-02T03:36:30.178589Z

Reference resolution

45 of 45 outbound references displayed

  • verified exact8
  • verified fuzzy16
  • unresolved21
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 368e8d34-37cf-4fbd-be4a-dd40b7593262 · outbound

This paper cites write newline.

High-Fidelity Simultaneous Speech-To-Speech Translation write newline

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-09T05:01:55.990379Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T05:01:55.990379Z digest=sha256:e7d28b233df597513a1b62903043d6b6bb183a7ab44231252bac390ebfa65151

Observation 5f71584e-53a2-4c4c-ac65-e7da7007146c · outbound

This paper cites A benchmark for evaluating machine translation metrics on dialects without standard orthography.

High-Fidelity Simultaneous Speech-To-Speech Translation A benchmark for evaluating machine translation metrics on dialects without standard orthography

Reference 2

Resolution
verified exact
doi, observed 2026-08-09T05:01:56.452529Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-09T05:01:55.997586Z digest=sha256:feb8c91ed0fa3c3246722d43538475ebe88137512573519c271e8800954112c4

Observation db04918e-1588-4b62-94e2-e2ab7ccfb8f7 · outbound

This paper cites Seamless: Multilingual Expressive and Streaming Speech Translation.

High-Fidelity Simultaneous Speech-To-Speech Translation Seamless: Multilingual Expressive and Streaming Speech Translation

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-09T05:01:56.004378Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T05:01:56.004378Z digest=sha256:92d12e8144616a26a869b1cb784e0fb814e4c390bd8c086e96f35a51eac0c245

Observation d4e9f196-c9d2-47b9-8c35-67858336164d · outbound

This paper cites Listen and Translate: A Proof of Concept for End-to-End Speech-to-Text Translation.

High-Fidelity Simultaneous Speech-To-Speech Translation Listen and Translate: A Proof of Concept for End-to-End Speech-to-Text Translation

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-09T05:01:56.011237Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T05:01:56.011237Z digest=sha256:778cf610c9fb6a0a6e50ad0781ebb30f593b6ed0ac680b3d0460446d562fe340

Observation 9330bbfe-c1c0-4efa-ba2d-4a29429e849f · outbound

This paper cites Audiolm: A language modeling approach to audio generation.

High-Fidelity Simultaneous Speech-To-Speech Translation Audiolm: A language modeling approach to audio generation

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T05:01:57.542379Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-09T05:01:56.016973Z digest=sha256:9eddd55c929435d29f681193a4e0dee026231414ec4652e8ae9bfa1d7b9b514f

Observation 34b01009-3b96-47c9-b488-b0d8b9ace1f1 · outbound

This paper cites Wavlm: Large-scale self-supervised pre-training for full stack speech processing.

High-Fidelity Simultaneous Speech-To-Speech Translation Wavlm: Large-scale self-supervised pre-training for full stack speech processing

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T05:01:57.523292Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-09T05:01:56.022240Z digest=sha256:29c9a45051a3d177121ae79deacfd0609b408618cda84ef9529114dc42577b86

Observation d89801ef-80e0-4155-a811-69c8263fe578 · outbound

This paper cites Simple and controllable music generation.

High-Fidelity Simultaneous Speech-To-Speech Translation Simple and controllable music generation

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T05:01:57.504119Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-09T05:01:56.027289Z digest=sha256:b55e382fe58930a08b4c3cf95a7a1650411ce50d87f90d1d76313f50cf4e519c

Observation 4ddc70aa-6e8b-4598-b434-cfbfa288e96d · outbound

This paper cites Moshi: a speech-text foundation model for real-time dialogue.

High-Fidelity Simultaneous Speech-To-Speech Translation Moshi: a speech-text foundation model for real-time dialogue

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-09T05:01:56.033220Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T05:01:56.033220Z digest=sha256:3656f153dd6de74f0a3b3f89b938adf8ae46a32a9fcc90f092b68de8d3ce8e0b

Observation 7faa02a1-24d9-4656-abfc-2d13a9480396 · outbound

This paper cites Daspeech: Directed acyclic transformer for fast and high-quality speech-to-speech translation.

High-Fidelity Simultaneous Speech-To-Speech Translation Daspeech: Directed acyclic transformer for fast and high-quality speech-to-speech translation

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T05:01:57.482088Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-09T05:01:56.038928Z digest=sha256:7b25e875555f1e154fb70c600f944d4d78630747292375135150659e67bf72e3

Observation 08518d04-c4ad-44c0-8467-9002f0bddf52 · outbound

This paper cites Gaussian Error Linear Units (GELUs).

High-Fidelity Simultaneous Speech-To-Speech Translation Gaussian Error Linear Units (GELUs)

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-09T05:01:56.044841Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T05:01:56.044841Z digest=sha256:7c5dcae81c89b9c1429912f7663dd8fa3cb0a0ce1991b71bf856a41a1316a2bf

Observation 50749da1-2985-47e5-b8e5-8f0c8d5d4887 · outbound

This paper cites U nit Y : Two-pass direct speech-to-speech translation with discrete units.

High-Fidelity Simultaneous Speech-To-Speech Translation U nit Y : Two-pass direct speech-to-speech translation with discrete units

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-09T05:01:56.050626Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T05:01:56.050626Z digest=sha256:4342f959a8e2ae85bbd1902a5a3e24b0cb57c03c6050aeda44852a8905455ec1

Observation 5b5976ed-61f7-4af7-a975-f8ebbc10f3ca · outbound

This paper cites N., McNair, A.

High-Fidelity Simultaneous Speech-To-Speech Translation N., McNair, A

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T05:01:57.464660Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-09T05:01:56.056181Z digest=sha256:b973ed3001848acd4a61aeefdd416088b960b57766c8423d347fc5d7bcdf9286

Observation 48cf3ee9-048c-46ad-8c88-689e1443a03f · outbound

This paper cites J., Biadsy, F., Macherey, W., Johnson, M., Chen, Z., and Wu, Y.

High-Fidelity Simultaneous Speech-To-Speech Translation J., Biadsy, F., Macherey, W., Johnson, M., Chen, Z., and Wu, Y

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-09T05:01:56.060726Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T05:01:56.060726Z digest=sha256:1279b60d75d82d51c75efc76e6befbce434b6805bee284bca390cfb133cabf32

Observation 9beecdd9-e835-4a2e-8a16-b6c0802f5206 · outbound

This paper cites T., Remez, T., and Pomerantz, R.

High-Fidelity Simultaneous Speech-To-Speech Translation T., Remez, T., and Pomerantz, R

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T05:01:57.446687Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-09T05:01:56.065521Z digest=sha256:34e8ec52c399f34131b3d837a37b3e81590db7470850f75eae5032c46165e9d9

Observation 05095930-e182-4e7f-962a-c527ea50e010 · outbound

This paper cites T., Wang, Q., and Zen, H.

High-Fidelity Simultaneous Speech-To-Speech Translation T., Wang, Q., and Zen, H

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T05:01:57.429257Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-09T05:01:56.070340Z digest=sha256:158c7ab01140544506f64ccb49614092a013069fdeb84a15ae76229c8c5ad55f

Observation e6ebff19-ee89-41c4-9881-4bc9efd3ca52 · outbound

This paper cites CTRL: A Conditional Transformer Language Model for Controllable Generation.

High-Fidelity Simultaneous Speech-To-Speech Translation CTRL: A Conditional Transformer Language Model for Controllable Generation

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-09T05:01:56.074836Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T05:01:56.074836Z digest=sha256:2341df25bf64d16f465703ade2e8587231e03aca7d0650fa7188ccd61b3aac40

Observation 24255d2b-6aa9-46a7-a490-196b1072a740 · outbound

This paper cites V., Buckley, C.

High-Fidelity Simultaneous Speech-To-Speech Translation V., Buckley, C

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T05:01:57.409573Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-09T05:01:56.079896Z digest=sha256:bfc0e35f45fa36a60eda72fac481ab78a2e1f2541a70ec7660718ddb1162bfa3

Observation ce7b1ffc-5375-44e8-abcc-6c5082a82716 · outbound

This paper cites Audiogen: Textually guided audio generation.

High-Fidelity Simultaneous Speech-To-Speech Translation Audiogen: Textually guided audio generation

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-09T05:01:56.084617Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T05:01:56.084617Z digest=sha256:f8126ba162f45b59780e2a308b456853a37049ca73b0b8d6bdcc69d27a5c1eb9

Observation ca2ee28e-bcc2-4b4c-b305-ca89eb803cc9 · outbound

This paper cites MADLAD-400: A multilingual and document-level large audited dataset.

High-Fidelity Simultaneous Speech-To-Speech Translation MADLAD-400: A multilingual and document-level large audited dataset

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T05:01:57.368523Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-09T05:01:56.089320Z digest=sha256:14b8582efe57feafb73ef235caa4a715bb04d74c25e5c4e1324587adf0d85560

Observation 9ec4525b-5246-484a-a059-e6eae5da9c5b · outbound

This paper cites Direct speech-to-speech translation with discrete units.

High-Fidelity Simultaneous Speech-To-Speech Translation Direct speech-to-speech translation with discrete units

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-09T05:01:56.093998Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T05:01:56.093998Z digest=sha256:092a15f4374893343f18844b5e01675a4f87759a73d075af990073db87db4bb3

Observation 1fd3affd-56e0-4277-be05-ca9550fe94fc · outbound

This paper cites Autoregressive image generation using residual quantization.

High-Fidelity Simultaneous Speech-To-Speech Translation Autoregressive image generation using residual quantization

Reference 21

Resolution
verified exact
arxiv_id_nonexistent, observed 2026-08-09T05:01:57.124988Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-09T05:01:56.098766Z digest=sha256:eff85bf8eb239355cbae0e24d7786e95c77a9234d807e93d9e028cb89d4019b6

Observation f92fafe5-50bb-47a0-9ed5-ab1b415a590b · outbound

This paper cites and Hutter, F.

High-Fidelity Simultaneous Speech-To-Speech Translation and Hutter, F

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T05:01:57.350884Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-09T05:01:56.103516Z digest=sha256:352298acd200ff260f617469e6392b4eb987ff68290a3e2c847529499575af81

Observation d1939de2-4210-4e52-a5ca-382d1616690a · outbound

This paper cites whisper-timestamped.

High-Fidelity Simultaneous Speech-To-Speech Translation whisper-timestamped

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T05:01:57.329358Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-09T05:01:56.108609Z digest=sha256:9ed7bfe3a933aedea40edc0a7a987d937408cf2b1ad04ed856aa76e85e4ece7d

Observation 60b1f997-233b-4e5f-be6a-4472f42a01f8 · outbound

This paper cites J., Koehn, P., and Pino, J.

High-Fidelity Simultaneous Speech-To-Speech Translation J., Koehn, P., and Pino, J

Reference 24

Resolution
verified exact
arxiv_id_nonexistent, observed 2026-08-09T05:01:56.923215Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-09T05:01:56.113420Z digest=sha256:2b3c46090e5b67612439bf5e0f56d93288914107aa9cdb81e08878f19178f368

Observation 7891340f-a044-49e7-93d2-97500fbd21be · outbound

This paper cites The ATR multilingual speech-to-speech translation system.

High-Fidelity Simultaneous Speech-To-Speech Translation The ATR multilingual speech-to-speech translation system

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T05:01:57.309907Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-09T05:01:56.118124Z digest=sha256:4f63525da74bce9c03e586e7e028aa8b0c8c700b9db7ec1e751f7602147118f2

Observation 57f85f41-24c6-476c-bf93-504b9269e077 · outbound

This paper cites Over-Generation Cannot Be Rewarded: Length-Adaptive Average Lagging for Simultaneous Speech Translation.

High-Fidelity Simultaneous Speech-To-Speech Translation Over-Generation Cannot Be Rewarded: Length-Adaptive Average Lagging for Simultaneous Speech Translation

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-09T05:01:56.122670Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T05:01:56.122670Z digest=sha256:966b965285267fbab5bdb2c7f89b2854ab8f579bc95323696f2edebbb058ba06

Observation 903523d6-6300-4dcb-b1f3-22c82a8d28ff · outbound

This paper cites A call for clarity in reporting BLEU scores.

High-Fidelity Simultaneous Speech-To-Speech Translation A call for clarity in reporting BLEU scores

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-09T05:01:56.127955Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T05:01:56.127955Z digest=sha256:79583b720f9ce8215d5abe7d75823410864408776fcf398b9f4e16ab48cdb45c

Observation c6daa7a8-5227-4a5e-95ac-5cbefc82e74e · outbound

This paper cites W., Xu, T., Brockman, G., McLeavey, C., and Sutskever, I.

High-Fidelity Simultaneous Speech-To-Speech Translation W., Xu, T., Brockman, G., McLeavey, C., and Sutskever, I

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T05:01:57.287455Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-09T05:01:56.132916Z digest=sha256:d86835626ecaae17c906ec4c6cc51da7cf77f162ff1da64665d769b11914f66e

Observation 6edda002-bfde-4916-ac13-a2239a5ad4e8 · outbound

This paper cites Generating diverse high-fidelity images with vq-vae-2.

High-Fidelity Simultaneous Speech-To-Speech Translation Generating diverse high-fidelity images with vq-vae-2

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-09T05:01:56.137836Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T05:01:56.137836Z digest=sha256:fe3ba35a5b14779009ced08b4048dc71f343f1b68c5eeabbba1608deb5d62acc

Observation 6d378446-2f56-4326-839d-ba18092464e6 · outbound

This paper cites S imul S peech: End-to-end simultaneous speech to text translation.

High-Fidelity Simultaneous Speech-To-Speech Translation S imul S peech: End-to-end simultaneous speech to text translation

Reference 30

Resolution
verified exact
doi, observed 2026-08-09T05:01:56.338499Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-09T05:01:56.142393Z digest=sha256:49c6b3541acb50a8fdbc418a0d014fea448cf75167edac8d6b530ae25ae54c28

Observation 8156c230-6d5b-4196-81e1-391c6ab067b3 · outbound

This paper cites AudioPaLM: A Large Language Model That Can Speak and Listen.

High-Fidelity Simultaneous Speech-To-Speech Translation AudioPaLM: A Large Language Model That Can Speak and Listen

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-09T05:01:56.148897Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T05:01:56.148897Z digest=sha256:f39677890d52c7b529670b67c4fcebdbc5f53f4b71ceb956170254462d0e5eba

Observation 0e000f0b-d714-4bd1-9d26-2353563fc308 · outbound

This paper cites PySBD: Pragmatic Sentence Boundary Disambiguation.

High-Fidelity Simultaneous Speech-To-Speech Translation PySBD: Pragmatic Sentence Boundary Disambiguation

Reference 32

Resolution
verified exact
local_arxiv, observed 2026-08-09T05:01:56.730713Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-09T05:01:56.154220Z digest=sha256:6f73478ddd456bcafb2db93474c258dcb83e6143723d54c9e2ef882dedc2d470

Observation f51f16b1-a65a-4bca-8c9e-bfae812a60f2 · outbound

This paper cites GLU Variants Improve Transformer.

High-Fidelity Simultaneous Speech-To-Speech Translation GLU Variants Improve Transformer

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-09T05:01:56.159706Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T05:01:56.159706Z digest=sha256:477521b306737b8295c4406d57903d01a3e77219760b3b3a283e39e678c66b22

Observation d9fdb950-94ea-46d3-9e15-0baa1977ca41 · outbound

This paper cites N., and and, L.

High-Fidelity Simultaneous Speech-To-Speech Translation N., and and, L

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T05:01:57.258491Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-09T05:01:56.165010Z digest=sha256:bffac4d1ef207f7e17a18cef0e462ac7122a2f6c2483f81a1496a29c3947522d

Observation 3c3af593-c798-465e-942f-de4f1d60516f · outbound

This paper cites Verbmobil: Foundations of speech-to-speech translation.

High-Fidelity Simultaneous Speech-To-Speech Translation Verbmobil: Foundations of speech-to-speech translation

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T05:01:57.240688Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-09T05:01:56.170325Z digest=sha256:13fda8ef21e4ab1c21122693798c10e7796ae949fd395ae2e0507d8c907711d8

Observation 91613f5d-e64e-475e-b9b5-78b1301e43e2 · outbound

This paper cites Fairseq S 2 T : Fast speech-to-text modeling with fairseq.

High-Fidelity Simultaneous Speech-To-Speech Translation Fairseq S 2 T : Fast speech-to-text modeling with fairseq

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-09T05:01:56.175353Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T05:01:56.175353Z digest=sha256:bd498e1dc9267eeb569df0a8e8e84c9661a6abfb2b7d08221c0462020432f288

Observation 28bb8ab5-b258-4122-b959-8943babad454 · outbound

This paper cites M., and Dupoux, E.

High-Fidelity Simultaneous Speech-To-Speech Translation M., and Dupoux, E

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-09T05:01:56.180159Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T05:01:56.180159Z digest=sha256:45701a83b0e7dd8d5914d4ae90c1e3f8d4ae3bbcd57c3d0185fd8d0d20dac1ce

Observation a21a9aa8-5682-4f4d-964e-55c2a8a53bf0 · outbound

This paper cites Covost 2 and massively multilingual speech translation.

High-Fidelity Simultaneous Speech-To-Speech Translation Covost 2 and massively multilingual speech translation

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-09T05:01:56.185031Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T05:01:56.185031Z digest=sha256:c0e69a7c463c6561bb5f923e19aeb52dbeb944dc37715f783d2cdccc60b57bca

Observation bcd0aecf-2797-4f6f-81af-9765b1c5006c · outbound

This paper cites J., Chorowski, J., Jaitly, N., Wu, Y., and Chen, Z.

High-Fidelity Simultaneous Speech-To-Speech Translation J., Chorowski, J., Jaitly, N., Wu, Y., and Chen, Z

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-09T05:01:56.190014Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T05:01:56.190014Z digest=sha256:3c0cb5bf59931a764db8570b45b5206ff465ff192a7fc1f5491d182af39eaf25

Observation e1ad2248-1f26-4077-b858-4f3daf0ffb90 · outbound

This paper cites UniAudio: An Audio Foundation Model Toward Universal Audio Generation.

High-Fidelity Simultaneous Speech-To-Speech Translation UniAudio: An Audio Foundation Model Toward Universal Audio Generation

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-09T05:01:56.194777Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T05:01:56.194777Z digest=sha256:8b172c390926c00ffaf4fafd38ec87e78e1326f19b3134b50811bdb0c289fdfc

Observation bf300c78-a469-47f9-a984-78049e551e54 · outbound

This paper cites Soundstream: An end-to-end neural audio codec.

High-Fidelity Simultaneous Speech-To-Speech Translation Soundstream: An end-to-end neural audio codec

Reference 41

Resolution
verified exact
arxiv_id_nonexistent, observed 2026-08-09T05:01:56.666912Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-09T05:01:56.199848Z digest=sha256:ed11cc04795dd2a8c8a0ed75c313adffc27c8edcf52ee380c8198ccc274013f6

Observation 8cf7ae20-c6a1-44f5-900f-6a3bcafa2fe1 · outbound

This paper cites Realtrans: End-to-end simultaneous speech translation with convolutional weighted-shrinking transformer.

High-Fidelity Simultaneous Speech-To-Speech Translation Realtrans: End-to-end simultaneous speech translation with convolutional weighted-shrinking transformer

Reference 42

Resolution
verified exact
doi, observed 2026-08-09T05:01:56.271331Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-09T05:01:56.204690Z digest=sha256:4def339bedc69ca8b59393350f5191c0c81ca16f77f6e2f839f18e2a3373c3fe

Observation 37dce851-5b88-42c3-824d-2a94cbe06235 · outbound

This paper cites Streamspeech: Simultaneous speech-to-speech translation with multi-task learning.

High-Fidelity Simultaneous Speech-To-Speech Translation Streamspeech: Simultaneous speech-to-speech translation with multi-task learning

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-09T05:01:56.209804Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T05:01:56.209804Z digest=sha256:42d988b1d7e7f3e4b4807ff7b0dea13e92d5dfe7260eeb7efa561427c380c6da

Observation bdbce54b-8e80-4095-8595-534cee55da31 · outbound

This paper cites Speechtokenizer: Unified speech tokenizer for speech language models.

High-Fidelity Simultaneous Speech-To-Speech Translation Speechtokenizer: Unified speech tokenizer for speech language models

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T05:01:57.208909Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-09T05:01:56.214556Z digest=sha256:41d3ffeaf076f780fceadec9b06ca5b956eff8d2f8bf30785e8d8e2aa5563ce7

Observation b55d94bf-86ef-4ebc-ab18-0c2c607ccbc7 · outbound

This paper cites Textless Streaming Speech-to-Speech Translation using Semantic Speech Tokens.

High-Fidelity Simultaneous Speech-To-Speech Translation Textless Streaming Speech-to-Speech Translation using Semantic Speech Tokens

Reference 45

Resolution
verified exact
local_arxiv, observed 2026-08-09T05:01:56.476954Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-09T05:01:56.219187Z digest=sha256:53554c6efd3ef537b25565610429d8f2268313184c8e05e712d8e9c8ebd4bc38

Pith citing papers

Observation e1bcde44-36e0-4eca-ae94-34ec2c0fdf4f · inbound

Breaking the Barriers of Text-Hungry and Audio-Deficient AI cites this paper.

Breaking the Barriers of Text-Hungry and Audio-Deficient AI High-Fidelity Simultaneous Speech-To-Speech Translation

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-07T11:28:52.222146Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:28:52.222146Z digest=sha256:19ef79600ef1d29dfaa32d9742c7eda4a527ea4bbd785c543c71c732754f08e2

Observation 9893ded1-943e-47cd-b054-faab046ad98c · inbound

Streaming Endpointer for Spoken Dialogue using Neural Audio Codecs and Label-Delayed Training cites this paper.

Streaming Endpointer for Spoken Dialogue using Neural Audio Codecs and Label-Delayed Training High-Fidelity Simultaneous Speech-To-Speech Translation

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T05:47:12.941618Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:47:12.941618Z digest=sha256:36fcd86f202bfedbf184b140066c18f6110bbcb42814d22711c1f765493b3b15

Observation 541cf954-a92b-44d0-ac2b-02138a4b836a · inbound

DiffSoundStream: Efficient Speech Tokenization via Diffusion Decoding cites this paper.

DiffSoundStream: Efficient Speech Tokenization via Diffusion Decoding High-Fidelity Simultaneous Speech-To-Speech Translation

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-06T22:11:58.428052Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:11:58.428052Z digest=sha256:086cf666e05013c72780f55a5ace347a8f526c2329eb20adbf0f24e1cd1225e1

Observation 8337e72c-d692-417b-90ce-912298b5ca36 · inbound

Regularized Entropy Information Adaptation with Temporal-Awareness Networks for Simultaneous Speech Translation cites this paper.

Regularized Entropy Information Adaptation with Temporal-Awareness Networks for Simultaneous Speech Translation High-Fidelity Simultaneous Speech-To-Speech Translation

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-05-11T07:20:59.322747Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-10T17:12:56.433339Z digest=sha256:15c0a8d8eddffb7948c616b405c1eceff6348031085fe98665bedf57fdce245e

Observation 433a5b46-cec5-4904-b6a8-20cf5ea6305c · inbound

A Pocket Offline Model for Simultaneous Speech Translation as CUNI Submission to IWSLT 2026 cites this paper.

A Pocket Offline Model for Simultaneous Speech Translation as CUNI Submission to IWSLT 2026 High-Fidelity Simultaneous Speech-To-Speech Translation

Reference 7

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T03:36:30.180383Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-06-28T09:49:20.431240Z digest=sha256:6b6db09db67c84ff4a057ff3a7f7e01e915a5bd0c158ab83317f6da29adfe303