Pith. sign in

Paper Citation Record · LEDGER

Length Aware Speech Translation for Video Dubbing

As of 14 August 2026, this Paper Citation Record lists 24 of 24 outbound references and 1 inbound Pith citation observation for arXiv:2506.00740.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.00740 v1

Coverage vector

measured 24 of 24 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T12:02:22.477354Z

measured 25 of 25 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-14T06:32:32.682623+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T12:02:20.273561Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-07T12:02:22.672497Z

Reference resolution

24 of 24 outbound references displayed

  • verified exact1
  • verified fuzzy21
  • unresolved2
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 50be5103-5baf-4227-a80e-2b28d4433ed9 · outbound

This paper cites Length Aware Speech Translation for Video Dubbing.

Length Aware Speech Translation for Video Dubbing Length Aware Speech Translation for Video Dubbing

Reference 1

Resolution
verified exact
local_arxiv, observed 2026-08-07T12:02:22.724435Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T12:02:20.273561Z digest=sha256:870702131e1d796919239a547de01f7e8e88101ca017fb5919b55d5e1d9dc11e

Observation d6cd933b-2236-41c7-b09a-e344d07b06c1 · outbound

This paper cites The duration of translated audio is influ- enced by: (a) the length of the translated text, and (b) the dura- tion model within the text-to-speech (TTS) system.

Length Aware Speech Translation for Video Dubbing The duration of translated audio is influ- enced by: (a) the length of the translated text, and (b) the dura- tion model within the text-to-speech (TTS) system

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:02:26.774957Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T12:02:20.320178Z digest=sha256:73853cf96677f24c7378cf6b6b143115d3e5b5190e573a80918b4a9d26d6da3c

Observation 55c0d34f-7e87-4ae5-8f9e-480ebab415d9 · outbound

This paper cites Model and Data The ST model used in our experiments is multilingual and jointly trained on Spanish (ES) and Korean (KO) data.

Length Aware Speech Translation for Video Dubbing Model and Data The ST model used in our experiments is multilingual and jointly trained on Spanish (ES) and Korean (KO) data

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:02:26.482554Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T12:02:20.493478Z digest=sha256:eba8f47f639fb292ea15a13bcf368335d1383e18705052383ca8628a2093e1cb

Observation 9336c538-4a13-4902-a877-f02753dfcf8f · outbound

This paper cites an unresolved cited work.

Length Aware Speech Translation for Video Dubbing Unresolved cited work

Reference 4

Resolution
unresolved
raw_fallback, observed 2026-08-07T12:02:26.610354Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T12:02:20.403503Z digest=sha256:d0e5508d3b201ea72ec7efbaf12f9079c6977ceff1e1a7010b159334a2d90296

Observation e5019cc5-db24-48f8-85f7-84aae49d17e8 · outbound

This paper cites an unresolved cited work.

Length Aware Speech Translation for Video Dubbing Unresolved cited work

Reference 5

Resolution
unresolved
raw_fallback, observed 2026-08-07T12:02:26.355715Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T12:02:20.581943Z digest=sha256:83e7c9959d51448caae22d765950b6e765c779caeb70486c4900a62cc1bf6f28

Observation 2af8228e-6317-4b13-8f9d-f1e6e8ee17af · outbound

This paper cites Our approach leverages predefined length control tokens to generate transla- tions of varying lengths—short, normal, and long—while main- taining high translation quality.

Length Aware Speech Translation for Video Dubbing Our approach leverages predefined length control tokens to generate transla- tions of varying lengths—short, normal, and long—while main- taining high translation quality

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:02:26.158753Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T12:02:20.678246Z digest=sha256:b836bc3476b9fac70b8f94cf60d5a572b7c77205c14ef7d4031fac8f4a5d7420

Observation d44a2f01-7e25-4985-b926-9ae24a81c927 · outbound

This paper cites Leveraging weakly supervised data to improve end-to-end speech-to-text translation,.

Length Aware Speech Translation for Video Dubbing Leveraging weakly supervised data to improve end-to-end speech-to-text translation,

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:02:25.962999Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T12:02:20.739344Z digest=sha256:89a7b2158534cb4541f577ee749ca20d8a18851e1979b99b8d5923a91a3dc215

Observation e4c0e5fa-9e0c-4500-8491-37fb26954bee · outbound

This paper cites Large-scale stream- ing end-to-end speech translation with neural transducers,.

Length Aware Speech Translation for Video Dubbing Large-scale stream- ing end-to-end speech translation with neural transducers,

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:02:25.805883Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T12:02:20.844471Z digest=sha256:211dc15ba3154c45d72ea643acb50cd6dcc40a2ea39ae098395fcd8686c758b5

Observation 5f522bd4-46cd-490f-911b-a374363298d9 · outbound

This paper cites Revisiting end-to-end speech-to-text translation from scratch,.

Length Aware Speech Translation for Video Dubbing Revisiting end-to-end speech-to-text translation from scratch,

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:02:25.677659Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T12:02:20.972785Z digest=sha256:ed62d70bec8d0412ad9190a31fd652a05138224f430669064c25e842d53233dc

Observation 80d6cb22-325f-485a-a43f-4405985b0eea · outbound

This paper cites Videodubber: machine translation with speech-aware length control for video dubbing,.

Length Aware Speech Translation for Video Dubbing Videodubber: machine translation with speech-aware length control for video dubbing,

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:02:25.469350Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T12:02:21.090318Z digest=sha256:dcf782e2b9a29c0270b0132ca8680e48677ceb4f05159df88874e41ed85aad03

Observation 6f8690b8-3172-4649-860e-18a7fc8f9f60 · outbound

This paper cites Controlling machine translation for multiple attributes with additive interven- tions,.

Length Aware Speech Translation for Video Dubbing Controlling machine translation for multiple attributes with additive interven- tions,

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:02:25.307976Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T12:02:21.144276Z digest=sha256:9220e9906e1fc6bf07088808c4bae5ea4d9d4ca400d7e2ca51e063c8b91de3f5

Observation c891b55d-409f-43c4-9d31-54d915e32e42 · outbound

This paper cites Is 42 the answer to ev- erything in subtitling-oriented speech translation?.

Length Aware Speech Translation for Video Dubbing Is 42 the answer to ev- erything in subtitling-oriented speech translation?

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:02:25.122444Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T12:02:21.231253Z digest=sha256:a85b7f9e8b615be71414ef4f436cdd6f1b58378339dc5705005150ffe6cf907d

Observation 77e30b5e-5581-4f16-8001-2021b968ed32 · outbound

This paper cites HW-TSC’s participa- tion in the IWSLT 2022 isometric spoken language translation,.

Length Aware Speech Translation for Video Dubbing HW-TSC’s participa- tion in the IWSLT 2022 isometric spoken language translation,

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:02:24.977450Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T12:02:21.321075Z digest=sha256:7c00ecf2e7dfa68ee2160ef9e5e5d76a924ad5baaa48c3bbb45de13b8b64bd82

Observation 786a7a35-b98e-4a3d-b333-e4cdf9587e0e · outbound

This paper cites Adapting end-to-end speech recognition for readable subtitles,.

Length Aware Speech Translation for Video Dubbing Adapting end-to-end speech recognition for readable subtitles,

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:02:24.782957Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T12:02:21.444140Z digest=sha256:a7bdc4dcfd8437c023ac5ddea082a0b5d5f1ce0dd7a7c3e288994ed700108ab7

Observation 94eab385-7f58-4ef7-a5a4-0a720508dfa6 · outbound

This paper cites Length- aware NMT and adaptive duration for automatic dubbing,.

Length Aware Speech Translation for Video Dubbing Length- aware NMT and adaptive duration for automatic dubbing,

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:02:24.579489Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T12:02:21.522466Z digest=sha256:3302a82ffd9acf5b43ffc074957cc0c11ff58b3a185e4c29e34e6059c33b1974

Observation 63532acc-da9b-4a64-8027-a88178be48d5 · outbound

This paper cites Machine translation verbosity control for automatic dubbing,.

Length Aware Speech Translation for Video Dubbing Machine translation verbosity control for automatic dubbing,

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:02:24.232143Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T12:02:21.641051Z digest=sha256:43060b2354fc7968688ca4bfde7c4edb6d745a05ba01b4155f8e27db9d8a24a3

Observation 04a7a3fd-aef3-43e6-bca2-b6d024091488 · outbound

This paper cites Isochrony-aware neural machine translation for automatic dub- bing,.

Length Aware Speech Translation for Video Dubbing Isochrony-aware neural machine translation for automatic dub- bing,

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:02:24.056077Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T12:02:21.720649Z digest=sha256:99f606ee8f70ead1667073917c15733b5cdf3e6fcf38d23e86aba04456065ede

Observation 1e40c525-7a80-446e-9302-88fa9f79df3c · outbound

This paper cites Isometric neural machine translation us- ing phoneme count ratio reward-based reinforcement learning,.

Length Aware Speech Translation for Video Dubbing Isometric neural machine translation us- ing phoneme count ratio reward-based reinforcement learning,

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:02:23.923490Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T12:02:21.815996Z digest=sha256:616dff582638281d176e568498c368c76e0149b18a71052e848943a4c3ec6584

Observation c164e01c-07f4-4b11-ab56-17cadb7bef10 · outbound

This paper cites FastSpeech 2: Fast and high-quality end-to-end text to speech,.

Length Aware Speech Translation for Video Dubbing FastSpeech 2: Fast and high-quality end-to-end text to speech,

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:02:23.764767Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T12:02:21.932443Z digest=sha256:7b09359dfec0ab5eedc6a974c6361a2e15c413d1d032288484d287881fadbf3a

Observation 16f9d6ea-13d9-43c2-a79f-a6925629b3d5 · outbound

This paper cites Conformer: Convolution-augmented transformer for speech recognition,.

Length Aware Speech Translation for Video Dubbing Conformer: Convolution-augmented transformer for speech recognition,

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:02:23.548438Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T12:02:22.031806Z digest=sha256:7b1c7831a55319b68a75c2bc86bf8f328fd7c87aebe52f311de48c5daeb126b4

Observation 1c399852-892d-4cbc-b3d9-4033bca30e9c · outbound

This paper cites Hybrid CTC/attention architecture for end-to-end speech recog- nition,.

Length Aware Speech Translation for Video Dubbing Hybrid CTC/attention architecture for end-to-end speech recog- nition,

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:02:23.409250Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T12:02:22.090541Z digest=sha256:8ae0064c7c5fc02936702347ba0a98839aaeab6f50f845691dc672b984e1ed9e

Observation 0b7d8599-7d6a-4c05-a552-fd47b114cc47 · outbound

This paper cites Fleurs: Few-shot learning evaluation of universal representations of speech,.

Length Aware Speech Translation for Video Dubbing Fleurs: Few-shot learning evaluation of universal representations of speech,

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:02:23.206909Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T12:02:22.186666Z digest=sha256:acb093484c728b7c1a910a4b4c7ff7a5487ec32a771302502625506ef9e68dcd

Observation b412060a-49bc-47aa-a46f-c407124203a7 · outbound

This paper cites LeanSpeech: The Microsoft lightweight speech synthesis system for limmits challenge 2023,.

Length Aware Speech Translation for Video Dubbing LeanSpeech: The Microsoft lightweight speech synthesis system for limmits challenge 2023,

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:02:23.010430Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T12:02:22.337542Z digest=sha256:572f0f75015f2c215e4f933b1605a0e1c44eb8fbd23177989e835c5011ac04db

Observation 930721e1-96fd-4a3e-a177-4d0affc1bfeb · outbound

This paper cites A call for clarity in reporting BLEU scores,.

Length Aware Speech Translation for Video Dubbing A call for clarity in reporting BLEU scores,

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:02:22.869228Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T12:02:22.477354Z digest=sha256:39b3bc04e4bb0547eefde78ec38f7bef0bae772c6613e546fd399f342950d9d2

Pith citing papers

Observation 50be5103-5baf-4227-a80e-2b28d4433ed9 · inbound

Length Aware Speech Translation for Video Dubbing cites this paper.

Length Aware Speech Translation for Video Dubbing Length Aware Speech Translation for Video Dubbing

Reference 1

Resolution
verified exact
local_arxiv, observed 2026-08-07T12:02:22.724435Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T12:02:20.273561Z digest=sha256:870702131e1d796919239a547de01f7e8e88101ca017fb5919b55d5e1d9dc11e