Pith. sign in

Paper Citation Record · LEDGER

Improving Lip-synchrony in Direct Audio-Visual Speech-to-Speech Translation

As of 21 August 2026, this Paper Citation Record lists 38 of 38 outbound references and 1 inbound Pith citation observation for arXiv:2412.16530.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2412.16530 v1

Coverage vector

measured 38 of 38 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-11T10:35:09.332508Z

measured 39 of 39 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-21T06:32:19.484+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-11T10:35:09.064293Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-11T10:35:09.599954Z

Reference resolution

38 of 38 outbound references displayed

  • verified exact3
  • verified fuzzy28
  • unresolved5
  • parse uncertain0
  • malformed identifier2
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 48f01f24-b522-42a5-a818-78d0f2e95868 · outbound

This paper cites However, improving lip synchrony should not compromise translation quality and natural- ness [4, 5].

Improving Lip-synchrony in Direct Audio-Visual Speech-to-Speech Translation However, improving lip synchrony should not compromise translation quality and natural- ness [4, 5]

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T10:35:10.426069Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-11T10:35:09.052485Z digest=sha256:62a3b1683a325d528d9f1c83b42cb9e764a0109488b3155f197b48b16c2090f3

Observation 09b13862-6ead-4263-94e1-2b23d9679301 · outbound

This paper cites Improving Lip-synchrony in Direct Audio-Visual Speech-to-Speech Translation.

Improving Lip-synchrony in Direct Audio-Visual Speech-to-Speech Translation Improving Lip-synchrony in Direct Audio-Visual Speech-to-Speech Translation

Reference 2

Resolution
verified exact
local_arxiv, observed 2026-08-11T10:35:09.611871Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-11T10:35:09.064293Z digest=sha256:1a07039a2ca5ed2722cb4cab7ab4372e92cc76235d7cfe02975dabef51b6f16c

Observation 1cc8f621-11e5-4b54-b384-6f20809688a0 · outbound

This paper cites length pre- dictor.

Improving Lip-synchrony in Direct Audio-Visual Speech-to-Speech Translation length pre- dictor

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T10:35:10.407359Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-11T10:35:09.071584Z digest=sha256:5e7d47edd331879e1772fbdcf41a30ea1fbff0262eb189c249622296427a9bba

Observation bd53b506-d9e8-459e-ad3c-e1cdb2554c6c · outbound

This paper cites Dataset We leverage LRS3 [28] which is a large-scale video data consisting of thousands of spoken sentences collected from TED talks.

Improving Lip-synchrony in Direct Audio-Visual Speech-to-Speech Translation Dataset We leverage LRS3 [28] which is a large-scale video data consisting of thousands of spoken sentences collected from TED talks

Reference 4

Resolution
malformed identifier
raw_fallback, observed 2026-08-11T10:35:10.383773Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-11T10:35:09.077958Z digest=sha256:684af3656a0095967cb07ea90f6551a1c30d5434cf9f55697e75dfe007447e68

Observation 0e87283f-f2f5-4279-a515-02381e8255ba · outbound

This paper cites Baselines We use the latest work of A V2A V [6] as a strong baseline for this work.

Improving Lip-synchrony in Direct Audio-Visual Speech-to-Speech Translation Baselines We use the latest work of A V2A V [6] as a strong baseline for this work

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T10:35:10.357561Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-11T10:35:09.087546Z digest=sha256:2153a359b6d16ea6e7aec75da354bb0dffe33ffa8a58234c701e76560743e439

Observation 8157bb37-63b4-4cf8-8d26-e21d8a365cc2 · outbound

This paper cites not trading off lip-synchrony improvements over speech translation quality and naturalness.

Improving Lip-synchrony in Direct Audio-Visual Speech-to-Speech Translation not trading off lip-synchrony improvements over speech translation quality and naturalness

Reference 6

Resolution
malformed identifier
raw_fallback, observed 2026-08-11T10:35:10.331549Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-11T10:35:09.095939Z digest=sha256:89ffd538f36c5e771d035e78cbceebad5d1e7a4840b198918544e19171463dda

Observation b3da65a4-4462-492d-972e-228dc9eb1f3f · outbound

This paper cites Our A VS2S framework incorporates lip-synchrony and duration loss to enhance the alignment between speech and lip movements in audio-visual translation models.

Improving Lip-synchrony in Direct Audio-Visual Speech-to-Speech Translation Our A VS2S framework incorporates lip-synchrony and duration loss to enhance the alignment between speech and lip movements in audio-visual translation models

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T10:35:10.310981Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-11T10:35:09.104037Z digest=sha256:203f3663479832c855248978235586421f244ae6289a35707969ee6edd081fe9

Observation 5b4d7f32-4bc9-47d3-ac65-249b21f5439f · outbound

This paper cites an unresolved cited work.

Improving Lip-synchrony in Direct Audio-Visual Speech-to-Speech Translation Unresolved cited work

Reference 8

Resolution
unresolved
raw_fallback, observed 2026-08-11T10:35:10.279147Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-11T10:35:09.118415Z digest=sha256:b98c6d4d4b915bde6cd8dac71325e554769c9581cd7759c35ae4018e1a480037

Observation c2e5ce6e-1e54-4ae0-85fa-ba9292f71906 · outbound

This paper cites Neural dubber: Dubbing for videos according to scripts,.

Improving Lip-synchrony in Direct Audio-Visual Speech-to-Speech Translation Neural dubber: Dubbing for videos according to scripts,

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T10:35:10.254149Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-11T10:35:09.131802Z digest=sha256:d567d274157dc4fb1f086e54b1ce5e343592d9210e6b0a76f357070a997d4ac1

Observation d87a7425-51f0-442f-b41f-85034527586e · outbound

This paper cites Neural style-preserving visual dubbing,.

Improving Lip-synchrony in Direct Audio-Visual Speech-to-Speech Translation Neural style-preserving visual dubbing,

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T10:35:10.226639Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-11T10:35:09.139335Z digest=sha256:9523b7aad4c47cfcf372877190a33b2e4be09770a315df99f63320edb613e142

Observation 7bcbdedd-8edd-4c30-980d-4f51a097cae4 · outbound

This paper cites An empirical take on the dubbing vs. subtitling debate: An eye movement study,.

Improving Lip-synchrony in Direct Audio-Visual Speech-to-Speech Translation An empirical take on the dubbing vs. subtitling debate: An eye movement study,

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T10:35:10.202474Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-11T10:35:09.146123Z digest=sha256:8ce3211bd715d51ec25fb66aa1d12dd5043959c416a45495c28d6191c75134dd

Observation db58254e-3b8d-41e7-af47-0ea539b880bb · outbound

This paper cites Dub- bing in practice: A large scale study of human localization with insights for automatic dubbing,.

Improving Lip-synchrony in Direct Audio-Visual Speech-to-Speech Translation Dub- bing in practice: A large scale study of human localization with insights for automatic dubbing,

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T10:35:10.176027Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-11T10:35:09.153496Z digest=sha256:0fe0e214b3d5ccf15b7166313bd6950a789b0f55519e1177ee2943e00a6a45d7

Observation ae9b1e37-533a-463e-8761-e5a042b3c78b · outbound

This paper cites Av2av: Direct audio-visual speech to audio-visual speech translation with unified audio-visual speech representation,.

Improving Lip-synchrony in Direct Audio-Visual Speech-to-Speech Translation Av2av: Direct audio-visual speech to audio-visual speech translation with unified audio-visual speech representation,

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T10:35:10.144852Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-11T10:35:09.160001Z digest=sha256:c8655eae098a21fde351141a6ffb39d0e464c728a2aa2d97e4dbc9eea408b0ba

Observation 5fc1aa79-cb70-4753-9aa0-ebeb9854336c · outbound

This paper cites A lip sync expert is all you need for speech to lip generation in the wild,.

Improving Lip-synchrony in Direct Audio-Visual Speech-to-Speech Translation A lip sync expert is all you need for speech to lip generation in the wild,

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T10:35:10.125322Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-11T10:35:09.166546Z digest=sha256:ca0466ab8b3176899e0f118f04f40d358c17c5340e574d7024be228fb619ebf4

Observation b30dd3ac-d018-4241-95a4-fc97368a5577 · outbound

This paper cites Exposing lip-syncing deepfakes from mouth inconsistencies,.

Improving Lip-synchrony in Direct Audio-Visual Speech-to-Speech Translation Exposing lip-syncing deepfakes from mouth inconsistencies,

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T10:35:10.091239Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-11T10:35:09.172234Z digest=sha256:42c080d025aa1d539135970ff7e6c8910b8d2ed04d81ee8b7712e1cfa89e9ac9

Observation 4efb3245-a2e9-4dbc-b3bd-501912ea1623 · outbound

This paper cites The DeepFake Detection Challenge (DFDC) Dataset.

Improving Lip-synchrony in Direct Audio-Visual Speech-to-Speech Translation The DeepFake Detection Challenge (DFDC) Dataset

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-11T10:35:09.178173Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T10:35:09.178173Z digest=sha256:5f7cf6fb4eef7b70cc3bcd4925f03071deab24f053af0ebe93b23f53760b8a22

Observation 3d2f24ab-3b92-4199-98df-c3cefd8febbc · outbound

This paper cites Faces of the future: How generative ai is redefining likeness and identity in the age of artificial intelligence,.

Improving Lip-synchrony in Direct Audio-Visual Speech-to-Speech Translation Faces of the future: How generative ai is redefining likeness and identity in the age of artificial intelligence,

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T10:35:10.064467Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-11T10:35:09.184553Z digest=sha256:dc7379f45fda5e7eed6f21134f6e4dfd2aa12d47876a35873459704f8fdfdd09

Observation 597d37e1-ed56-4bad-b3ce-5a10760bc06c · outbound

This paper cites Face/off: Changing the face of movies with deepfakes,.

Improving Lip-synchrony in Direct Audio-Visual Speech-to-Speech Translation Face/off: Changing the face of movies with deepfakes,

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T10:35:10.036165Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-11T10:35:09.194328Z digest=sha256:4d8572e139f965fefadfa3f03d93fbaf90dc60f20984a181e2f466f91be14b9f

Observation 5ca9d8c6-0665-4be9-bb34-cad046ac5217 · outbound

This paper cites Regu- lating deep fakes: legal and ethical considerations,.

Improving Lip-synchrony in Direct Audio-Visual Speech-to-Speech Translation Regu- lating deep fakes: legal and ethical considerations,

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T10:35:10.007227Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-11T10:35:09.202548Z digest=sha256:1bc4ea0f640013250192c960bf32d371b2155bd5d0da9c5bcd08cdfead41e6dc

Observation e5449a1a-143f-490d-8929-bfc5e69fd43a · outbound

This paper cites VASA-1: Lifelike Audio-Driven Talking Faces Generated in Real Time.

Improving Lip-synchrony in Direct Audio-Visual Speech-to-Speech Translation VASA-1: Lifelike Audio-Driven Talking Faces Generated in Real Time

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-11T10:35:09.209823Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T10:35:09.209823Z digest=sha256:089e50d58072595f8441ee446b16619f12a829e3d2f802a3427e3893edcdb95a

Observation 968d1cab-5d56-4bec-bd12-ca9ff9c6e8d7 · outbound

This paper cites AV-TranSpeech: Audio-Visual Robust Speech-to-Speech Translation.

Improving Lip-synchrony in Direct Audio-Visual Speech-to-Speech Translation AV-TranSpeech: Audio-Visual Robust Speech-to-Speech Translation

Reference 21

Resolution
verified exact
local_arxiv, observed 2026-08-11T10:35:09.513252Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-11T10:35:09.217279Z digest=sha256:cf85c5b7c40dfb6704ceccf818f08544daadd36b6cef3eaa64f40f2f5438f6ad

Observation 2bcf0595-4ea8-48f0-b2a0-60dea3d0c6bf · outbound

This paper cites Mixspeech: Cross-modality self-learning with audio- visual stream mixup for visual speech translation and recogni- tion,.

Improving Lip-synchrony in Direct Audio-Visual Speech-to-Speech Translation Mixspeech: Cross-modality self-learning with audio- visual stream mixup for visual speech translation and recogni- tion,

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T10:35:09.983616Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-11T10:35:09.224683Z digest=sha256:a8c8c8c5da06e4688a8efcc919f591452e93cbed5209cb1a20e71228738555bc

Observation 7856dcaa-b336-4025-9c3f-d5ab29dd9ded · outbound

This paper cites Isometric MT: Neural Machine Translation for Automatic Dubbing.

Improving Lip-synchrony in Direct Audio-Visual Speech-to-Speech Translation Isometric MT: Neural Machine Translation for Automatic Dubbing

Reference 23

Resolution
verified exact
local_arxiv, observed 2026-08-11T10:35:09.473359Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-11T10:35:09.231449Z digest=sha256:7be5e979985e60a724d1e9fa44edf91df39eaeb52250ac7352fda9084a8a6677

Observation a2b0c858-646d-44f2-85e3-a560cbd37672 · outbound

This paper cites Duration modeling of neural tts for au- tomatic dubbing,.

Improving Lip-synchrony in Direct Audio-Visual Speech-to-Speech Translation Duration modeling of neural tts for au- tomatic dubbing,

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T10:35:09.959827Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-11T10:35:09.238692Z digest=sha256:b720036ebee0d353a3ba9cda1f1879accfd5dc2fc8e415535263119dc6eeafc2

Observation 5a59b774-35e0-47c9-9aef-d94b44097f21 · outbound

This paper cites Prosodic alignment for off-screen automatic dubbing,.

Improving Lip-synchrony in Direct Audio-Visual Speech-to-Speech Translation Prosodic alignment for off-screen automatic dubbing,

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T10:35:09.935445Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-11T10:35:09.245664Z digest=sha256:5c16d1149992fb833e341fdd4ec25658096992cd02171841406e275423e046db

Observation 8f9d8b29-748a-4cc4-9cd1-16e032190486 · outbound

This paper cites Improving isochronous machine translation with target factors and auxil- iary counters,.

Improving Lip-synchrony in Direct Audio-Visual Speech-to-Speech Translation Improving isochronous machine translation with target factors and auxil- iary counters,

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T10:35:09.911417Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-11T10:35:09.251652Z digest=sha256:0c88652fc493c6fe4de7878e2f42602d2a56cb38980d2ed0d20b5ebf6f74e118

Observation 2d7bcfbd-78ec-4504-ad66-1ba22fd1f0b5 · outbound

This paper cites Jointly optimizing translations and speech timing to improve isochrony in automatic dubbing,.

Improving Lip-synchrony in Direct Audio-Visual Speech-to-Speech Translation Jointly optimizing translations and speech timing to improve isochrony in automatic dubbing,

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T10:35:09.887262Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-11T10:35:09.258509Z digest=sha256:e5e7d309b1d9141cc65b589f3e080095b24655edbe7c28177b97ec5df120b3e7

Observation 45d4c327-5039-489d-bbee-273c17c64754 · outbound

This paper cites A lip sync expert is all you need for speech to lip generation in the wild,.

Improving Lip-synchrony in Direct Audio-Visual Speech-to-Speech Translation A lip sync expert is all you need for speech to lip generation in the wild,

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T10:35:09.865088Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-11T10:35:09.264839Z digest=sha256:6b3fb9c390de7bd0a4368c5d3502c40635fef3725277ed90d7ab0fb7fe0bcd3d

Observation 2f4f0e13-2e69-486b-b98c-53a90407e7b7 · outbound

This paper cites Towards realistic visual dubbing with heterogeneous sources,.

Improving Lip-synchrony in Direct Audio-Visual Speech-to-Speech Translation Towards realistic visual dubbing with heterogeneous sources,

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T10:35:09.838034Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-11T10:35:09.270885Z digest=sha256:317c7783eaefebf387acecf82aed21c9497850bf9bd4f1061be2286fb40c8119

Observation 10b682fe-da9f-4fb1-8504-5b4ebe74edad · outbound

This paper cites Learning audio-visual speech representation by masked multimodal cluster prediction,.

Improving Lip-synchrony in Direct Audio-Visual Speech-to-Speech Translation Learning audio-visual speech representation by masked multimodal cluster prediction,

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T10:35:09.804108Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-11T10:35:09.277607Z digest=sha256:c5e687cbabe4180ddf38205d3757e573cab154fa1ecec1c6d4a90b090e390df8

Observation d2425594-730b-45eb-bff1-6d52f5006d3c · outbound

This paper cites Textless Unit-to-Unit training for Many-to-Many Multilingual Speech-to-Speech Translation.

Improving Lip-synchrony in Direct Audio-Visual Speech-to-Speech Translation Textless Unit-to-Unit training for Many-to-Many Multilingual Speech-to-Speech Translation

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-11T10:35:09.283881Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T10:35:09.283881Z digest=sha256:d514878c69fb54f0661b300f56800dbc6196a5687bf91ceb2eafd190cb1383fa

Observation 5c5933dd-6f8f-49d9-8433-a7b144d0d95d · outbound

This paper cites Hifi-gan: Generative adversarial networks for efficient and high fidelity speech synthesis,.

Improving Lip-synchrony in Direct Audio-Visual Speech-to-Speech Translation Hifi-gan: Generative adversarial networks for efficient and high fidelity speech synthesis,

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T10:35:09.777830Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-11T10:35:09.291223Z digest=sha256:f7ac737821f82c039884179f471404f1e47c1e9310776bce15f15a875b3ece1f

Observation 2c217f1f-2490-4f0c-aaac-ef9188e734cd · outbound

This paper cites Direct speech-to-speech translation with discrete units.

Improving Lip-synchrony in Direct Audio-Visual Speech-to-Speech Translation Direct speech-to-speech translation with discrete units

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-11T10:35:09.298096Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T10:35:09.298096Z digest=sha256:dc42dde3ae511e2849fa3ee29eea0d72aa38d82ff0cff0db6120a7832738d834

Observation b8162e89-b6ec-434d-870d-2ca6ec764b82 · outbound

This paper cites Out of time: automated lip sync in the wild,.

Improving Lip-synchrony in Direct Audio-Visual Speech-to-Speech Translation Out of time: automated lip sync in the wild,

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T10:35:09.750054Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-11T10:35:09.305579Z digest=sha256:a28e7dbe4deb67fa51001bb6afc99904528fae3cc033415cc5cf830bce9ebe6e

Observation c1020e8c-7ec4-4652-9238-bbd7e7c3dfbd · outbound

This paper cites Deep audio-visual speech recognition,.

Improving Lip-synchrony in Direct Audio-Visual Speech-to-Speech Translation Deep audio-visual speech recognition,

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T10:35:09.721959Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-11T10:35:09.311869Z digest=sha256:524c11c8e800a097df6bb20be22e4fffce890f6f84226ae4ced5d7978057d1aa

Observation 15901fcc-e929-4b2a-8533-f827d0fe3a59 · outbound

This paper cites Seamlessm4t: Massively multilin- gual and multimodal machine translation,.

Improving Lip-synchrony in Direct Audio-Visual Speech-to-Speech Translation Seamlessm4t: Massively multilin- gual and multimodal machine translation,

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T10:35:09.696514Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-11T10:35:09.318632Z digest=sha256:f4450dfff6229ebbeeb1a6705859d94087ab16e28ca025ed3fa20c0541077dba

Observation 03b8eb3a-bd75-4700-9285-ad17a4e449cc · outbound

This paper cites Bleu: a method for automatic evaluation of machine translation,.

Improving Lip-synchrony in Direct Audio-Visual Speech-to-Speech Translation Bleu: a method for automatic evaluation of machine translation,

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T10:35:09.669365Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-11T10:35:09.325077Z digest=sha256:b01b4a2eb9f1c7dc0d7c7891def51508f32c014fe7c1798314a415d8598382f5

Observation c974c953-a91b-400b-86cd-7f22352d87eb · outbound

This paper cites Decoupled weight decay regularization,.

Improving Lip-synchrony in Direct Audio-Visual Speech-to-Speech Translation Decoupled weight decay regularization,

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T10:35:09.643732Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-11T10:35:09.332508Z digest=sha256:e90e573be32b785efce3536b0d8b524c86b6d36560172b5a331e263ca0475cd1

Pith citing papers

Observation 09b13862-6ead-4263-94e1-2b23d9679301 · inbound

Improving Lip-synchrony in Direct Audio-Visual Speech-to-Speech Translation cites this paper.

Improving Lip-synchrony in Direct Audio-Visual Speech-to-Speech Translation Improving Lip-synchrony in Direct Audio-Visual Speech-to-Speech Translation

Reference 2

Resolution
verified exact
local_arxiv, observed 2026-08-11T10:35:09.611871Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-11T10:35:09.064293Z digest=sha256:1a07039a2ca5ed2722cb4cab7ab4372e92cc76235d7cfe02975dabef51b6f16c