Pith. sign in

Paper Citation Record · LEDGER

Multimodal Large Language Model-Enabled Video Translation: A Role-Oriented Survey

As of 21 August 2026, this Paper Citation Record lists 100 of 176 outbound references and 0 inbound Pith citation observations for arXiv:2604.11283.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2604.11283 v2

Coverage vector

measured 100 of 176 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-07-12T22:04:31.302192Z

measured 100 of 100 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-21T06:32:19.484+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

100 of 176 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved100
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 10768c91-4c8f-48fa-a272-7d1ab7561600 · outbound

This paper cites Deep neural networks for acoustic modeling in speech recognition: The shared views of four research groups,.

Multimodal Large Language Model-Enabled Video Translation: A Role-Oriented Survey Deep neural networks for acoustic modeling in speech recognition: The shared views of four research groups,

Reference 1

Resolution
unresolved
no resolver link, observed 2026-07-12T22:04:31.302192Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T22:04:31.302192Z digest=sha256:72b2a5c9e845b17da06fba71ce98509c32b9edfe0aae90a1035fb975fcde3984

Observation 160cea28-2304-4f49-b1da-d2adfc3268ab · outbound

This paper cites Attention is all you need,.

Multimodal Large Language Model-Enabled Video Translation: A Role-Oriented Survey Attention is all you need,

Reference 2

Resolution
unresolved
no resolver link, observed 2026-07-12T22:04:31.302192Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T22:04:31.302192Z digest=sha256:4b1ae878ba6f86ea7a0b1d24c3ff1747d6577ac922273ab1cc98b210df38db42

Observation 9fc18f9d-2a66-4c58-8660-89aecf196f2d · outbound

This paper cites Tacotron: Towards end-to-end speech synthesis,.

Multimodal Large Language Model-Enabled Video Translation: A Role-Oriented Survey Tacotron: Towards end-to-end speech synthesis,

Reference 3

Resolution
unresolved
no resolver link, observed 2026-07-12T22:04:31.302192Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T22:04:31.302192Z digest=sha256:4921e5013a3b51d4436ac49500a2b78b5b0ca3dac549dbb80c5654ca3f74050b

Observation 3dcdd481-d9f2-4978-bef9-da71125d0718 · outbound

This paper cites A lip sync expert is all you need for speech to lip generation in the wild,.

Multimodal Large Language Model-Enabled Video Translation: A Role-Oriented Survey A lip sync expert is all you need for speech to lip generation in the wild,

Reference 4

Resolution
unresolved
no resolver link, observed 2026-07-12T22:04:31.302192Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T22:04:31.302192Z digest=sha256:0392ec5db9aca96d7547da4d315bdc0c3b984fe2ec15d184ef8e5ed7a1dd9fc7

Observation e5303075-5d96-42b2-98ae-d02318d4c7a1 · outbound

This paper cites Direct speech-to-speech translation with a sequence-to- sequence model,.

Multimodal Large Language Model-Enabled Video Translation: A Role-Oriented Survey Direct speech-to-speech translation with a sequence-to- sequence model,

Reference 5

Resolution
unresolved
no resolver link, observed 2026-07-12T22:04:31.302192Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T22:04:31.302192Z digest=sha256:4ab181fdcdb31b8c4725666565edbcb97982679747f1f415dfb1eedb2dea80de

Observation 55e557c7-a83c-461d-982a-04df1c1f32ee · outbound

This paper cites Learning transferable visual models from natural language supervi- sion,.

Multimodal Large Language Model-Enabled Video Translation: A Role-Oriented Survey Learning transferable visual models from natural language supervi- sion,

Reference 6

Resolution
unresolved
no resolver link, observed 2026-07-12T22:04:31.302192Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T22:04:31.302192Z digest=sha256:ff4be951814d18e867386095ab814334e5b0586557dd34d54c5e3105ee133d15

Observation d9a5b663-85ae-46c3-9c1d-b639bcbaf8a1 · outbound

This paper cites Flamingo: A visual language model for few-shot learning,.

Multimodal Large Language Model-Enabled Video Translation: A Role-Oriented Survey Flamingo: A visual language model for few-shot learning,

Reference 7

Resolution
unresolved
no resolver link, observed 2026-07-12T22:04:31.302192Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T22:04:31.302192Z digest=sha256:3f59d26e60c7f66173f81a556c82cff16cbf7af1ae8f3eccd966435381d2bfae

Observation f8270dea-65ae-4719-bcf5-da0ce077fb42 · outbound

This paper cites Blip-2: Bootstrapping language- image pre-training with frozen image encoders and large language models,.

Multimodal Large Language Model-Enabled Video Translation: A Role-Oriented Survey Blip-2: Bootstrapping language- image pre-training with frozen image encoders and large language models,

Reference 8

Resolution
unresolved
no resolver link, observed 2026-07-12T22:04:31.302192Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T22:04:31.302192Z digest=sha256:82e6589260600a044f8239e9c8a8bf719d4813d4e44657c2747021952398b267

Observation a10440fb-602f-488a-8a5f-a99d21359d6f · outbound

This paper cites Visual instruction tuning,.

Multimodal Large Language Model-Enabled Video Translation: A Role-Oriented Survey Visual instruction tuning,

Reference 9

Resolution
unresolved
no resolver link, observed 2026-07-12T22:04:31.302192Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T22:04:31.302192Z digest=sha256:d5f8ab3b2b11ab0b4f2fffa2421418a62615b8601faa725f9dcd16183b22fab2

Observation f4dd3ce0-a03e-43fc-92bc-5c1ff2c46ad4 · outbound

This paper cites mplug-owl2: Revolutionizing multi-modal large language model with modality collaboration,.

Multimodal Large Language Model-Enabled Video Translation: A Role-Oriented Survey mplug-owl2: Revolutionizing multi-modal large language model with modality collaboration,

Reference 10

Resolution
unresolved
no resolver link, observed 2026-07-12T22:04:31.302192Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T22:04:31.302192Z digest=sha256:6c05059e23cab9441d19792f466d722c0fb44eeb2063839b560f74a89a9e251c

Observation c52d8c46-8963-4dcc-a36d-3f38695b628a · outbound

This paper cites Video-chatgpt: Towards detailed video understanding via large vision and language models,.

Multimodal Large Language Model-Enabled Video Translation: A Role-Oriented Survey Video-chatgpt: Towards detailed video understanding via large vision and language models,

Reference 11

Resolution
unresolved
no resolver link, observed 2026-07-12T22:04:31.302192Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T22:04:31.302192Z digest=sha256:adf7f4d7aaff4e83e4175ed0b1d391dd40d4c5eb68596d5d6925b5bc79c2bf79

Observation 90cb84c3-4bd3-4607-a61a-152f361024b2 · outbound

This paper cites Video-llama: An instruction-tuned audio-visual language model for video understanding,.

Multimodal Large Language Model-Enabled Video Translation: A Role-Oriented Survey Video-llama: An instruction-tuned audio-visual language model for video understanding,

Reference 12

Resolution
unresolved
no resolver link, observed 2026-07-12T22:04:31.302192Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T22:04:31.302192Z digest=sha256:b36cfe62eeace239c0712141063d0a54b78f9e3df114cbf4cb509beafc7007bf

Observation 64c1ce9a-3c23-4695-9a24-c3e21dd9b449 · outbound

This paper cites Llama-vid: An image is worth 2 tokens in large language models,.

Multimodal Large Language Model-Enabled Video Translation: A Role-Oriented Survey Llama-vid: An image is worth 2 tokens in large language models,

Reference 13

Resolution
unresolved
no resolver link, observed 2026-07-12T22:04:31.302192Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T22:04:31.302192Z digest=sha256:2e36da7c1f82e86531c794ee6eb789f1ef076f3773050252b75c02086313d5bc

Observation c54f1da5-4857-4b49-a020-fc06cc90dc37 · outbound

This paper cites Internvideo2: Scaling foundation models for multimodal video understanding,.

Multimodal Large Language Model-Enabled Video Translation: A Role-Oriented Survey Internvideo2: Scaling foundation models for multimodal video understanding,

Reference 14

Resolution
unresolved
no resolver link, observed 2026-07-12T22:04:31.302192Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T22:04:31.302192Z digest=sha256:ee40934c7aed96255923d5e50f7a8519b07175a91dbd37b555e73e17ddb192e9

Observation a4899045-5f27-4698-9c8e-33b41a8884b0 · outbound

This paper cites CosyVoice 2: Scalable Streaming Speech Synthesis with Large Language Models.

Multimodal Large Language Model-Enabled Video Translation: A Role-Oriented Survey CosyVoice 2: Scalable Streaming Speech Synthesis with Large Language Models

Reference 15

Resolution
unresolved
no resolver link, observed 2026-07-12T22:04:31.302192Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T22:04:31.302192Z digest=sha256:89029b4076bd06608ff90211701fbb166ca3e6b5e9a7cdce3bf7301d7e06812f

Observation ea2a24ff-9863-478d-81c4-6342b93c7be4 · outbound

This paper cites F5-tts: A fairytaler that fakes fluent and faithful speech with flow matching,.

Multimodal Large Language Model-Enabled Video Translation: A Role-Oriented Survey F5-tts: A fairytaler that fakes fluent and faithful speech with flow matching,

Reference 16

Resolution
unresolved
no resolver link, observed 2026-07-12T22:04:31.302192Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T22:04:31.302192Z digest=sha256:62c4aab3f51ebfcdb8f9361532708791a6e1b2b0af67b5aa7973cf9dd55d2ae5

Observation 275fcc5a-7699-43c7-962b-7237ba203c8c · outbound

This paper cites A survey on video diffusion models,.

Multimodal Large Language Model-Enabled Video Translation: A Role-Oriented Survey A survey on video diffusion models,

Reference 17

Resolution
unresolved
no resolver link, observed 2026-07-12T22:04:31.302192Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T22:04:31.302192Z digest=sha256:a2404019ae281c082bf49aefb050926d0e4cd0ae9ae6f2bd11d9213648aa4844

Observation dfc92ba4-c667-4b87-a752-7c922f749208 · outbound

This paper cites Omnisync: Towards universal lip synchronization via diffusion transformers,.

Multimodal Large Language Model-Enabled Video Translation: A Role-Oriented Survey Omnisync: Towards universal lip synchronization via diffusion transformers,

Reference 18

Resolution
unresolved
no resolver link, observed 2026-07-12T22:04:31.302192Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T22:04:31.302192Z digest=sha256:0071fa3701cf2de7df36bbb81dd59b149438b4b8c69d0298883722da864831c6

Observation e706ad03-e91b-4dd9-85ae-4b5d210013d1 · outbound

This paper cites A Survey on Multi-modal Machine Translation: Tasks, Methods and Challenges.

Multimodal Large Language Model-Enabled Video Translation: A Role-Oriented Survey A Survey on Multi-modal Machine Translation: Tasks, Methods and Challenges

Reference 19

Resolution
unresolved
no resolver link, observed 2026-07-12T22:04:31.302192Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T22:04:31.302192Z digest=sha256:40c9200c3031ec646783bf29a99773faa61852d9d6a4b6a8959c9f66ebfbf711

Observation 66d72910-9195-4787-bb78-39e8c1b4dd81 · outbound

This paper cites Video-guided machine translation: A survey of models, datasets, and challenges,.

Multimodal Large Language Model-Enabled Video Translation: A Role-Oriented Survey Video-guided machine translation: A survey of models, datasets, and challenges,

Reference 20

Resolution
unresolved
no resolver link, observed 2026-07-12T22:04:31.302192Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T22:04:31.302192Z digest=sha256:2de9855d29c632bf876d0d5989d41fe81f9b8937a4f232c248f644e39cdc2cae

Observation f8f38568-bab9-4527-9df5-4febdffb687f · outbound

This paper cites A survey on multimodal large language models,.

Multimodal Large Language Model-Enabled Video Translation: A Role-Oriented Survey A survey on multimodal large language models,

Reference 21

Resolution
unresolved
no resolver link, observed 2026-07-12T22:04:31.302192Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T22:04:31.302192Z digest=sha256:e776b39d9abf250378bd90e7b6428e3a1c86693c899407d7d51019b1662ca421

Observation db826b02-7e8f-420c-a01c-f7a3abafbed3 · outbound

This paper cites Video understanding with large language models: A survey,.

Multimodal Large Language Model-Enabled Video Translation: A Role-Oriented Survey Video understanding with large language models: A survey,

Reference 22

Resolution
unresolved
no resolver link, observed 2026-07-12T22:04:31.302192Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T22:04:31.302192Z digest=sha256:b984fc46511550e6d3a59259f5ac8fa470917c01d05ab8e50e79fea9014dab14

Observation 26eccb96-1dd0-44e9-b82f-316a3a735192 · outbound

This paper cites A survey on video temporal grounding with multimodal large language model,.

Multimodal Large Language Model-Enabled Video Translation: A Role-Oriented Survey A survey on video temporal grounding with multimodal large language model,

Reference 23

Resolution
unresolved
no resolver link, observed 2026-07-12T22:04:31.302192Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T22:04:31.302192Z digest=sha256:d6713a6489632bc3303c337ef1eda24f95c19375192dc9eb57419648259a9a26

Observation cc0aed22-173e-4283-a805-bdd0a25edd93 · outbound

This paper cites Direct Speech-to-Speech Neural Machine Translation: A Survey.

Multimodal Large Language Model-Enabled Video Translation: A Role-Oriented Survey Direct Speech-to-Speech Neural Machine Translation: A Survey

Reference 24

Resolution
unresolved
no resolver link, observed 2026-07-12T22:04:31.302192Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T22:04:31.302192Z digest=sha256:2dcf8322cb66eb5b683ebee9fef44a4fe267586d05975a8ce85969ae9b7da234

Observation ce5f19c4-0f83-4a90-87fa-0e9f9e34ac11 · outbound

This paper cites Towards controllable speech synthesis in the era of large language models: A systematic survey,.

Multimodal Large Language Model-Enabled Video Translation: A Role-Oriented Survey Towards controllable speech synthesis in the era of large language models: A systematic survey,

Reference 25

Resolution
unresolved
no resolver link, observed 2026-07-12T22:04:31.302192Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T22:04:31.302192Z digest=sha256:b2540f4c3253b26579f99209c6c3d948a70a4c27907153629e6f97944d405839

Observation 5fa792ce-96a7-448d-ad97-39496589d1bd · outbound

This paper cites Advancing talking head generation: A comprehensive survey of multi-modal methodologies, datasets, evaluation metrics, and loss functions,.

Multimodal Large Language Model-Enabled Video Translation: A Role-Oriented Survey Advancing talking head generation: A comprehensive survey of multi-modal methodologies, datasets, evaluation metrics, and loss functions,

Reference 26

Resolution
unresolved
no resolver link, observed 2026-07-12T22:04:31.302192Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T22:04:31.302192Z digest=sha256:b56e4d67f9bee46f47d78cf4d2b884021999cee64d6d6c25ae9f6929edff69d1

Observation fa4cfb08-33cb-4d22-bc6b-a9b6e4760899 · outbound

This paper cites Rabiner and B.-H.

Multimodal Large Language Model-Enabled Video Translation: A Role-Oriented Survey Rabiner and B.-H

Reference 27

Resolution
unresolved
no resolver link, observed 2026-07-12T22:04:31.302192Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T22:04:31.302192Z digest=sha256:6c139ea6c75a147ae012c551329ca2ff96e8fcb534066d89ef7bc3785bbec518

Observation 63e0517a-cf0a-4063-b2f1-c93093946845 · outbound

This paper cites Connection- ist temporal classification: Labelling unsegmented sequence data with recurrent neural networks,.

Multimodal Large Language Model-Enabled Video Translation: A Role-Oriented Survey Connection- ist temporal classification: Labelling unsegmented sequence data with recurrent neural networks,

Reference 28

Resolution
unresolved
no resolver link, observed 2026-07-12T22:04:31.302192Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T22:04:31.302192Z digest=sha256:dc53a8081aff31d5f5631b55eb7c05440fb05c5213f74db7796d3c62091cb83a

Observation 80b95312-4df9-4ba7-af79-1b6243102680 · outbound

This paper cites Deep Speech: Scaling up end-to-end speech recognition.

Multimodal Large Language Model-Enabled Video Translation: A Role-Oriented Survey Deep Speech: Scaling up end-to-end speech recognition

Reference 29

Resolution
unresolved
no resolver link, observed 2026-07-12T22:04:31.302192Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T22:04:31.302192Z digest=sha256:c7f86151ef4fe4eccea8a2219da2c02afe00ec7dcb04ea07a610ee64d65e6a25

Observation fe51874d-d962-4993-9687-15d22b3cea4a · outbound

This paper cites Listen, attend and spell,.

Multimodal Large Language Model-Enabled Video Translation: A Role-Oriented Survey Listen, attend and spell,

Reference 30

Resolution
unresolved
no resolver link, observed 2026-07-12T22:04:31.302192Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T22:04:31.302192Z digest=sha256:3a6fd25bd9327f13886a4507255d634c8f0d1e3d95ce4456d1e16a7104aaf10e

Observation 82758b61-a320-4cef-9ec0-d8e837e138c9 · outbound

This paper cites The mathematics of statistical machine translation: Parameter estimation,.

Multimodal Large Language Model-Enabled Video Translation: A Role-Oriented Survey The mathematics of statistical machine translation: Parameter estimation,

Reference 31

Resolution
unresolved
no resolver link, observed 2026-07-12T22:04:31.302192Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T22:04:31.302192Z digest=sha256:052b4a66dd626838ccba555d73c5ffbaf2ad44ca3573b2733e9c4b61fe48d1db

Observation ad93a695-7de4-415e-b647-666d5f3d0846 · outbound

This paper cites Statistical phrase-based transla- tion,.

Multimodal Large Language Model-Enabled Video Translation: A Role-Oriented Survey Statistical phrase-based transla- tion,

Reference 32

Resolution
unresolved
no resolver link, observed 2026-07-12T22:04:31.302192Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T22:04:31.302192Z digest=sha256:76e1c64aa3e0bb6e1993701eb87f268df1b8819f6878edff21fbb8fd9e174b6a

Observation 9cfd52df-9c8c-4911-b3e0-31fadbe1e67f · outbound

This paper cites Sequence to sequence learning with neural networks,.

Multimodal Large Language Model-Enabled Video Translation: A Role-Oriented Survey Sequence to sequence learning with neural networks,

Reference 33

Resolution
unresolved
no resolver link, observed 2026-07-12T22:04:31.302192Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T22:04:31.302192Z digest=sha256:61a6cdc42dc88d1bd1d3bce52eb0ad2da00a24f53871783c314f1350bc5ff8ed

Observation a8c4693f-be0c-4201-87c0-7809eeb7524e · outbound

This paper cites Speech parameter generation algorithms for hmm-based speech syn- thesis,.

Multimodal Large Language Model-Enabled Video Translation: A Role-Oriented Survey Speech parameter generation algorithms for hmm-based speech syn- thesis,

Reference 34

Resolution
unresolved
no resolver link, observed 2026-07-12T22:04:31.302192Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T22:04:31.302192Z digest=sha256:0456534aa3bfb110f9a1c56bc2f00ef65469d0bef7eb3b09c1d56ef67670c9f9

Observation 8b1a67ab-e196-4150-a9e6-3b888ee3815b · outbound

This paper cites Statistical parametric speech synthesis,.

Multimodal Large Language Model-Enabled Video Translation: A Role-Oriented Survey Statistical parametric speech synthesis,

Reference 35

Resolution
unresolved
no resolver link, observed 2026-07-12T22:04:31.302192Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T22:04:31.302192Z digest=sha256:38e44a9547a691d7e5fe8f38d5cf14cb7ddd2bcb63046056c996728fd97d0955

Observation d11733ce-2223-4cbb-9f7d-d36caa5c0f0b · outbound

This paper cites WaveNet: A Generative Model for Raw Audio.

Multimodal Large Language Model-Enabled Video Translation: A Role-Oriented Survey WaveNet: A Generative Model for Raw Audio

Reference 36

Resolution
unresolved
no resolver link, observed 2026-07-12T22:04:31.302192Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T22:04:31.302192Z digest=sha256:fc4fcab810b37b72014fb1dddc65b31c6bb4741190c7f6a79553520da18f3729

Observation d2be9e07-127d-46ac-a2b5-6052be1cb955 · outbound

This paper cites Photo-realistic talking-heads from image samples,.

Multimodal Large Language Model-Enabled Video Translation: A Role-Oriented Survey Photo-realistic talking-heads from image samples,

Reference 37

Resolution
unresolved
no resolver link, observed 2026-07-12T22:04:31.302192Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T22:04:31.302192Z digest=sha256:a33e9d2a8f29da6ac990aea7611cf2d73aebb78aa9c145d59d856951e2471c3c

Observation 9379d2f9-a24f-47dd-af73-53137e5254d2 · outbound

This paper cites Lip reading sentences in the wild,.

Multimodal Large Language Model-Enabled Video Translation: A Role-Oriented Survey Lip reading sentences in the wild,

Reference 38

Resolution
unresolved
no resolver link, observed 2026-07-12T22:04:31.302192Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T22:04:31.302192Z digest=sha256:2684b0df9e1b47f7126149dc18a968e3abb6e6b60c5fc88db8f7bd32145142e9

Observation b3f0000a-9e5b-480a-8ecc-bb5ad6a9b90a · outbound

This paper cites Few-shot adversarial learning of realistic neural talking head models,.

Multimodal Large Language Model-Enabled Video Translation: A Role-Oriented Survey Few-shot adversarial learning of realistic neural talking head models,

Reference 39

Resolution
unresolved
no resolver link, observed 2026-07-12T22:04:31.302192Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T22:04:31.302192Z digest=sha256:56a5aa591cc18b6c2938450e02632274aece92ba93152814d5dfc152b8ac64fd

Observation 5ce64f5f-d3fa-49fc-92df-656f8c3a125e · outbound

This paper cites Onellm: One framework to align all modalities with language,.

Multimodal Large Language Model-Enabled Video Translation: A Role-Oriented Survey Onellm: One framework to align all modalities with language,

Reference 40

Resolution
unresolved
no resolver link, observed 2026-07-12T22:04:31.302192Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T22:04:31.302192Z digest=sha256:7f1b53c13cc20210c511fa962d5667b012b57f9fcdb145ded676ee64a834f843

Observation 7d73ef55-7391-4aea-bce7-041f6685bbf2 · outbound

This paper cites An image is worth 16x16 words: Transformers for image recognition at scale,.

Multimodal Large Language Model-Enabled Video Translation: A Role-Oriented Survey An image is worth 16x16 words: Transformers for image recognition at scale,

Reference 41

Resolution
unresolved
no resolver link, observed 2026-07-12T22:04:31.302192Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T22:04:31.302192Z digest=sha256:e890f390547ae56621e50f3f6cfe0abc6c218ad7c5f93ba70e5edc7a912f46ea

Observation daf14975-2cad-485a-93ad-6bf2230210c0 · outbound

This paper cites Scaling up visual and vision-language representation learning with noisy text supervision,.

Multimodal Large Language Model-Enabled Video Translation: A Role-Oriented Survey Scaling up visual and vision-language representation learning with noisy text supervision,

Reference 42

Resolution
unresolved
no resolver link, observed 2026-07-12T22:04:31.302192Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T22:04:31.302192Z digest=sha256:171c7cedde07cf9136aecf46c41dc05f50d93d03f444078424da9aa9a08b825b

Observation a8ff4e49-2cc0-47ea-9c56-24f88d298d42 · outbound

This paper cites Videobert: A joint model for video and language representation learning,.

Multimodal Large Language Model-Enabled Video Translation: A Role-Oriented Survey Videobert: A joint model for video and language representation learning,

Reference 43

Resolution
unresolved
no resolver link, observed 2026-07-12T22:04:31.302192Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T22:04:31.302192Z digest=sha256:562a7caebee83b958d50d6543d5ee2f59ba36d207b787d3cfefd84150257cad6

Observation b35132a1-d16d-4fe1-91de-d9e115b87974 · outbound

This paper cites Actbert: Learning global-local video-text repre- sentations,.

Multimodal Large Language Model-Enabled Video Translation: A Role-Oriented Survey Actbert: Learning global-local video-text repre- sentations,

Reference 44

Resolution
unresolved
no resolver link, observed 2026-07-12T22:04:31.302192Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T22:04:31.302192Z digest=sha256:d9490e101d3d961cf297a02f356d2ea36aecf80709ede2c1f217def30d3e158f

Observation c09d45f8-535c-453f-b463-64b97345b01b · outbound

This paper cites Merlot: Multimodal neural script knowledge models,.

Multimodal Large Language Model-Enabled Video Translation: A Role-Oriented Survey Merlot: Multimodal neural script knowledge models,

Reference 45

Resolution
unresolved
no resolver link, observed 2026-07-12T22:04:31.302192Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T22:04:31.302192Z digest=sha256:d17b5851842b189c9a5fa6e85673d4e07ef589413a89a924509b2e99c8d7abd3

Observation 483d256a-4ef9-4304-af45-e3e508f8d609 · outbound

This paper cites Frozen in time: A joint video and image encoder for end-to-end retrieval,.

Multimodal Large Language Model-Enabled Video Translation: A Role-Oriented Survey Frozen in time: A joint video and image encoder for end-to-end retrieval,

Reference 46

Resolution
unresolved
no resolver link, observed 2026-07-12T22:04:31.302192Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T22:04:31.302192Z digest=sha256:570bbf0b2e2b3e71e9e9fd4a7ed60ca0ecfb6cefe6002b82823d306ca94d9f43

Observation 2dfdf966-6e4c-43ee-aaf5-0669857d0008 · outbound

This paper cites Violet: End-to-end video-language transformers with masked visual- token modeling,.

Multimodal Large Language Model-Enabled Video Translation: A Role-Oriented Survey Violet: End-to-end video-language transformers with masked visual- token modeling,

Reference 47

Resolution
unresolved
no resolver link, observed 2026-07-12T22:04:31.302192Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T22:04:31.302192Z digest=sha256:1f302fdc8ebeb43fe84aca3c2e62aa16f935ae1bbc486930eb30f53c5a80578e

Observation fa371314-5ff1-4d14-aaf2-f7878025e093 · outbound

This paper cites Omnivl: One foundation model for image- language and video-language tasks,.

Multimodal Large Language Model-Enabled Video Translation: A Role-Oriented Survey Omnivl: One foundation model for image- language and video-language tasks,

Reference 48

Resolution
unresolved
no resolver link, observed 2026-07-12T22:04:31.302192Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T22:04:31.302192Z digest=sha256:171484c720913c2e798d443d3c2cd14c68feac4a92491ecb770742f9c100c60e

Observation 80c999f0-f809-4892-80c9-05670df423ee · outbound

This paper cites InternVideo: General Video Foundation Models via Generative and Discriminative Learning.

Multimodal Large Language Model-Enabled Video Translation: A Role-Oriented Survey InternVideo: General Video Foundation Models via Generative and Discriminative Learning

Reference 49

Resolution
unresolved
no resolver link, observed 2026-07-12T22:04:31.302192Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T22:04:31.302192Z digest=sha256:b0814c2616ddd9ff96302e614a7c3e04bf84837cf7c2faff28f7aa1d7680dddf

Observation 4507d018-e051-4a22-b52d-49fbc6b6d26b · outbound

This paper cites Unival: Unified model for image, video, audio and language tasks,.

Multimodal Large Language Model-Enabled Video Translation: A Role-Oriented Survey Unival: Unified model for image, video, audio and language tasks,

Reference 50

Resolution
unresolved
no resolver link, observed 2026-07-12T22:04:31.302192Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T22:04:31.302192Z digest=sha256:3ee908d6a29442d56c1982b6f519fa1bb3871de77d5f02b0e7fb981ac9aaaa39

Observation 6004dae6-7026-4b8e-bafc-53a5143f010a · outbound

This paper cites Unified-io 2: Scaling autoregressive multimodal models with vision language audio and action,.

Multimodal Large Language Model-Enabled Video Translation: A Role-Oriented Survey Unified-io 2: Scaling autoregressive multimodal models with vision language audio and action,

Reference 51

Resolution
unresolved
no resolver link, observed 2026-07-12T22:04:31.302192Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T22:04:31.302192Z digest=sha256:84d93fadad943f770122319e79645f44387f8d2117ebad6f4dc16c7e2083411c

Observation 2fa3e50f-30de-4fa4-980a-fbbf94b42c88 · outbound

This paper cites PaLI-X: On Scaling up a Multilingual Vision and Language Model.

Multimodal Large Language Model-Enabled Video Translation: A Role-Oriented Survey PaLI-X: On Scaling up a Multilingual Vision and Language Model

Reference 52

Resolution
unresolved
no resolver link, observed 2026-07-12T22:04:31.302192Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T22:04:31.302192Z digest=sha256:49580eb1cc021684a4ce2b4cf86b1dc747d10e0ca206a3de6f03c0736274cace

Observation 33b91d41-a5c1-4b18-9ac9-70901ebbfe73 · outbound

This paper cites Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond.

Multimodal Large Language Model-Enabled Video Translation: A Role-Oriented Survey Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond

Reference 53

Resolution
unresolved
no resolver link, observed 2026-07-12T22:04:31.302192Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T22:04:31.302192Z digest=sha256:0ef15f3c24167870f1d7ffd37a78a459fefcd1cd67ac0d74eeac771a44562095

Observation 5b60ed54-80be-45aa-82d0-3f1d280504ee · outbound

This paper cites VideoLLaMA 3: Frontier Multimodal Foundation Models for Image and Video Understanding.

Multimodal Large Language Model-Enabled Video Translation: A Role-Oriented Survey VideoLLaMA 3: Frontier Multimodal Foundation Models for Image and Video Understanding

Reference 54

Resolution
unresolved
no resolver link, observed 2026-07-12T22:04:31.302192Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T22:04:31.302192Z digest=sha256:ba627aafae884008d5195679cc1c31cd51e701df1e9b55039d34b9240c5993bc

Observation 50f1f047-b84d-4a46-95c5-dbe1724e49d1 · outbound

This paper cites VideoChat: Chat-Centric Video Understanding.

Multimodal Large Language Model-Enabled Video Translation: A Role-Oriented Survey VideoChat: Chat-Centric Video Understanding

Reference 55

Resolution
unresolved
no resolver link, observed 2026-07-12T22:04:31.302192Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T22:04:31.302192Z digest=sha256:50fc0f3818c8b1a19df5c9ae084d69b8231b1fe3000708ba5d1eefb4d9c89ec2

Observation 6d2e3842-bbb3-4e92-84ad-800e932d1c89 · outbound

This paper cites Valley: Video Assistant with Large Language model Enhanced abilitY.

Multimodal Large Language Model-Enabled Video Translation: A Role-Oriented Survey Valley: Video Assistant with Large Language model Enhanced abilitY

Reference 56

Resolution
unresolved
no resolver link, observed 2026-07-12T22:04:31.302192Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T22:04:31.302192Z digest=sha256:e2eacea39478471fb5bf47251b31f3668641a2bfacf4fc57d0ce8b8c3bfbea65

Observation 5d22707e-a9a9-45ad-be56-2c7f62545229 · outbound

This paper cites Mvbench: A comprehensive multi- modal video understanding benchmark,.

Multimodal Large Language Model-Enabled Video Translation: A Role-Oriented Survey Mvbench: A comprehensive multi- modal video understanding benchmark,

Reference 57

Resolution
unresolved
no resolver link, observed 2026-07-12T22:04:31.302192Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T22:04:31.302192Z digest=sha256:922bb5ff5e6d06afebdd7fa3abae676857751cacba019feb5c5f7c638e5c8a80

Observation be48bc47-fefc-4306-8f73-aa1e24715988 · outbound

This paper cites Groundinggpt: Language enhanced multi-modal grounding model,.

Multimodal Large Language Model-Enabled Video Translation: A Role-Oriented Survey Groundinggpt: Language enhanced multi-modal grounding model,

Reference 58

Resolution
unresolved
no resolver link, observed 2026-07-12T22:04:31.302192Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T22:04:31.302192Z digest=sha256:0fa08446a0277ed4eae5a6b22c171e7dc7bd9e06c74c327f1a13108a7a275d35

Observation a5ca39dc-8743-4b51-afb6-6072a4cf7452 · outbound

This paper cites Moviechat: From dense token to sparse memory for long video understanding,.

Multimodal Large Language Model-Enabled Video Translation: A Role-Oriented Survey Moviechat: From dense token to sparse memory for long video understanding,

Reference 59

Resolution
unresolved
no resolver link, observed 2026-07-12T22:04:31.302192Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T22:04:31.302192Z digest=sha256:cd0bd1a1030bfc437357d967bbd6fd6b9ee278414ac5feec2c523d165c82fce8

Observation 5dde9bc3-9171-4c61-bc3a-f7ffe8c86a11 · outbound

This paper cites DreamFrame: Enhancing Video Understanding via Automatically Generated QA and Style-Consistent Keyframes.

Multimodal Large Language Model-Enabled Video Translation: A Role-Oriented Survey DreamFrame: Enhancing Video Understanding via Automatically Generated QA and Style-Consistent Keyframes

Reference 60

Resolution
unresolved
no resolver link, observed 2026-07-12T22:04:31.302192Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T22:04:31.302192Z digest=sha256:8a21c88b75fc6328d672e35edce0ba861b3fa3ff9f75316aa3d6063e632a9c35

Observation 54af9688-e32a-4b0d-83c8-0854bf176e86 · outbound

This paper cites Longvlm: Efficient long video understanding via large language models,.

Multimodal Large Language Model-Enabled Video Translation: A Role-Oriented Survey Longvlm: Efficient long video understanding via large language models,

Reference 61

Resolution
unresolved
no resolver link, observed 2026-07-12T22:04:31.302192Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T22:04:31.302192Z digest=sha256:3b0337e79d248df00628cc19d4187da22dd8b07f54c49f1fefbd69cafb222c4c

Observation 98f16a7d-ef0c-4731-9d87-6bfb2f9429cf · outbound

This paper cites Timechat: A time-sensitive multimodal large language model for long video understanding,.

Multimodal Large Language Model-Enabled Video Translation: A Role-Oriented Survey Timechat: A time-sensitive multimodal large language model for long video understanding,

Reference 62

Resolution
unresolved
no resolver link, observed 2026-07-12T22:04:31.302192Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T22:04:31.302192Z digest=sha256:8f36440de8567d0a78a7e6abfc161565faf048a90a06c0104d95167e49a68830

Observation 7b7076a6-220a-4457-b723-32c353d2f7c0 · outbound

This paper cites Vtimellm: Empower llm to grasp video moments,.

Multimodal Large Language Model-Enabled Video Translation: A Role-Oriented Survey Vtimellm: Empower llm to grasp video moments,

Reference 63

Resolution
unresolved
no resolver link, observed 2026-07-12T22:04:31.302192Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T22:04:31.302192Z digest=sha256:45de3564f39472e319c523601206f7e2e8333e50e79d85828e39dbc0e5492407

Observation b14dd2fa-b7d8-4e7a-afc2-dc5e2c459a8d · outbound

This paper cites Momentor: Advancing Video Large Language Model with Fine-Grained Temporal Reasoning.

Multimodal Large Language Model-Enabled Video Translation: A Role-Oriented Survey Momentor: Advancing Video Large Language Model with Fine-Grained Temporal Reasoning

Reference 64

Resolution
unresolved
no resolver link, observed 2026-07-12T22:04:31.302192Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T22:04:31.302192Z digest=sha256:b22ec35b478f44ab46f2e35aee4c5b30835f9bd6b2a583e6668b87790fec8cc7

Observation 43c77e3f-1d5a-4a5a-b119-3c912e7a53c2 · outbound

This paper cites Videollm-online: Online video large language model for streaming video,.

Multimodal Large Language Model-Enabled Video Translation: A Role-Oriented Survey Videollm-online: Online video large language model for streaming video,

Reference 65

Resolution
unresolved
no resolver link, observed 2026-07-12T22:04:31.302192Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T22:04:31.302192Z digest=sha256:aea6e8692ef08bfa571c791cb923858d240b97930e7c0cb449062b514f7186cd

Observation 1b1bcd79-5687-4d88-a1fc-899c71622054 · outbound

This paper cites Lita: Language instructed temporal-localization assistant,.

Multimodal Large Language Model-Enabled Video Translation: A Role-Oriented Survey Lita: Language instructed temporal-localization assistant,

Reference 66

Resolution
unresolved
no resolver link, observed 2026-07-12T22:04:31.302192Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T22:04:31.302192Z digest=sha256:cc6da5e7aa1e7376e3702701d96210f90c7b84f002d1a3109bce39d244fdd0ef

Observation 5398d75c-2f60-4d21-8592-3986f2665dbd · outbound

This paper cites Vtg-llm: Integrating timestamp knowledge into video llms for enhanced video temporal grounding,.

Multimodal Large Language Model-Enabled Video Translation: A Role-Oriented Survey Vtg-llm: Integrating timestamp knowledge into video llms for enhanced video temporal grounding,

Reference 67

Resolution
unresolved
no resolver link, observed 2026-07-12T22:04:31.302192Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T22:04:31.302192Z digest=sha256:bf429c11fdfc81a0f308bfde11e03676fd21a0f9a1d4a9804e06fb75b724bfef

Observation 2ffa213a-6bef-424f-935e-45aa0041ff81 · outbound

This paper cites Video-xl: Extra-long vision language model for hour-scale video understanding,.

Multimodal Large Language Model-Enabled Video Translation: A Role-Oriented Survey Video-xl: Extra-long vision language model for hour-scale video understanding,

Reference 68

Resolution
unresolved
no resolver link, observed 2026-07-12T22:04:31.302192Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T22:04:31.302192Z digest=sha256:3616858b9abe29b895fc1fe5ce7750b4eaaa80d1b59ad576feabbe9a6e46dc66

Observation 5e8ebaf9-3410-4ae5-bd62-0a945189ece1 · outbound

This paper cites Time-R1: Post-Training Large Vision Language Model for Temporal Video Grounding.

Multimodal Large Language Model-Enabled Video Translation: A Role-Oriented Survey Time-R1: Post-Training Large Vision Language Model for Temporal Video Grounding

Reference 69

Resolution
unresolved
no resolver link, observed 2026-07-12T22:04:31.302192Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T22:04:31.302192Z digest=sha256:da65c1ddef3f72216f08b247a2559367318923ede50acc16c434f369a394c8c9

Observation be03c237-266f-43f1-9f75-ee4376385482 · outbound

This paper cites SeamlessM4T: Massively Multilingual & Multimodal Machine Translation.

Multimodal Large Language Model-Enabled Video Translation: A Role-Oriented Survey SeamlessM4T: Massively Multilingual & Multimodal Machine Translation

Reference 70

Resolution
unresolved
no resolver link, observed 2026-07-12T22:04:31.302192Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T22:04:31.302192Z digest=sha256:f84be8f8544e953cd5601b74de68d472638a90e09b9ae96cf2fbef02e489b370

Observation 37ef2916-36f7-4037-9372-f0039b9b1239 · outbound

This paper cites Av2av: Direct audio- visual speech to audio-visual speech translation with unified audio- visual speech representation,.

Multimodal Large Language Model-Enabled Video Translation: A Role-Oriented Survey Av2av: Direct audio- visual speech to audio-visual speech translation with unified audio- visual speech representation,

Reference 71

Resolution
unresolved
no resolver link, observed 2026-07-12T22:04:31.302192Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T22:04:31.302192Z digest=sha256:909253f05dcd82feb42e2eaa69dd6a66f3005ac83f8117188b329f07e9607375

Observation 9d4d6e7a-9647-421a-9d92-bd5f93a675bf · outbound

This paper cites Adaptive inner speech-text alignment for llm-based speech trans- lation,.

Multimodal Large Language Model-Enabled Video Translation: A Role-Oriented Survey Adaptive inner speech-text alignment for llm-based speech trans- lation,

Reference 72

Resolution
unresolved
no resolver link, observed 2026-07-12T22:04:31.302192Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T22:04:31.302192Z digest=sha256:949a7305132936674b310629de38e2caecd9f9f0dd856e87124cae71c98cd17a

Observation a8f12530-b7c1-41d5-92f2-7e5223a631f3 · outbound

This paper cites Inimagetrans: Multimodal llm-based text image machine translation,.

Multimodal Large Language Model-Enabled Video Translation: A Role-Oriented Survey Inimagetrans: Multimodal llm-based text image machine translation,

Reference 73

Resolution
unresolved
no resolver link, observed 2026-07-12T22:04:31.302192Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T22:04:31.302192Z digest=sha256:6aa83b4f62dab57b5cc40bbdb5799f0a447668430bed3a449667b011622030b0

Observation 60853e80-729e-4c6a-8be1-e44a06e385f5 · outbound

This paper cites Llama-adapter: Efficient fine-tuning of language models with zero-init attention,.

Multimodal Large Language Model-Enabled Video Translation: A Role-Oriented Survey Llama-adapter: Efficient fine-tuning of language models with zero-init attention,

Reference 74

Resolution
unresolved
no resolver link, observed 2026-07-12T22:04:31.302192Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T22:04:31.302192Z digest=sha256:700537c6a24b9fac9d48c40c8f4750975d872816b68a813eb568f2cc7211df03

Observation 2f6a9045-e26e-4717-b525-43abcce74042 · outbound

This paper cites Bt-adapter: Video conversation is feasible without video instruction tuning,.

Multimodal Large Language Model-Enabled Video Translation: A Role-Oriented Survey Bt-adapter: Video conversation is feasible without video instruction tuning,

Reference 75

Resolution
unresolved
no resolver link, observed 2026-07-12T22:04:31.302192Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T22:04:31.302192Z digest=sha256:884c7338af571ccdeae7f1a5294913be374f2a16216e12ef31b1731ae94cf7f4

Observation aa46b280-4fb8-47fe-b285-5ffe9559d136 · outbound

This paper cites Otter: A multi-modal model with in-context instruction tuning,.

Multimodal Large Language Model-Enabled Video Translation: A Role-Oriented Survey Otter: A multi-modal model with in-context instruction tuning,

Reference 76

Resolution
unresolved
no resolver link, observed 2026-07-12T22:04:31.302192Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T22:04:31.302192Z digest=sha256:cb6ff031559cfe247ce64f37e7b5ac8cfe2c7cdb232f8c763bf86a167b8eefa9

Observation 5f3ad4f5-f0cb-481c-a1e4-4ea5f04d58bc · outbound

This paper cites PLLaVA : Parameter-free LLaVA Extension from Images to Videos for Video Dense Captioning.

Multimodal Large Language Model-Enabled Video Translation: A Role-Oriented Survey PLLaVA : Parameter-free LLaVA Extension from Images to Videos for Video Dense Captioning

Reference 77

Resolution
unresolved
no resolver link, observed 2026-07-12T22:04:31.302192Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T22:04:31.302192Z digest=sha256:9b123084f63f775b874f46575e155a669e61b408b52b978347a155f5a3cbeabb

Observation 7dd0a2ad-7068-4bb3-a3dc-a6079849cc68 · outbound

This paper cites From Image to Video, what do we need in multimodal LLMs?.

Multimodal Large Language Model-Enabled Video Translation: A Role-Oriented Survey From Image to Video, what do we need in multimodal LLMs?

Reference 78

Resolution
unresolved
no resolver link, observed 2026-07-12T22:04:31.302192Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T22:04:31.302192Z digest=sha256:35c38f651fd51354eade3e8fb1f8838e553c554c502e7eb6181fb9d07b6bc06b

Observation 82dee9e6-900f-4f14-84ba-277bccc87674 · outbound

This paper cites Reef: Relevance-aware and efficient llm adapter for video understanding,.

Multimodal Large Language Model-Enabled Video Translation: A Role-Oriented Survey Reef: Relevance-aware and efficient llm adapter for video understanding,

Reference 79

Resolution
unresolved
no resolver link, observed 2026-07-12T22:04:31.302192Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T22:04:31.302192Z digest=sha256:741c9621317ea6a98410cfc1ec68bbf3f9bf9e299e08806c611432c0490cd991

Observation d0072e5b-df64-4ba2-be21-c4407ea5164a · outbound

This paper cites Vlog: Video-language models by generative retrieval of narration vocabulary,.

Multimodal Large Language Model-Enabled Video Translation: A Role-Oriented Survey Vlog: Video-language models by generative retrieval of narration vocabulary,

Reference 80

Resolution
unresolved
no resolver link, observed 2026-07-12T22:04:31.302192Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T22:04:31.302192Z digest=sha256:207a932f82fd189e00bd13763db0fe69454c7cac6769e1c495ac2bbab1e79370

Observation a4d8b360-d86c-4a6e-b5a3-ad3d28189a0b · outbound

This paper cites Vall-e r: Robust and efficient zero-shot text-to-speech synthesis via monotonic alignment,.

Multimodal Large Language Model-Enabled Video Translation: A Role-Oriented Survey Vall-e r: Robust and efficient zero-shot text-to-speech synthesis via monotonic alignment,

Reference 81

Resolution
unresolved
no resolver link, observed 2026-07-12T22:04:31.302192Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T22:04:31.302192Z digest=sha256:45d122804922afe1bd8073ace1dad1cc4ab99d2dbde0ecbc43844df07e95f9c4

Observation 66f06165-fdfa-496c-8376-a8fe887675b3 · outbound

This paper cites CosyVoice: A Scalable Multilingual Zero-shot Text-to-speech Synthesizer based on Supervised Semantic Tokens.

Multimodal Large Language Model-Enabled Video Translation: A Role-Oriented Survey CosyVoice: A Scalable Multilingual Zero-shot Text-to-speech Synthesizer based on Supervised Semantic Tokens

Reference 82

Resolution
unresolved
no resolver link, observed 2026-07-12T22:04:31.302192Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T22:04:31.302192Z digest=sha256:07141449624abc6bc9c700d38c2e42e1e0355f5318e046ff7d02093c1dda6126

Observation f348cb99-400d-4a35-b18c-dc165d39c950 · outbound

This paper cites Fish-Speech: Leveraging Large Language Models for Advanced Multilingual Text-to-Speech Synthesis.

Multimodal Large Language Model-Enabled Video Translation: A Role-Oriented Survey Fish-Speech: Leveraging Large Language Models for Advanced Multilingual Text-to-Speech Synthesis

Reference 83

Resolution
unresolved
no resolver link, observed 2026-07-12T22:04:31.302192Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T22:04:31.302192Z digest=sha256:cc359a971434381a8c03ff00d2ae182dd1309a3c136409b59c2d95239bbbe294

Observation 148aecbb-bd2a-4eb1-b797-1cb8b87f8337 · outbound

This paper cites HALL-E: Hierarchical Neural Codec Language Model for Minute-Long Zero-Shot Text-to-Speech Synthesis.

Multimodal Large Language Model-Enabled Video Translation: A Role-Oriented Survey HALL-E: Hierarchical Neural Codec Language Model for Minute-Long Zero-Shot Text-to-Speech Synthesis

Reference 84

Resolution
unresolved
no resolver link, observed 2026-07-12T22:04:31.302192Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T22:04:31.302192Z digest=sha256:997d6c1bb4e6bf4ce8e70bb64eaf8b2aea962f5e3388013aa65296a583a92c74

Observation aa24220d-3b3d-4b45-8d3f-e9dc9e36fabd · outbound

This paper cites V oxinstruct: Expressive human instruction-to-speech generation with unified multilingual codec language modelling,.

Multimodal Large Language Model-Enabled Video Translation: A Role-Oriented Survey V oxinstruct: Expressive human instruction-to-speech generation with unified multilingual codec language modelling,

Reference 85

Resolution
unresolved
no resolver link, observed 2026-07-12T22:04:31.302192Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T22:04:31.302192Z digest=sha256:e484461ee970e631cf3f3eb476d1e5f7d96af246882b5d28067f6461fb1a8bd2

Observation 5bb3292c-5288-40fd-be9f-aa3f4257990a · outbound

This paper cites Spark-TTS: An Efficient LLM-Based Text-to-Speech Model with Single-Stream Decoupled Speech Tokens.

Multimodal Large Language Model-Enabled Video Translation: A Role-Oriented Survey Spark-TTS: An Efficient LLM-Based Text-to-Speech Model with Single-Stream Decoupled Speech Tokens

Reference 86

Resolution
unresolved
no resolver link, observed 2026-07-12T22:04:31.302192Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T22:04:31.302192Z digest=sha256:43c0002a72c7bbf1c729a82fed1a062d6918a04f951f92786ab3b3097475e271

Observation 03fc689e-c71d-42b2-a81e-7ef85ef8b1af · outbound

This paper cites Mega-tts 2: Boosting prompting mechanisms for zero-shot speech synthesis,.

Multimodal Large Language Model-Enabled Video Translation: A Role-Oriented Survey Mega-tts 2: Boosting prompting mechanisms for zero-shot speech synthesis,

Reference 87

Resolution
unresolved
no resolver link, observed 2026-07-12T22:04:31.302192Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T22:04:31.302192Z digest=sha256:91b6480162b0a40302f7affdd50b7cb26f0df2c56a467449dee899b2b383c53b

Observation a4c39120-64ca-43b2-befd-ce6d88be8442 · outbound

This paper cites MegaTTS 3: Sparse Alignment Enhanced Latent Diffusion Transformer for Zero-Shot Speech Synthesis.

Multimodal Large Language Model-Enabled Video Translation: A Role-Oriented Survey MegaTTS 3: Sparse Alignment Enhanced Latent Diffusion Transformer for Zero-Shot Speech Synthesis

Reference 88

Resolution
unresolved
no resolver link, observed 2026-07-12T22:04:31.302192Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T22:04:31.302192Z digest=sha256:a8de90af97b98686ea6835317d809534ace920419e2b3e24a2e5b711dc87957a

Observation a6a1f961-7ef8-4cc2-ac74-94e3e8085089 · outbound

This paper cites Seed-TTS: A Family of High-Quality Versatile Speech Generation Models.

Multimodal Large Language Model-Enabled Video Translation: A Role-Oriented Survey Seed-TTS: A Family of High-Quality Versatile Speech Generation Models

Reference 89

Resolution
unresolved
no resolver link, observed 2026-07-12T22:04:31.302192Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T22:04:31.302192Z digest=sha256:b6b83fb39283dea3793b10135e54890c85bcbde53a9c17d1e25d040b97119f41

Observation a850a614-1c5e-462a-93bb-3ff0952954d1 · outbound

This paper cites PromptTTS 2: Describing and Generating Voices with Text Prompt.

Multimodal Large Language Model-Enabled Video Translation: A Role-Oriented Survey PromptTTS 2: Describing and Generating Voices with Text Prompt

Reference 91

Resolution
unresolved
no resolver link, observed 2026-07-12T22:04:31.302192Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T22:04:31.302192Z digest=sha256:61090b2cb0bc9c8db0c6029656a75b946bd631349980327f3b9822f546bde7d6

Observation d7dca130-2869-4c72-9d71-5b43ab92708f · outbound

This paper cites InstructTTS: Modelling Expressive TTS in Discrete Latent Space with Natural Language Style Prompt.

Multimodal Large Language Model-Enabled Video Translation: A Role-Oriented Survey InstructTTS: Modelling Expressive TTS in Discrete Latent Space with Natural Language Style Prompt

Reference 92

Resolution
unresolved
no resolver link, observed 2026-07-12T22:04:31.302192Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T22:04:31.302192Z digest=sha256:4ced102178cf4b4220d4a753069a87abedf80f4c3c1c1f7e1840fa4d3a8ff809

Observation 06085be6-89e8-469c-8f4a-485aece58c1b · outbound

This paper cites Styletts 2: Towards human-level text-to-speech through style diffu- sion and adversarial training with large speech language models,.

Multimodal Large Language Model-Enabled Video Translation: A Role-Oriented Survey Styletts 2: Towards human-level text-to-speech through style diffu- sion and adversarial training with large speech language models,

Reference 93

Resolution
unresolved
no resolver link, observed 2026-07-12T22:04:31.302192Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T22:04:31.302192Z digest=sha256:53852fd995a0fdc0de88c73faa0a13a22dc5996f6796b3b68f1c835ab83152ef

Observation 008ddd17-1a0e-4145-bed2-d3466210c9d4 · outbound

This paper cites Controlspeech: Towards simultaneous and independent zero- shot speaker cloning and zero-shot language style control,.

Multimodal Large Language Model-Enabled Video Translation: A Role-Oriented Survey Controlspeech: Towards simultaneous and independent zero- shot speaker cloning and zero-shot language style control,

Reference 94

Resolution
unresolved
no resolver link, observed 2026-07-12T22:04:31.302192Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T22:04:31.302192Z digest=sha256:15133803581086bdbcf69c63c81a0cc632efdffc6063cc212ec7a8df6c02a7b6

Observation de70324d-a926-4f93-ab8b-cc9bd948aa7b · outbound

This paper cites SC VALL-E: Style-Controllable Zero-Shot Text to Speech Synthesizer.

Multimodal Large Language Model-Enabled Video Translation: A Role-Oriented Survey SC VALL-E: Style-Controllable Zero-Shot Text to Speech Synthesizer

Reference 95

Resolution
unresolved
no resolver link, observed 2026-07-12T22:04:31.302192Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T22:04:31.302192Z digest=sha256:4ba26c87928bb20d4a3d91ff1709cc94d879dd940f7c1fe54a68ab151c73016b

Observation 1674fc9b-f029-4588-b108-cfe923adc7ce · outbound

This paper cites E2 TTS: Embarrassingly Easy Fully Non-Autoregressive Zero-Shot TTS.

Multimodal Large Language Model-Enabled Video Translation: A Role-Oriented Survey E2 TTS: Embarrassingly Easy Fully Non-Autoregressive Zero-Shot TTS

Reference 96

Resolution
unresolved
no resolver link, observed 2026-07-12T22:04:31.302192Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T22:04:31.302192Z digest=sha256:9821b15f0725a3ada9c12ae46031a32df8b2c0e3c90579fa584eb48015c659b6

Observation c3e6503c-fabf-4965-ab72-5abcbb1424b8 · outbound

This paper cites V oicebox: Text- guided multilingual universal speech generation at scale,.

Multimodal Large Language Model-Enabled Video Translation: A Role-Oriented Survey V oicebox: Text- guided multilingual universal speech generation at scale,

Reference 97

Resolution
unresolved
no resolver link, observed 2026-07-12T22:04:31.302192Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T22:04:31.302192Z digest=sha256:22a9aad915eeb1a3cfbe7bd326755bf5247ad52a566231649f5338b1a7c26001

Observation 46c4a353-49c8-4094-9e8c-fc49dce02c09 · outbound

This paper cites NaturalSpeech 2: Latent Diffusion Models are Natural and Zero-Shot Speech and Singing Synthesizers.

Multimodal Large Language Model-Enabled Video Translation: A Role-Oriented Survey NaturalSpeech 2: Latent Diffusion Models are Natural and Zero-Shot Speech and Singing Synthesizers

Reference 98

Resolution
unresolved
no resolver link, observed 2026-07-12T22:04:31.302192Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T22:04:31.302192Z digest=sha256:749f6cbb13b7b7264277ab89177fc1048ab0b5b829a2fb027a8dc1963328a895

Observation 20d62fa4-d936-4e00-9ab3-b8291178187d · outbound

This paper cites NaturalSpeech 3: Zero-Shot Speech Synthesis with Factorized Codec and Diffusion Models.

Multimodal Large Language Model-Enabled Video Translation: A Role-Oriented Survey NaturalSpeech 3: Zero-Shot Speech Synthesis with Factorized Codec and Diffusion Models

Reference 99

Resolution
unresolved
no resolver link, observed 2026-07-12T22:04:31.302192Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T22:04:31.302192Z digest=sha256:f91cd245795cb26f4bc7ca499dd6b01d21efc9b0af581d0df5e35b5ad3a07a1f

Observation 0e3aaad8-48ce-4b1e-a5c5-1a1eb2beeeec · outbound

This paper cites DiTTo-TTS: Diffusion Transformers for Scalable Text-to-Speech without Domain-Specific Factors.

Multimodal Large Language Model-Enabled Video Translation: A Role-Oriented Survey DiTTo-TTS: Diffusion Transformers for Scalable Text-to-Speech without Domain-Specific Factors

Reference 100

Resolution
unresolved
no resolver link, observed 2026-07-12T22:04:31.302192Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T22:04:31.302192Z digest=sha256:f398986af0e4c9b7c1bd657699f41abfd98c9e627999a88384346a1991507ee5

Observation fae34acb-80fa-4cd5-a6f1-5152c13aca10 · outbound

This paper cites MaskGCT: Zero-Shot Text-to-Speech with Masked Generative Codec Transformer.

Multimodal Large Language Model-Enabled Video Translation: A Role-Oriented Survey MaskGCT: Zero-Shot Text-to-Speech with Masked Generative Codec Transformer

Reference 101

Resolution
unresolved
no resolver link, observed 2026-07-12T22:04:31.302192Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T22:04:31.302192Z digest=sha256:794bb008fa7b0ce36a99a9de3461b6cca6f78c1c93f551d16ea5332954cdd72b

Pith citing papers

No inbound Pith citation observations are available.