Pith. sign in

Paper Citation Record · LEDGER

ViCocktail: Automated Multi-Modal Data Collection for Vietnamese Audio-Visual Speech Recognition

As of 8 August 2026, this Paper Citation Record lists 41 of 41 outbound references and 1 inbound Pith citation observation for arXiv:2506.04635.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.04635 v1

Coverage vector

measured 41 of 41 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T10:44:02.577587Z

measured 42 of 42 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T10:43:58.304998Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-07T10:44:02.849295Z

Reference resolution

41 of 41 outbound references displayed

  • verified exact0
  • verified fuzzy33
  • unresolved6
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch1

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 14534fc1-1de0-45a5-9a32-140aa25824c8 · outbound

This paper cites ViCocktail: Automated Multi-Modal Data Collection for Vietnamese Audio-Visual Speech Recognition.

ViCocktail: Automated Multi-Modal Data Collection for Vietnamese Audio-Visual Speech Recognition ViCocktail: Automated Multi-Modal Data Collection for Vietnamese Audio-Visual Speech Recognition

Reference 1

Resolution
metadata mismatch
local_arxiv, observed 2026-08-07T10:44:02.920063Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T10:43:58.304998Z digest=sha256:b4b85e72652accdedde887eaa9aa6d537a269278ef9442cdafa95e204d85c63e

Observation cda3b0d1-e221-46f6-8a58-e7c8b0932664 · outbound

This paper cites ViCocktail dataset In this section, we describe the multi-stage pipeline for auto- matically generating a dataset for A VSR model.

ViCocktail: Automated Multi-Modal Data Collection for Vietnamese Audio-Visual Speech Recognition ViCocktail dataset In this section, we describe the multi-stage pipeline for auto- matically generating a dataset for A VSR model

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:44:09.037825Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T10:43:58.395010Z digest=sha256:dd34dfee78730eee79d086579c1d6aac8b06b73d6c4d565edded69531a47351f

Observation af47041d-62e8-4410-ac52-cee99ec66afa · outbound

This paper cites The first model uses a Conformer-based encoder [31] with a CTC/Attention de- coder [32].

ViCocktail: Automated Multi-Modal Data Collection for Vietnamese Audio-Visual Speech Recognition The first model uses a Conformer-based encoder [31] with a CTC/Attention de- coder [32]

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:44:08.739735Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T10:43:58.659123Z digest=sha256:a2158d467d0ab81fed3f6c80b080b8724b04c32ad3c6e3703df60e338a9868eb

Observation c55a85f0-9deb-42c5-b28f-c3c9109eb22d · outbound

This paper cites an unresolved cited work.

ViCocktail: Automated Multi-Modal Data Collection for Vietnamese Audio-Visual Speech Recognition Unresolved cited work

Reference 4

Resolution
malformed identifier
raw_fallback, observed 2026-08-07T10:44:08.461469Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T10:43:58.897221Z digest=sha256:399da3a532c67ddd086edac241a816b2ca19ef1be2115da72105fa25ae1fac32

Observation 34ae3160-da82-41d3-b579-faa8a6eb7395 · outbound

This paper cites natural”, “music.

ViCocktail: Automated Multi-Modal Data Collection for Vietnamese Audio-Visual Speech Recognition natural”, “music

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:44:08.603449Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T10:43:58.762305Z digest=sha256:4133df5dfa9ebdd057c5d21166204b29675b168a34cd329c59b0d7617d796751

Observation a8e7cc0e-08b4-4f64-b6c7-a5b6f96bd24f · outbound

This paper cites LRS3-TED: a large-scale dataset for visual speech recognition.

ViCocktail: Automated Multi-Modal Data Collection for Vietnamese Audio-Visual Speech Recognition LRS3-TED: a large-scale dataset for visual speech recognition

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T10:43:59.852640Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:43:59.852640Z digest=sha256:2e6d88ba500f281d314fd172124a61c92f2c8bce27ba1b5834d233da33d0210e

Observation 31069dcb-5d22-4b6d-8860-ed195367982f · outbound

This paper cites Our data collection and preparation process is fully automated and can be extended to other languages.

ViCocktail: Automated Multi-Modal Data Collection for Vietnamese Audio-Visual Speech Recognition Our data collection and preparation process is fully automated and can be extended to other languages

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:44:08.315209Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T10:43:59.065158Z digest=sha256:f1902d037bec1b7ae8c0fd10f2c9edd124a8fac77eaec58c45863dd33ba94bb8

Observation 0b9620e7-3120-47fc-83fe-813600c7f8d6 · outbound

This paper cites Hearing lips and seeing voices,.

ViCocktail: Automated Multi-Modal Data Collection for Vietnamese Audio-Visual Speech Recognition Hearing lips and seeing voices,

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T10:43:59.187872Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:43:59.187872Z digest=sha256:792aea7d951fa3d978c2af669b76f607d127bc1c6b1b0adf07527ddd80ee4d7e

Observation e025039d-30c9-482c-9372-2276f42acc19 · outbound

This paper cites Integration of acous- tic and visual speech signals using neural networks,.

ViCocktail: Automated Multi-Modal Data Collection for Vietnamese Audio-Visual Speech Recognition Integration of acous- tic and visual speech signals using neural networks,

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:44:08.158035Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T10:43:59.325929Z digest=sha256:d25430b420e3eb47df2ec71e8bb309c20a3e7456efe38df1d594c27f8d99a202

Observation 66ef00be-2e43-4000-8ffb-1a7c9f6da7fe · outbound

This paper cites See me, hear me: inte- grating automatic speech recognition and lip-reading,.

ViCocktail: Automated Multi-Modal Data Collection for Vietnamese Audio-Visual Speech Recognition See me, hear me: inte- grating automatic speech recognition and lip-reading,

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:44:08.028221Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T10:43:59.466040Z digest=sha256:7daeccd4adbb8f2e44e8040bc2d0cb2b75f0fe60c58bca9c688e47dcdf9291a8

Observation cf6772e2-062b-4b11-a5f5-d8f0eaa910dd · outbound

This paper cites Multimodal interfaces,.

ViCocktail: Automated Multi-Modal Data Collection for Vietnamese Audio-Visual Speech Recognition Multimodal interfaces,

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:44:07.885954Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T10:43:59.589762Z digest=sha256:8fae869ed2c81b847330302d3d7273f9adbd264b3b54486dc12d25304076d094

Observation ddaa6b4d-bd68-4df0-81f8-e6212c4d5f7f · outbound

This paper cites Lip read- ing sentences in the wild,.

ViCocktail: Automated Multi-Modal Data Collection for Vietnamese Audio-Visual Speech Recognition Lip read- ing sentences in the wild,

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:44:07.750127Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T10:43:59.712319Z digest=sha256:7fbec26b4a642048acee320a90c2a9a6e7a2deb37b2046479f4a700c1f3fca7a

Observation 1abe6788-433b-426a-b964-dccf5c0adfcd · outbound

This paper cites RUSA VIC corpus: Russian audio-visual speech in cars,.

ViCocktail: Automated Multi-Modal Data Collection for Vietnamese Audio-Visual Speech Recognition RUSA VIC corpus: Russian audio-visual speech in cars,

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:44:06.869597Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T10:44:00.799174Z digest=sha256:42869dad66b6aa6b57439f6c80beaf10972a78d10253608c29e935b2872ce8af

Observation 97db2532-7544-4c47-9e69-7132363b267b · outbound

This paper cites Large-scale visual speech recognition,.

ViCocktail: Automated Multi-Modal Data Collection for Vietnamese Audio-Visual Speech Recognition Large-scale visual speech recognition,

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:44:07.601661Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T10:43:59.968516Z digest=sha256:d48b52f0be746ae9a75f40e8433d5e71886cc737c5affb36bbdfd4d4f2a1f396

Observation 80da9692-e488-4ef2-8012-8ff559f16835 · outbound

This paper cites Avas: Speech database for multimodal recognition applications,.

ViCocktail: Automated Multi-Modal Data Collection for Vietnamese Audio-Visual Speech Recognition Avas: Speech database for multimodal recognition applications,

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:44:07.478794Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T10:44:00.102673Z digest=sha256:a6d7330813d38e31c630db06cb41d6e9a1a9a8022b7d4380ad49982730daa563

Observation 16c16b9d-f679-400a-88c3-c35c431f57ec · outbound

This paper cites An arabic visual dataset for visual speech recognition,.

ViCocktail: Automated Multi-Modal Data Collection for Vietnamese Audio-Visual Speech Recognition An arabic visual dataset for visual speech recognition,

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:44:07.328730Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T10:44:00.250298Z digest=sha256:a89c301280b63bc48ddc8b27845049ab22507c106eae864e6c8caa7ca2e02018

Observation a85b86a1-6850-4757-a564-b82e1ddc902d · outbound

This paper cites Cn-cvs: A mandarin audio- visual dataset for large vocabulary continuous visual to speech synthesis,.

ViCocktail: Automated Multi-Modal Data Collection for Vietnamese Audio-Visual Speech Recognition Cn-cvs: A mandarin audio- visual dataset for large vocabulary continuous visual to speech synthesis,

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:44:07.159675Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T10:44:00.380608Z digest=sha256:371f766e1caf4bb400c6d69a3bd61dd490343aa7130e9cbe3a85edd68a618862

Observation caf27eae-3b0e-4f6a-b147-4f5a72c66176 · outbound

This paper cites Lipreading with densenet and resbi-lstm,.

ViCocktail: Automated Multi-Modal Data Collection for Vietnamese Audio-Visual Speech Recognition Lipreading with densenet and resbi-lstm,

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:44:07.028858Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T10:44:00.508386Z digest=sha256:dd8e092cc7d4c675e006d97599637a79725e8551d2333b55c10ca72c776b4c92

Observation c5fb88b5-9090-425a-afc2-c04eba3c98d2 · outbound

This paper cites A cascade sequence- to-sequence model for chinese mandarin lip reading,.

ViCocktail: Automated Multi-Modal Data Collection for Vietnamese Audio-Visual Speech Recognition A cascade sequence- to-sequence model for chinese mandarin lip reading,

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T10:44:00.574357Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:44:00.574357Z digest=sha256:594df3a95d7f92f720bcb05d63d8189809b39118873ab7612b7f1938e6ceb4d3

Observation 4a6fb9e2-5068-4426-a192-7c4b4f6122bd · outbound

This paper cites Auto-avsr: Audio-visual speech recognition with automatic labels,.

ViCocktail: Automated Multi-Modal Data Collection for Vietnamese Audio-Visual Speech Recognition Auto-avsr: Audio-visual speech recognition with automatic labels,

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:44:05.859305Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T10:44:01.557071Z digest=sha256:41cd349ff2d161e88d8413a9a6f283aad271ca1340bcbe2a09cc46c654cee740

Observation 82c188c4-f436-4a9b-8934-fe8c410b2b7b · outbound

This paper cites Havrus corpus: High-speed recordings of audio-visual russian speech,.

ViCocktail: Automated Multi-Modal Data Collection for Vietnamese Audio-Visual Speech Recognition Havrus corpus: High-speed recordings of audio-visual russian speech,

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:44:06.720953Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T10:44:00.958153Z digest=sha256:d0394da59aefe69b3c04580b33f8b83ba52584a09037efcb5e8330c4417e8d56

Observation bef02caf-cb17-4f85-b162-160266766334 · outbound

This paper cites Towards es- timating the upper bound of visual-speech recognition: The visual lip-reading feasibility database,.

ViCocktail: Automated Multi-Modal Data Collection for Vietnamese Audio-Visual Speech Recognition Towards es- timating the upper bound of visual-speech recognition: The visual lip-reading feasibility database,

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:44:06.600363Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T10:44:01.140474Z digest=sha256:cb0466c774e90d940a91e02695a7cb962ff49aec9c30d8edf166f40c4c2251ab

Observation 42e40033-5729-4718-bbf1-5cdc395bb8ae · outbound

This paper cites Visual lip reading dataset in turkish,.

ViCocktail: Automated Multi-Modal Data Collection for Vietnamese Audio-Visual Speech Recognition Visual lip reading dataset in turkish,

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:44:06.491929Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T10:44:01.293218Z digest=sha256:b7bfa8b9d8dd765fd682a174d9a4d26560cc0521468a4e3aecd4dba3881e65e5

Observation 6b858014-4c99-4a7e-8f23-27711c5e815f · outbound

This paper cites The ASD model relies on both audio and visual features to deter- mine when a person in the video is actually speaking.

ViCocktail: Automated Multi-Modal Data Collection for Vietnamese Audio-Visual Speech Recognition The ASD model relies on both audio and visual features to deter- mine when a person in the video is actually speaking

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:44:08.900652Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T10:43:58.506495Z digest=sha256:55a8fcc13ccc977d57f1e2382151df9986d99730b04c3b39cf8848188e742c1c

Observation 760ad800-c7ad-4ebc-8cc5-b6a609744b58 · outbound

This paper cites Lip reading in the wild,.

ViCocktail: Automated Multi-Modal Data Collection for Vietnamese Audio-Visual Speech Recognition Lip reading in the wild,

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:44:06.383881Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T10:44:01.352567Z digest=sha256:ad3910849403d1f412fa5137a2f87b6118da2573ceb836cd700a15c0bf03b203

Observation 7b42453e-0765-48e7-947b-7acba01e11d4 · outbound

This paper cites Robust self-supervised audio-visual speech recognition,.

ViCocktail: Automated Multi-Modal Data Collection for Vietnamese Audio-Visual Speech Recognition Robust self-supervised audio-visual speech recognition,

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:44:06.199784Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T10:44:01.425993Z digest=sha256:f58dd41b091ca8cfd6320a3318370c9c1dea6703b93029209be5cd37936ee8ae

Observation d456a7ff-1ee3-4d35-b82c-aaed08e9dc57 · outbound

This paper cites Whisper-flamingo: Integrating visual fea- tures into whisper for audio-visual speech recognition and trans- lation,.

ViCocktail: Automated Multi-Modal Data Collection for Vietnamese Audio-Visual Speech Recognition Whisper-flamingo: Integrating visual fea- tures into whisper for audio-visual speech recognition and trans- lation,

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:44:06.022408Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T10:44:01.499122Z digest=sha256:d761db4a4879dcd3003719886d0e23811093cab78d83f8428d6d3104e8a86792

Observation 471ae582-469e-466f-8343-a189ba9faf54 · outbound

This paper cites Xls-r: Self-supervised cross-lingual speech representation learning at scale,.

ViCocktail: Automated Multi-Modal Data Collection for Vietnamese Audio-Visual Speech Recognition Xls-r: Self-supervised cross-lingual speech representation learning at scale,

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T10:44:01.615895Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:44:01.615895Z digest=sha256:32b2a5a15f586c1429e2449774e4745cbf910e33c40d89cf63e86c1d7577b441

Observation 24ef293d-e845-4682-afbc-87b2ad79bfc6 · outbound

This paper cites Simple and effective zero-shot cross-lingual phoneme recognition,.

ViCocktail: Automated Multi-Modal Data Collection for Vietnamese Audio-Visual Speech Recognition Simple and effective zero-shot cross-lingual phoneme recognition,

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T10:44:01.671995Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:44:01.671995Z digest=sha256:2b14b4d540286eb001efe6ebeb94d81453e89f19fd4d96915f7e3efdab4d3314

Observation 2849a824-84b2-42d5-8a38-e82436792501 · outbound

This paper cites S3fd: Single shot scale-invariant face detector,.

ViCocktail: Automated Multi-Modal Data Collection for Vietnamese Audio-Visual Speech Recognition S3fd: Single shot scale-invariant face detector,

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:44:05.697665Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T10:44:01.726648Z digest=sha256:ca1273219d60098c977529a394f314e75b430a3bb398392c0324e2a49cfd1fb4

Observation abbf2c26-b893-40cc-a4bf-c1e08487b3ea · outbound

This paper cites A light weight model for active speaker detection,.

ViCocktail: Automated Multi-Modal Data Collection for Vietnamese Audio-Visual Speech Recognition A light weight model for active speaker detection,

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:44:05.532314Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T10:44:01.800743Z digest=sha256:f95effe897ef9b7e78e410abf028fbaf09d48e2d1a15c5ca64781b0bb5817153

Observation fedafee6-38e3-4709-9e24-ab31cf6c6faa · outbound

This paper cites Out of time: Automated lip sync in the wild,.

ViCocktail: Automated Multi-Modal Data Collection for Vietnamese Audio-Visual Speech Recognition Out of time: Automated lip sync in the wild,

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:44:05.413907Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T10:44:01.862618Z digest=sha256:40ee9a7db2881d8accfbd05cef58fd7f57c62fffb0f684130cbd5097b8cc0f11

Observation 53c7fd6b-edc6-4728-81b2-38816c36a69c · outbound

This paper cites Dlib-ml: A machine learning toolkit,.

ViCocktail: Automated Multi-Modal Data Collection for Vietnamese Audio-Visual Speech Recognition Dlib-ml: A machine learning toolkit,

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:44:05.230983Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T10:44:01.938616Z digest=sha256:6a1c8a32f8ecf7c311b84c56e601721775861ce830cff2b330e251e8676cc33e

Observation 8258663b-4295-4237-949e-95f1acb13e81 · outbound

This paper cites Vietnamese end-to-end speech recognition using wav2vec 2.0,.

ViCocktail: Automated Multi-Modal Data Collection for Vietnamese Audio-Visual Speech Recognition Vietnamese end-to-end speech recognition using wav2vec 2.0,

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:44:04.910813Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T10:44:01.986945Z digest=sha256:4abbfe0cf5cefd5f1653fca599ccf87c1f9f5f0eb147fe7301c9ea244b055aaa

Observation 06e23409-5baf-4e7e-a8e4-c44496ccd595 · outbound

This paper cites Synthetic conversations improve multi-talker asr,.

ViCocktail: Automated Multi-Modal Data Collection for Vietnamese Audio-Visual Speech Recognition Synthetic conversations improve multi-talker asr,

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:44:04.554124Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T10:44:02.068194Z digest=sha256:0f5c2aa4ee219c745314bd76c3cc5133eb51432af7eaaa1b70808d25a6799a39

Observation b3e1757f-8f27-4af3-8d94-6012dfc76184 · outbound

This paper cites Msa-asr: Efficient multilingual speaker attribution with frozen asr models,.

ViCocktail: Automated Multi-Modal Data Collection for Vietnamese Audio-Visual Speech Recognition Msa-asr: Efficient multilingual speaker attribution with frozen asr models,

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:44:04.202199Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T10:44:02.192619Z digest=sha256:40bac9fbfd98070ce4c815c26392199a84bae85c62855ec6db80c91c26f17e7d

Observation bb1324eb-2107-40b9-8c39-0e37d2b80fcf · outbound

This paper cites Phowhisper: Automatic speech recognition for vietnamese,.

ViCocktail: Automated Multi-Modal Data Collection for Vietnamese Audio-Visual Speech Recognition Phowhisper: Automatic speech recognition for vietnamese,

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:44:03.940516Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T10:44:02.277536Z digest=sha256:6fcd471963b6929213b9a346bdf8bc96bfcee1f996f23ebc8afe994c5587133c

Observation d5ceb486-d7ba-4b2a-af87-238661caba7d · outbound

This paper cites End-to-end audio-visual speech recognition with conformers,.

ViCocktail: Automated Multi-Modal Data Collection for Vietnamese Audio-Visual Speech Recognition End-to-end audio-visual speech recognition with conformers,

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:44:03.678059Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T10:44:02.360780Z digest=sha256:d0329fbcc6c58818719c63c98fb9a3656cec932839408d4dce72eeeab5616e20

Observation 56da5b47-bc35-473c-a8de-6f583bad848d · outbound

This paper cites Hy- brid ctc/attention architecture for end-to-end speech recognition,.

ViCocktail: Automated Multi-Modal Data Collection for Vietnamese Audio-Visual Speech Recognition Hy- brid ctc/attention architecture for end-to-end speech recognition,

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-07T10:44:02.427128Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:44:02.427128Z digest=sha256:3c8f8108f6d5784dda4888421fc75964d22b7a466651dc6d04bf81152e6159d0

Observation d4d35d24-1f69-4048-856e-89f12b758e9b · outbound

This paper cites Subword regularization: Improving neural network translation models with multiple subword candidates,.

ViCocktail: Automated Multi-Modal Data Collection for Vietnamese Audio-Visual Speech Recognition Subword regularization: Improving neural network translation models with multiple subword candidates,

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:44:03.293287Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T10:44:02.518850Z digest=sha256:56b673871bceaa546b46086549a4db5a7b59f956fe77278f9fce7ef24746a4ad

Observation 5dc24ec7-717d-4f4b-8817-26656b7bf6cd · outbound

This paper cites Muavic: A multilingual audio-visual corpus for robust speech recognition and robust speech-to-text translation,.

ViCocktail: Automated Multi-Modal Data Collection for Vietnamese Audio-Visual Speech Recognition Muavic: A multilingual audio-visual corpus for robust speech recognition and robust speech-to-text translation,

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:44:03.080745Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T10:44:02.577587Z digest=sha256:dc51dbe303af4f1371ec9c855ee1f1235e0eab20e0a5b9a94b17d7e7c43aae98

Pith citing papers

Observation 14534fc1-1de0-45a5-9a32-140aa25824c8 · inbound

ViCocktail: Automated Multi-Modal Data Collection for Vietnamese Audio-Visual Speech Recognition cites this paper.

ViCocktail: Automated Multi-Modal Data Collection for Vietnamese Audio-Visual Speech Recognition ViCocktail: Automated Multi-Modal Data Collection for Vietnamese Audio-Visual Speech Recognition

Reference 1

Resolution
metadata mismatch
local_arxiv, observed 2026-08-07T10:44:02.920063Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T10:43:58.304998Z digest=sha256:b4b85e72652accdedde887eaa9aa6d537a269278ef9442cdafa95e204d85c63e