Pith. sign in

Paper Citation Record · LEDGER

ViCocktail: Automated Multi-Modal Data Collection for Vietnamese Audio-Visual Speech Recognition

As of 8 August 2026, this Paper Citation Record lists 41 of 41 outbound references and 1 inbound Pith citation observation for arXiv:2506.04635.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.04635 v1

Coverage vector

measured 41 of 41 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T10:44:02.577587Z

measured 42 of 42 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T10:43:58.304998Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-07T10:44:02.849295Z

Reference resolution

41 of 41 outbound references displayed

  • verified exact0
  • verified fuzzy33
  • unresolved6
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch1

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 14534fc1-1de0-45a5-9a32-140aa25824c8 · outbound

This paper cites ViCocktail: Automated Multi-Modal Data Collection for Vietnamese Audio-Visual Speech Recognition.

ViCocktail: Automated Multi-Modal Data Collection for Vietnamese Audio-Visual Speech Recognition ViCocktail: Automated Multi-Modal Data Collection for Vietnamese Audio-Visual Speech Recognition

Reference 1

Resolution
metadata mismatch
local_arxiv, observed 2026-08-07T10:44:02.920063Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T10:43:58.304998Z digest=sha256:ee63a165f8aae98f6e47e80263aecc134ed041b654934ab0b707ad450e75641b

Observation cda3b0d1-e221-46f6-8a58-e7c8b0932664 · outbound

This paper cites ViCocktail dataset In this section, we describe the multi-stage pipeline for auto- matically generating a dataset for A VSR model.

ViCocktail: Automated Multi-Modal Data Collection for Vietnamese Audio-Visual Speech Recognition ViCocktail dataset In this section, we describe the multi-stage pipeline for auto- matically generating a dataset for A VSR model

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:44:09.037825Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T10:43:58.395010Z digest=sha256:434b4a6e2f9dd7d455b5eeb8bfd2ea85f92e99fdf71f0da7bbe38abe7c2b79ea

Observation af47041d-62e8-4410-ac52-cee99ec66afa · outbound

This paper cites The first model uses a Conformer-based encoder [31] with a CTC/Attention de- coder [32].

ViCocktail: Automated Multi-Modal Data Collection for Vietnamese Audio-Visual Speech Recognition The first model uses a Conformer-based encoder [31] with a CTC/Attention de- coder [32]

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:44:08.739735Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T10:43:58.659123Z digest=sha256:bc7f9fd9a87208c0dc861ba1b3433589e70b061a16d0fd996c7e9c808c3d4d05

Observation c55a85f0-9deb-42c5-b28f-c3c9109eb22d · outbound

This paper cites an unresolved cited work.

ViCocktail: Automated Multi-Modal Data Collection for Vietnamese Audio-Visual Speech Recognition Unresolved cited work

Reference 4

Resolution
malformed identifier
raw_fallback, observed 2026-08-07T10:44:08.461469Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T10:43:58.897221Z digest=sha256:2541b68a6ea9799ab2a641a90eb774f92747fa4800a0b3097f311502d010a23c

Observation 34ae3160-da82-41d3-b579-faa8a6eb7395 · outbound

This paper cites natural”, “music.

ViCocktail: Automated Multi-Modal Data Collection for Vietnamese Audio-Visual Speech Recognition natural”, “music

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:44:08.603449Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T10:43:58.762305Z digest=sha256:d384434f39b85e365626e7eb76a95845870400e723fe785f45a904d5b2513d6a

Observation a8e7cc0e-08b4-4f64-b6c7-a5b6f96bd24f · outbound

This paper cites LRS3-TED: a large-scale dataset for visual speech recognition.

ViCocktail: Automated Multi-Modal Data Collection for Vietnamese Audio-Visual Speech Recognition LRS3-TED: a large-scale dataset for visual speech recognition

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T10:43:59.852640Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:43:59.852640Z digest=sha256:2e6d88ba500f281d314fd172124a61c92f2c8bce27ba1b5834d233da33d0210e

Observation 31069dcb-5d22-4b6d-8860-ed195367982f · outbound

This paper cites Our data collection and preparation process is fully automated and can be extended to other languages.

ViCocktail: Automated Multi-Modal Data Collection for Vietnamese Audio-Visual Speech Recognition Our data collection and preparation process is fully automated and can be extended to other languages

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:44:08.315209Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T10:43:59.065158Z digest=sha256:a54a35d4d15f3e271ac78d544716956e99310db950593466a7ba2608d01af239

Observation 0b9620e7-3120-47fc-83fe-813600c7f8d6 · outbound

This paper cites Hearing lips and seeing voices,.

ViCocktail: Automated Multi-Modal Data Collection for Vietnamese Audio-Visual Speech Recognition Hearing lips and seeing voices,

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T10:43:59.187872Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:43:59.187872Z digest=sha256:792aea7d951fa3d978c2af669b76f607d127bc1c6b1b0adf07527ddd80ee4d7e

Observation e025039d-30c9-482c-9372-2276f42acc19 · outbound

This paper cites Integration of acous- tic and visual speech signals using neural networks,.

ViCocktail: Automated Multi-Modal Data Collection for Vietnamese Audio-Visual Speech Recognition Integration of acous- tic and visual speech signals using neural networks,

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:44:08.158035Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T10:43:59.325929Z digest=sha256:e96ea5c332628f20aa2fb6830033b1a919b3589524c5a4a8c8a193923673a593

Observation 66ef00be-2e43-4000-8ffb-1a7c9f6da7fe · outbound

This paper cites See me, hear me: inte- grating automatic speech recognition and lip-reading,.

ViCocktail: Automated Multi-Modal Data Collection for Vietnamese Audio-Visual Speech Recognition See me, hear me: inte- grating automatic speech recognition and lip-reading,

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:44:08.028221Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T10:43:59.466040Z digest=sha256:dae31a903d678d7baa5df3f07781e6734a1299304e581d2990426cc15822b565

Observation cf6772e2-062b-4b11-a5f5-d8f0eaa910dd · outbound

This paper cites Multimodal interfaces,.

ViCocktail: Automated Multi-Modal Data Collection for Vietnamese Audio-Visual Speech Recognition Multimodal interfaces,

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:44:07.885954Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T10:43:59.589762Z digest=sha256:098f1e186b5b5f4a75a01331a5adb22679482677c8d9923c4e777b033fef3674

Observation ddaa6b4d-bd68-4df0-81f8-e6212c4d5f7f · outbound

This paper cites Lip read- ing sentences in the wild,.

ViCocktail: Automated Multi-Modal Data Collection for Vietnamese Audio-Visual Speech Recognition Lip read- ing sentences in the wild,

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:44:07.750127Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T10:43:59.712319Z digest=sha256:3c82ff10bdb05b10193217cca2763c5752e9aee5c0159559776617bf649c0005

Observation 1abe6788-433b-426a-b964-dccf5c0adfcd · outbound

This paper cites RUSA VIC corpus: Russian audio-visual speech in cars,.

ViCocktail: Automated Multi-Modal Data Collection for Vietnamese Audio-Visual Speech Recognition RUSA VIC corpus: Russian audio-visual speech in cars,

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:44:06.869597Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T10:44:00.799174Z digest=sha256:63eda2ce2f14f8289ced449b33271bfa39b63421fbf9178543c50ae8467992cb

Observation 97db2532-7544-4c47-9e69-7132363b267b · outbound

This paper cites Large-scale visual speech recognition,.

ViCocktail: Automated Multi-Modal Data Collection for Vietnamese Audio-Visual Speech Recognition Large-scale visual speech recognition,

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:44:07.601661Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T10:43:59.968516Z digest=sha256:2d49e10d1620b53f2b0847ca5cd73032fa1041e6ace60e95e8c62a2feb281dc5

Observation 80da9692-e488-4ef2-8012-8ff559f16835 · outbound

This paper cites Avas: Speech database for multimodal recognition applications,.

ViCocktail: Automated Multi-Modal Data Collection for Vietnamese Audio-Visual Speech Recognition Avas: Speech database for multimodal recognition applications,

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:44:07.478794Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T10:44:00.102673Z digest=sha256:0aac42949ff6903cf328ccade1abdcbfc05579fd3c18710855bf0b52c9ec20f9

Observation 16c16b9d-f679-400a-88c3-c35c431f57ec · outbound

This paper cites An arabic visual dataset for visual speech recognition,.

ViCocktail: Automated Multi-Modal Data Collection for Vietnamese Audio-Visual Speech Recognition An arabic visual dataset for visual speech recognition,

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:44:07.328730Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T10:44:00.250298Z digest=sha256:b98bb0df14379e2291d0807d1132b2842f5088661d51faa767b33559cdfe0400

Observation a85b86a1-6850-4757-a564-b82e1ddc902d · outbound

This paper cites Cn-cvs: A mandarin audio- visual dataset for large vocabulary continuous visual to speech synthesis,.

ViCocktail: Automated Multi-Modal Data Collection for Vietnamese Audio-Visual Speech Recognition Cn-cvs: A mandarin audio- visual dataset for large vocabulary continuous visual to speech synthesis,

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:44:07.159675Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T10:44:00.380608Z digest=sha256:5b6d2460639a6b859ca8ec7f9f754287fb809dbf7477e1848db5446d8981f51d

Observation caf27eae-3b0e-4f6a-b147-4f5a72c66176 · outbound

This paper cites Lipreading with densenet and resbi-lstm,.

ViCocktail: Automated Multi-Modal Data Collection for Vietnamese Audio-Visual Speech Recognition Lipreading with densenet and resbi-lstm,

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:44:07.028858Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T10:44:00.508386Z digest=sha256:13a2edea4a9b6e0a9031cc48dfad3c3e7c3e756ed640752eefd09fc206e33466

Observation c5fb88b5-9090-425a-afc2-c04eba3c98d2 · outbound

This paper cites A cascade sequence- to-sequence model for chinese mandarin lip reading,.

ViCocktail: Automated Multi-Modal Data Collection for Vietnamese Audio-Visual Speech Recognition A cascade sequence- to-sequence model for chinese mandarin lip reading,

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T10:44:00.574357Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:44:00.574357Z digest=sha256:594df3a95d7f92f720bcb05d63d8189809b39118873ab7612b7f1938e6ceb4d3

Observation 4a6fb9e2-5068-4426-a192-7c4b4f6122bd · outbound

This paper cites Auto-avsr: Audio-visual speech recognition with automatic labels,.

ViCocktail: Automated Multi-Modal Data Collection for Vietnamese Audio-Visual Speech Recognition Auto-avsr: Audio-visual speech recognition with automatic labels,

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:44:05.859305Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T10:44:01.557071Z digest=sha256:4d8686dfdab0e3d4d3020d3808464c76844dbfd464af99cb9bdea873b6d68aa9

Observation 82c188c4-f436-4a9b-8934-fe8c410b2b7b · outbound

This paper cites Havrus corpus: High-speed recordings of audio-visual russian speech,.

ViCocktail: Automated Multi-Modal Data Collection for Vietnamese Audio-Visual Speech Recognition Havrus corpus: High-speed recordings of audio-visual russian speech,

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:44:06.720953Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T10:44:00.958153Z digest=sha256:4b7ced36186169fa494695e05171d4ed854743cd03172d11246141c3675f724c

Observation bef02caf-cb17-4f85-b162-160266766334 · outbound

This paper cites Towards es- timating the upper bound of visual-speech recognition: The visual lip-reading feasibility database,.

ViCocktail: Automated Multi-Modal Data Collection for Vietnamese Audio-Visual Speech Recognition Towards es- timating the upper bound of visual-speech recognition: The visual lip-reading feasibility database,

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:44:06.600363Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T10:44:01.140474Z digest=sha256:4ed8a1620cb2ccb9f667d55ad0156e9c76444f3223b1e185366e4b92c8fa57fe

Observation 42e40033-5729-4718-bbf1-5cdc395bb8ae · outbound

This paper cites Visual lip reading dataset in turkish,.

ViCocktail: Automated Multi-Modal Data Collection for Vietnamese Audio-Visual Speech Recognition Visual lip reading dataset in turkish,

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:44:06.491929Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T10:44:01.293218Z digest=sha256:c97f12ece0973ca084d697ec824175499ee7dd877237c933925e48204de22cf5

Observation 6b858014-4c99-4a7e-8f23-27711c5e815f · outbound

This paper cites The ASD model relies on both audio and visual features to deter- mine when a person in the video is actually speaking.

ViCocktail: Automated Multi-Modal Data Collection for Vietnamese Audio-Visual Speech Recognition The ASD model relies on both audio and visual features to deter- mine when a person in the video is actually speaking

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:44:08.900652Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T10:43:58.506495Z digest=sha256:5e29d9c62a2ffec2653ec980a104697bc5ecd7380dba14f22cc3fd905ab1558e

Observation 760ad800-c7ad-4ebc-8cc5-b6a609744b58 · outbound

This paper cites Lip reading in the wild,.

ViCocktail: Automated Multi-Modal Data Collection for Vietnamese Audio-Visual Speech Recognition Lip reading in the wild,

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:44:06.383881Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T10:44:01.352567Z digest=sha256:97100a9059e5c8b296c0b059437ac09d5ddcf55baad521119689fc1b7d42bd01

Observation 7b42453e-0765-48e7-947b-7acba01e11d4 · outbound

This paper cites Robust self-supervised audio-visual speech recognition,.

ViCocktail: Automated Multi-Modal Data Collection for Vietnamese Audio-Visual Speech Recognition Robust self-supervised audio-visual speech recognition,

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:44:06.199784Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T10:44:01.425993Z digest=sha256:8201536ca3597a2addb3dca9a743c3bee6d52280dfc0ec8345fe072643b67f28

Observation d456a7ff-1ee3-4d35-b82c-aaed08e9dc57 · outbound

This paper cites Whisper-flamingo: Integrating visual fea- tures into whisper for audio-visual speech recognition and trans- lation,.

ViCocktail: Automated Multi-Modal Data Collection for Vietnamese Audio-Visual Speech Recognition Whisper-flamingo: Integrating visual fea- tures into whisper for audio-visual speech recognition and trans- lation,

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:44:06.022408Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T10:44:01.499122Z digest=sha256:305d55a740877dfbc83d8091a3d644c8faa62beb7d115e5c9d57b3d698b64c8b

Observation 471ae582-469e-466f-8343-a189ba9faf54 · outbound

This paper cites Xls-r: Self-supervised cross-lingual speech representation learning at scale,.

ViCocktail: Automated Multi-Modal Data Collection for Vietnamese Audio-Visual Speech Recognition Xls-r: Self-supervised cross-lingual speech representation learning at scale,

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T10:44:01.615895Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:44:01.615895Z digest=sha256:32b2a5a15f586c1429e2449774e4745cbf910e33c40d89cf63e86c1d7577b441

Observation 24ef293d-e845-4682-afbc-87b2ad79bfc6 · outbound

This paper cites Simple and effective zero-shot cross-lingual phoneme recognition,.

ViCocktail: Automated Multi-Modal Data Collection for Vietnamese Audio-Visual Speech Recognition Simple and effective zero-shot cross-lingual phoneme recognition,

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T10:44:01.671995Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:44:01.671995Z digest=sha256:2b14b4d540286eb001efe6ebeb94d81453e89f19fd4d96915f7e3efdab4d3314

Observation 2849a824-84b2-42d5-8a38-e82436792501 · outbound

This paper cites S3fd: Single shot scale-invariant face detector,.

ViCocktail: Automated Multi-Modal Data Collection for Vietnamese Audio-Visual Speech Recognition S3fd: Single shot scale-invariant face detector,

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:44:05.697665Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T10:44:01.726648Z digest=sha256:6933cf9e865e105425632bae0424461655caeb417a2dcfdd852aee9af051cf16

Observation abbf2c26-b893-40cc-a4bf-c1e08487b3ea · outbound

This paper cites A light weight model for active speaker detection,.

ViCocktail: Automated Multi-Modal Data Collection for Vietnamese Audio-Visual Speech Recognition A light weight model for active speaker detection,

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:44:05.532314Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T10:44:01.800743Z digest=sha256:48d901ea6e1442dd90584e84ae50d116f5d9d35c7a5765c90bbf6f16a19144ef

Observation fedafee6-38e3-4709-9e24-ab31cf6c6faa · outbound

This paper cites Out of time: Automated lip sync in the wild,.

ViCocktail: Automated Multi-Modal Data Collection for Vietnamese Audio-Visual Speech Recognition Out of time: Automated lip sync in the wild,

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:44:05.413907Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T10:44:01.862618Z digest=sha256:cdd05774f3cc23da53a3a6d581e2dc4f5f4ce8e1f80025b7316243c8e9cf5764

Observation 53c7fd6b-edc6-4728-81b2-38816c36a69c · outbound

This paper cites Dlib-ml: A machine learning toolkit,.

ViCocktail: Automated Multi-Modal Data Collection for Vietnamese Audio-Visual Speech Recognition Dlib-ml: A machine learning toolkit,

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:44:05.230983Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T10:44:01.938616Z digest=sha256:1ac51eec94fc010e60507a9b6fa07a91be102178c8efba8a1400d1285eca7060

Observation 8258663b-4295-4237-949e-95f1acb13e81 · outbound

This paper cites Vietnamese end-to-end speech recognition using wav2vec 2.0,.

ViCocktail: Automated Multi-Modal Data Collection for Vietnamese Audio-Visual Speech Recognition Vietnamese end-to-end speech recognition using wav2vec 2.0,

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:44:04.910813Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T10:44:01.986945Z digest=sha256:0d018e8c5b28e985af6a75269257bef8474f46f6ac46ff9ba9bd2beba8668de5

Observation 06e23409-5baf-4e7e-a8e4-c44496ccd595 · outbound

This paper cites Synthetic conversations improve multi-talker asr,.

ViCocktail: Automated Multi-Modal Data Collection for Vietnamese Audio-Visual Speech Recognition Synthetic conversations improve multi-talker asr,

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:44:04.554124Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T10:44:02.068194Z digest=sha256:7daf043bfa5f0caef138b3d43f7c404ba4397c5a1738e9dfb249c8ed20c795bd

Observation b3e1757f-8f27-4af3-8d94-6012dfc76184 · outbound

This paper cites Msa-asr: Efficient multilingual speaker attribution with frozen asr models,.

ViCocktail: Automated Multi-Modal Data Collection for Vietnamese Audio-Visual Speech Recognition Msa-asr: Efficient multilingual speaker attribution with frozen asr models,

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:44:04.202199Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T10:44:02.192619Z digest=sha256:493bf5975ca1d9ec294555f5ffff077f59e19d095e46ac2a1537f4d8a155d6e6

Observation bb1324eb-2107-40b9-8c39-0e37d2b80fcf · outbound

This paper cites Phowhisper: Automatic speech recognition for vietnamese,.

ViCocktail: Automated Multi-Modal Data Collection for Vietnamese Audio-Visual Speech Recognition Phowhisper: Automatic speech recognition for vietnamese,

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:44:03.940516Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T10:44:02.277536Z digest=sha256:e7ad9601838730e8d3d9cc090ebdbd37a3219e4c48131eb88523ef54f6aa3eef

Observation d5ceb486-d7ba-4b2a-af87-238661caba7d · outbound

This paper cites End-to-end audio-visual speech recognition with conformers,.

ViCocktail: Automated Multi-Modal Data Collection for Vietnamese Audio-Visual Speech Recognition End-to-end audio-visual speech recognition with conformers,

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:44:03.678059Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T10:44:02.360780Z digest=sha256:cde84e153ae90d247def2c4cb6fbc9d63bf49c1013cadb0c142f9f4c576de2fe

Observation 56da5b47-bc35-473c-a8de-6f583bad848d · outbound

This paper cites Hy- brid ctc/attention architecture for end-to-end speech recognition,.

ViCocktail: Automated Multi-Modal Data Collection for Vietnamese Audio-Visual Speech Recognition Hy- brid ctc/attention architecture for end-to-end speech recognition,

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-07T10:44:02.427128Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:44:02.427128Z digest=sha256:3c8f8108f6d5784dda4888421fc75964d22b7a466651dc6d04bf81152e6159d0

Observation d4d35d24-1f69-4048-856e-89f12b758e9b · outbound

This paper cites Subword regularization: Improving neural network translation models with multiple subword candidates,.

ViCocktail: Automated Multi-Modal Data Collection for Vietnamese Audio-Visual Speech Recognition Subword regularization: Improving neural network translation models with multiple subword candidates,

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:44:03.293287Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T10:44:02.518850Z digest=sha256:7e5a28f169ef7ed66026c2a669df2def7e8f8a92a3b90fa8f21d974d804a02d2

Observation 5dc24ec7-717d-4f4b-8817-26656b7bf6cd · outbound

This paper cites Muavic: A multilingual audio-visual corpus for robust speech recognition and robust speech-to-text translation,.

ViCocktail: Automated Multi-Modal Data Collection for Vietnamese Audio-Visual Speech Recognition Muavic: A multilingual audio-visual corpus for robust speech recognition and robust speech-to-text translation,

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:44:03.080745Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T10:44:02.577587Z digest=sha256:be358634126d6201d1af43097d24e80e110939ea69e4d22918796c8a3f29b68f

Pith citing papers

Observation 14534fc1-1de0-45a5-9a32-140aa25824c8 · inbound

ViCocktail: Automated Multi-Modal Data Collection for Vietnamese Audio-Visual Speech Recognition cites this paper.

ViCocktail: Automated Multi-Modal Data Collection for Vietnamese Audio-Visual Speech Recognition ViCocktail: Automated Multi-Modal Data Collection for Vietnamese Audio-Visual Speech Recognition

Reference 1

Resolution
metadata mismatch
local_arxiv, observed 2026-08-07T10:44:02.920063Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T10:43:58.304998Z digest=sha256:ee63a165f8aae98f6e47e80263aecc134ed041b654934ab0b707ad450e75641b