Pith. sign in

Paper Citation Record · LEDGER

Fusion of Discrete Representations and Self-Augmented Representations for Multilingual Automatic Speech Recognition

As of 13 August 2026, this Paper Citation Record lists 40 of 40 outbound references and 1 inbound Pith citation observation for arXiv:2411.18107.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2411.18107 v1

Coverage vector

measured 40 of 40 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-12T11:34:43.082458Z

measured 41 of 41 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-13T06:32:02.005865+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-12T11:34:42.898199Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-12T11:34:43.135889Z

Reference resolution

40 of 40 outbound references displayed

  • verified exact1
  • verified fuzzy36
  • unresolved2
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation addbb92a-e430-416c-840f-0a6916870228 · outbound

This paper cites Fusion of Discrete Representations and Self-Augmented Representations for Multilingual Automatic Speech Recognition.

Fusion of Discrete Representations and Self-Augmented Representations for Multilingual Automatic Speech Recognition Fusion of Discrete Representations and Self-Augmented Representations for Multilingual Automatic Speech Recognition

Reference 1

Resolution
verified exact
local_arxiv, observed 2026-08-12T11:34:43.143055Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T11:34:42.898199Z digest=sha256:5a41db7d6e044b4f49fc635d2cc6c49cf53b97155f6e0cb7d2fad80bc9239ce2

Observation 7778cee4-a5f6-4558-93e2-96a7bcf9caba · outbound

This paper cites Discretization process Figure 1 provides a high-level overview of our fusion pipeline.

Fusion of Discrete Representations and Self-Augmented Representations for Multilingual Automatic Speech Recognition Discretization process Figure 1 provides a high-level overview of our fusion pipeline

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:34:43.716932Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T11:34:42.904143Z digest=sha256:84ace8157e9ce827abadce4e09097cc34db533ebfffeb7e219cee537b91bd221

Observation 40d72017-11ef-4591-84a0-9c18c8aca1f1 · outbound

This paper cites an unresolved cited work.

Fusion of Discrete Representations and Self-Augmented Representations for Multilingual Automatic Speech Recognition Unresolved cited work

Reference 3

Resolution
unresolved
raw_fallback, observed 2026-08-12T11:34:43.702238Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T11:34:42.909090Z digest=sha256:549aa9557e6517b9b1466db84388563b47d34f779658e50aa4c2b33ac350d021

Observation 5c85dc06-3d76-4957-b5a5-be28f69811f9 · outbound

This paper cites Dataset We evaluated the proposed method on LibriSpeech-100h [29] and ML-SUPERB [12].

Fusion of Discrete Representations and Self-Augmented Representations for Multilingual Automatic Speech Recognition Dataset We evaluated the proposed method on LibriSpeech-100h [29] and ML-SUPERB [12]

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:34:43.687931Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T11:34:42.913865Z digest=sha256:787b3497008e7d9757512de6b3aae9562448a84d8e0214de8b14a2b37ddbbc63

Observation 41f686be-3a5c-4de4-9a6f-89846c00aa29 · outbound

This paper cites Quantitative result ASR results.

Fusion of Discrete Representations and Self-Augmented Representations for Multilingual Automatic Speech Recognition Quantitative result ASR results

Reference 5

Resolution
malformed identifier
raw_fallback, observed 2026-08-12T11:34:43.673010Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T11:34:42.918993Z digest=sha256:0790f77e5b36a9921578b7b6c0f97077b2d29a602e8bddb84678fa233ba3ea08

Observation 48d47749-f251-40bf-bc93-c1d89e78286e · outbound

This paper cites self-augmented.

Fusion of Discrete Representations and Self-Augmented Representations for Multilingual Automatic Speech Recognition self-augmented

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:34:43.658034Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T11:34:42.924135Z digest=sha256:f7eb2ab71b1bceb6b66705ecd91dc7dbe4ad3e158f3f59d0daa541a896723591

Observation 4d42857b-202d-4640-bf32-919a286e4c11 · outbound

This paper cites HuBERT: Self-supervised speech rep- resentation learning by masked prediction of hidden units,.

Fusion of Discrete Representations and Self-Augmented Representations for Multilingual Automatic Speech Recognition HuBERT: Self-supervised speech rep- resentation learning by masked prediction of hidden units,

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:34:43.642511Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T11:34:42.928858Z digest=sha256:edfaa7051dfaa84d99f17e34ffe95037b4c01977558f19901a217670a9e4fcb3

Observation 2ccf15b9-a944-4d7e-8834-1aa0bfd65702 · outbound

This paper cites A Survey of Multi- lingual Models for Automatic Speech Recognition,.

Fusion of Discrete Representations and Self-Augmented Representations for Multilingual Automatic Speech Recognition A Survey of Multi- lingual Models for Automatic Speech Recognition,

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:34:43.627743Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T11:34:42.933675Z digest=sha256:c8368cfe5071783ecebfdbdea3dd4da14a7f7146f84e62c8ae260e8a011016f1

Observation f9ded218-8b25-4537-b6d8-4f548066e6af · outbound

This paper cites Cross-lingual Automatic Speech Recog- nition Exploiting Articulatory Features,.

Fusion of Discrete Representations and Self-Augmented Representations for Multilingual Automatic Speech Recognition Cross-lingual Automatic Speech Recog- nition Exploiting Articulatory Features,

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:34:43.612015Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T11:34:42.938373Z digest=sha256:9655a0c3c1387654355b6b283a4564f4194452e909e96704f2c19b741ab304f9

Observation 2431b646-89c6-4a6f-a99b-070af72cca2b · outbound

This paper cites wav2vec 2.0: A framework for self- supervised learning of speech representations,.

Fusion of Discrete Representations and Self-Augmented Representations for Multilingual Automatic Speech Recognition wav2vec 2.0: A framework for self- supervised learning of speech representations,

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:34:43.596629Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T11:34:42.943156Z digest=sha256:3b192b7a30e350bcf5bd7ca7b2d2d796fa961c0772351acfc77f7f504606eed3

Observation 95077b77-3dfd-4519-8262-985d41698746 · outbound

This paper cites XLS-R: Self-supervised Cross-lingual Speech Representation Learning at Scale,.

Fusion of Discrete Representations and Self-Augmented Representations for Multilingual Automatic Speech Recognition XLS-R: Self-supervised Cross-lingual Speech Representation Learning at Scale,

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:34:43.582162Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T11:34:42.948404Z digest=sha256:2420aecc31d11cca0d5229deb02289f4c7db39e01b33c45ded6031960460dd1e

Observation 347b95c1-9056-47b1-97dd-117cf5aef508 · outbound

This paper cites W2v-BERT: Combining contrastive learning and masked language modeling for self-supervised speech pre-training,.

Fusion of Discrete Representations and Self-Augmented Representations for Multilingual Automatic Speech Recognition W2v-BERT: Combining contrastive learning and masked language modeling for self-supervised speech pre-training,

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:34:43.567358Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T11:34:42.952834Z digest=sha256:7219cf526bdb142c8782e15a593b787d6ee29f430f826e5e2a805098faf50f36

Observation c34e7edf-e82d-4c8d-aa10-45eaae192b7a · outbound

This paper cites Wavlm: Large-scale self-supervised pre- training for full stack speech processing,.

Fusion of Discrete Representations and Self-Augmented Representations for Multilingual Automatic Speech Recognition Wavlm: Large-scale self-supervised pre- training for full stack speech processing,

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:34:43.552377Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T11:34:42.958911Z digest=sha256:cf1580549edada25b17f5d1ff032e31c57408f2de3d38870cc16236ffa63ed33

Observation 95714afb-2bb4-4efc-8a4c-6bc0841a6e58 · outbound

This paper cites Self-supervised speech repre- sentation learning: A review,.

Fusion of Discrete Representations and Self-Augmented Representations for Multilingual Automatic Speech Recognition Self-supervised speech repre- sentation learning: A review,

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:34:43.537588Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T11:34:42.963844Z digest=sha256:c2be17299557792520ed3bbeeda3a4f6f80b1554523107c5663d288493ecc461

Observation 9dfbe30f-78d1-40f9-a069-5ec9779f8083 · outbound

This paper cites Multi-resolution huBERT: Multi-resolution speech self-supervised learning with masked unit prediction,.

Fusion of Discrete Representations and Self-Augmented Representations for Multilingual Automatic Speech Recognition Multi-resolution huBERT: Multi-resolution speech self-supervised learning with masked unit prediction,

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:34:43.522813Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T11:34:42.968484Z digest=sha256:adb6976b168b1a601590c4225e98179bb790b9a228c20e76ae6ab08201766fed

Observation a945e848-6a0d-49b3-b28d-ba67249f793f · outbound

This paper cites Scaling speech technology to 1,000+ languages,.

Fusion of Discrete Representations and Self-Augmented Representations for Multilingual Automatic Speech Recognition Scaling speech technology to 1,000+ languages,

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:34:43.508388Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T11:34:42.973148Z digest=sha256:3c3839294a06d851e661f83b410794ec6552d034c267e0a3ed412bc7d705d9ae

Observation 812838b2-3ebd-496e-8fd9-71e18ce10da1 · outbound

This paper cites SUPERB: Speech Processing Universal PERformance Benchmark,.

Fusion of Discrete Representations and Self-Augmented Representations for Multilingual Automatic Speech Recognition SUPERB: Speech Processing Universal PERformance Benchmark,

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:34:43.493670Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T11:34:42.977682Z digest=sha256:6f1642ca88af5f48888575acff06255bddd1608057335924f5db606a7ddd4300

Observation ab5e4686-c59d-4fde-b42c-18f259eb17ce · outbound

This paper cites ML-SUPERB: Multilingual Speech Uni- versal PERformance Benchmark,.

Fusion of Discrete Representations and Self-Augmented Representations for Multilingual Automatic Speech Recognition ML-SUPERB: Multilingual Speech Uni- versal PERformance Benchmark,

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:34:43.478182Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T11:34:42.982163Z digest=sha256:a40538c8231b3ee472ef5cae0eab935ac7360f2e7ff3eebc5342e36ae91cf88d

Observation f5bcc383-6ce3-4b16-83ab-0532e19ffd77 · outbound

This paper cites Exploring speech recognition, transla- tion, and understanding with discrete speech units: A compar- ative study,.

Fusion of Discrete Representations and Self-Augmented Representations for Multilingual Automatic Speech Recognition Exploring speech recognition, transla- tion, and understanding with discrete speech units: A compar- ative study,

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:34:43.462846Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T11:34:42.986449Z digest=sha256:43caed3499f4a03defb7825875394710594dee04747879ac4cd926c690e02049

Observation 5d281b8c-764a-42e6-8108-d988429f4f0d · outbound

This paper cites An Exploration of Prompt Tuning on Generative Spoken Language Model for Speech Processing Tasks,.

Fusion of Discrete Representations and Self-Augmented Representations for Multilingual Automatic Speech Recognition An Exploration of Prompt Tuning on Generative Spoken Language Model for Speech Processing Tasks,

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:34:43.447264Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T11:34:42.991048Z digest=sha256:7b4e2c2db79b60265327787eeaf0951effde66e104636b8f327fe193b6489abd

Observation 5b73ff31-a19a-476a-9370-8edcc61265af · outbound

This paper cites Towards universal speech discrete tokens: A case study for asr and tts,.

Fusion of Discrete Representations and Self-Augmented Representations for Multilingual Automatic Speech Recognition Towards universal speech discrete tokens: A case study for asr and tts,

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:34:43.432134Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T11:34:42.995775Z digest=sha256:441332e8b3ecc890bb189105d093365522b0873f11b64faea547072cd7c8fb94

Observation 7b89fbaf-c25d-4936-bbf2-50ce11887f5a · outbound

This paper cites Acoustic bpe for speech generation with dis- crete tokens,.

Fusion of Discrete Representations and Self-Augmented Representations for Multilingual Automatic Speech Recognition Acoustic bpe for speech generation with dis- crete tokens,

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:34:43.417656Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T11:34:43.000017Z digest=sha256:b4d71253961d6ded26259a645e38f8d540863796df21d9b387ac736c40120fca

Observation a70f3d72-a00d-41c6-b89e-914edc1694e8 · outbound

This paper cites V oxtlm: Unified decoder-only models for consolidating speech recognition, synthesis and speech, text continuation tasks,.

Fusion of Discrete Representations and Self-Augmented Representations for Multilingual Automatic Speech Recognition V oxtlm: Unified decoder-only models for consolidating speech recognition, synthesis and speech, text continuation tasks,

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:34:43.402217Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T11:34:43.004358Z digest=sha256:a48a3134a294e62f1c0f8c406b0aa3d61e9e5723ef5180820011d4d0569164ea

Observation 66479bd1-a262-479e-aa8c-f980edee6fd6 · outbound

This paper cites TokenSplit: Using Discrete Speech Representations for Direct, Refined, and Transcript- Conditioned Speech Separation and Recognition,.

Fusion of Discrete Representations and Self-Augmented Representations for Multilingual Automatic Speech Recognition TokenSplit: Using Discrete Speech Representations for Direct, Refined, and Transcript- Conditioned Speech Separation and Recognition,

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:34:43.387524Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T11:34:43.008752Z digest=sha256:359956ed515aad457179a297232a4e93d915e54aeef0b37d9ebcaed4e56a6e86

Observation 52ebb251-95ee-4eb4-90de-9411abd50b6d · outbound

This paper cites Akvsr: Audio knowledge empowered visual speech recognition by compressing audio knowledge of a pretrained model,.

Fusion of Discrete Representations and Self-Augmented Representations for Multilingual Automatic Speech Recognition Akvsr: Audio knowledge empowered visual speech recognition by compressing audio knowledge of a pretrained model,

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:34:43.372982Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T11:34:43.013179Z digest=sha256:15db5f19eecf8776bbba75f2d907fd62c4bbdebaaf464da2a78eed6306cb3033

Observation f0cbf6cc-5289-4c96-8589-a5dc10c30fe1 · outbound

This paper cites Lip reading for low-resource lan- guages by learning and combining general speech knowl- edge and language-specific knowledge,.

Fusion of Discrete Representations and Self-Augmented Representations for Multilingual Automatic Speech Recognition Lip reading for low-resource lan- guages by learning and combining general speech knowl- edge and language-specific knowledge,

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:34:43.358350Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T11:34:43.017917Z digest=sha256:2d271796f0ddfb55ff81c276873e2f40874cdcec60fb579ba30648ff4fa27160

Observation aa65cded-ebcc-482e-90bc-5f2238d1202e · outbound

This paper cites Exploration of Efficient End-to-End ASR using Discretized Input from Self-Supervised Learning,.

Fusion of Discrete Representations and Self-Augmented Representations for Multilingual Automatic Speech Recognition Exploration of Efficient End-to-End ASR using Discretized Input from Self-Supervised Learning,

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:34:43.344252Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T11:34:43.022835Z digest=sha256:c31bf930a9787315130e26dd8b3738bf1df06c23cc0478bc06fed1bd388109bc

Observation 7e60bd7e-2e94-4034-abc3-7680c251df96 · outbound

This paper cites EFFUSE: Efficient self-supervised fea- ture fusion for e2e asr in multilingual and low resource scenar- ios,.

Fusion of Discrete Representations and Self-Augmented Representations for Multilingual Automatic Speech Recognition EFFUSE: Efficient self-supervised fea- ture fusion for e2e asr in multilingual and low resource scenar- ios,

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:34:43.329653Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T11:34:43.027434Z digest=sha256:156d9ee5c5aff77abd3fab0b91c10bd85e9126e4951186c253f82fa4dd58f22c

Observation ed665449-56b6-4c1c-95b4-45594bde2e13 · outbound

This paper cites FeaRLESS: Feature Refinement Loss for Ensembling Self-Supervised Learning Features in Robust End-to-end Speech Recognition,.

Fusion of Discrete Representations and Self-Augmented Representations for Multilingual Automatic Speech Recognition FeaRLESS: Feature Refinement Loss for Ensembling Self-Supervised Learning Features in Robust End-to-end Speech Recognition,

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:34:43.315058Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T11:34:43.032257Z digest=sha256:18fe81a426e6afa90ca35518207b9833aa5d4166b35b0f2fb981fbcbc5ebd045

Observation 758dd497-1111-4d7e-8ce3-052c1661eb16 · outbound

This paper cites Combining spectral and self-supervised features for low resource speech recognition and translation,.

Fusion of Discrete Representations and Self-Augmented Representations for Multilingual Automatic Speech Recognition Combining spectral and self-supervised features for low resource speech recognition and translation,

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:34:43.300591Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T11:34:43.036753Z digest=sha256:4b915e386b6d93099986b74cfab7ed0878baa2c74d2588a426e1509b5365c72b

Observation c08db803-9d7c-4e75-a2c2-0497a54eee67 · outbound

This paper cites Many-to-many spoken language translation via unified speech and text representation learning with unit- to-unit translation,.

Fusion of Discrete Representations and Self-Augmented Representations for Multilingual Automatic Speech Recognition Many-to-many spoken language translation via unified speech and text representation learning with unit- to-unit translation,

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:34:43.285599Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T11:34:43.041184Z digest=sha256:072075d3af0d15c8d7d94fdfdd84ed66d1a0309534708bc0b98fce795bd23b38

Observation e619195d-d4ef-4add-a9e3-734a93d3ec92 · outbound

This paper cites Intelligible Lip-to-Speech Synthesis with Speech Units,.

Fusion of Discrete Representations and Self-Augmented Representations for Multilingual Automatic Speech Recognition Intelligible Lip-to-Speech Synthesis with Speech Units,

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:34:43.270423Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T11:34:43.045634Z digest=sha256:8aa588e6a25a4335d06add04e91c66de305840aa10c4f00935d6e38862199d41

Observation 18648f34-54b8-402d-b5dc-afcece88096e · outbound

This paper cites Tmt: Tri-modal translation between speech, image, and text by processing different modalities as different languages,.

Fusion of Discrete Representations and Self-Augmented Representations for Multilingual Automatic Speech Recognition Tmt: Tri-modal translation between speech, image, and text by processing different modalities as different languages,

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:34:43.255045Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T11:34:43.050228Z digest=sha256:85f7b2b6bf7e4f9e19373dc1affc6e8fc941ffab5f0af8e4625a2274523b64b0

Observation 190ceaf9-0dd3-47b3-bb07-d98b69d21cd7 · outbound

This paper cites Towards practical and efficient image- to-speech captioning with vision-language pre-training and multi-modal tokens,.

Fusion of Discrete Representations and Self-Augmented Representations for Multilingual Automatic Speech Recognition Towards practical and efficient image- to-speech captioning with vision-language pre-training and multi-modal tokens,

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:34:43.239436Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T11:34:43.054853Z digest=sha256:fd3f7fc193bd8c5bf4f157580b3385d1764472c5c3ce0978db2255f31dc954b8

Observation 210d6f74-d2a7-4b09-92db-e31384f746cd · outbound

This paper cites Librispeech: An asr corpus based on public domain audio books,.

Fusion of Discrete Representations and Self-Augmented Representations for Multilingual Automatic Speech Recognition Librispeech: An asr corpus based on public domain audio books,

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:34:43.222619Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T11:34:43.059407Z digest=sha256:cf6cd5db46bc6de889495db4ccb5be0e82b278a9ebe124c6163077db7d7170a2

Observation 4af76440-2cae-4eaf-b3b3-5917da8e5201 · outbound

This paper cites The interspeech 2024 challenge on speech processing using discrete units,.

Fusion of Discrete Representations and Self-Augmented Representations for Multilingual Automatic Speech Recognition The interspeech 2024 challenge on speech processing using discrete units,

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:34:43.206840Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T11:34:43.063839Z digest=sha256:2344baf2b869ffd28998fad19506b264973114df6d51b04c68a9c1fe099f1bec

Observation e45b46d1-df8a-4939-b472-01b756ae291d · outbound

This paper cites Attention is all you need,.

Fusion of Discrete Representations and Self-Augmented Representations for Multilingual Automatic Speech Recognition Attention is all you need,

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:34:43.191353Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T11:34:43.068873Z digest=sha256:ee6253de1ae2fda92523677259e6c58a87313ff62f395860c9fec4a500885fe1

Observation 66439a34-ee79-42e3-bf6d-e5e31190b3ae · outbound

This paper cites ESPnet: End-to-end speech processing toolkit,.

Fusion of Discrete Representations and Self-Augmented Representations for Multilingual Automatic Speech Recognition ESPnet: End-to-end speech processing toolkit,

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:34:43.175704Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T11:34:43.073261Z digest=sha256:8b0bc8bc266aebc9a8ea6dd7345d11f961698284a75e099854b0cc7f2af9ecd4

Observation 82db58c5-2ddc-4dfa-8f23-43ea7faea4c8 · outbound

This paper cites librosa: Audio and music signal analy- sis in python,.

Fusion of Discrete Representations and Self-Augmented Representations for Multilingual Automatic Speech Recognition librosa: Audio and music signal analy- sis in python,

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:34:43.159399Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T11:34:43.077938Z digest=sha256:472def1cdecc4c5abea7bc573f7c8fe41ad77615e914a04b718a66352c5b1961

Observation e8ff56df-eb2a-424f-a1d6-d3f96d7c5660 · outbound

This paper cites Adam: A Method for Stochastic Optimization.

Fusion of Discrete Representations and Self-Augmented Representations for Multilingual Automatic Speech Recognition Adam: A Method for Stochastic Optimization

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-12T11:34:43.082458Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:34:43.082458Z digest=sha256:e7a09344524adf5289aed420b2355a3b924f8753962553c09602b5cf110b0ac7

Pith citing papers

Observation addbb92a-e430-416c-840f-0a6916870228 · inbound

Fusion of Discrete Representations and Self-Augmented Representations for Multilingual Automatic Speech Recognition cites this paper.

Fusion of Discrete Representations and Self-Augmented Representations for Multilingual Automatic Speech Recognition Fusion of Discrete Representations and Self-Augmented Representations for Multilingual Automatic Speech Recognition

Reference 1

Resolution
verified exact
local_arxiv, observed 2026-08-12T11:34:43.143055Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T11:34:42.898199Z digest=sha256:5a41db7d6e044b4f49fc635d2cc6c49cf53b97155f6e0cb7d2fad80bc9239ce2