Pith. sign in

Paper Citation Record · LEDGER

Fusion of Discrete Representations and Self-Augmented Representations for Multilingual Automatic Speech Recognition

As of 12 August 2026, this Paper Citation Record lists 40 of 40 outbound references and 1 inbound Pith citation observation for arXiv:2411.18107.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2411.18107 v1

Coverage vector

measured 40 of 40 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-12T11:34:43.082458Z

measured 41 of 41 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-12T06:34:41.77262+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-12T11:34:42.898199Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-12T11:34:43.135889Z

Reference resolution

40 of 40 outbound references displayed

  • verified exact1
  • verified fuzzy36
  • unresolved2
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation addbb92a-e430-416c-840f-0a6916870228 · outbound

This paper cites Fusion of Discrete Representations and Self-Augmented Representations for Multilingual Automatic Speech Recognition.

Fusion of Discrete Representations and Self-Augmented Representations for Multilingual Automatic Speech Recognition Fusion of Discrete Representations and Self-Augmented Representations for Multilingual Automatic Speech Recognition

Reference 1

Resolution
verified exact
local_arxiv, observed 2026-08-12T11:34:43.143055Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T11:34:42.898199Z digest=sha256:d19c7508de218fd051c536780d59e13f5f1ccc8266b569f3a858e61fca215635

Observation 7778cee4-a5f6-4558-93e2-96a7bcf9caba · outbound

This paper cites Discretization process Figure 1 provides a high-level overview of our fusion pipeline.

Fusion of Discrete Representations and Self-Augmented Representations for Multilingual Automatic Speech Recognition Discretization process Figure 1 provides a high-level overview of our fusion pipeline

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:34:43.716932Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T11:34:42.904143Z digest=sha256:0d36bb79d4c3e900ea3d57678697c4e36fa47246a7dd2fdee3a0be09647dbcb1

Observation 40d72017-11ef-4591-84a0-9c18c8aca1f1 · outbound

This paper cites an unresolved cited work.

Fusion of Discrete Representations and Self-Augmented Representations for Multilingual Automatic Speech Recognition Unresolved cited work

Reference 3

Resolution
unresolved
raw_fallback, observed 2026-08-12T11:34:43.702238Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T11:34:42.909090Z digest=sha256:79ae703215b1056dedf87e590992d83991d90ea2bd8aa77f81727e79b66054c9

Observation 5c85dc06-3d76-4957-b5a5-be28f69811f9 · outbound

This paper cites Dataset We evaluated the proposed method on LibriSpeech-100h [29] and ML-SUPERB [12].

Fusion of Discrete Representations and Self-Augmented Representations for Multilingual Automatic Speech Recognition Dataset We evaluated the proposed method on LibriSpeech-100h [29] and ML-SUPERB [12]

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:34:43.687931Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T11:34:42.913865Z digest=sha256:35349e411532e465e738fa0b532058d3477a627a8c35ab63705f313fdbece5e4

Observation 41f686be-3a5c-4de4-9a6f-89846c00aa29 · outbound

This paper cites Quantitative result ASR results.

Fusion of Discrete Representations and Self-Augmented Representations for Multilingual Automatic Speech Recognition Quantitative result ASR results

Reference 5

Resolution
malformed identifier
raw_fallback, observed 2026-08-12T11:34:43.673010Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T11:34:42.918993Z digest=sha256:5cc97ba753b6b1a76d1f205abcf69a3d44a1208516be6b9277b7b6a0117c793a

Observation 48d47749-f251-40bf-bc93-c1d89e78286e · outbound

This paper cites self-augmented.

Fusion of Discrete Representations and Self-Augmented Representations for Multilingual Automatic Speech Recognition self-augmented

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:34:43.658034Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T11:34:42.924135Z digest=sha256:5bd64af617f8ddec831456a5d0d9011325fc61d12a88bd1d566c840b819df3f6

Observation 4d42857b-202d-4640-bf32-919a286e4c11 · outbound

This paper cites HuBERT: Self-supervised speech rep- resentation learning by masked prediction of hidden units,.

Fusion of Discrete Representations and Self-Augmented Representations for Multilingual Automatic Speech Recognition HuBERT: Self-supervised speech rep- resentation learning by masked prediction of hidden units,

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:34:43.642511Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T11:34:42.928858Z digest=sha256:534687ee1df71603b28f4f81f07a0a30d1df633591217c73bf6d6137051e8210

Observation 2ccf15b9-a944-4d7e-8834-1aa0bfd65702 · outbound

This paper cites A Survey of Multi- lingual Models for Automatic Speech Recognition,.

Fusion of Discrete Representations and Self-Augmented Representations for Multilingual Automatic Speech Recognition A Survey of Multi- lingual Models for Automatic Speech Recognition,

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:34:43.627743Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T11:34:42.933675Z digest=sha256:741332c1b7421de9f6a7dbd40d79223738eaa201a2fc3e686213539a0b0d231b

Observation f9ded218-8b25-4537-b6d8-4f548066e6af · outbound

This paper cites Cross-lingual Automatic Speech Recog- nition Exploiting Articulatory Features,.

Fusion of Discrete Representations and Self-Augmented Representations for Multilingual Automatic Speech Recognition Cross-lingual Automatic Speech Recog- nition Exploiting Articulatory Features,

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:34:43.612015Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T11:34:42.938373Z digest=sha256:2cd135f671b0901cd63b2955f961bd2c44351b0ca60807eae400973ddba2d5c4

Observation 2431b646-89c6-4a6f-a99b-070af72cca2b · outbound

This paper cites wav2vec 2.0: A framework for self- supervised learning of speech representations,.

Fusion of Discrete Representations and Self-Augmented Representations for Multilingual Automatic Speech Recognition wav2vec 2.0: A framework for self- supervised learning of speech representations,

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:34:43.596629Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T11:34:42.943156Z digest=sha256:366f9067e3ae0f0687a90f43cc6e90c64ab3b778e1e4c62d226aa8ef554f7ef6

Observation 95077b77-3dfd-4519-8262-985d41698746 · outbound

This paper cites XLS-R: Self-supervised Cross-lingual Speech Representation Learning at Scale,.

Fusion of Discrete Representations and Self-Augmented Representations for Multilingual Automatic Speech Recognition XLS-R: Self-supervised Cross-lingual Speech Representation Learning at Scale,

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:34:43.582162Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T11:34:42.948404Z digest=sha256:03e99deab022e6114f46c0c3be07711ab6f1ce827580505a8e94a407c59154c6

Observation 347b95c1-9056-47b1-97dd-117cf5aef508 · outbound

This paper cites W2v-BERT: Combining contrastive learning and masked language modeling for self-supervised speech pre-training,.

Fusion of Discrete Representations and Self-Augmented Representations for Multilingual Automatic Speech Recognition W2v-BERT: Combining contrastive learning and masked language modeling for self-supervised speech pre-training,

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:34:43.567358Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T11:34:42.952834Z digest=sha256:ca42247d66aaa1f14c6e79c2ca363b3c380f9a8cc36aca82014f94f849afaa75

Observation c34e7edf-e82d-4c8d-aa10-45eaae192b7a · outbound

This paper cites Wavlm: Large-scale self-supervised pre- training for full stack speech processing,.

Fusion of Discrete Representations and Self-Augmented Representations for Multilingual Automatic Speech Recognition Wavlm: Large-scale self-supervised pre- training for full stack speech processing,

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:34:43.552377Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T11:34:42.958911Z digest=sha256:247ffc1254dd0934856b76544baa58d4d44b83ccd79c808c868b2cbf1fe5d610

Observation 95714afb-2bb4-4efc-8a4c-6bc0841a6e58 · outbound

This paper cites Self-supervised speech repre- sentation learning: A review,.

Fusion of Discrete Representations and Self-Augmented Representations for Multilingual Automatic Speech Recognition Self-supervised speech repre- sentation learning: A review,

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:34:43.537588Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T11:34:42.963844Z digest=sha256:d98c6ca4452cc267aaa110d87a40925cf39682b4ee8315f58fb81b63b1617211

Observation 9dfbe30f-78d1-40f9-a069-5ec9779f8083 · outbound

This paper cites Multi-resolution huBERT: Multi-resolution speech self-supervised learning with masked unit prediction,.

Fusion of Discrete Representations and Self-Augmented Representations for Multilingual Automatic Speech Recognition Multi-resolution huBERT: Multi-resolution speech self-supervised learning with masked unit prediction,

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:34:43.522813Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T11:34:42.968484Z digest=sha256:94624547f550292989cc0fc32ada13ad2d12b2c2cff461036d88a145cbcbd344

Observation a945e848-6a0d-49b3-b28d-ba67249f793f · outbound

This paper cites Scaling speech technology to 1,000+ languages,.

Fusion of Discrete Representations and Self-Augmented Representations for Multilingual Automatic Speech Recognition Scaling speech technology to 1,000+ languages,

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:34:43.508388Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T11:34:42.973148Z digest=sha256:fa818633865fe423e29d5232312038e673d250824b88efb3d3c276897be46ab5

Observation 812838b2-3ebd-496e-8fd9-71e18ce10da1 · outbound

This paper cites SUPERB: Speech Processing Universal PERformance Benchmark,.

Fusion of Discrete Representations and Self-Augmented Representations for Multilingual Automatic Speech Recognition SUPERB: Speech Processing Universal PERformance Benchmark,

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:34:43.493670Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T11:34:42.977682Z digest=sha256:eddf2696febe42a9ebc24763292eb1264bd0d9c1a51bcdaf8b77f1b5cd642ec3

Observation ab5e4686-c59d-4fde-b42c-18f259eb17ce · outbound

This paper cites ML-SUPERB: Multilingual Speech Uni- versal PERformance Benchmark,.

Fusion of Discrete Representations and Self-Augmented Representations for Multilingual Automatic Speech Recognition ML-SUPERB: Multilingual Speech Uni- versal PERformance Benchmark,

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:34:43.478182Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T11:34:42.982163Z digest=sha256:10093ce0f0ea02977f84af5de4928621fda666b4b3cacbe6e708ba08892fb860

Observation f5bcc383-6ce3-4b16-83ab-0532e19ffd77 · outbound

This paper cites Exploring speech recognition, transla- tion, and understanding with discrete speech units: A compar- ative study,.

Fusion of Discrete Representations and Self-Augmented Representations for Multilingual Automatic Speech Recognition Exploring speech recognition, transla- tion, and understanding with discrete speech units: A compar- ative study,

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:34:43.462846Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T11:34:42.986449Z digest=sha256:8a1fcf1cd2b0d2a1e3e65c4500010598a3d0328d77b57350c2c73abc6b88037b

Observation 5d281b8c-764a-42e6-8108-d988429f4f0d · outbound

This paper cites An Exploration of Prompt Tuning on Generative Spoken Language Model for Speech Processing Tasks,.

Fusion of Discrete Representations and Self-Augmented Representations for Multilingual Automatic Speech Recognition An Exploration of Prompt Tuning on Generative Spoken Language Model for Speech Processing Tasks,

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:34:43.447264Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T11:34:42.991048Z digest=sha256:96d449225b87fa5852a1be9f95883d3b6f1d7033d2064e69885d9da161a1fa05

Observation 5b73ff31-a19a-476a-9370-8edcc61265af · outbound

This paper cites Towards universal speech discrete tokens: A case study for asr and tts,.

Fusion of Discrete Representations and Self-Augmented Representations for Multilingual Automatic Speech Recognition Towards universal speech discrete tokens: A case study for asr and tts,

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:34:43.432134Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T11:34:42.995775Z digest=sha256:4b5d9d4475bed8cb285c761bc80f2f7975e9d167430f5e6b93fed64ffacad7a9

Observation 7b89fbaf-c25d-4936-bbf2-50ce11887f5a · outbound

This paper cites Acoustic bpe for speech generation with dis- crete tokens,.

Fusion of Discrete Representations and Self-Augmented Representations for Multilingual Automatic Speech Recognition Acoustic bpe for speech generation with dis- crete tokens,

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:34:43.417656Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T11:34:43.000017Z digest=sha256:439da67f1b2473f678ad501ee41393752a619d1a7a3ec982972ac0f5dd9323ab

Observation a70f3d72-a00d-41c6-b89e-914edc1694e8 · outbound

This paper cites V oxtlm: Unified decoder-only models for consolidating speech recognition, synthesis and speech, text continuation tasks,.

Fusion of Discrete Representations and Self-Augmented Representations for Multilingual Automatic Speech Recognition V oxtlm: Unified decoder-only models for consolidating speech recognition, synthesis and speech, text continuation tasks,

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:34:43.402217Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T11:34:43.004358Z digest=sha256:7a2ec259441d4b9b5bcd2e55b36b64691e9e5c1ab2948961a1f808c64852417e

Observation 66479bd1-a262-479e-aa8c-f980edee6fd6 · outbound

This paper cites TokenSplit: Using Discrete Speech Representations for Direct, Refined, and Transcript- Conditioned Speech Separation and Recognition,.

Fusion of Discrete Representations and Self-Augmented Representations for Multilingual Automatic Speech Recognition TokenSplit: Using Discrete Speech Representations for Direct, Refined, and Transcript- Conditioned Speech Separation and Recognition,

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:34:43.387524Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T11:34:43.008752Z digest=sha256:cfdc6c3758eee3358dd8c2a7419204103380896a03a2105103a27772ff3d469b

Observation 52ebb251-95ee-4eb4-90de-9411abd50b6d · outbound

This paper cites Akvsr: Audio knowledge empowered visual speech recognition by compressing audio knowledge of a pretrained model,.

Fusion of Discrete Representations and Self-Augmented Representations for Multilingual Automatic Speech Recognition Akvsr: Audio knowledge empowered visual speech recognition by compressing audio knowledge of a pretrained model,

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:34:43.372982Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T11:34:43.013179Z digest=sha256:03267d522c38358a2e024853eda39a03ebd9787c205167b06472c7913748f402

Observation f0cbf6cc-5289-4c96-8589-a5dc10c30fe1 · outbound

This paper cites Lip reading for low-resource lan- guages by learning and combining general speech knowl- edge and language-specific knowledge,.

Fusion of Discrete Representations and Self-Augmented Representations for Multilingual Automatic Speech Recognition Lip reading for low-resource lan- guages by learning and combining general speech knowl- edge and language-specific knowledge,

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:34:43.358350Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T11:34:43.017917Z digest=sha256:cf3b28f2e960d90b5718e93f85ea836c3c836d0231b8bee81a9d9943bea53988

Observation aa65cded-ebcc-482e-90bc-5f2238d1202e · outbound

This paper cites Exploration of Efficient End-to-End ASR using Discretized Input from Self-Supervised Learning,.

Fusion of Discrete Representations and Self-Augmented Representations for Multilingual Automatic Speech Recognition Exploration of Efficient End-to-End ASR using Discretized Input from Self-Supervised Learning,

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:34:43.344252Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T11:34:43.022835Z digest=sha256:27e2b923a044765e38410b3c74dcbe28350a2260adace351a13060a1275d84cc

Observation 7e60bd7e-2e94-4034-abc3-7680c251df96 · outbound

This paper cites EFFUSE: Efficient self-supervised fea- ture fusion for e2e asr in multilingual and low resource scenar- ios,.

Fusion of Discrete Representations and Self-Augmented Representations for Multilingual Automatic Speech Recognition EFFUSE: Efficient self-supervised fea- ture fusion for e2e asr in multilingual and low resource scenar- ios,

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:34:43.329653Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T11:34:43.027434Z digest=sha256:321bee6f828d18848b2831b8a8d991f5faa09f97ab217b4152346476ba2ab33d

Observation ed665449-56b6-4c1c-95b4-45594bde2e13 · outbound

This paper cites FeaRLESS: Feature Refinement Loss for Ensembling Self-Supervised Learning Features in Robust End-to-end Speech Recognition,.

Fusion of Discrete Representations and Self-Augmented Representations for Multilingual Automatic Speech Recognition FeaRLESS: Feature Refinement Loss for Ensembling Self-Supervised Learning Features in Robust End-to-end Speech Recognition,

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:34:43.315058Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T11:34:43.032257Z digest=sha256:cf26d82ffc05c50f134395f608482603a64b76df0017652903d3841fc1f4f49e

Observation 758dd497-1111-4d7e-8ce3-052c1661eb16 · outbound

This paper cites Combining spectral and self-supervised features for low resource speech recognition and translation,.

Fusion of Discrete Representations and Self-Augmented Representations for Multilingual Automatic Speech Recognition Combining spectral and self-supervised features for low resource speech recognition and translation,

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:34:43.300591Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T11:34:43.036753Z digest=sha256:d944e98778da9666ee4f83d0f900bcde165d651204adde533d2c125ab586629a

Observation c08db803-9d7c-4e75-a2c2-0497a54eee67 · outbound

This paper cites Many-to-many spoken language translation via unified speech and text representation learning with unit- to-unit translation,.

Fusion of Discrete Representations and Self-Augmented Representations for Multilingual Automatic Speech Recognition Many-to-many spoken language translation via unified speech and text representation learning with unit- to-unit translation,

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:34:43.285599Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T11:34:43.041184Z digest=sha256:e1718097e4da913a880b49bc17ed6b2f8a384f817a5b129d7544c40cda4eddee

Observation e619195d-d4ef-4add-a9e3-734a93d3ec92 · outbound

This paper cites Intelligible Lip-to-Speech Synthesis with Speech Units,.

Fusion of Discrete Representations and Self-Augmented Representations for Multilingual Automatic Speech Recognition Intelligible Lip-to-Speech Synthesis with Speech Units,

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:34:43.270423Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T11:34:43.045634Z digest=sha256:ca80993bd3f4588349cad17277afe7970abc99c88c8fa395ce67fae504a19ede

Observation 18648f34-54b8-402d-b5dc-afcece88096e · outbound

This paper cites Tmt: Tri-modal translation between speech, image, and text by processing different modalities as different languages,.

Fusion of Discrete Representations and Self-Augmented Representations for Multilingual Automatic Speech Recognition Tmt: Tri-modal translation between speech, image, and text by processing different modalities as different languages,

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:34:43.255045Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T11:34:43.050228Z digest=sha256:61abd0471ed42504dc4b6c9119fa6cf69763913a2824b93f0499d1e8b1ecbd6b

Observation 190ceaf9-0dd3-47b3-bb07-d98b69d21cd7 · outbound

This paper cites Towards practical and efficient image- to-speech captioning with vision-language pre-training and multi-modal tokens,.

Fusion of Discrete Representations and Self-Augmented Representations for Multilingual Automatic Speech Recognition Towards practical and efficient image- to-speech captioning with vision-language pre-training and multi-modal tokens,

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:34:43.239436Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T11:34:43.054853Z digest=sha256:5fdfbc390877b7d17e426bad1550f4daec3e3b66baa8f4b3c127e8635e408749

Observation 210d6f74-d2a7-4b09-92db-e31384f746cd · outbound

This paper cites Librispeech: An asr corpus based on public domain audio books,.

Fusion of Discrete Representations and Self-Augmented Representations for Multilingual Automatic Speech Recognition Librispeech: An asr corpus based on public domain audio books,

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:34:43.222619Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T11:34:43.059407Z digest=sha256:2fa80bf33e0d6770524f8753c5324cb8d2cdf67fc3d0de74fdc304a2f8f55440

Observation 4af76440-2cae-4eaf-b3b3-5917da8e5201 · outbound

This paper cites The interspeech 2024 challenge on speech processing using discrete units,.

Fusion of Discrete Representations and Self-Augmented Representations for Multilingual Automatic Speech Recognition The interspeech 2024 challenge on speech processing using discrete units,

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:34:43.206840Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T11:34:43.063839Z digest=sha256:14f3801df993b428ed3fdd84b669ca58da2047051c8fa2052c92ed4694f2e5e6

Observation e45b46d1-df8a-4939-b472-01b756ae291d · outbound

This paper cites Attention is all you need,.

Fusion of Discrete Representations and Self-Augmented Representations for Multilingual Automatic Speech Recognition Attention is all you need,

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:34:43.191353Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T11:34:43.068873Z digest=sha256:c006a60b5d6e02e25fec5be78af50d35c764c48f905ca09ad151642e68a93062

Observation 66439a34-ee79-42e3-bf6d-e5e31190b3ae · outbound

This paper cites ESPnet: End-to-end speech processing toolkit,.

Fusion of Discrete Representations and Self-Augmented Representations for Multilingual Automatic Speech Recognition ESPnet: End-to-end speech processing toolkit,

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:34:43.175704Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T11:34:43.073261Z digest=sha256:7e6b8491ed5fe1acf539d55bcca7592cf6128436926e00d86d31a40290a0d9c6

Observation 82db58c5-2ddc-4dfa-8f23-43ea7faea4c8 · outbound

This paper cites librosa: Audio and music signal analy- sis in python,.

Fusion of Discrete Representations and Self-Augmented Representations for Multilingual Automatic Speech Recognition librosa: Audio and music signal analy- sis in python,

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:34:43.159399Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T11:34:43.077938Z digest=sha256:77d473d7afb75bbdbd11e2427692d4f3dcba426a5fcc09bc290ca0879d519dec

Observation e8ff56df-eb2a-424f-a1d6-d3f96d7c5660 · outbound

This paper cites Adam: A Method for Stochastic Optimization.

Fusion of Discrete Representations and Self-Augmented Representations for Multilingual Automatic Speech Recognition Adam: A Method for Stochastic Optimization

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-12T11:34:43.082458Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:34:43.082458Z digest=sha256:e7a09344524adf5289aed420b2355a3b924f8753962553c09602b5cf110b0ac7

Pith citing papers

Observation addbb92a-e430-416c-840f-0a6916870228 · inbound

Fusion of Discrete Representations and Self-Augmented Representations for Multilingual Automatic Speech Recognition cites this paper.

Fusion of Discrete Representations and Self-Augmented Representations for Multilingual Automatic Speech Recognition Fusion of Discrete Representations and Self-Augmented Representations for Multilingual Automatic Speech Recognition

Reference 1

Resolution
verified exact
local_arxiv, observed 2026-08-12T11:34:43.143055Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T11:34:42.898199Z digest=sha256:d19c7508de218fd051c536780d59e13f5f1ccc8266b569f3a858e61fca215635