Pith. sign in

Paper Citation Record · LEDGER

Conformer: Convolution-augmented Transformer for Speech Recognition

As of 16 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 100 inbound Pith citation observations for arXiv:2005.08100.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2005.08100 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 100 of 100 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-16T06:30:59.297886+00:00

measured 100 of 125 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-16T12:38:53.919600Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-05T02:28:24.338817Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

385
pith, observed 2026-08-05T02:28:24.338817Z

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 04016ebf-6b02-4846-90c0-323f59c02367 · inbound

Scaling Autoregressive Models for Content-Rich Text-to-Image Generation cites this paper.

Scaling Autoregressive Models for Content-Rich Text-to-Image Generation Conformer: Convolution-augmented Transformer for Speech Recognition

Reference 23

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T04:49:31.123367Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-12T04:49:30.873360Z digest=sha256:ecbcec8e70437daf2a14916804099f6b0dca1da98c31aa77a3945ac769be1a9c

Observation 5f2ea2a7-c11f-4bf2-bdcf-26290f640ceb · inbound

FlashAttention-3: Fast and Accurate Attention with Asynchrony and Low-precision cites this paper.

FlashAttention-3: Fast and Accurate Attention with Asynchrony and Low-precision Conformer: Convolution-augmented Transformer for Speech Recognition

Reference 23

Resolution
metadata mismatch
arxiv_id, observed 2026-05-20T19:45:36.424652Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-20T19:45:36.337956Z digest=sha256:28ebdef37e9a2dd2f8a722f7fcca096c835f1daa99f8d991de7fb3057f25badf

Observation cfab2fda-ac46-43bf-980d-589f7f3de15c · inbound

An End-To-End Stuttering Detection Method Based On Conformer And BILSTM cites this paper.

An End-To-End Stuttering Detection Method Based On Conformer And BILSTM Conformer: Convolution-augmented Transformer for Speech Recognition

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-12T20:38:27.297576Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T20:38:27.297576Z digest=sha256:721ad8e93b7d627c631a49396978623af876a3013aaccb3a186fb1c8ab8325ed

Observation 98a14354-bb3c-4ea4-9491-3ce727bcc453 · inbound

Brain-to-Text Decoding with Context-Aware Neural Representations and Large Language Models cites this paper.

Brain-to-Text Decoding with Context-Aware Neural Representations and Large Language Models Conformer: Convolution-augmented Transformer for Speech Recognition

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-12T19:31:53.147344Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T19:31:53.147344Z digest=sha256:4dde4e0dadcfeec12d8cb7ab40a61d420104f8c455f814b1dbffc72bb0074ba9

Observation ce81ee95-2193-4b00-a4f9-27085fd51afe · inbound

DGSNA: Dynamic Generative Scene-based Noise Addition method cites this paper.

DGSNA: Dynamic Generative Scene-based Noise Addition method Conformer: Convolution-augmented Transformer for Speech Recognition

Reference 52

Resolution
verified exact
arxiv_id, observed 2026-05-23T17:45:46.264977Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-23T17:43:47.086524Z digest=sha256:179031ae3ba9b34a2d8a22ab9ac02fa1ae03f1bbf369e6604b9f02eabd857790

Observation 6a70d22d-6fec-4429-8511-d2bc765167c3 · inbound

TPCNet: Representation learning for HI mapping cites this paper.

TPCNet: Representation learning for HI mapping Conformer: Convolution-augmented Transformer for Speech Recognition

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-12T16:39:03.635883Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T16:39:03.635883Z digest=sha256:1ca7494b60a52660befc12248c68c4554ed9ca65480ceb3aa52a1f477a836767

Observation e25f21ab-818f-4d35-ba98-6f46038e3b9f · inbound

Tiny-Align: Bridging Automatic Speech Recognition and Large Language Model on the Edge cites this paper.

Tiny-Align: Bridging Automatic Speech Recognition and Large Language Model on the Edge Conformer: Convolution-augmented Transformer for Speech Recognition

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-12T15:58:19.912750Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:58:19.912750Z digest=sha256:9ae04202ed983d08ed798f76dc17afa7187d11c8156311ad3251724bc8175f37

Observation 4f747a62-0642-4d58-9448-abbc1ee71cb4 · inbound

X-CrossNet: A complex spectral mapping approach to target speaker extraction with cross attention speaker embedding fusion cites this paper.

X-CrossNet: A complex spectral mapping approach to target speaker extraction with cross attention speaker embedding fusion Conformer: Convolution-augmented Transformer for Speech Recognition

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-12T15:53:11.092180Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:53:11.092180Z digest=sha256:3b325dc7caf86b613425421d60a70270f94b77e86e8125bff083b39a00467409

Observation 72c49616-a6f7-4f0c-bab0-19185faae6fd · inbound

CAIP: Detecting Router Misconfigurations with Context-Aware Iterative Prompting of LLMs cites this paper.

CAIP: Detecting Router Misconfigurations with Context-Aware Iterative Prompting of LLMs Conformer: Convolution-augmented Transformer for Speech Recognition

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-12T15:24:37.520943Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:24:37.520943Z digest=sha256:0af158962a4242ee3c2676489c9ed2116d10bb4c9953eb6d1f6e478797ff69db

Observation de702da3-e804-42f5-9342-b6696ed0b4fb · inbound

Simplifying CLIP: Unleashing the Power of Large-Scale Models on Consumer-level Computers cites this paper.

Simplifying CLIP: Unleashing the Power of Large-Scale Models on Consumer-level Computers Conformer: Convolution-augmented Transformer for Speech Recognition

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-12T14:59:52.219465Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:59:52.219465Z digest=sha256:d20610b9c69144e09f1cb10b284a34e9e35c64946e7fdb1735831b3878d914d3

Observation 19a5796f-f185-47c5-899f-c9f5460e71e7 · inbound

The SVASR System for Text-dependent Speaker Verification (TdSV) AAIC Challenge 2024 cites this paper.

The SVASR System for Text-dependent Speaker Verification (TdSV) AAIC Challenge 2024 Conformer: Convolution-augmented Transformer for Speech Recognition

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-12T13:27:06.802778Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T13:27:06.802778Z digest=sha256:fde9401221a96947775df9d3e29186a3071eba9206c233178fd37043ef78c4d3

Observation 97c55f22-a88e-4c00-a1d6-689ffa66afb3 · inbound

Sample adaptive data augmentation with progressive scheduling cites this paper.

Sample adaptive data augmentation with progressive scheduling Conformer: Convolution-augmented Transformer for Speech Recognition

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-12T05:27:56.712197Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T05:27:56.712197Z digest=sha256:ff84170095ceca39242668a484695fd2d4d4b8220878af570f0a1d0cc1be5cf2

Observation 86ac59d4-dba7-4676-93ca-6e7923d2ae11 · inbound

From Audio Deepfake Detection to AI-Generated Music Detection -- A Pathway and Overview cites this paper.

From Audio Deepfake Detection to AI-Generated Music Detection -- A Pathway and Overview Conformer: Convolution-augmented Transformer for Speech Recognition

Reference 133

Resolution
unresolved
no resolver link, observed 2026-08-12T05:15:07.718194Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T05:15:07.718194Z digest=sha256:eaf0ce768b570d57cd35f02893fa0ac73088b98d5fe74ab306ae1d08dc777fe4

Observation f7d8372a-6714-47ac-b3f5-77c74528a2ee · inbound

Complexity boosted adaptive training for better low resource ASR performance cites this paper.

Complexity boosted adaptive training for better low resource ASR performance Conformer: Convolution-augmented Transformer for Speech Recognition

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-12T04:59:13.263015Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:59:13.263015Z digest=sha256:f21bb94c8c405184c89370ab9e72d81043f0292095166c44fc79c448852197c6

Observation 1b6aa9a6-81ba-4226-811d-ba94547946eb · inbound

Privacy-Preserving Gesture Tracking System Utilizing Frequency-Hopping RFID Signals cites this paper.

Privacy-Preserving Gesture Tracking System Utilizing Frequency-Hopping RFID Signals Conformer: Convolution-augmented Transformer for Speech Recognition

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-11T21:55:58.410553Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:55:58.410553Z digest=sha256:5363cf62a404b5a36b3c277e0a0de6d3e9ab8ac020e5cc4dbd311c0973668a39

Observation 468ba433-b7d5-4625-9837-42621a8a68e5 · inbound

WhisperFlow: speech foundation models in real time cites this paper.

WhisperFlow: speech foundation models in real time Conformer: Convolution-augmented Transformer for Speech Recognition

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-11T15:11:57.788464Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:11:57.788464Z digest=sha256:5562aa89227713aea5bfb16e552d568d500222de915062e58070af001a2f487c

Observation 5f9036c2-d73d-4c88-a99c-b9688de14ed4 · inbound

Leveraging User-Generated Metadata of Online Videos for Cover Song Identification cites this paper.

Leveraging User-Generated Metadata of Online Videos for Cover Song Identification Conformer: Convolution-augmented Transformer for Speech Recognition

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-11T14:36:52.547219Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T14:36:52.547219Z digest=sha256:43921e68693b687adeffff8d4a0a1ad0c54091f7bf5eb47a642667780e60e884

Observation b983785e-31dd-4196-8fb4-43bf11a9bc9b · inbound

Efficient Speech Command Recognition Leveraging Spiking Neural Network and Curriculum Learning-based Knowledge Distillation cites this paper.

Efficient Speech Command Recognition Leveraging Spiking Neural Network and Curriculum Learning-based Knowledge Distillation Conformer: Convolution-augmented Transformer for Speech Recognition

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-11T13:44:44.021314Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T13:44:44.021314Z digest=sha256:fe657707abc85a19f4d5b52b6f380c79c3caa33e688fce3f255e278e43cdff75

Observation f4494061-488f-40dd-81dd-7e3c85f032c4 · inbound

Model Decides How to Tokenize: Adaptive DNA Sequence Tokenization with MxDNA cites this paper.

Model Decides How to Tokenize: Adaptive DNA Sequence Tokenization with MxDNA Conformer: Convolution-augmented Transformer for Speech Recognition

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-11T12:55:40.510711Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:55:40.510711Z digest=sha256:ccf49e56023b69becac245ce0a2de0db0784cb1f114ce74852c58ac042954057

Observation 3b012534-5e9b-4b99-9a68-c9ae19e56324 · inbound

Open Universal Arabic ASR Leaderboard cites this paper.

Open Universal Arabic ASR Leaderboard Conformer: Convolution-augmented Transformer for Speech Recognition

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-11T12:51:45.534532Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:51:45.534532Z digest=sha256:9f059e9e0d02e68549df679d41739803396b56be485bb659ee65a2ef1dab015e

Observation c29e5446-e879-48c3-ae36-b9cf25835a2d · inbound

A Decade of Deep Learning: A Survey on The Magnificent Seven cites this paper.

A Decade of Deep Learning: A Survey on The Magnificent Seven Conformer: Convolution-augmented Transformer for Speech Recognition

Reference 2015

Resolution
unresolved
no resolver link, observed 2026-08-11T16:10:24.412313Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T16:10:24.412313Z digest=sha256:3b46244470e1d880836085964fdf4f07a1326b371419f78ec2447c87ef82b36b

Observation 89d46f2c-2d74-4a29-b42c-dbd73854f43c · inbound

Unity is Strength: Unifying Convolutional and Transformeral Features for Better Person Re-Identification cites this paper.

Unity is Strength: Unifying Convolutional and Transformeral Features for Better Person Re-Identification Conformer: Convolution-augmented Transformer for Speech Recognition

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-11T05:45:17.493612Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T05:45:17.493612Z digest=sha256:ef16e2187558ee46989e22efb6aa8b1d6da75e7075a8aa2a82c8e0de8c449fd1

Observation 6e43c16b-898b-4600-99cb-bc3ae141eba5 · inbound

An Attention-based Framework with Multistation Information for Earthquake Early Warnings cites this paper.

An Attention-based Framework with Multistation Information for Earthquake Early Warnings Conformer: Convolution-augmented Transformer for Speech Recognition

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-11T05:06:43.616829Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T05:06:43.616829Z digest=sha256:d9f34ab6a9bf9ad0ae54eb0b8fa86f44739f7616f1ce96358080bbc0f808d25c

Observation 86d44d48-9026-4a54-aa1e-5988254c7054 · inbound

Next Token Prediction Towards Multimodal Intelligence: A Comprehensive Survey cites this paper.

Next Token Prediction Towards Multimodal Intelligence: A Comprehensive Survey Conformer: Convolution-augmented Transformer for Speech Recognition

Reference 141

Resolution
unresolved
no resolver link, observed 2026-08-11T14:59:01.753245Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T14:59:01.753245Z digest=sha256:41d347928585c282779d76b48d97234e4874f9bb3fc124fb8adc73951cf3a047

Observation adeb4dd4-d4c7-412c-a215-0322aac9d917 · inbound

MuQ: Self-Supervised Music Representation Learning with Mel Residual Vector Quantization cites this paper.

MuQ: Self-Supervised Music Representation Learning with Mel Residual Vector Quantization Conformer: Convolution-augmented Transformer for Speech Recognition

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-10T22:40:22.034251Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T22:40:22.034251Z digest=sha256:cf6cb796759501708faaea6de6edcfc46308c513dded514b04f725e0b7388562

Observation d2223582-9090-4ad8-b19c-fc326d1abe6f · inbound

On the Robustness of Cover Version Identification Models: A Study Using Cover Versions from YouTube cites this paper.

On the Robustness of Cover Version Identification Models: A Study Using Cover Versions from YouTube Conformer: Convolution-augmented Transformer for Speech Recognition

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-10T22:35:19.068918Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:35:19.068918Z digest=sha256:0e2c937597c86e59a9ab12b0b01d66619e9a055cdce34f3e3a1d9f9984dc7be7

Observation dba533ba-213b-44b6-b407-465ed4750d52 · inbound

A Non-autoregressive Model for Joint STT and TTS cites this paper.

A Non-autoregressive Model for Joint STT and TTS Conformer: Convolution-augmented Transformer for Speech Recognition

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-10T20:15:28.480440Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:15:28.480440Z digest=sha256:24fe6a4ead66b48dd68a79dd1753db71d8dd3f1ebad779373981aa987b2fc3d8

Observation 478bd955-ecd5-49f8-bc58-12503b7b8c1a · inbound

Let SSMs be ConvNets: State-space Modeling with Optimal Tensor Contractions cites this paper.

Let SSMs be ConvNets: State-space Modeling with Optimal Tensor Contractions Conformer: Convolution-augmented Transformer for Speech Recognition

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-10T16:26:51.213613Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T16:26:51.213613Z digest=sha256:0b46d0e3bd8cfd8661686319611279e91d3d456c2bfe045d3a592177d000714d

Observation 2810c41d-1183-413b-a937-fa3ecd9ce378 · inbound

Efficient Audiovisual Speech Processing via MUTUD: Multimodal Training and Unimodal Deployment cites this paper.

Efficient Audiovisual Speech Processing via MUTUD: Multimodal Training and Unimodal Deployment Conformer: Convolution-augmented Transformer for Speech Recognition

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-10T00:35:49.015294Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T00:35:49.015294Z digest=sha256:14c44470504d178ebc53232745311091f36e3aab96c95d207267b44539f5d614

Observation 08e73f46-7912-462b-97c7-b3159b4031ef · inbound

Aligner-Encoders: Self-Attention Transformers Can Be Self-Transducers cites this paper.

Aligner-Encoders: Self-Attention Transformers Can Be Self-Transducers Conformer: Convolution-augmented Transformer for Speech Recognition

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-08T22:31:38.784878Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T22:31:38.784878Z digest=sha256:2d9ba35fa1cba16ff580e2d3da760aa6ca4efb07406f5baf2d703d88870d0917

Observation 2812bdb9-7c65-4922-b502-a03c0395e5fd · inbound

Synergistic Effects of Knowledge Distillation and Structured Pruning for Self-Supervised Speech Models cites this paper.

Synergistic Effects of Knowledge Distillation and Structured Pruning for Self-Supervised Speech Models Conformer: Convolution-augmented Transformer for Speech Recognition

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-08T17:49:20.093965Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T17:49:20.093965Z digest=sha256:634815940b17c23e6e306f9eae72caefc25046c5088b79ea0bd25808ebfae78f

Observation 8795af33-1621-412d-bb13-fb6533b37fbb · inbound

Advancing Arabic Speech Recognition Through Large-Scale Weakly Supervised Learning cites this paper.

Advancing Arabic Speech Recognition Through Large-Scale Weakly Supervised Learning Conformer: Convolution-augmented Transformer for Speech Recognition

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-16T12:38:53.919600Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T12:38:53.919600Z digest=sha256:3a0cc8d6aeec683e66a1e33d2e9109f2b8165daacd66903458a103eaacacdb12

Observation c06933bf-c638-48a4-bc5d-18da64b3c0dd · inbound

Pretraining Large Brain Language Model for Active BCI: Silent Speech cites this paper.

Pretraining Large Brain Language Model for Active BCI: Silent Speech Conformer: Convolution-augmented Transformer for Speech Recognition

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-16T05:14:57.770317Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:14:57.770317Z digest=sha256:6fb1e98ebc52fb3db3bd25e62a719e70d403a219164aada20b3ef701858cce80

Observation f1e85b49-285f-41be-aedc-04332efe825a · inbound

Remote Rowhammer Attack using Adversarial Observations on Federated Learning Clients cites this paper.

Remote Rowhammer Attack using Adversarial Observations on Federated Learning Clients Conformer: Convolution-augmented Transformer for Speech Recognition

Reference 37

Resolution
metadata mismatch
arxiv_id, observed 2026-05-22T15:24:57.799931Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-22T15:22:59.827762Z digest=sha256:ff131fbf062c5e3999b043d17431a333f9595bab1c9a825bb6d2edf60ce10427

Observation e0acadb9-521d-4ae5-ad88-dff1dd8f2fad · inbound

Survey of End-to-End Multi-Speaker Automatic Speech Recognition for Monaural Audio cites this paper.

Survey of End-to-End Multi-Speaker Automatic Speech Recognition for Monaural Audio Conformer: Convolution-augmented Transformer for Speech Recognition

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-15T21:03:18.003997Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:03:18.003997Z digest=sha256:3222da00650800cd652092574dabc1248091ab4464ae7395afaff80f77e64377

Observation 4c8d1bc7-1d2d-41b1-8955-627e553c55c1 · inbound

The Computation of Generalized Embeddings for Underwater Acoustic Target Recognition using Contrastive Learning cites this paper.

The Computation of Generalized Embeddings for Underwater Acoustic Target Recognition using Contrastive Learning Conformer: Convolution-augmented Transformer for Speech Recognition

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-15T20:28:07.078243Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:28:07.078243Z digest=sha256:0213378f28b9b9c10eace4d3977bbf1fe68c3389fcb5436bcbd890f4e2790d5b

Observation 75484c64-f2a6-4f31-9cc0-97337f4de5e5 · inbound

Calm-Whisper: Reduce Whisper Hallucination On Non-Speech By Calming Crazy Heads Down cites this paper.

Calm-Whisper: Reduce Whisper Hallucination On Non-Speech By Calming Crazy Heads Down Conformer: Convolution-augmented Transformer for Speech Recognition

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-15T20:27:12.004662Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:27:12.004662Z digest=sha256:0ec64e6cb70915f230138e72ff746df53a653895a087def3ea7b01fdb9d5e11a

Observation 7ed58234-60da-43f8-ad8a-aafdf344953b · inbound

Cross-modal Knowledge Transfer Learning as Graph Matching Based on Optimal Transport for ASR cites this paper.

Cross-modal Knowledge Transfer Learning as Graph Matching Based on Optimal Transport for ASR Conformer: Convolution-augmented Transformer for Speech Recognition

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-15T20:26:15.047857Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:26:15.047857Z digest=sha256:8e21a0721049bc8d10f817c17425097aaa2a71579195666df80b73de58cfb0ad

Observation 9cb69022-47e4-49ed-ad31-9af27db019a2 · inbound

ClapFM-EVC: High-Fidelity and Flexible Emotional Voice Conversion with Dual Control from Natural Language and Speech cites this paper.

ClapFM-EVC: High-Fidelity and Flexible Emotional Voice Conversion with Dual Control from Natural Language and Speech Conformer: Convolution-augmented Transformer for Speech Recognition

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-15T20:13:45.889714Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:13:45.889714Z digest=sha256:51ea2ef7ede0200c91e03515790f0301c16af69fcac1ed59dcd8473fcf85ea39

Observation 75b3023e-f670-4324-8b08-aae5b29f9a25 · inbound

The Multimodal Information Based Speech Processing (MISP) 2025 Challenge: Audio-Visual Diarization and Recognition cites this paper.

The Multimodal Information Based Speech Processing (MISP) 2025 Challenge: Audio-Visual Diarization and Recognition Conformer: Convolution-augmented Transformer for Speech Recognition

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T15:43:03.146074Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:43:03.146074Z digest=sha256:9f0eb08137a29a54c1f7a8ca2efa9c4a036fa7d9d74c9b4879776528714d3a9b

Observation 86ab1fd2-8887-4e89-943f-98ef29dd12c7 · inbound

QUADS: QUAntized Distillation Framework for Efficient Speech Language Understanding cites this paper.

QUADS: QUAntized Distillation Framework for Efficient Speech Language Understanding Conformer: Convolution-augmented Transformer for Speech Recognition

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-15T20:21:35.938764Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:21:35.938764Z digest=sha256:3e9b9cccde789ff6f6309ee9d84fc1e244ef1f0b7dbe362698f072eb8208d89a

Observation 6ed2d261-021f-4fbb-94ef-b7a55cd3ded7 · inbound

In-Context Learning Boosts Speech Recognition via Human-like Adaptation to Speakers and Language Varieties cites this paper.

In-Context Learning Boosts Speech Recognition via Human-like Adaptation to Speakers and Language Varieties Conformer: Convolution-augmented Transformer for Speech Recognition

Reference 2020

Resolution
unresolved
no resolver link, observed 2026-08-07T15:30:35.814004Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:30:35.814004Z digest=sha256:cdbd1ffe66851c1fc8a12b4cc08b27ac93fd56672b983bf6376a4e29b9c2aaa8

Observation b0c18d76-84ba-45a3-a517-3d8140f15e51 · inbound

MPE-TTS: Customized Emotion Zero-Shot Text-To-Speech Using Multi-Modal Prompt cites this paper.

MPE-TTS: Customized Emotion Zero-Shot Text-To-Speech Using Multi-Modal Prompt Conformer: Convolution-augmented Transformer for Speech Recognition

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T14:33:40.174375Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:33:40.174375Z digest=sha256:0da27ffeb8079fa597cdd15ef60333faaf99228f8a3e9899d7f73ed376bdb2c0

Observation bda5f11f-5787-4aae-94a2-bb11cfdf7a1a · inbound

Test-Time Adaptation with Binary Feedback cites this paper.

Test-Time Adaptation with Binary Feedback Conformer: Convolution-augmented Transformer for Speech Recognition

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-07T14:35:32.658211Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:35:32.658211Z digest=sha256:8208baa5c4a3044657d8d871ac0cea429ab8ec3cc8c96fbd02cdd4b452a6bc3f

Observation 1e9dd8c2-89ec-41fb-9003-a684f8c85455 · inbound

DuRep: Dual-Mode Speech Representation Learning via ASR-Aware Distillation cites this paper.

DuRep: Dual-Mode Speech Representation Learning via ASR-Aware Distillation Conformer: Convolution-augmented Transformer for Speech Recognition

Reference 2020

Resolution
unresolved
no resolver link, observed 2026-08-07T14:12:48.881846Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:12:48.881846Z digest=sha256:27fb9366a542379d7badb2a4adb0c54f9381a23dfbe612699b32ccd864dc04aa

Observation 169549f2-9d70-4f14-882e-aa95a3b8fb7f · inbound

Music Audio-Visual Question Answering Requires Specialized Multimodal Designs cites this paper.

Music Audio-Visual Question Answering Requires Specialized Multimodal Designs Conformer: Convolution-augmented Transformer for Speech Recognition

Reference 5

Resolution
metadata mismatch
arxiv_id, observed 2026-05-19T14:12:22.089771Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-19T14:11:52.011626Z digest=sha256:b2000f3673b738854397f38f7cbfd4c5d9bfb59c55c4c8e017a79d78edc30c21

Observation 823371dd-2fe8-497f-bf8c-80c1fb16fbab · inbound

Leveraging LLM for Stuttering Speech: A Unified Architecture Bridging Recognition and Event Detection cites this paper.

Leveraging LLM for Stuttering Speech: A Unified Architecture Bridging Recognition and Event Detection Conformer: Convolution-augmented Transformer for Speech Recognition

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T13:21:55.806950Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:21:55.806950Z digest=sha256:62f620358074adf1dece5703d73919e5acfc5723874c760d5dda3aaf57edd888

Observation 6bb3fe55-5ae3-4a56-86bf-0c853d741b88 · inbound

Delayed-KD: Delayed Knowledge Distillation based CTC for Low-Latency Streaming ASR cites this paper.

Delayed-KD: Delayed Knowledge Distillation based CTC for Low-Latency Streaming ASR Conformer: Convolution-augmented Transformer for Speech Recognition

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T13:24:29.745390Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:24:29.745390Z digest=sha256:a6d2c48ed198b33250ad2c27574af572b8134fa57b1575a364d49bddb5027035

Observation 61506fae-6150-4f64-9c98-12aa2c30bd3e · inbound

Towards disentangling the contributions of articulation and acoustics in multimodal phoneme recognition cites this paper.

Towards disentangling the contributions of articulation and acoustics in multimodal phoneme recognition Conformer: Convolution-augmented Transformer for Speech Recognition

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T12:39:46.387249Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:39:46.387249Z digest=sha256:19ea17635a2f062dcd2625881dbc0c7bbe1b811bd65bc0672fe97ebab5df7e1a

Observation 26300d89-a224-4daa-a4b5-8383f6b86d6f · inbound

Transformers Are Universally Consistent cites this paper.

Transformers Are Universally Consistent Conformer: Convolution-augmented Transformer for Speech Recognition

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T12:35:03.644947Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:35:03.644947Z digest=sha256:e5ffa876a21b2df42c5a11f19bbd417bfd3c44c615d4379ced8fbfda1f2ebea4

Observation 9d099ae3-f46a-4463-bb50-feaf6ed86116 · inbound

Pureformer-VC: Non-parallel Voice Conversion with Pure Stylized Transformer Blocks and Triplet Discriminative Training cites this paper.

Pureformer-VC: Non-parallel Voice Conversion with Pure Stylized Transformer Blocks and Triplet Discriminative Training Conformer: Convolution-augmented Transformer for Speech Recognition

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T05:20:56.451555Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:20:56.451555Z digest=sha256:9d3a09fce02c49ac0b829baceff15e44ada62b6b95644b7da146cf2102c40494

Observation 41a1e8f9-7454-46d1-955a-b62514b7c4c1 · inbound

Recognizing Every Voice: Towards Inclusive ASR for Rural Bhojpuri Women cites this paper.

Recognizing Every Voice: Towards Inclusive ASR for Rural Bhojpuri Women Conformer: Convolution-augmented Transformer for Speech Recognition

Reference 2020

Resolution
unresolved
no resolver link, observed 2026-08-07T04:51:25.472023Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:51:25.472023Z digest=sha256:af6422094662d28385ef28ee05b116930d7737aafd2c7e6c015d8e64105a7396

Observation 540381c8-97f5-4e5d-ab92-9a47a0d66f1c · inbound

FairASR: Fair Audio Contrastive Learning for Automatic Speech Recognition cites this paper.

FairASR: Fair Audio Contrastive Learning for Automatic Speech Recognition Conformer: Convolution-augmented Transformer for Speech Recognition

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T04:24:25.967159Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:24:25.967159Z digest=sha256:73e8f8aea793e33b8b145a803d2b0b611527cd3d38d5cc1716f4148196ae2398

Observation 7d1701fd-bfef-49fa-abcc-e7495a5d0181 · inbound

StreamMel: Real-Time Zero-shot Text-to-Speech via Interleaved Continuous Autoregressive Modeling cites this paper.

StreamMel: Real-Time Zero-shot Text-to-Speech via Interleaved Continuous Autoregressive Modeling Conformer: Convolution-augmented Transformer for Speech Recognition

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T00:52:03.732161Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:52:03.732161Z digest=sha256:486f38acf1d8e1a808b0f52a8bba78754db61998356b1b1eb7ba079715b326db

Observation 0fa73f01-3f40-4b90-a87c-e5527b34ac03 · inbound

SC-SOT: Conditioning the Decoder on Diarized Speaker Information for End-to-End Overlapped Speech Recognition cites this paper.

SC-SOT: Conditioning the Decoder on Diarized Speaker Information for End-to-End Overlapped Speech Recognition Conformer: Convolution-augmented Transformer for Speech Recognition

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-07T00:50:17.256472Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:50:17.256472Z digest=sha256:c7bba7beb5d7a741af4296364512adf58650139d6d709413abaf8550c02622f2

Observation bd5a5ea5-4b6a-49a9-85d0-37d29fff8901 · inbound

Pushing the Performance of Synthetic Speech Detection with Kolmogorov-Arnold Networks and Self-Supervised Learning Models cites this paper.

Pushing the Performance of Synthetic Speech Detection with Kolmogorov-Arnold Networks and Self-Supervised Learning Models Conformer: Convolution-augmented Transformer for Speech Recognition

Reference 2020

Resolution
unresolved
no resolver link, observed 2026-08-07T00:22:24.974764Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:22:24.974764Z digest=sha256:c39ccc2a62be5b49d9ebcf80cd2a640c358b4e14d5d90b43aa8bc1d1cc564452

Observation f9f662a8-0597-4ae8-9df8-e171d0b0d4d2 · inbound

Early Attentive Sparsification Accelerates Neural Speech Transcription cites this paper.

Early Attentive Sparsification Accelerates Neural Speech Transcription Conformer: Convolution-augmented Transformer for Speech Recognition

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-06T23:50:41.136057Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T23:50:41.136057Z digest=sha256:2c51174838ec8aa75ca0110e5a77ef015381d7b55e1c9c4a12071a90d5961be9

Observation da49e256-227c-4898-9ed6-4c9792f3051c · inbound

Instituto de Telecomunica\c{c}\~oes at IWSLT 2025: Aligning Small-Scale Speech and Language Models for Speech-to-Text Learning cites this paper.

Instituto de Telecomunica\c{c}\~oes at IWSLT 2025: Aligning Small-Scale Speech and Language Models for Speech-to-Text Learning Conformer: Convolution-augmented Transformer for Speech Recognition

Reference 2020

Resolution
unresolved
no resolver link, observed 2026-08-06T23:34:07.813102Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:34:07.813102Z digest=sha256:8c6642de7ed19b29f85bbe19dad7f2ef3e329de6f6e5a4b2a38e4b7fa23dd4c2

Observation 9c4f57d1-2c8e-4f24-a2a5-58267cb3db99 · inbound

A foundation model with multi-variate parallel attention to generate neuronal activity cites this paper.

A foundation model with multi-variate parallel attention to generate neuronal activity Conformer: Convolution-augmented Transformer for Speech Recognition

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-06T22:55:46.416417Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:55:46.416417Z digest=sha256:32e22b602f2358fa2ee76d4b19415f69030b9e8033dd94d74c5c68c9ef013e39

Observation 37fe7df5-fad6-4e18-9363-f78955c3aa74 · inbound

WTFormer: A Wavelet Conformer Network for MIMO Speech Enhancement with Spatial Cues Peservation cites this paper.

WTFormer: A Wavelet Conformer Network for MIMO Speech Enhancement with Spatial Cues Peservation Conformer: Convolution-augmented Transformer for Speech Recognition

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-06T22:26:17.482190Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:26:17.482190Z digest=sha256:b543ca52fdca053a6af4ded06c3c3625dda6ed151e310f3e2698fa3710a798d2

Observation 9c492056-dde3-4364-b09b-f71204d78236 · inbound

MuteSwap: Visual-informed Silent Video Identity Conversion cites this paper.

MuteSwap: Visual-informed Silent Video Identity Conversion Conformer: Convolution-augmented Transformer for Speech Recognition

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-06T21:19:51.504954Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T21:19:51.504954Z digest=sha256:3b051f449fe01968532a1b02944af99ed50139190cee457faac69ad8496149b9

Observation 38f6d3e4-7233-4e86-911c-998220d2cfdd · inbound

OMAR-RQ: Open Music Audio Representation Model Trained with Multi-Feature Masked Token Prediction cites this paper.

OMAR-RQ: Open Music Audio Representation Model Trained with Multi-Feature Masked Token Prediction Conformer: Convolution-augmented Transformer for Speech Recognition

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-06T20:14:10.664092Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:14:10.664092Z digest=sha256:f60cdb0d9024ec93d760efa931c1f53ebdcb4e0ee96965bd7558bbd5734dd37e

Observation ecbfb686-1f96-4a78-be4c-5361c2c574dd · inbound

Causal Foundation Models: Disentangling Physics from Instrument Properties cites this paper.

Causal Foundation Models: Disentangling Physics from Instrument Properties Conformer: Convolution-augmented Transformer for Speech Recognition

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-06T19:32:22.134628Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:32:22.134628Z digest=sha256:66ba5680296c14f5b4682133f66c1a2b86699960c0a377a1697c8791d5f99661

Observation 57f2890f-83b5-4b77-9885-6887a56d5653 · inbound

STARS: A Unified Framework for Singing Transcription, Alignment, and Refined Style Annotation cites this paper.

STARS: A Unified Framework for Singing Transcription, Alignment, and Refined Style Annotation Conformer: Convolution-augmented Transformer for Speech Recognition

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-06T19:01:27.510652Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T19:01:27.510652Z digest=sha256:eb8d9b542e19ea36d5c2c371ef6dcd23b2545bffaee1296527728678b400042b

Observation 581e39d2-a2e8-4a33-9fa6-4007ec3be640 · inbound

SpecASR: Accelerating LLM-based Automatic Speech Recognition via Speculative Decoding cites this paper.

SpecASR: Accelerating LLM-based Automatic Speech Recognition via Speculative Decoding Conformer: Convolution-augmented Transformer for Speech Recognition

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-06T14:42:56.113650Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:42:56.113650Z digest=sha256:f0bfc3d7ee23b5a87bb1485f34f6da8612f6c35698e26715595fcc892d39fed4

Observation 19912d0a-9069-4a27-96cd-5287d68f9453 · inbound

Self-Improvement for Audio Large Language Model using Unlabeled Speech cites this paper.

Self-Improvement for Audio Large Language Model using Unlabeled Speech Conformer: Convolution-augmented Transformer for Speech Recognition

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-15T17:52:49.739127Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:52:49.739127Z digest=sha256:d5ebda890934d36e5638d168926705714ef143ccac57a4ce764fa7473195fc33

Observation ecd4dfe7-a943-40bb-ac73-a8eb964e1f91 · inbound

Non-Intrusive Automatic Speech Recognition Refinement: A Survey cites this paper.

Non-Intrusive Automatic Speech Recognition Refinement: A Survey Conformer: Convolution-augmented Transformer for Speech Recognition

Reference 3

Resolution
metadata mismatch
arxiv_id, observed 2026-05-21T23:35:46.452979Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-21T23:34:52.249143Z digest=sha256:9d61dab941d88bf541b804a9d2f7621901594925b9156f1f92d4fca906aae571

Observation 9d6e07af-e3de-4174-9279-b39644d2194d · inbound

A Signer-Invariant Conformer and Multi-Scale Fusion Transformer for Continuous Sign Language Recognition cites this paper.

A Signer-Invariant Conformer and Multi-Scale Fusion Transformer for Continuous Sign Language Recognition Conformer: Convolution-augmented Transformer for Speech Recognition

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-05T21:12:01.925265Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T21:12:01.925265Z digest=sha256:1d8b5fe539bb94055b8cd743b125b03ac5b158a8a4d63a3ba665af89d54735aa

Observation 12c4f2be-3549-4235-b773-52986618c110 · inbound

The Sound of Risk: A Multimodal Physics-Informed Acoustic Model for Forecasting Market Volatility and Enhancing Market Interpretability cites this paper.

The Sound of Risk: A Multimodal Physics-Informed Acoustic Model for Forecasting Market Volatility and Enhancing Market Interpretability Conformer: Convolution-augmented Transformer for Speech Recognition

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-05T16:24:13.925390Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T16:24:13.925390Z digest=sha256:1a0c0a4b49e7f819289aa62f06bf88881137baf7008cad54cf8bd1d57e78b353

Observation 2bd2cf12-0a45-40ce-bee6-41ffad99057b · inbound

Fundamentals of Data-Driven Approaches to Acoustic Signal Detection, Filtering, and Transformation cites this paper.

Fundamentals of Data-Driven Approaches to Acoustic Signal Detection, Filtering, and Transformation Conformer: Convolution-augmented Transformer for Speech Recognition

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-05T14:24:31.055305Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T14:24:31.055305Z digest=sha256:6f755435b2688d04c986b53ebade2ed47df2ad3239f1f6c83ba5a7141a51885d

Observation 9deb145b-5d8f-44c6-a778-135df11c596b · inbound

Automatic Pronunciation Error Detection and Correction of the Holy Quran's Learners Using Deep Learning cites this paper.

Automatic Pronunciation Error Detection and Correction of the Holy Quran's Learners Using Deep Learning Conformer: Convolution-augmented Transformer for Speech Recognition

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-05T15:22:20.578116Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T15:22:20.578116Z digest=sha256:1e2999f26f4c439c5dec02c4a901979eeeb003f1d985beb93d71fc5884c0d29a

Observation 71d59571-a5e8-4a80-920c-49ba09d4ac7a · inbound

Entropy-based Coarse and Compressed Semantic Speech Representation Learning cites this paper.

Entropy-based Coarse and Compressed Semantic Speech Representation Learning Conformer: Convolution-augmented Transformer for Speech Recognition

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-05T13:36:04.660557Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T13:36:04.660557Z digest=sha256:4e6f313073c795b0bac2853b5e6ee077e2d8d74ef44a4c6d66b7b8b05f77374d

Observation 13982067-8fed-4bfc-95ad-6293d88d9d1e · inbound

NADI 2025: The First Multidialectal Arabic Speech Processing Shared Task cites this paper.

NADI 2025: The First Multidialectal Arabic Speech Processing Shared Task Conformer: Convolution-augmented Transformer for Speech Recognition

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-05T12:00:47.991578Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T12:00:47.991578Z digest=sha256:071bf931dd89a0cf8ccd4194238e0583f928e2cc8711a4cfd1383b979837fdcf

Observation dd0b3353-cd8a-4cff-9ef2-c00dd16d36e1 · inbound

AudioCodecBench: A Comprehensive Benchmark for Audio Codec Evaluation cites this paper.

AudioCodecBench: A Comprehensive Benchmark for Audio Codec Evaluation Conformer: Convolution-augmented Transformer for Speech Recognition

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-05T11:40:59.970746Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T11:40:59.970746Z digest=sha256:7b34a3fba3a36082b9f37fe8f57ea1f3da8238bdfb6cdb03ea1ac3f7f50eb037

Observation 89d46fd3-6fb4-4019-9d62-d02e9e973939 · inbound

Contextualized Token Discrimination for Speech Search Query Correction cites this paper.

Contextualized Token Discrimination for Speech Search Query Correction Conformer: Convolution-augmented Transformer for Speech Recognition

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-05T10:16:11.974460Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T10:16:11.974460Z digest=sha256:3d696cad0245f4036568bc8dd6f1d9b51ce55c171211a0f0a177cf7c623fc12b

Observation 109d7505-fcb0-4f70-9f06-f3381a8736ef · inbound

Unified Learnable 2D Convolutional Feature Extraction for ASR cites this paper.

Unified Learnable 2D Convolutional Feature Extraction for ASR Conformer: Convolution-augmented Transformer for Speech Recognition

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-04T18:20:02.004092Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T18:20:02.004092Z digest=sha256:abf1e1566762501977164e7bbb565334eee8307df073f6705df06e07be91d01c

Observation 963e2091-2ee9-4915-83de-4b9c17528fac · inbound

HRTFformer: A Spatially-Aware Transformer for Individual HRTF Upsampling in Immersive Audio Rendering cites this paper.

HRTFformer: A Spatially-Aware Transformer for Individual HRTF Upsampling in Immersive Audio Rendering Conformer: Convolution-augmented Transformer for Speech Recognition

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-04T12:50:26.217151Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T12:50:26.217151Z digest=sha256:1a292e2644199dcea78290bb805b31990e33d24b6e9d9e89e0e06729d28c9741

Observation df9c363b-3daf-4ef6-9ea0-617f78190a81 · inbound

How to Evaluate Speech Translation with Source-Aware Neural MT Metrics cites this paper.

How to Evaluate Speech Translation with Source-Aware Neural MT Metrics Conformer: Convolution-augmented Transformer for Speech Recognition

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-05-18T01:40:37.282806Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-18T01:36:53.294436Z digest=sha256:0a065a926f73a41c67b5ef25a324ea4f7d8cf0cd0521344936473f7832bd9787

Observation a62e6233-4ecd-4cb0-ab4c-74f06e339abb · inbound

Physics of Language Models: Part 4.1, Architecture Design and the Magic of Canon Layers cites this paper.

Physics of Language Models: Part 4.1, Architecture Design and the Magic of Canon Layers Conformer: Convolution-augmented Transformer for Speech Recognition

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-03T15:22:49.687539Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T15:22:49.687539Z digest=sha256:21918c07b60d0f8e812056315ec94657f7745110f2b5177d7eb345b1f2834838

Observation eae177ea-07c3-4cd9-9be1-0870d84ec4f6 · inbound

KernelEvolve: Scaling Agentic Kernel Coding for Heterogeneous AI Accelerators at Meta cites this paper.

KernelEvolve: Scaling Agentic Kernel Coding for Heterogeneous AI Accelerators at Meta Conformer: Convolution-augmented Transformer for Speech Recognition

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-03T13:45:23.606608Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T13:45:23.606608Z digest=sha256:ce3e10b2843a43041af6d818e7994463d2d972ae4e0bb29c7921250893c2a08f

Observation 52dae236-5a3d-4b27-9707-13a82a1b72d5 · inbound

FastSLM: Hierarchical Temporal Abstraction for Efficient Long-Form Speech Adaptation cites this paper.

FastSLM: Hierarchical Temporal Abstraction for Efficient Long-Form Speech Adaptation Conformer: Convolution-augmented Transformer for Speech Recognition

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-03T12:02:00.595452Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T12:02:00.595452Z digest=sha256:0b5d9c1cffbecde1bbcc64df3f7dd314b9dfa8402224ceecd7864981a9dcda7a

Observation 6799d9a5-7ce7-48e8-8d67-09f0c696cfcd · inbound

Noise-Robust Contrastive Learning with an MFCC-Conformer For Coronary Artery Disease Detection cites this paper.

Noise-Robust Contrastive Learning with an MFCC-Conformer For Coronary Artery Disease Detection Conformer: Convolution-augmented Transformer for Speech Recognition

Reference 16

Resolution
metadata mismatch
arxiv_id, observed 2026-05-16T11:27:48.195077Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-16T11:26:15.596039Z digest=sha256:23ff6236c7353238be87574ba2e5563f4e5b68c04991859a45ff3b74cae2b322

Observation c3d2a669-ff9a-4efd-a10c-0b7692264241 · inbound

Escaping the BLEU Trap: A Signal-Grounded Framework with Decoupled Semantic Guidance for EEG-to-Text Decoding cites this paper.

Escaping the BLEU Trap: A Signal-Grounded Framework with Decoupled Semantic Guidance for EEG-to-Text Decoding Conformer: Convolution-augmented Transformer for Speech Recognition

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-03T03:27:34.432446Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T03:27:34.432446Z digest=sha256:093f7a4c4bfa1a08986e76cb4c53b4191e0fef253785df7e7e40a4375cedee05

Observation bfd8d8b4-0d06-45a6-8bf3-3feaa79919e9 · inbound

Evolution Strategy-Based Calibration for Low-Bit Quantization of Speech Models cites this paper.

Evolution Strategy-Based Calibration for Low-Bit Quantization of Speech Models Conformer: Convolution-augmented Transformer for Speech Recognition

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-02T18:35:50.419650Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T18:35:50.419650Z digest=sha256:a90ee1a0b243c1b453801f577a4e5698101971444985b92245bc87c6274ce7c5

Observation 7fd691fd-e381-4a10-b00e-311eb6e9bbcd · inbound

SEMamba++: A General Speech Restoration Framework Leveraging Global, Local, and Periodic Spectral Patterns cites this paper.

SEMamba++: A General Speech Restoration Framework Leveraging Global, Local, and Periodic Spectral Patterns Conformer: Convolution-augmented Transformer for Speech Recognition

Reference 22

Resolution
unresolved
no resolver link, observed 2026-07-14T22:43:13.611383Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T22:43:13.611383Z digest=sha256:ba7b1a318cbaa78a77f615789855e60d251cad3c6c0fc13139397e0cb7099ab1

Observation 0c892742-1fd4-400a-bd8a-c6751a917048 · inbound

Making Separation-First Multi-Stream Audio Watermarking Feasible via Joint Training cites this paper.

Making Separation-First Multi-Stream Audio Watermarking Feasible via Joint Training Conformer: Convolution-augmented Transformer for Speech Recognition

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-02T18:03:41.525577Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T18:03:41.525577Z digest=sha256:02a5075f76e9bcdc131397a1c0723b4e8d9e0d2e8f3c1e298a409ef6bb371a51

Observation 25ebb840-3fa1-4c2e-a9b3-c8c223cac999 · inbound

Sharp spectral estimates for free boundary problems arising in plasma physics cites this paper.

Sharp spectral estimates for free boundary problems arising in plasma physics Conformer: Convolution-augmented Transformer for Speech Recognition

Reference 33

Resolution
unresolved
no resolver link, observed 2026-07-14T19:56:33.793167Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T19:56:33.793167Z digest=sha256:1c4f2532ba5414f007a005cf097a8ff439b64ab74d2ef2d99f393a9688ccb13e

Observation f98ad145-9767-4ece-9a06-260d23469722 · inbound

FastTurn: Unifying Acoustic and Streaming Semantic Cues for Low-Latency and Robust Turn Detection cites this paper.

FastTurn: Unifying Acoustic and Streaming Semantic Cues for Low-Latency and Robust Turn Detection Conformer: Convolution-augmented Transformer for Speech Recognition

Reference 33

Resolution
metadata mismatch
arxiv_id, observed 2026-05-13T20:58:15.996997Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-13T20:55:09.141906Z digest=sha256:5731f52130db5e833b50735261b39e3d61462907a39f53ccf866608b7f587609

Observation 70fa7b1f-716f-42bf-992d-0f62b9440657 · inbound

MALEFA: Multi-grAnularity Learning and Effective False Alarm Suppression for Zero-shot Keyword Spotting cites this paper.

MALEFA: Multi-grAnularity Learning and Effective False Alarm Suppression for Zero-shot Keyword Spotting Conformer: Convolution-augmented Transformer for Speech Recognition

Reference 19

Resolution
metadata mismatch
arxiv_id, observed 2026-05-13T17:23:02.525280Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-13T17:20:18.787074Z digest=sha256:cd297358777646b3023952b33ae1cb5671f6e9f1c2f4bb692a7a9bfb7b42a4c6

Observation 27e726b8-4a17-4182-85ad-ec707cbfaa1d · inbound

Rethinking Entropy Allocation in LLM-based ASR: Understanding the Dynamics between Speech Encoders and LLMs cites this paper.

Rethinking Entropy Allocation in LLM-based ASR: Understanding the Dynamics between Speech Encoders and LLMs Conformer: Convolution-augmented Transformer for Speech Recognition

Reference 13

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T05:30:57.621115Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-10T18:06:50.408402Z digest=sha256:b839aab754fae21465fd807cc61e60a919a6f6e0976d2625cde0aaf6c7a64ff5

Observation 65366e2f-cbbe-4ce2-adcd-1e79f012d250 · inbound

NIM4-ASR: Towards Efficient, Robust, and Customizable Real-Time LLM-Based ASR cites this paper.

NIM4-ASR: Towards Efficient, Robust, and Customizable Real-Time LLM-Based ASR Conformer: Convolution-augmented Transformer for Speech Recognition

Reference 8

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T12:21:06.024806Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-10T03:48:14.211240Z digest=sha256:7bdecd3e2c9566dcbef25dc7cfc4297e67e60f50268d51b5103fe79a5ddb3ea6

Observation 39e80b68-f2a1-4a0b-a324-29a11dafecc3 · inbound

NIM4-ASR: Towards Efficient, Robust, and Customizable Real-Time LLM-Based ASR cites this paper.

NIM4-ASR: Towards Efficient, Robust, and Customizable Real-Time LLM-Based ASR Conformer: Convolution-augmented Transformer for Speech Recognition

Reference 8

Resolution
metadata mismatch
local_arxiv, observed 2026-07-05T13:21:06.270505Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-07-05T13:15:50.794969Z digest=sha256:5d31f24439d7aabf56507c1acfdd6be55b1506201a6985ed5f279e387193fc27

Observation 41a58e19-f876-4a08-afd5-02077e5c2d07 · inbound

Learning Posterior Predictive Distributions for Node Classification from Synthetic Graph Priors cites this paper.

Learning Posterior Predictive Distributions for Node Classification from Synthetic Graph Priors Conformer: Convolution-augmented Transformer for Speech Recognition

Reference 230

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T12:21:04.509497Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-05-10T03:50:44.626261Z digest=sha256:121fe0df21476911a5369b360cb0037c2670d2ad8e0221c884289722a0929758

Observation 99efac63-c8cc-4bb1-838e-ec9722303421 · inbound

ResAF-Net: An Anchor-Free Attention-Based Network for Tree Detection and Agricultural Mapping in Palestine cites this paper.

ResAF-Net: An Anchor-Free Attention-Based Network for Tree Detection and Agricultural Mapping in Palestine Conformer: Convolution-augmented Transformer for Speech Recognition

Reference 31

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T21:06:13.351597Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-08T06:51:54.321037Z digest=sha256:565a7f19e9d32b97cea1ba847540b81d9eacfd4ad5540c01fdecd7cd83854cbd

Observation b514e865-dcd0-4746-871f-a3e83e9a65c8 · inbound

AccLock: Unlocking Identity with Heartbeat Using In-Ear Accelerometers cites this paper.

AccLock: Unlocking Identity with Heartbeat Using In-Ear Accelerometers Conformer: Convolution-augmented Transformer for Speech Recognition

Reference 25

Resolution
verified exact
arxiv_id, observed 2026-05-13T05:27:18.633700Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-13T05:24:04.546820Z digest=sha256:e5a1979aae7e6f930853750b8a29464e454dfefada99b9d703dacd8792a2a829

Observation df2d7a77-9d60-400e-ac7f-7e97e2084e00 · inbound

MedASR: An Open-Source Model for High-Accuracy Medical Dictation cites this paper.

MedASR: An Open-Source Model for High-Accuracy Medical Dictation Conformer: Convolution-augmented Transformer for Speech Recognition

Reference 17

Resolution
metadata mismatch
arxiv_id, observed 2026-05-19T21:22:48.080788Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-19T21:18:56.084055Z digest=sha256:a2699a5f1800d13acd90ddf8beb6ebc5110d82ca6830ed827c4271cc179dc092

Observation a8e13d0d-7ece-4114-a837-e5bdd89d3948 · inbound

Multilingual and Multimodal LLMs in the Wild: Building for Low-Resource Languages cites this paper.

Multilingual and Multimodal LLMs in the Wild: Building for Low-Resource Languages Conformer: Convolution-augmented Transformer for Speech Recognition

Reference 242

Resolution
metadata mismatch
arxiv_id, observed 2026-05-20T14:38:21.743997Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-05-20T14:33:36.100966Z digest=sha256:4344a282ff146f11bc1f74a12bd6b9115742deaad40bcda3aed2e28b520fc9ba

Observation 321c05f4-83f6-4fb6-9d51-64546cd41c36 · inbound

Executable Boundary Contracts for Sound Event Traces cites this paper.

Executable Boundary Contracts for Sound Event Traces Conformer: Convolution-augmented Transformer for Speech Recognition

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-05-20T02:12:58.511264Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-20T02:08:44.066554Z digest=sha256:df6b90c27a1945ace7664be5ff023491342dd21ba256b6a8f5ab4aca2edb224e

Observation 1ab776c2-1163-4f25-a61c-21b3550f82c6 · inbound

Data-Efficient On-Policy Distillation for Automatic Speech Recognition cites this paper.

Data-Efficient On-Policy Distillation for Automatic Speech Recognition Conformer: Convolution-augmented Transformer for Speech Recognition

Reference 2

Resolution
metadata mismatch
arxiv_id, observed 2026-06-29T12:03:23.188660Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-06-29T12:02:35.332633Z digest=sha256:cb086b98bda057e585927e8cac39a00dfa444d099e0e96a1fedbe2fb81d42784

Observation 87626e57-fa71-489b-9e89-f75dca9e397f · inbound

ConTrans: Learning Text-enhanced Local-global Temporal Representations for Zero-shot Temporal Action Localization cites this paper.

ConTrans: Learning Text-enhanced Local-global Temporal Representations for Zero-shot Temporal Action Localization Conformer: Convolution-augmented Transformer for Speech Recognition

Reference 30

Resolution
metadata mismatch
arxiv_id, observed 2026-06-28T23:32:47.449856Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-06-28T23:27:25.975576Z digest=sha256:2bc6db51d6539d0a0f9fe2dfa937fff2ad4e31bbcaa6606094ff4eed651f2064