Pith. sign in

Paper Citation Record · LEDGER

Prepending or Cross-Attention for Speech-to-Text? An Empirical Comparison

As of 12 August 2026, this Paper Citation Record lists 73 of 73 outbound references and 0 inbound Pith citation observations for arXiv:2501.02370.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2501.02370 v3

Coverage vector

measured 73 of 73 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-10T22:18:01.876712Z

measured 73 of 73 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-11T06:34:44.6726+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

73 of 73 outbound references displayed

  • verified exact9
  • verified fuzzy0
  • unresolved63
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch1

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 0f3b9c67-05fa-4f7a-8bf9-5a49f7184a13 · outbound

This paper cites online" 'onlinestring :=.

Prepending or Cross-Attention for Speech-to-Text? An Empirical Comparison online" 'onlinestring :=

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-10T22:18:01.448832Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T22:18:01.448832Z digest=sha256:1ee35bf099b1a9f24c20d25bd333d5bb253f4e043c7715ac09fe5652c7b1c757

Observation e6c487ac-5620-44b7-8e09-33629c904bcd · outbound

This paper cites write newline.

Prepending or Cross-Attention for Speech-to-Text? An Empirical Comparison write newline

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-10T22:18:01.455015Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T22:18:01.455015Z digest=sha256:3c1ecd510b1504aa82163f75b29f683d56b932767bf749a28ef7b09b10775f25

Observation 0b4a17ef-b74b-4526-98e5-a84dd5879d8e · outbound

This paper cites an unresolved cited work.

Prepending or Cross-Attention for Speech-to-Text? An Empirical Comparison Unresolved cited work

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-10T22:18:01.461178Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T22:18:01.461178Z digest=sha256:9b23eadd5fad5e0de893e6793287d4f71f1b4b4f3cb4fe9cb73b21d6c7e433a8

Observation c5fb2eb8-c6ac-495a-b8c4-797fc17f1434 · outbound

This paper cites SpeechT5: Unified-Modal Encoder-Decoder Pre-Training for Spoken Language Processing.

Prepending or Cross-Attention for Speech-to-Text? An Empirical Comparison SpeechT5: Unified-Modal Encoder-Decoder Pre-Training for Spoken Language Processing

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-10T22:18:01.466667Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T22:18:01.466667Z digest=sha256:1b5f99ab924d6df8e0f29c5c1f2e04256c847aba87cf46682300821b1f5c1e6e

Observation 00681776-ef60-4a60-8c94-5e2a3a3247d6 · outbound

This paper cites an unresolved cited work.

Prepending or Cross-Attention for Speech-to-Text? An Empirical Comparison Unresolved cited work

Reference 5

Resolution
unresolved
raw_fallback, observed 2026-08-10T22:18:03.519305Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-10T22:18:01.472850Z digest=sha256:41029d301c72fe3ba6f7ccef0c47e3cdbaf5d3cc0851b7a9bc9eb17865f95919

Observation 2cbddd3e-636f-447e-9c0b-2b184ddb2405 · outbound

This paper cites SeamlessM4T: Massively Multilingual & Multimodal Machine Translation.

Prepending or Cross-Attention for Speech-to-Text? An Empirical Comparison SeamlessM4T: Massively Multilingual & Multimodal Machine Translation

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-10T22:18:01.478377Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T22:18:01.478377Z digest=sha256:dab8fda569e877dda0d855036c37eef2bc0734bdc59ad52a572fa32c8bfba299

Observation 2a26895a-ffd2-4510-8ed0-23822735c55b · outbound

This paper cites an unresolved cited work.

Prepending or Cross-Attention for Speech-to-Text? An Empirical Comparison Unresolved cited work

Reference 7

Resolution
unresolved
raw_fallback, observed 2026-08-10T22:18:03.500520Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-10T22:18:01.484075Z digest=sha256:7c6991da47ea719a835892e91e747a5f1308517ebdb4e4efb6513222e3084578

Observation 9d04f292-3889-4c34-8ced-6301b562e8fd · outbound

This paper cites Listen and Translate: A Proof of Concept for End-to-End Speech-to-Text Translation.

Prepending or Cross-Attention for Speech-to-Text? An Empirical Comparison Listen and Translate: A Proof of Concept for End-to-End Speech-to-Text Translation

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-10T22:18:01.489299Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T22:18:01.489299Z digest=sha256:e0c9c9e7d4bdf07af25980c744c022c93669c4dfad1726d7152b3f0b8c92e4d0

Observation 74368e12-e5ec-4696-8e83-edd0f44976db · outbound

This paper cites Language Models are Few-Shot Learners.

Prepending or Cross-Attention for Speech-to-Text? An Empirical Comparison Language Models are Few-Shot Learners

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-10T22:18:01.495695Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T22:18:01.495695Z digest=sha256:1ac1f87844e0c37efd1e7463f4ecd391f88334108e8eaf280d63df596aba65dc

Observation 98577954-92cd-496f-af45-0363a881fa0d · outbound

This paper cites an unresolved cited work.

Prepending or Cross-Attention for Speech-to-Text? An Empirical Comparison Unresolved cited work

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-10T22:18:01.500795Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T22:18:01.500795Z digest=sha256:9b288d6f60d88837c1337a61e86608bca77f34bbf4b0c1f68b188423e908b862

Observation c57b96bd-cf23-4998-bd38-e087f706bb27 · outbound

This paper cites Listen, Attend and Spell.

Prepending or Cross-Attention for Speech-to-Text? An Empirical Comparison Listen, Attend and Spell

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-10T22:18:01.505886Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T22:18:01.505886Z digest=sha256:69836c98c500ca60f8f804e8d15eebbb0fa088a6235f543cc357eac58ded3c9e

Observation 97e5cbe5-2c95-4895-86a2-30664262a0a3 · outbound

This paper cites X-LLM: Bootstrapping Advanced Large Language Models by Treating Multi-Modalities as Foreign Languages.

Prepending or Cross-Attention for Speech-to-Text? An Empirical Comparison X-LLM: Bootstrapping Advanced Large Language Models by Treating Multi-Modalities as Foreign Languages

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-10T22:18:01.511482Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T22:18:01.511482Z digest=sha256:cebf2b351d7daf45859c14a762a8bf3aab465ff3f6b9a1e646e9628a85801110

Observation 307c7a12-9d78-4dda-8972-ec0762f8bffb · outbound

This paper cites an unresolved cited work.

Prepending or Cross-Attention for Speech-to-Text? An Empirical Comparison Unresolved cited work

Reference 13

Resolution
verified exact
doi, observed 2026-08-10T22:18:02.297511Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-10T22:18:01.517081Z digest=sha256:4ea47b6c5aece27ab21e2523fe4cd380f30eda6e96c059107714090e91525675

Observation 1969d1a8-c8df-4690-8436-94d3999077bb · outbound

This paper cites an unresolved cited work.

Prepending or Cross-Attention for Speech-to-Text? An Empirical Comparison Unresolved cited work

Reference 14

Resolution
unresolved
raw_fallback, observed 2026-08-10T22:18:03.484626Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-10T22:18:01.522137Z digest=sha256:967390f89aabc59b2c12af8a58e39614542a30e8770a19b4b7d8f9ee28e9dfbe

Observation 1880c362-05f2-4d71-a28b-ff22d270cc57 · outbound

This paper cites BESTOW: Efficient and Streamable Speech Language Model with the Best of Two Worlds in GPT and T5.

Prepending or Cross-Attention for Speech-to-Text? An Empirical Comparison BESTOW: Efficient and Streamable Speech Language Model with the Best of Two Worlds in GPT and T5

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-10T22:18:01.526989Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T22:18:01.526989Z digest=sha256:f78c2a0768284ee28818ddec49dd13e86e341ddb81973f6048dfa08bc73fda32

Observation 72d40374-923b-48ca-8cad-cc33ea7bb60c · outbound

This paper cites Qwen-Audio: Advancing Universal Audio Understanding via Unified Large-Scale Audio-Language Models.

Prepending or Cross-Attention for Speech-to-Text? An Empirical Comparison Qwen-Audio: Advancing Universal Audio Understanding via Unified Large-Scale Audio-Language Models

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-10T22:18:01.532170Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T22:18:01.532170Z digest=sha256:7f4c6ad54b879c628391c462fe5d62e4ae361fc302716624e3db3e80cbf06d52

Observation a02aebfe-1914-42a6-af68-14ac5762e034 · outbound

This paper cites No Language Left Behind: Scaling Human-Centered Machine Translation.

Prepending or Cross-Attention for Speech-to-Text? An Empirical Comparison No Language Left Behind: Scaling Human-Centered Machine Translation

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-10T22:18:01.537186Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T22:18:01.537186Z digest=sha256:5eea784c62182bbd7ee176b03fc98a098453ebaf422172b8ba067b7663d6f86e

Observation f0674089-193f-4071-9c26-11a92076e2ae · outbound

This paper cites Di Gangi, Roldano Cattoni, Luisa Bentivogli, Matteo Negri, and Marco Turchi.

Prepending or Cross-Attention for Speech-to-Text? An Empirical Comparison Di Gangi, Roldano Cattoni, Luisa Bentivogli, Matteo Negri, and Marco Turchi

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-10T22:18:01.542521Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T22:18:01.542521Z digest=sha256:bd747ef5f916ebeb13c4982c61ca3a17b4c4033dbcce822cb98905c71f52a22b

Observation ff73b565-cb79-408f-973b-3990f54efd73 · outbound

This paper cites an unresolved cited work.

Prepending or Cross-Attention for Speech-to-Text? An Empirical Comparison Unresolved cited work

Reference 19

Resolution
unresolved
raw_fallback, observed 2026-08-10T22:18:03.467338Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-10T22:18:01.547880Z digest=sha256:dd9d07fe81782a3c36a00a96d336283326edb41f6f1437b544ca4af0a0b9e160

Observation 4f598268-324a-4286-8dd8-5d40c1c8453a · outbound

This paper cites an unresolved cited work.

Prepending or Cross-Attention for Speech-to-Text? An Empirical Comparison Unresolved cited work

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-10T22:18:01.553570Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T22:18:01.553570Z digest=sha256:e816ff9b5ce7fabec536f063345a07d4457b754fd449891f2dd5fdbc1b9e531d

Observation 63aed7e9-80d5-461c-85a1-03212c2050dd · outbound

This paper cites an unresolved cited work.

Prepending or Cross-Attention for Speech-to-Text? An Empirical Comparison Unresolved cited work

Reference 21

Resolution
unresolved
raw_fallback, observed 2026-08-10T22:18:03.450486Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-10T22:18:01.558447Z digest=sha256:4082acdb6924bb3965aacb969ab71c25bfecc72b80d96799e70cbce0273ccecb

Observation b1c261b1-4c9d-4a98-a3ba-f6ff39dbf4ae · outbound

This paper cites an unresolved cited work.

Prepending or Cross-Attention for Speech-to-Text? An Empirical Comparison Unresolved cited work

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-10T22:18:01.563366Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T22:18:01.563366Z digest=sha256:5791dcac85ab44c4fdff0095ffe2a5306f839b6c97b4001a0cb47e8981544c96

Observation 5434c9e8-0a74-4150-b062-866bcce2a019 · outbound

This paper cites an unresolved cited work.

Prepending or Cross-Attention for Speech-to-Text? An Empirical Comparison Unresolved cited work

Reference 23

Resolution
verified exact
doi, observed 2026-08-10T22:18:02.216980Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-10T22:18:01.568325Z digest=sha256:dee64574a93f9b78914141d5a12c55d0116e8c2da15acce3e4db60547aa657c2

Observation 19cf8ee4-d37b-4e0e-9258-6277cd896d8e · outbound

This paper cites an unresolved cited work.

Prepending or Cross-Attention for Speech-to-Text? An Empirical Comparison Unresolved cited work

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-10T22:18:01.574114Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T22:18:01.574114Z digest=sha256:efb78b67059372d7d78a10d14f4b03eb4da53ef90dd01e76771d7482a86a6b9c

Observation 244d2d17-71da-4dfa-a609-1ebf6f162b4c · outbound

This paper cites an unresolved cited work.

Prepending or Cross-Attention for Speech-to-Text? An Empirical Comparison Unresolved cited work

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-10T22:18:01.579165Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T22:18:01.579165Z digest=sha256:15f8a828f367218ec193cdd74f94f018ba50eb3350281eef2b797743eb819cbf

Observation f8cd1f14-14f0-4fc8-928e-f052921a25ff · outbound

This paper cites an unresolved cited work.

Prepending or Cross-Attention for Speech-to-Text? An Empirical Comparison Unresolved cited work

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-10T22:18:01.584589Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T22:18:01.584589Z digest=sha256:96dd9ff20d51d2aa0b1e9cda6aff11446c365bcf4ca8f92c0d559f1aef2a8d4e

Observation 6da82df2-b3aa-46bb-91c9-5b3008b8003f · outbound

This paper cites an unresolved cited work.

Prepending or Cross-Attention for Speech-to-Text? An Empirical Comparison Unresolved cited work

Reference 27

Resolution
verified exact
doi, observed 2026-08-10T22:18:02.157931Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-10T22:18:01.589891Z digest=sha256:b7623433eba30d492d9f118c75a6054105afa2caadbef4f9c3ce4061e236e8a0

Observation e0f6b5d2-e2d3-495f-b902-18f169f5a594 · outbound

This paper cites an unresolved cited work.

Prepending or Cross-Attention for Speech-to-Text? An Empirical Comparison Unresolved cited work

Reference 28

Resolution
verified exact
doi, observed 2026-08-10T22:18:02.131824Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-10T22:18:01.594610Z digest=sha256:b999c7e3c9caf16fd1c79149599c7547dac9046a33259134a28552be95c773cb

Observation df371e8f-507e-43ae-8b4e-3424983f3466 · outbound

This paper cites WavLLM: Towards Robust and Adaptive Speech Large Language Model.

Prepending or Cross-Attention for Speech-to-Text? An Empirical Comparison WavLLM: Towards Robust and Adaptive Speech Large Language Model

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-10T22:18:01.599424Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T22:18:01.599424Z digest=sha256:ead6b6d9b063813832f12d545fdc265f182cc1ae64a21d11e70e58518b5c3c28

Observation 92e0aa0a-c532-4b8b-b187-eb395f6f779d · outbound

This paper cites Investigating Decoder-only Large Language Models for Speech-to-text Translation.

Prepending or Cross-Attention for Speech-to-Text? An Empirical Comparison Investigating Decoder-only Large Language Models for Speech-to-text Translation

Reference 30

Resolution
verified exact
local_arxiv, observed 2026-08-10T22:18:02.937954Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-10T22:18:01.604423Z digest=sha256:8b385fb6e198a4e340995715185d88bab9f42ad98f0d31535d44503c9ee3006f

Observation 4e97a6a2-584e-4b67-b218-fafbb27dba81 · outbound

This paper cites an unresolved cited work.

Prepending or Cross-Attention for Speech-to-Text? An Empirical Comparison Unresolved cited work

Reference 31

Resolution
unresolved
raw_fallback, observed 2026-08-10T22:18:03.424437Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-10T22:18:01.609511Z digest=sha256:fd17059ff6408344dda76cd7ace628e76168446fbf0a8653889bbfb6cb844558

Observation 8c847965-af42-4aef-bb62-18c628730887 · outbound

This paper cites an unresolved cited work.

Prepending or Cross-Attention for Speech-to-Text? An Empirical Comparison Unresolved cited work

Reference 32

Resolution
verified exact
doi, observed 2026-08-10T22:18:02.106711Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-10T22:18:01.615827Z digest=sha256:9d368cadfaec10b4aa1e7918063040c5112947a769f9ac93e7464186f5ebbcee

Observation ae5c7f6d-a245-4092-b0d6-ece5797f5c75 · outbound

This paper cites Mistral 7B.

Prepending or Cross-Attention for Speech-to-Text? An Empirical Comparison Mistral 7B

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-10T22:18:01.621418Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T22:18:01.621418Z digest=sha256:0b71697364d0e79af4d8885be16a1a224d68c8c3b5531bf826574d45c7aeabce

Observation 7856850b-884c-4fc6-8192-6dbbb2df1237 · outbound

This paper cites an unresolved cited work.

Prepending or Cross-Attention for Speech-to-Text? An Empirical Comparison Unresolved cited work

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-10T22:18:01.626619Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T22:18:01.626619Z digest=sha256:a741d628089e0ef453babd73e76e9f4e55cdbc98397d47a25e82aa929e798b4c

Observation fcca62ab-ab44-41ce-843d-d24e3073b5a5 · outbound

This paper cites an unresolved cited work.

Prepending or Cross-Attention for Speech-to-Text? An Empirical Comparison Unresolved cited work

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-10T22:18:01.632173Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T22:18:01.632173Z digest=sha256:e7fdc940c854b1933bd3013ff22d65ae91f981dbc5cdc721d30d7c0f892d4a9d

Observation 3b01ca19-9c4c-40e5-80de-73d21c0b79b1 · outbound

This paper cites an unresolved cited work.

Prepending or Cross-Attention for Speech-to-Text? An Empirical Comparison Unresolved cited work

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-10T22:18:01.637237Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T22:18:01.637237Z digest=sha256:64e1b89aa3623dbb3f3818a529137ac23da1e2f42555b159d8f60e115c78b2a3

Observation 8616c69c-db18-47df-aa44-57a6322e35ae · outbound

This paper cites an unresolved cited work.

Prepending or Cross-Attention for Speech-to-Text? An Empirical Comparison Unresolved cited work

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-10T22:18:01.643157Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T22:18:01.643157Z digest=sha256:5ab885fa0123eb4c1bd80eb68d496b28e613b5fe951cfc79c305033c440d2ef5

Observation d8ec79a6-45d1-4b9c-b645-645dca211bcf · outbound

This paper cites an unresolved cited work.

Prepending or Cross-Attention for Speech-to-Text? An Empirical Comparison Unresolved cited work

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-10T22:18:01.648424Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T22:18:01.648424Z digest=sha256:2c10e1de5d9824490dfeecbe7e903eb9ac10aa50c02345c358a196781a4ca6f4

Observation 0115e4cf-721c-4f46-9eb9-de17847d53eb · outbound

This paper cites an unresolved cited work.

Prepending or Cross-Attention for Speech-to-Text? An Empirical Comparison Unresolved cited work

Reference 39

Resolution
unresolved
raw_fallback, observed 2026-08-10T22:18:03.381856Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-10T22:18:01.653492Z digest=sha256:00f95bf381ff286bd6c6a472865fe54085608f4d6fa3de52a1ab1833e3b5efa4

Observation 8799987d-4c64-45a9-be1c-410a45cd3d9c · outbound

This paper cites Sparks of Large Audio Models: A Survey and Outlook.

Prepending or Cross-Attention for Speech-to-Text? An Empirical Comparison Sparks of Large Audio Models: A Survey and Outlook

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-10T22:18:01.658865Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T22:18:01.658865Z digest=sha256:13abe2f1389a4b1d841795e74808d5171defbbb1ae81acd9b50a4b507e7a74f7

Observation 256b2189-67b7-47ef-a118-8261b5c7fd61 · outbound

This paper cites Speech-MASSIVE: A Multilingual Speech Dataset for SLU and Beyond.

Prepending or Cross-Attention for Speech-to-Text? An Empirical Comparison Speech-MASSIVE: A Multilingual Speech Dataset for SLU and Beyond

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-10T22:18:01.663850Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T22:18:01.663850Z digest=sha256:491ee3b50a06b80351eaea05731cefdc927358d16b3256bfbab4233587f5de30

Observation 806e12d1-8eb8-47a9-aed8-cb59242914d9 · outbound

This paper cites an unresolved cited work.

Prepending or Cross-Attention for Speech-to-Text? An Empirical Comparison Unresolved cited work

Reference 42

Resolution
unresolved
raw_fallback, observed 2026-08-10T22:18:03.363918Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-10T22:18:01.669010Z digest=sha256:f960c814afbf7405bfec7826ebcefce7c7abed5c3d89d6e29cdbef8d2472fa0a

Observation e3cd9b70-8b15-47fd-b296-a2514ce960d1 · outbound

This paper cites Bridging the Modality Gap for Speech-to-Text Translation.

Prepending or Cross-Attention for Speech-to-Text? An Empirical Comparison Bridging the Modality Gap for Speech-to-Text Translation

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-10T22:18:01.676462Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T22:18:01.676462Z digest=sha256:65ce89c8ec75e7cf28d2d6806962d7332021ebf411258a1acb868d30d3473f41

Observation cdb43061-e459-472f-b3f2-da6fab3196c9 · outbound

This paper cites an unresolved cited work.

Prepending or Cross-Attention for Speech-to-Text? An Empirical Comparison Unresolved cited work

Reference 44

Resolution
metadata mismatch
raw_fallback, observed 2026-08-10T22:18:02.824558Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-10T22:18:01.682072Z digest=sha256:63a05d84931221e25bf06c8813d6e6761f7fc8de8b5406b4d95488f929b32b65

Observation 466f143f-f364-478c-b65d-e4e593f55972 · outbound

This paper cites an unresolved cited work.

Prepending or Cross-Attention for Speech-to-Text? An Empirical Comparison Unresolved cited work

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-10T22:18:01.687235Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T22:18:01.687235Z digest=sha256:25513a724e44b98354155cdecbff856937d2410d4a66ea625dc9796b3f66710b

Observation 36cd5e72-c0b7-4a5e-96fc-e1f6397f3ac3 · outbound

This paper cites COSMIC: Data Efficient Instruction-tuning For Speech In-Context Learning.

Prepending or Cross-Attention for Speech-to-Text? An Empirical Comparison COSMIC: Data Efficient Instruction-tuning For Speech In-Context Learning

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-10T22:18:01.693105Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T22:18:01.693105Z digest=sha256:b0a7fab119cf393a48eaa091e81231b0228842d0d47c9ef3f969ee29e2fb2cf0

Observation 246de955-aec3-410f-a4e0-3c73c6c6d1a6 · outbound

This paper cites an unresolved cited work.

Prepending or Cross-Attention for Speech-to-Text? An Empirical Comparison Unresolved cited work

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-10T22:18:01.698618Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T22:18:01.698618Z digest=sha256:085898ac24cb88aab27acb6940e5804681af3746043e0579ec5a7274d31e7eab

Observation e4057a89-793b-438e-a24d-9f910f5a11fc · outbound

This paper cites How "Real" is Your Real-Time Simultaneous Speech-to-Text Translation System?.

Prepending or Cross-Attention for Speech-to-Text? An Empirical Comparison How "Real" is Your Real-Time Simultaneous Speech-to-Text Translation System?

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-10T22:18:01.704779Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T22:18:01.704779Z digest=sha256:aac969fa9b4c76de46e158b67d9f089a0ace761816c49a6f0e44f7e24cb871d8

Observation f719ed18-019d-4485-ba60-dfa1a8cbb67e · outbound

This paper cites Park, William Chan, Yu Zhang, Chung-Cheng Chiu, Barret Zoph, Ekin D.

Prepending or Cross-Attention for Speech-to-Text? An Empirical Comparison Park, William Chan, Yu Zhang, Chung-Cheng Chiu, Barret Zoph, Ekin D

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-10T22:18:01.710676Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T22:18:01.710676Z digest=sha256:84b5ffedc8dbaf89bdec17eadaa551e00e992d6bd11f4f037b10d8d9f9c03e17

Observation 3209d94a-e7e1-45ff-859a-b2e20fa5ceba · outbound

This paper cites an unresolved cited work.

Prepending or Cross-Attention for Speech-to-Text? An Empirical Comparison Unresolved cited work

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-10T22:18:01.718564Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T22:18:01.718564Z digest=sha256:c14ef5cb7aae088d621aed4d6155de27062b6fca277ffeebe8c6e0c300355bfa

Observation 4d804dca-fb92-44b3-8cc1-071290405279 · outbound

This paper cites an unresolved cited work.

Prepending or Cross-Attention for Speech-to-Text? An Empirical Comparison Unresolved cited work

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-10T22:18:01.724578Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T22:18:01.724578Z digest=sha256:3b38809e8287ebb8126811cf4342fd016340a049524aee147f2f65f3c0cb0b86

Observation 69a02b05-5d6d-4b4e-b1ee-e0c005a3b919 · outbound

This paper cites an unresolved cited work.

Prepending or Cross-Attention for Speech-to-Text? An Empirical Comparison Unresolved cited work

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-10T22:18:01.730703Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T22:18:01.730703Z digest=sha256:433a6781f3ef58d0472fad799772e61bae286ab461ca8cce26ad2d6d9d15a37e

Observation edcbf2c9-c3c0-4f04-bd45-b7937242395b · outbound

This paper cites an unresolved cited work.

Prepending or Cross-Attention for Speech-to-Text? An Empirical Comparison Unresolved cited work

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-10T22:18:01.739051Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T22:18:01.739051Z digest=sha256:c761cc98ded65cbe7f562890460a3d7bc39afe98cc4d0deb39ac6188aaa64de9

Observation 4da58c74-06dc-4733-93b5-c4d7593c3af8 · outbound

This paper cites an unresolved cited work.

Prepending or Cross-Attention for Speech-to-Text? An Empirical Comparison Unresolved cited work

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-10T22:18:01.749335Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T22:18:01.749335Z digest=sha256:161f1dff5ce72bba678737199c6ebca9982d0d95260724fe2dc10835534c684a

Observation 1fba9d1e-d065-4ba1-949c-e8de8356654b · outbound

This paper cites an unresolved cited work.

Prepending or Cross-Attention for Speech-to-Text? An Empirical Comparison Unresolved cited work

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-10T22:18:01.759295Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T22:18:01.759295Z digest=sha256:b39b180640f672b883323ea32b6c1fed72214b8649ed223dd05590079e8e7876

Observation 6a6f444c-2183-44cf-ab36-322bbcb87733 · outbound

This paper cites Chameleon: Mixed-Modal Early-Fusion Foundation Models.

Prepending or Cross-Attention for Speech-to-Text? An Empirical Comparison Chameleon: Mixed-Modal Early-Fusion Foundation Models

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-10T22:18:01.765430Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T22:18:01.765430Z digest=sha256:5113a241e0440633e881590c95bb02a7a4ad8bff832495cd7599a00458180874

Observation 4ed36dde-8573-4f60-8e79-f392e325972d · outbound

This paper cites Gemma: Open Models Based on Gemini Research and Technology.

Prepending or Cross-Attention for Speech-to-Text? An Empirical Comparison Gemma: Open Models Based on Gemini Research and Technology

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-10T22:18:01.773692Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T22:18:01.773692Z digest=sha256:bc83cde4e184c7a8da32873e7b35649498226a324f87a26e341ca757e6845533

Observation 46f9c9c7-73d9-47be-b23f-91df747c0546 · outbound

This paper cites Llama 2: Open Foundation and Fine-Tuned Chat Models.

Prepending or Cross-Attention for Speech-to-Text? An Empirical Comparison Llama 2: Open Foundation and Fine-Tuned Chat Models

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-10T22:18:01.779841Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T22:18:01.779841Z digest=sha256:66e89e40b529e937410a9507195bb56e3902c2834ec6049c50bd6ae4c357a960

Observation 203e68d0-e43c-453f-94bb-ab6fe7c00206 · outbound

This paper cites G \'a llego, Jos \'e A.

Prepending or Cross-Attention for Speech-to-Text? An Empirical Comparison G \'a llego, Jos \'e A

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-10T22:18:01.786265Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T22:18:01.786265Z digest=sha256:209425990056b86ea329edc9c4c40ae4fac15b13f3ad296e8cd6b00924cfbb2b

Observation 7a2a09cc-e600-4eb8-af59-8b27d1c38858 · outbound

This paper cites Gállego, José A.

Prepending or Cross-Attention for Speech-to-Text? An Empirical Comparison Gállego, José A

Reference 60

Resolution
verified exact
doi, observed 2026-08-10T22:18:01.970986Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-10T22:18:01.791567Z digest=sha256:b35ac4cc3ae8bb1c37a5269d90df481b08cc90c9b531c6828a2e28747c57ec44

Observation e53bfefd-adc3-4cb6-9abb-cdd0ff7e0f80 · outbound

This paper cites Decoder-only Architecture for Speech Recognition with CTC Prompts and Text Data Augmentation.

Prepending or Cross-Attention for Speech-to-Text? An Empirical Comparison Decoder-only Architecture for Speech Recognition with CTC Prompts and Text Data Augmentation

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-10T22:18:01.797969Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T22:18:01.797969Z digest=sha256:3c02ef940461998518301c3ac97d5e29dd12c07e828ab3695b24d4a199d30f81

Observation 968d45f1-92ec-4447-9021-36170cbc3b77 · outbound

This paper cites Decoder-only Architecture for Streaming End-to-end Speech Recognition.

Prepending or Cross-Attention for Speech-to-Text? An Empirical Comparison Decoder-only Architecture for Streaming End-to-end Speech Recognition

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-10T22:18:01.804040Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T22:18:01.804040Z digest=sha256:d0b177be3bafee46cb06980841311d20909714508e6b1f5cb2cf4140b990dd4b

Observation c7c732f0-f88c-4ff2-8108-13a27d9f7fbe · outbound

This paper cites an unresolved cited work.

Prepending or Cross-Attention for Speech-to-Text? An Empirical Comparison Unresolved cited work

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-10T22:18:01.811218Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T22:18:01.811218Z digest=sha256:febd35d74895ad25f41af1f5bc41d9800d0162b56fede04688b252b1613cc589

Observation 0b180781-470c-4722-afd9-cc36cdc536ee · outbound

This paper cites How to Connect Speech Foundation Models and Large Language Models? What Matters and What Does Not.

Prepending or Cross-Attention for Speech-to-Text? An Empirical Comparison How to Connect Speech Foundation Models and Large Language Models? What Matters and What Does Not

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-10T22:18:01.816674Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T22:18:01.816674Z digest=sha256:a62c1dedb67af290cb8df14c304925d43ec781eef15e67672103e2fc42080e5c

Observation 644d796c-cb4b-48ad-a834-32c75b6bac46 · outbound

This paper cites an unresolved cited work.

Prepending or Cross-Attention for Speech-to-Text? An Empirical Comparison Unresolved cited work

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-10T22:18:01.824647Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T22:18:01.824647Z digest=sha256:357707f69046eb529a6f60ec4989dbcc6238efbaf7e04f4cb44d15e21e4bbc40

Observation 50f35035-05eb-4bf6-8366-b0ebf5ae317a · outbound

This paper cites an unresolved cited work.

Prepending or Cross-Attention for Speech-to-Text? An Empirical Comparison Unresolved cited work

Reference 66

Resolution
unresolved
raw_fallback, observed 2026-08-10T22:18:03.265035Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-10T22:18:01.831175Z digest=sha256:68cbe0bae86e767a7badd5af837633f67dd5699b380bf8dc22dbb1280679d41d

Observation 750102ff-069e-4043-b423-4e8dbce94f38 · outbound

This paper cites SLM: Bridge the thin gap between speech and text foundation models.

Prepending or Cross-Attention for Speech-to-Text? An Empirical Comparison SLM: Bridge the thin gap between speech and text foundation models

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-10T22:18:01.839297Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T22:18:01.839297Z digest=sha256:0f8b8d5553b893cd8e04d54ecb817b4535c5443fdb4a5cc253e1e10ceb79b3fc

Observation 0fccce70-8841-4921-8626-f3f84e2f7aa7 · outbound

This paper cites Sequence-to-Sequence Models Can Directly Translate Foreign Speech.

Prepending or Cross-Attention for Speech-to-Text? An Empirical Comparison Sequence-to-Sequence Models Can Directly Translate Foreign Speech

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-10T22:18:01.845871Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T22:18:01.845871Z digest=sha256:cf21e2838cd20b72e28ae9b83f2f0326e1256beda1808f029b4f16f38758d20e

Observation 393cfcda-844f-4c8b-8c67-a64940f5100e · outbound

This paper cites an unresolved cited work.

Prepending or Cross-Attention for Speech-to-Text? An Empirical Comparison Unresolved cited work

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-10T22:18:01.851106Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T22:18:01.851106Z digest=sha256:b499ffdf6b189319eedd28af627d980f167b586a3e2790ceea44c901624f6ad3

Observation b3d35543-044d-4a02-98fb-3b5d6cea7074 · outbound

This paper cites an unresolved cited work.

Prepending or Cross-Attention for Speech-to-Text? An Empirical Comparison Unresolved cited work

Reference 70

Resolution
verified exact
doi, observed 2026-08-10T22:18:01.937062Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-10T22:18:01.857584Z digest=sha256:f8796f60a4a8cfe99053bc2e109189a3f4492193298a51c5d9bcec0bef8c8280

Observation facb8f58-7fed-4995-8aa3-8c9320ef2013 · outbound

This paper cites EMMeTT: Efficient Multimodal Machine Translation Training.

Prepending or Cross-Attention for Speech-to-Text? An Empirical Comparison EMMeTT: Efficient Multimodal Machine Translation Training

Reference 71

Resolution
verified exact
local_arxiv, observed 2026-08-10T22:18:02.395738Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-10T22:18:01.863187Z digest=sha256:ddb2d9f75f3c788b34a26c952934bed25d1d7055c9007f281f63df58d5396f74

Observation 88b24748-0963-43be-b838-277307e7072f · outbound

This paper cites Understanding Knowledge Distillation in Non-autoregressive Machine Translation.

Prepending or Cross-Attention for Speech-to-Text? An Empirical Comparison Understanding Knowledge Distillation in Non-autoregressive Machine Translation

Reference 72

Resolution
unresolved
no resolver link, observed 2026-08-10T22:18:01.868319Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T22:18:01.868319Z digest=sha256:233a655babd7a43abbd733c0eab35bc597afa4f0a1c74cc6ab9f1815266ed800

Observation 90e60d0c-b511-4a28-bb23-4298f8d61ce6 · outbound

This paper cites Contrastive Learning for Task-Independent SpeechLLM-Pretraining.

Prepending or Cross-Attention for Speech-to-Text? An Empirical Comparison Contrastive Learning for Task-Independent SpeechLLM-Pretraining

Reference 73

Resolution
unresolved
no resolver link, observed 2026-08-10T22:18:01.876712Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T22:18:01.876712Z digest=sha256:99100dab67345d7fae33b596a7adc5ea0315f114296bb131c7ef36fe162f3eeb

Pith citing papers

No inbound Pith citation observations are available.