Pith. sign in

Paper Citation Record · LEDGER

Instituto de Telecomunica\c{c}\~oes at IWSLT 2025: Aligning Small-Scale Speech and Language Models for Speech-to-Text Learning

As of 7 August 2026, this Paper Citation Record lists 27 of 27 outbound references and 0 inbound Pith citation observations for arXiv:2506.17019.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.17019 v1

Coverage vector

measured 27 of 27 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T23:34:09.446488Z

measured 27 of 27 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

27 of 27 outbound references displayed

  • verified exact0
  • verified fuzzy7
  • unresolved20
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation f723c553-dd42-4d7d-b2bf-1f1b8bea7139 · outbound

This paper cites Phi-4-Mini Technical Report: Compact yet Powerful Multimodal Language Models via Mixture-of-LoRAs.

Instituto de Telecomunica\c{c}\~oes at IWSLT 2025: Aligning Small-Scale Speech and Language Models for Speech-to-Text Learning Phi-4-Mini Technical Report: Compact yet Powerful Multimodal Language Models via Mixture-of-LoRAs

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-06T23:34:06.447360Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:34:06.447360Z digest=sha256:8c0e0a484b7eacbadb8a242de2aa0fc12035247e7bd1e956ff459004c25316cc

Observation 88e5864e-a2e5-4be5-a0b6-af6a54601ce3 · outbound

This paper cites arXiv preprint arXiv:2503.10620.

Instituto de Telecomunica\c{c}\~oes at IWSLT 2025: Aligning Small-Scale Speech and Language Models for Speech-to-Text Learning arXiv preprint arXiv:2503.10620

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-06T23:34:06.624600Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:34:06.624600Z digest=sha256:1d01e64da37a8deb0408337627b513a7b624c227851192d995d25c03a102119b

Observation 7e6a1902-06e1-4559-98fb-e208df2133fc · outbound

This paper cites Qwen2-Audio Technical Report.

Instituto de Telecomunica\c{c}\~oes at IWSLT 2025: Aligning Small-Scale Speech and Language Models for Speech-to-Text Learning Qwen2-Audio Technical Report

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-06T23:34:06.960891Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:34:06.960891Z digest=sha256:0195b9bf97d7711fb5d25315a25fec03431c65acbd08f1db36dbc272c4715075

Observation 483e7c21-5dfc-4b52-bfac-70f87284c879 · outbound

This paper cites Qwen-Audio: Advancing Universal Audio Understanding via Unified Large-Scale Audio-Language Models.

Instituto de Telecomunica\c{c}\~oes at IWSLT 2025: Aligning Small-Scale Speech and Language Models for Speech-to-Text Learning Qwen-Audio: Advancing Universal Audio Understanding via Unified Large-Scale Audio-Language Models

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-06T23:34:07.072917Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:34:07.072917Z digest=sha256:fb11fd88373903e5b69ea9b23fff6297a4439e11107718bdcc2538c5a77b87b1

Observation a3ff050c-a152-471c-940b-a48d59925cba · outbound

This paper cites In 2022 IEEE Spoken Language Technology Workshop (SLT), pages 798–805.

Instituto de Telecomunica\c{c}\~oes at IWSLT 2025: Aligning Small-Scale Speech and Language Models for Speech-to-Text Learning In 2022 IEEE Spoken Language Technology Workshop (SLT), pages 798–805

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:34:11.169768Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T23:34:07.129512Z digest=sha256:2190da43ea9de799efe075512e737cb8d4a269fac20063f2930557cc73e9c7b0

Observation 3e69669d-9fc9-4cac-bb93-3d48cdd6be69 · outbound

This paper cites Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models.

Instituto de Telecomunica\c{c}\~oes at IWSLT 2025: Aligning Small-Scale Speech and Language Models for Speech-to-Text Learning Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-06T23:34:07.278388Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:34:07.278388Z digest=sha256:17bb9f5159bb064b2cc48d48441907a8a1f48ab5c316822c829e33501f6f3bee

Observation 6dcc4476-94e2-48e5-b951-dbf7ab1d4ac7 · outbound

This paper cites Gemma 2: Improving Open Language Models at a Practical Size.

Instituto de Telecomunica\c{c}\~oes at IWSLT 2025: Aligning Small-Scale Speech and Language Models for Speech-to-Text Learning Gemma 2: Improving Open Language Models at a Practical Size

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-06T23:34:07.562021Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:34:07.562021Z digest=sha256:e44cecf2955b80293bcca1f5e6c23fc212c57a4fb11de5cb51affe445e0b375b

Observation 36ee1ed8-a703-4db4-8cbf-4966191d8a52 · outbound

This paper cites The Llama 3 Herd of Models.

Instituto de Telecomunica\c{c}\~oes at IWSLT 2025: Aligning Small-Scale Speech and Language Models for Speech-to-Text Learning The Llama 3 Herd of Models

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-06T23:34:07.676909Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:34:07.676909Z digest=sha256:016270129563780dae1872806812e40dd3f886c741a14e83c99ce03df157762a

Observation c7cc3634-92b5-4eea-86da-8264703a0e80 · outbound

This paper cites In Findings of the Asso- ciation for Computational Linguistics: EMNLP 2024, pages 4552–4572, Miami, Florida, USA.

Instituto de Telecomunica\c{c}\~oes at IWSLT 2025: Aligning Small-Scale Speech and Language Models for Speech-to-Text Learning In Findings of the Asso- ciation for Computational Linguistics: EMNLP 2024, pages 4552–4572, Miami, Florida, USA

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:34:11.046606Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T23:34:07.937565Z digest=sha256:9b24a622231d208a70315b6c0d278e87d7a764b0b21a8908d91d2b69f393bfe7

Observation d1bc3b29-c5cb-4d4e-9183-02054e959fba · outbound

This paper cites Speech Translation with Large Language Models: An Industrial Practice.

Instituto de Telecomunica\c{c}\~oes at IWSLT 2025: Aligning Small-Scale Speech and Language Models for Speech-to-Text Learning Speech Translation with Large Language Models: An Industrial Practice

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-06T23:34:08.081674Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:34:08.081674Z digest=sha256:c4d018baa08669df13d3bf70d90b58cbc50b984f4873e154dc1cdc3f1bdf51dc

Observation c2cc3502-ee96-4697-b803-f288712f68e8 · outbound

This paper cites In ICASSP 2024-2024 IEEE Inter- national Conference on Acoustics, Speech and Signal Processing (ICASSP), pages 13326–13330.

Instituto de Telecomunica\c{c}\~oes at IWSLT 2025: Aligning Small-Scale Speech and Language Models for Speech-to-Text Learning In ICASSP 2024-2024 IEEE Inter- national Conference on Acoustics, Speech and Signal Processing (ICASSP), pages 13326–13330

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:34:10.732535Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T23:34:08.247414Z digest=sha256:2baf5f1b3c578e49d9e3b77900c9bdcb53e81cd7a985fa9f88a149d74b7154d7

Observation 912abaa8-18dd-420b-b3a2-6c4210d234ec · outbound

This paper cites EuroLLM: Multilingual Language Models for Europe.

Instituto de Telecomunica\c{c}\~oes at IWSLT 2025: Aligning Small-Scale Speech and Language Models for Speech-to-Text Learning EuroLLM: Multilingual Language Models for Europe

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-06T23:34:08.333794Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:34:08.333794Z digest=sha256:02e38847a7e9d4daf45c08ef460e95c0fd57c3466f0526f06f8998e4c899aa70

Observation 52b9ffae-d251-4592-8138-9f2843071aff · outbound

This paper cites an unresolved cited work.

Instituto de Telecomunica\c{c}\~oes at IWSLT 2025: Aligning Small-Scale Speech and Language Models for Speech-to-Text Learning Unresolved cited work

Reference 21

Resolution
unresolved
raw_fallback, observed 2026-08-06T23:34:10.304258Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T23:34:08.694602Z digest=sha256:8111015be94914d5aaa597b8487d74a0a366af8de5844ae1f6f1a3ea18f6e4db

Observation 8be90602-4665-47d1-9e30-66dedbaac60b · outbound

This paper cites In Proceedings of the Seventh Conference on Machine Translation (WMT) , pages 634–645, Abu Dhabi, United Arab Emirates (Hybrid).

Instituto de Telecomunica\c{c}\~oes at IWSLT 2025: Aligning Small-Scale Speech and Language Models for Speech-to-Text Learning In Proceedings of the Seventh Conference on Machine Translation (WMT) , pages 634–645, Abu Dhabi, United Arab Emirates (Hybrid)

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:34:10.105964Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T23:34:08.823192Z digest=sha256:2fb4af7b1f6397e4d12fa2e90986b0be8edf450422fa245458361ac7cf97f39b

Observation 3769e255-2bcd-4641-a8d3-7107add7efd4 · outbound

This paper cites AudioPaLM: A Large Language Model That Can Speak and Listen.

Instituto de Telecomunica\c{c}\~oes at IWSLT 2025: Aligning Small-Scale Speech and Language Models for Speech-to-Text Learning AudioPaLM: A Large Language Model That Can Speak and Listen

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-06T23:34:08.892983Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:34:08.892983Z digest=sha256:0bca7ae1294d9003275bbcecd95d3db2ecef537ea96e141d3e831feaa92ebc03

Observation ee251e65-6943-49ba-b351-b7e7d8353639 · outbound

This paper cites SALMONN: Towards Generic Hearing Abilities for Large Language Models.

Instituto de Telecomunica\c{c}\~oes at IWSLT 2025: Aligning Small-Scale Speech and Language Models for Speech-to-Text Learning SALMONN: Towards Generic Hearing Abilities for Large Language Models

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-06T23:34:08.989599Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:34:08.989599Z digest=sha256:3b7e8355029a84b001e9df847792cef598083d479a907528efc759b4fef09835

Observation 9d90560f-7ae2-47ac-a5cf-b99fa37eaee7 · outbound

This paper cites VoxPopuli: A Large-Scale Multilingual Speech Corpus for Representation Learning, Semi-Supervised Learning and Interpretation.

Instituto de Telecomunica\c{c}\~oes at IWSLT 2025: Aligning Small-Scale Speech and Language Models for Speech-to-Text Learning VoxPopuli: A Large-Scale Multilingual Speech Corpus for Representation Learning, Semi-Supervised Learning and Interpretation

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-06T23:34:09.173525Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:34:09.173525Z digest=sha256:8fc5b018940d7a8815831b8a164b2e2923254c0fef272727802dd81796c798d5

Observation 2cd2468b-8d69-4899-8a8f-edfd7ca0502b · outbound

This paper cites CoVoST 2 and Massively Multilingual Speech-to-Text Translation.

Instituto de Telecomunica\c{c}\~oes at IWSLT 2025: Aligning Small-Scale Speech and Language Models for Speech-to-Text Learning CoVoST 2 and Massively Multilingual Speech-to-Text Translation

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-06T23:34:09.324302Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:34:09.324302Z digest=sha256:63f0b0b55dce7277ce9f4d782ac47fb99f286bfe6919f8a44d5deb6e1b02f3be

Observation cc82fe11-6225-4091-9bae-df733af749d4 · outbound

This paper cites Qwen2.5 Technical Report.

Instituto de Telecomunica\c{c}\~oes at IWSLT 2025: Aligning Small-Scale Speech and Language Models for Speech-to-Text Learning Qwen2.5 Technical Report

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-06T23:34:09.446488Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:34:09.446488Z digest=sha256:413a8a345efca5c23790e7c10ca2a30fa01680a2bef445a0684a49cfedf59651

Observation 75a282bb-6d7c-4218-9015-e2b28c55ef58 · outbound

This paper cites In 2015 IEEE international conference on acoustics, speech and signal processing (ICASSP), pages 5206–5210.

Instituto de Telecomunica\c{c}\~oes at IWSLT 2025: Aligning Small-Scale Speech and Language Models for Speech-to-Text Learning In 2015 IEEE international conference on acoustics, speech and signal processing (ICASSP), pages 5206–5210

Reference 2015

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:34:10.497929Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T23:34:08.567919Z digest=sha256:b2a623ccf373a78c224a434539f5da650617157ad113af4ab1e87c605502882b

Observation 3cbfb931-1243-401a-9c4d-1451cfd44ed1 · outbound

This paper cites an unresolved cited work.

Instituto de Telecomunica\c{c}\~oes at IWSLT 2025: Aligning Small-Scale Speech and Language Models for Speech-to-Text Learning Unresolved cited work

Reference 2018

Resolution
unresolved
raw_fallback, observed 2026-08-06T23:34:10.891849Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T23:34:08.177645Z digest=sha256:47a0a76ff5953694d9ef6e9df94454e9750bffa39834ee84cd6e91f73943a0fd

Observation da49e256-227c-4898-9ed6-4c9792f3051c · outbound

This paper cites Conformer: Convolution-augmented Transformer for Speech Recognition.

Instituto de Telecomunica\c{c}\~oes at IWSLT 2025: Aligning Small-Scale Speech and Language Models for Speech-to-Text Learning Conformer: Convolution-augmented Transformer for Speech Recognition

Reference 2020

Resolution
unresolved
no resolver link, observed 2026-08-06T23:34:07.813102Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:34:07.813102Z digest=sha256:66765b11b045b9cb38b1f278d565734b771742cb3a95d6d6595bd48d1f18656a

Observation 82dc9500-d820-4672-8a3b-a1b4f3ff12c7 · outbound

This paper cites The People's Speech: A Large-Scale Diverse English Speech Recognition Dataset for Commercial Usage.

Instituto de Telecomunica\c{c}\~oes at IWSLT 2025: Aligning Small-Scale Speech and Language Models for Speech-to-Text Learning The People's Speech: A Large-Scale Diverse English Speech Recognition Dataset for Commercial Usage

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-06T23:34:07.432916Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:34:07.432916Z digest=sha256:f0cf174363f61a13a5da9baca5dbe294a0367fabbc29df3b408c097996443bef

Observation 97ce2ada-7205-422f-9073-2058f440d569 · outbound

This paper cites No Language Left Behind: Scaling Human-Centered Machine Translation.

Instituto de Telecomunica\c{c}\~oes at IWSLT 2025: Aligning Small-Scale Speech and Language Models for Speech-to-Text Learning No Language Left Behind: Scaling Human-Centered Machine Translation

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-06T23:34:08.449809Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:34:08.449809Z digest=sha256:f012f66d3cbc503ba2944a85f84acede1c89bc644d1f1f64a894b629d131b778

Observation 8117339d-05be-41c1-85d4-dc3e9f67baea · outbound

This paper cites Seamless: Multilingual Expressive and Streaming Speech Translation.

Instituto de Telecomunica\c{c}\~oes at IWSLT 2025: Aligning Small-Scale Speech and Language Models for Speech-to-Text Learning Seamless: Multilingual Expressive and Streaming Speech Translation

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-06T23:34:06.852611Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:34:06.852611Z digest=sha256:600a53350e660b5d0b1f6cf9d3fee4d1e357056ac9b577c239e2c14c4357e8ab

Observation cf7879b9-d990-47df-88d3-753e3a0f641b · outbound

This paper cites In Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing, pages 21318–21340, Miami, Florida, USA.

Instituto de Telecomunica\c{c}\~oes at IWSLT 2025: Aligning Small-Scale Speech and Language Models for Speech-to-Text Learning In Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing, pages 21318–21340, Miami, Florida, USA

Reference 2024

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:34:11.300091Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T23:34:06.715685Z digest=sha256:69aa8c2afefff1c41b38e90f5d04ebd31fe21daacdd1f04a918ec8268ff06ff9

Observation d21fd134-09ea-408a-994e-1be5c6334827 · outbound

This paper cites In Proceedings of the 22nd Interna- tional Conference on Spoken Language Translation (IWSLT 2025), Vienna, Austria (in-person and on- line).

Instituto de Telecomunica\c{c}\~oes at IWSLT 2025: Aligning Small-Scale Speech and Language Models for Speech-to-Text Learning In Proceedings of the 22nd Interna- tional Conference on Spoken Language Translation (IWSLT 2025), Vienna, Austria (in-person and on- line)

Reference 2025

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:34:11.419483Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T23:34:06.365008Z digest=sha256:7f68cd91f616bb685114d60e6d3c0e6bf8136022c611339aeb81c7b54ac43357

Pith citing papers

No inbound Pith citation observations are available.