Pith. sign in

Paper Citation Record · LEDGER

LLM can Read Spectrogram: Encoder-free Speech-Language Modeling

As of 4 August 2026, this Paper Citation Record lists 46 of 46 outbound references and 1 inbound Pith citation observation for arXiv:2606.10231.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2606.10231 v3

Coverage vector

measured 46 of 46 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-07-03T23:55:11.698728Z

measured 47 of 47 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-03T06:30:56.289259+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-06-27T14:40:40.673091Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-07-03T03:47:35.741508Z

Reference resolution

46 of 46 outbound references displayed

  • verified exact11
  • verified fuzzy34
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch1

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 6c7dbc0f-113f-46a8-a039-5381085db54b · outbound

This paper cites Prompting large language models with speech recognition abilities,.

LLM can Read Spectrogram: Encoder-free Speech-Language Modeling Prompting large language models with speech recognition abilities,

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-07-04T23:30:12.217537Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-07-03T23:55:11.698728Z digest=sha256:dbea819de2c2101e58afcba3c9179316b2251b54919b86ec2971e728ef9245c3

Observation 362812f5-2384-4cf9-a5a6-0cbfab55b69b · outbound

This paper cites Train short, infer long: Speech-llm enables zero-shot streamable joint asr and diarization on long audio,.

LLM can Read Spectrogram: Encoder-free Speech-Language Modeling Train short, infer long: Speech-llm enables zero-shot streamable joint asr and diarization on long audio,

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-07-04T23:30:12.236804Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-07-03T23:55:11.698728Z digest=sha256:81c0de136cfc79e53558605d615ed9a8e826704844148ba4cbc1aa6e8947456e

Observation f63f758d-1806-41b7-a231-a08cdb05976c · outbound

This paper cites Phi-4-Mini Technical Report: Compact yet Powerful Multimodal Language Models via Mixture-of-LoRAs.

LLM can Read Spectrogram: Encoder-free Speech-Language Modeling Phi-4-Mini Technical Report: Compact yet Powerful Multimodal Language Models via Mixture-of-LoRAs

Reference 3

Resolution
verified exact
local_arxiv, observed 2026-07-03T23:59:06.133918Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-07-03T23:55:11.698728Z digest=sha256:e24f2db859b24b5fff417590656238e323e89c4d93a3e1fa70210a04d32bcc1a

Observation 7f1e8f5b-0ac7-4b00-85b5-511704b01579 · outbound

This paper cites Robust speech recognition via large-scale weak super- vision,.

LLM can Read Spectrogram: Encoder-free Speech-Language Modeling Robust speech recognition via large-scale weak super- vision,

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-07-04T23:30:12.228477Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-07-03T23:55:11.698728Z digest=sha256:1bc494c78ca55363eaf81933d68c3acd39e98130753718e1caa594c1a9da7477

Observation 16a6441f-d2ab-486e-9524-56a2493356c6 · outbound

This paper cites Conformer: Convolution- augmented Transformer for speech recognition,.

LLM can Read Spectrogram: Encoder-free Speech-Language Modeling Conformer: Convolution- augmented Transformer for speech recognition,

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-07-04T23:30:12.267204Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-07-03T23:55:11.698728Z digest=sha256:47fa6e7d4e8c62837a8d3f59a671114c84af3d36f28c6667a783c8710e4c9877

Observation 633d82b5-ff05-4003-b51c-45826bcee60a · outbound

This paper cites Lora: Low-rank adaptation of large language models,.

LLM can Read Spectrogram: Encoder-free Speech-Language Modeling Lora: Low-rank adaptation of large language models,

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-07-04T23:30:12.241210Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-07-03T23:55:11.698728Z digest=sha256:27762e1b1fbbe2213ece70454268fe0f84bf4f7b540cd5d3946d6a5a0e0bb545

Observation 0c773364-73a7-4419-8000-8f07df4dee5b · outbound

This paper cites Fuyu-8B: A multimodal architecture for AI agents,.

LLM can Read Spectrogram: Encoder-free Speech-Language Modeling Fuyu-8B: A multimodal architecture for AI agents,

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-07-04T23:30:12.269052Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-07-03T23:55:11.698728Z digest=sha256:65f642cd58239b8dbad6bb1559e69ee78ccbd2006f32241673d2b5108e767eb7

Observation f8c075f6-cd2f-4784-9eaa-be8e92ef0bb3 · outbound

This paper cites Unveiling encoder-free vision- language models,.

LLM can Read Spectrogram: Encoder-free Speech-Language Modeling Unveiling encoder-free vision- language models,

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-07-04T23:30:12.224327Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-07-03T23:55:11.698728Z digest=sha256:9a7734be2fbbe992700821914c30dbd54fa0cc578dec52dff902b2bd3cbf1eef

Observation 3818f7cc-e2e1-40b3-a188-9ab78e46e9f9 · outbound

This paper cites Breaking the encoder barrier for seamless video-language understanding,.

LLM can Read Spectrogram: Encoder-free Speech-Language Modeling Breaking the encoder barrier for seamless video-language understanding,

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-07-04T23:30:12.270914Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-07-03T23:55:11.698728Z digest=sha256:d125b42fc1913d07ade32977bb45826b2fa37797aba033409816e6c17bbd593c

Observation d9042f0a-d363-4951-9f4d-4f4bee470f23 · outbound

This paper cites Open ASR Leaderboard: Towards Reproducible and Transparent Multilingual Speech Recognition Evaluation.

LLM can Read Spectrogram: Encoder-free Speech-Language Modeling Open ASR Leaderboard: Towards Reproducible and Transparent Multilingual Speech Recognition Evaluation

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-07-03T23:59:06.149805Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-07-03T23:55:11.698728Z digest=sha256:d3fd11f2f4aad051c568c53cb2bcac38d57daff0c2b499b1b0018699fa125a41

Observation 8eb00970-682f-4c67-8983-b2666ff1d1bf · outbound

This paper cites Joint audio and speech understanding,.

LLM can Read Spectrogram: Encoder-free Speech-Language Modeling Joint audio and speech understanding,

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-07-04T23:30:12.232480Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-07-03T23:55:11.698728Z digest=sha256:71a9b583118e8bba0ac97284d1ba6558b3a8b8006aebe1b788babe2b73669088

Observation 23532edd-5bb7-444c-9652-a3a9a0ed5c70 · outbound

This paper cites WavLLM: Towards robust and adaptive speech large language model,.

LLM can Read Spectrogram: Encoder-free Speech-Language Modeling WavLLM: Towards robust and adaptive speech large language model,

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-07-04T23:30:12.221450Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-07-03T23:55:11.698728Z digest=sha256:2b9b29885d37cf38e46521a22d23cb1d29cea2ff9cc2ceb93c758558a0060eb6

Observation 31d2e2a3-768f-4f47-8752-d29ec784e632 · outbound

This paper cites MMAU: A Massive Multi-Task Audio Understanding and Reasoning Benchmark.

LLM can Read Spectrogram: Encoder-free Speech-Language Modeling MMAU: A Massive Multi-Task Audio Understanding and Reasoning Benchmark

Reference 13

Resolution
verified exact
local_arxiv, observed 2026-07-03T23:59:06.146529Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-07-03T23:55:11.698728Z digest=sha256:e08d0fc2012919ef7609e674dce69b9d188f7f7fd8c487a545d3928aca9ce321

Observation 97abc9c8-d525-49fe-82e9-dddb473999ee · outbound

This paper cites Measuring massive multitask language understanding,.

LLM can Read Spectrogram: Encoder-free Speech-Language Modeling Measuring massive multitask language understanding,

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-07-04T23:30:12.208955Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-07-03T23:55:11.698728Z digest=sha256:42c265278522d870615b283ce378d79c2b04d0cd3f8f905f7add569f1bfa7b5f

Observation 8ec34912-5455-47e6-bf94-208642a8b258 · outbound

This paper cites SALMONN: Towards generic hearing abilities for large language models,.

LLM can Read Spectrogram: Encoder-free Speech-Language Modeling SALMONN: Towards generic hearing abilities for large language models,

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-07-04T23:30:12.243113Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-07-03T23:55:11.698728Z digest=sha256:e8f3b56512f4856f887d9187ffae7819ccc48bfe04bc88297daeefad795e6359

Observation c005f65f-b993-4f04-8b2f-d7d7e6c3a5c9 · outbound

This paper cites HuBERT: Self- supervised speech representation learning by masked prediction of hidden units,.

LLM can Read Spectrogram: Encoder-free Speech-Language Modeling HuBERT: Self- supervised speech representation learning by masked prediction of hidden units,

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-07-04T23:30:12.245066Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-07-03T23:55:11.698728Z digest=sha256:162c5fb6bf8c971ab3c5b4479c363a7e43dce1ba0a78295642a0c2e89d3b1dad

Observation f0a27b5a-5033-4ea5-82e0-4a60a0f47b41 · outbound

This paper cites SpeechGPT: Empowering large language models with intrinsic cross- modal conversational abilities,.

LLM can Read Spectrogram: Encoder-free Speech-Language Modeling SpeechGPT: Empowering large language models with intrinsic cross- modal conversational abilities,

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-07-04T23:30:12.265271Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-07-03T23:55:11.698728Z digest=sha256:b78a9f1b0ea498b637ca2f33fe69594380988e8c8c431c7f3f992d46762d5f4c

Observation b98001c5-f2d2-40db-92fd-26808ecb7e7a · outbound

This paper cites V oxtLM: Unified decoder-only models for consolidating speech recognition/synthesis and speech/text continuation tasks,.

LLM can Read Spectrogram: Encoder-free Speech-Language Modeling V oxtLM: Unified decoder-only models for consolidating speech recognition/synthesis and speech/text continuation tasks,

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-07-04T23:30:12.251718Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-07-03T23:55:11.698728Z digest=sha256:04354daea8fa719a8c08042105485f182f1110631efd633147fd57926a8811b1

Observation 5cc7e207-52eb-4a29-abee-57a4d35f7540 · outbound

This paper cites Spirit LM: Interleaved Spoken and Written Language Model.

LLM can Read Spectrogram: Encoder-free Speech-Language Modeling Spirit LM: Interleaved Spoken and Written Language Model

Reference 19

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T23:59:06.152286Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-07-03T23:55:11.698728Z digest=sha256:9c4f3be09e9626065bebe9eabef752371464e159e7d2b670e61f626da7a824a8

Observation 65fc4da2-194f-4003-9faf-ee978db382d8 · outbound

This paper cites dMel: Speech Tokenization made Simple.

LLM can Read Spectrogram: Encoder-free Speech-Language Modeling dMel: Speech Tokenization made Simple

Reference 20

Resolution
verified exact
arxiv_id, observed 2026-07-03T23:59:06.142179Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-07-03T23:55:11.698728Z digest=sha256:00d4ad9dc65f91d90621a8316d8c15fe0391f91014313f4c7cdd9fc1b2937e30

Observation d32b1a49-e06f-4e4e-83fa-4eecf6a6ceb8 · outbound

This paper cites Autoregressive speech synthesis without vector quantization,.

LLM can Read Spectrogram: Encoder-free Speech-Language Modeling Autoregressive speech synthesis without vector quantization,

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-07-04T23:30:12.261575Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-07-03T23:55:11.698728Z digest=sha256:64d11a0b25ef0d504f71d92f40e7295db8d4d17c6236a43c7dbb3cb0eefc5c84

Observation d428e616-bdc4-4573-8fdd-3db1e99ad2d2 · outbound

This paper cites VibeVoice Technical Report.

LLM can Read Spectrogram: Encoder-free Speech-Language Modeling VibeVoice Technical Report

Reference 22

Resolution
verified exact
arxiv_id, observed 2026-07-03T23:59:06.137863Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-07-03T23:55:11.698728Z digest=sha256:a7fad55b1ebb58986f1e7c41193868e11f144eb6623127ed40c267481ab6669f

Observation 602bb37a-9e30-4800-b9cb-b59167b4e5a2 · outbound

This paper cites HiFi-GAN: Generative adversarial net- works for efficient and high fidelity speech synthesis,.

LLM can Read Spectrogram: Encoder-free Speech-Language Modeling HiFi-GAN: Generative adversarial net- works for efficient and high fidelity speech synthesis,

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-07-04T23:30:12.256283Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-07-03T23:55:11.698728Z digest=sha256:66a5663d6e930afd20f6adcaf0f1a1eb13ad3203afd48b841d487079f73e33d1

Observation a62efd64-2a19-453b-a266-65cdb60508f8 · outbound

This paper cites UTMOS: UTokyo-SaruLab system for V oiceMOS chal- lenge 2022.

LLM can Read Spectrogram: Encoder-free Speech-Language Modeling UTMOS: UTokyo-SaruLab system for V oiceMOS chal- lenge 2022

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-07-04T23:30:12.230564Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-07-03T23:55:11.698728Z digest=sha256:7a5658b3fa9bcea5249e06248e33adbc5ae45bfae76a03ca6fd25f60fcee50c2

Observation 22b2f636-672c-4097-996f-adc3d54d0ec1 · outbound

This paper cites Simple and effective V AE training with calibrated decoders,.

LLM can Read Spectrogram: Encoder-free Speech-Language Modeling Simple and effective V AE training with calibrated decoders,

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-07-04T23:30:12.222434Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-07-03T23:55:11.698728Z digest=sha256:6fa8f68fa5ab10652688406ecfe6ad188bc0fd74bb5aad0f47d79ddc517b63f2

Observation b55fc75b-67a4-49dd-a29b-6a47fc00ccde · outbound

This paper cites Tuna-2: Pixel Embeddings Beat Vision Encoders for Multimodal Understanding and Generation.

LLM can Read Spectrogram: Encoder-free Speech-Language Modeling Tuna-2: Pixel Embeddings Beat Vision Encoders for Multimodal Understanding and Generation

Reference 26

Resolution
verified exact
local_arxiv, observed 2026-07-03T23:59:06.141726Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-07-03T23:55:11.698728Z digest=sha256:596f0b2464b4ef08a51b5b6cf57e56b1d22d40388ef142e5b10e0dab2713fd64

Observation 921868c2-f87e-478e-92ee-4b377756fa90 · outbound

This paper cites Representation Forcing for Bottleneck-Free Unified Multimodal Models.

LLM can Read Spectrogram: Encoder-free Speech-Language Modeling Representation Forcing for Bottleneck-Free Unified Multimodal Models

Reference 27

Resolution
verified exact
local_arxiv, observed 2026-07-03T23:59:06.125682Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-07-03T23:55:11.698728Z digest=sha256:9c609af45990df1c0b365057a1eff14a6ddaf3be281f2c7f28bfe5c6f6bf052b

Observation 364d29c4-d1b9-4b58-bcf5-2f1a3c2c3484 · outbound

This paper cites Toward Native Multimodal Modeling: A Roadmap.

LLM can Read Spectrogram: Encoder-free Speech-Language Modeling Toward Native Multimodal Modeling: A Roadmap

Reference 28

Resolution
verified exact
local_arxiv, observed 2026-07-03T23:59:06.145302Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-07-03T23:55:11.698728Z digest=sha256:4aae1ef0eb2711e5bf5ed40cb20d6fb793341e21c57c9f3e17d50c7ee8b9b5ec

Observation 91525c26-1a95-48b6-beb4-eab7ef6532ee · outbound

This paper cites MELA-TTS: Joint Transformer-diffusion model with rep- resentation alignment for speech synthesis,.

LLM can Read Spectrogram: Encoder-free Speech-Language Modeling MELA-TTS: Joint Transformer-diffusion model with rep- resentation alignment for speech synthesis,

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-07-04T23:30:12.256872Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-07-03T23:55:11.698728Z digest=sha256:754be7aee69424342712830a2493d7018493bce428edb96aa594c3c8a471009a

Observation 2f1c0eac-61fb-4a84-a34d-5fead643ddc3 · outbound

This paper cites MELD: Mel-Spectrogram-Based Speech Language Modeling with Discrete Latent Variables.

LLM can Read Spectrogram: Encoder-free Speech-Language Modeling MELD: Mel-Spectrogram-Based Speech Language Modeling with Discrete Latent Variables

Reference 30

Resolution
verified exact
local_arxiv, observed 2026-07-03T23:59:06.129686Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-07-03T23:55:11.698728Z digest=sha256:28a756ebde9793cdda4a097fac4cd7ee513f59ac512259cf8bdfd7ab48425cf7

Observation 316c97cf-0705-4025-bcf8-747a28f89ecb · outbound

This paper cites WavFlow: Audio Generation in Waveform Space.

LLM can Read Spectrogram: Encoder-free Speech-Language Modeling WavFlow: Audio Generation in Waveform Space

Reference 31

Resolution
verified exact
local_arxiv, observed 2026-07-03T23:59:06.154351Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-07-03T23:55:11.698728Z digest=sha256:c20998b693cc02ec0743c8cb22eaefb9a4bc378f9e6c0402cb6bb72438c1b755

Observation a640653e-73f3-4e2a-bf69-dea3768caee4 · outbound

This paper cites AlignFormer: Modality matching can achieve better zero-shot instruction-following speech-LLM,.

LLM can Read Spectrogram: Encoder-free Speech-Language Modeling AlignFormer: Modality matching can achieve better zero-shot instruction-following speech-LLM,

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-07-04T23:30:12.213302Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-07-03T23:55:11.698728Z digest=sha256:4460626a0665ccdae09d9a1876e90240c4333da34ec9e47a22b74796104f7b24

Observation 98f9e433-62b7-4192-bf1b-a7c0ab39a803 · outbound

This paper cites Towards efficient speech-text jointly decoding within one speech language model,.

LLM can Read Spectrogram: Encoder-free Speech-Language Modeling Towards efficient speech-text jointly decoding within one speech language model,

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-07-04T23:30:12.232996Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-07-03T23:55:11.698728Z digest=sha256:4b14b5d0605075cc99aa9524d669122daa27616fcc577799e48cbe174126d1b8

Observation 00602b0c-7b9a-47de-9d87-d76afee85d89 · outbound

This paper cites SLM-S2ST: A multimodal language model for direct speech-to-speech translation,.

LLM can Read Spectrogram: Encoder-free Speech-Language Modeling SLM-S2ST: A multimodal language model for direct speech-to-speech translation,

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-07-04T23:30:12.263409Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-07-03T23:55:11.698728Z digest=sha256:e4fb4772acc92e5eab83d343d24f38380cf162dfe4b737206728f763e45e664c

Observation b6e83f15-8f58-41bb-9e0e-05f1e7b7d919 · outbound

This paper cites Speech llms are contextual reasoning transcribers,.

LLM can Read Spectrogram: Encoder-free Speech-Language Modeling Speech llms are contextual reasoning transcribers,

Reference 35

Resolution
verified exact
arxiv_id, observed 2026-07-03T23:59:06.115752Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-07-03T23:55:11.698728Z digest=sha256:92bc9218496696be96b63cbbe78c62ddbcd1613c548e193af3c2197360ffc06f

Observation 2efeada1-7bb5-4578-ade6-d48da811fa11 · outbound

This paper cites LibriSpeech: An ASR corpus based on public domain audio books,.

LLM can Read Spectrogram: Encoder-free Speech-Language Modeling LibriSpeech: An ASR corpus based on public domain audio books,

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-07-04T23:30:12.215139Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-07-03T23:55:11.698728Z digest=sha256:1b73d9ec9c24ac29a31b6b724d116f7453b7d6407ef4ebdef39c357ba5be02a8

Observation dc4becae-bb29-46aa-b2f7-214c095c8661 · outbound

This paper cites LibriSpeech-PC: Benchmark for evaluation of punctuation and capitalization capabilities of end-to-end ASR models,.

LLM can Read Spectrogram: Encoder-free Speech-Language Modeling LibriSpeech-PC: Benchmark for evaluation of punctuation and capitalization capabilities of end-to-end ASR models,

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-07-04T23:30:12.249283Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-07-03T23:55:11.698728Z digest=sha256:2ad4a577021803d210352f0ab89f3e9e11e34044738ec5545794f2279521c187

Observation ebb12562-5a16-4232-9ae2-f90c82db2f79 · outbound

This paper cites GigaSpeech: An evolving, multi-domain ASR corpus with 10,000 hours of transcribed audio,.

LLM can Read Spectrogram: Encoder-free Speech-Language Modeling GigaSpeech: An evolving, multi-domain ASR corpus with 10,000 hours of transcribed audio,

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-07-04T23:30:12.254520Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-07-03T23:55:11.698728Z digest=sha256:daa8d30eb4a9ebd19aa0501d39998666a247a90e49eca08320909eb405d13a55

Observation 4b1f62fd-d442-4c0f-b03a-7cea879f530f · outbound

This paper cites MLS: A large-scale multilingual dataset for speech research,.

LLM can Read Spectrogram: Encoder-free Speech-Language Modeling MLS: A large-scale multilingual dataset for speech research,

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-07-04T23:30:12.258819Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-07-03T23:55:11.698728Z digest=sha256:c0268d51c89b76acea5a117a8418c1083f75bb922087b284b2d01c29dbae04b4

Observation 492e1106-ce4b-40c9-8369-edb966bdf933 · outbound

This paper cites SPGISpeech: 5,000 hours of transcribed financial audio for fully formatted end-to-end speech recognition,.

LLM can Read Spectrogram: Encoder-free Speech-Language Modeling SPGISpeech: 5,000 hours of transcribed financial audio for fully formatted end-to-end speech recognition,

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-07-04T23:30:12.240287Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-07-03T23:55:11.698728Z digest=sha256:87e1f5dd81c2cd10bfe0043136be98ce78cfa9b5582a5f13cfc86e32688c6ef4

Observation 7b1d7a7e-01ea-4736-ace2-a3e848b10ab1 · outbound

This paper cites Common V oice: A massively- multilingual speech corpus,.

LLM can Read Spectrogram: Encoder-free Speech-Language Modeling Common V oice: A massively- multilingual speech corpus,

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-07-04T23:30:12.249767Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-07-03T23:55:11.698728Z digest=sha256:b85a9b19638020bb5b4d45683eb4f1fe7baaad682dd3fb9b790971724aac4a34

Observation 89117a64-ea88-496d-a338-51ee90c19e8e · outbound

This paper cites V oxPopuli: A large-scale multilingual speech corpus for representation learning, semi-supervised learning and interpretation,.

LLM can Read Spectrogram: Encoder-free Speech-Language Modeling V oxPopuli: A large-scale multilingual speech corpus for representation learning, semi-supervised learning and interpretation,

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-07-04T23:30:12.253971Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-07-03T23:55:11.698728Z digest=sha256:b1ff44729d3ff3da54cab54d1ff6bfe96a4b9f7dea5bace51668e0e4d5c4320a

Observation 9e52c3f9-da2d-479e-94b7-d7be15153519 · outbound

This paper cites TED-LIUM 3: Twice as much data and corpus repartition for experi- ments on speaker adaptation,.

LLM can Read Spectrogram: Encoder-free Speech-Language Modeling TED-LIUM 3: Twice as much data and corpus repartition for experi- ments on speaker adaptation,

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-07-04T23:30:12.237577Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-07-03T23:55:11.698728Z digest=sha256:8e207ccf1204965ab90d93326b197fb1c4608a6e627881a0c0b41e3f50215876

Observation 9e55fff3-ec64-45ce-b4f4-110865708704 · outbound

This paper cites The AMI meeting corpus: A pre-announcement,.

LLM can Read Spectrogram: Encoder-free Speech-Language Modeling The AMI meeting corpus: A pre-announcement,

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-07-04T23:30:12.235426Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-07-03T23:55:11.698728Z digest=sha256:21e8939764383d6cc53c3130f5730f23f33379882be2b2754f59e2938d3aedcc

Observation 7d77cbd8-b221-4498-b2f8-54e189b1d620 · outbound

This paper cites Earnings-22: A practical benchmark for accents in the wild,.

LLM can Read Spectrogram: Encoder-free Speech-Language Modeling Earnings-22: A practical benchmark for accents in the wild,

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-07-04T23:30:12.230837Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-07-03T23:55:11.698728Z digest=sha256:be9c7af3fb5060d40260a03ae96cd2c092dbb4ae51526889f23dba7315745d1d

Observation 909e278b-f338-49a9-a4e4-df4c51425a3e · outbound

This paper cites FLEURS: Few-shot learning evaluation of universal representations of speech,.

LLM can Read Spectrogram: Encoder-free Speech-Language Modeling FLEURS: Few-shot learning evaluation of universal representations of speech,

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-07-04T23:30:12.219512Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-07-03T23:55:11.698728Z digest=sha256:cffd4b1728547ea305025b3d8d41f58baa42ceaba7653789ad3c0afb1de1c1cd

Pith citing papers

Observation 43da8cc7-5865-4229-a5de-de8f154919cc · inbound

LLM can Read Spectrogram: Encoder-free Speech-Language Modeling cites this paper.

LLM can Read Spectrogram: Encoder-free Speech-Language Modeling LLM can Read Spectrogram: Encoder-free Speech-Language Modeling

Reference 2

Resolution
metadata mismatch
local_arxiv, observed 2026-07-03T03:47:35.742884Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-06-27T14:40:40.673091Z digest=sha256:0aa8934b99ed727eb9123f059ff1ff86981ab023b0578f92db7984207fb10ca1