Pith. sign in

Paper Citation Record · LEDGER

Granite-speech: open-source speech-aware LLMs with strong English ASR capabilities

As of 18 August 2026, this Paper Citation Record lists 36 of 36 outbound references and 15 inbound Pith citation observations for arXiv:2505.08699.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.08699 v2

Coverage vector

measured 36 of 36 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-15T21:52:53.499480Z

measured 51 of 51 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-17T06:30:58.91139+00:00

measured 15 of 15 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-03T12:18:37.983799Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

36 of 36 outbound references displayed

  • verified exact0
  • verified fuzzy11
  • unresolved25
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

1
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

Observation f0ce88bc-771c-41df-abbe-9e8c4fe461e4 · outbound

This paper cites Gemini: A Family of Highly Capable Multimodal Models.

Granite-speech: open-source speech-aware LLMs with strong English ASR capabilities Gemini: A Family of Highly Capable Multimodal Models

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-15T21:52:53.346794Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:52:53.346794Z digest=sha256:65e54409fe2188dae9f26d3e0ae08bf894a147e6dffcf789cb5977b77ff67e3a

Observation 2ce94490-140a-4cdf-b031-aa9bb2ee6bb0 · outbound

This paper cites Mini-Omni: Language Models Can Hear, Talk While Thinking in Streaming.

Granite-speech: open-source speech-aware LLMs with strong English ASR capabilities Mini-Omni: Language Models Can Hear, Talk While Thinking in Streaming

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-15T21:52:53.351695Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:52:53.351695Z digest=sha256:32c39ab7b6a19b2d349702fc86b46f5e45101b7c9c98d715d2d76cb56bbe004c

Observation cec19f15-ca5a-41d5-9ea3-32c04a626a07 · outbound

This paper cites On The Landscape of Spoken Language Models: A Comprehensive Survey.

Granite-speech: open-source speech-aware LLMs with strong English ASR capabilities On The Landscape of Spoken Language Models: A Comprehensive Survey

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-15T21:52:53.355988Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:52:53.355988Z digest=sha256:39825dce780dba24af8224111b0280366d8bcd51822d5588e41ab165ad84c113

Observation 23f4e3f6-5db9-4bf7-81b7-63eb789b5f10 · outbound

This paper cites The Llama 3 Herd of Models.

Granite-speech: open-source speech-aware LLMs with strong English ASR capabilities The Llama 3 Herd of Models

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-15T21:52:53.360572Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:52:53.360572Z digest=sha256:d8d731920061e815410ebd6079422b6872792af039e8951eeec9f1e9da7ab7ea

Observation cb523cb1-2fc1-4d46-9972-a48ceda923f9 · outbound

This paper cites Phi-4-mini technical report: Compact yet powerful multimodal language models via mixture- of-loras,.

Granite-speech: open-source speech-aware LLMs with strong English ASR capabilities Phi-4-mini technical report: Compact yet powerful multimodal language models via mixture- of-loras,

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:52:54.072375Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T21:52:53.365324Z digest=sha256:52d20cdc1ef34ef829a6bdbd74e50d463af5d9197fbc0d236e2f57096879de71

Observation cb2d118d-d7ea-407c-8df0-8562a1164e0d · outbound

This paper cites An embarrassingly simple approach for llm with strong asr capacity,.

Granite-speech: open-source speech-aware LLMs with strong English ASR capabilities An embarrassingly simple approach for llm with strong asr capacity,

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:52:54.058518Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T21:52:53.369632Z digest=sha256:780b02a2a1215bb79a3fe5aeec502251cb42e5f7b42616a49b0eb59d069a1d70

Observation ce574837-53bc-4329-8cae-b3f73e5e1243 · outbound

This paper cites Salmonn: Towards generic hearing abilities for large language models,.

Granite-speech: open-source speech-aware LLMs with strong English ASR capabilities Salmonn: Towards generic hearing abilities for large language models,

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:52:54.045899Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T21:52:53.374119Z digest=sha256:88ebbaee5d7ee735625272ae6285798b9d132463ec9607caa82ef00376341bf0

Observation 43c88141-f0d6-4e02-91db-da23cde1c89d · outbound

This paper cites Qwen2-Audio Technical Report.

Granite-speech: open-source speech-aware LLMs with strong English ASR capabilities Qwen2-Audio Technical Report

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-15T21:52:53.378348Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:52:53.378348Z digest=sha256:8ff28f599d73bfbd1176dd75d04e59e5ffa617801b3bd5a6130b758b1ee18043

Observation 597bacf4-ffa5-44dc-b3f4-6f9b9e2d2807 · outbound

This paper cites MuTox: Universal MUltilingual Audio-based TOXicity Dataset and Zero-shot Detector.

Granite-speech: open-source speech-aware LLMs with strong English ASR capabilities MuTox: Universal MUltilingual Audio-based TOXicity Dataset and Zero-shot Detector

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-15T21:52:53.382546Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:52:53.382546Z digest=sha256:c2d6043cedf39d9191d4520d6d640b4e72c13535602e7c084d302e5887741acc

Observation 0b04b37b-d2f0-4758-a944-eb730058e271 · outbound

This paper cites AlignFormer: Modality Matching Can Achieve Better Zero-shot Instruction-Following Speech-LLM.

Granite-speech: open-source speech-aware LLMs with strong English ASR capabilities AlignFormer: Modality Matching Can Achieve Better Zero-shot Instruction-Following Speech-LLM

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-15T21:52:53.386966Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:52:53.386966Z digest=sha256:e26d3795026a67c84926e701955ecefeba2ad19b5978d5f765a188055b9861c2

Observation 73c36d68-d274-435c-88ca-f5db26321d45 · outbound

This paper cites MLS: A Large-Scale Multilingual Dataset for Speech Research.

Granite-speech: open-source speech-aware LLMs with strong English ASR capabilities MLS: A Large-Scale Multilingual Dataset for Speech Research

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-15T21:52:53.391559Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:52:53.391559Z digest=sha256:27533c0752fc562674097169d0795b3d64f6e3d42697cb6c26dbce0c551f89f9

Observation c62cb1a9-98d8-40f0-9cf9-153aa3c703e3 · outbound

This paper cites GigaSpeech: An Evolving, Multi-domain ASR Corpus with 10,000 Hours of Transcribed Audio.

Granite-speech: open-source speech-aware LLMs with strong English ASR capabilities GigaSpeech: An Evolving, Multi-domain ASR Corpus with 10,000 Hours of Transcribed Audio

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-15T21:52:53.397615Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:52:53.397615Z digest=sha256:2dde560032cfca8b2347b50746e74eb0cfdf26861ef9bf9484a9065b5a4908fa

Observation b2829192-216c-4a50-bb96-e906d00539d5 · outbound

This paper cites Common Voice: A Massively-Multilingual Speech Corpus.

Granite-speech: open-source speech-aware LLMs with strong English ASR capabilities Common Voice: A Massively-Multilingual Speech Corpus

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-15T21:52:53.402250Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:52:53.402250Z digest=sha256:a83f43fd3c5afea7bf54a7b78d5b51c3f86d87761cd218d450dced1795eea528

Observation 860b0fe1-93cb-4f3d-99e9-3ad8b988d53d · outbound

This paper cites Librispeech: an asr corpus based on public domain audio books,.

Granite-speech: open-source speech-aware LLMs with strong English ASR capabilities Librispeech: an asr corpus based on public domain audio books,

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-15T21:52:53.406720Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:52:53.406720Z digest=sha256:b378f8bb372fc38dfcb64a4d04bda7e89b2e14e1d34584f7f05b8ed3d7c72919

Observation 4dbce808-0594-4b3e-8ae8-6d608a832352 · outbound

This paper cites VoxPopuli: A Large-Scale Multilingual Speech Corpus for Representation Learning, Semi-Supervised Learning and Interpretation.

Granite-speech: open-source speech-aware LLMs with strong English ASR capabilities VoxPopuli: A Large-Scale Multilingual Speech Corpus for Representation Learning, Semi-Supervised Learning and Interpretation

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-15T21:52:53.410512Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:52:53.410512Z digest=sha256:5eeb7b059a5fc71640296de63e01429c12e30c40e305cd50a529bc8de70decf1

Observation 93f4c749-9c7b-42aa-a20a-7e3e81f12690 · outbound

This paper cites The ami meeting corpus,.

Granite-speech: open-source speech-aware LLMs with strong English ASR capabilities The ami meeting corpus,

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:52:54.022384Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T21:52:53.415132Z digest=sha256:521eb41af7e43c9f43aa5a35d751aad194c3b4c926f21a85795a4aa2bec933d5

Observation 10675d0f-1d20-4b90-878b-74fe856ec64b · outbound

This paper cites Yodas: Youtube-oriented dataset for audio and speech,.

Granite-speech: open-source speech-aware LLMs with strong English ASR capabilities Yodas: Youtube-oriented dataset for audio and speech,

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:52:54.008398Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T21:52:53.419728Z digest=sha256:012e9891a7bbceb0e9ae4d08125b9b5babd80e4a6038bfb55faa51aa14d9c5e6

Observation 2af46f37-0c48-4b83-99a6-ea7f5710996e · outbound

This paper cites SPGISpeech: 5,000 hours of transcribed financial audio for fully formatted end-to-end speech recognition.

Granite-speech: open-source speech-aware LLMs with strong English ASR capabilities SPGISpeech: 5,000 hours of transcribed financial audio for fully formatted end-to-end speech recognition

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-15T21:52:53.423620Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:52:53.423620Z digest=sha256:c82f100520aedceee27e1899dabd669ed36b9af86481fcf9c062f111a6af30c0

Observation 3afaf5a2-62b6-4f6b-a3c3-2ab947cbd7d6 · outbound

This paper cites Switchboard: Telephone speech corpus for research and development,.

Granite-speech: open-source speech-aware LLMs with strong English ASR capabilities Switchboard: Telephone speech corpus for research and development,

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:52:53.992358Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T21:52:53.427999Z digest=sha256:658974baf91dbd4ca8d4a04fb543a563d44b91df9e1269285c71d18d3540adca

Observation d0331f27-359c-43f5-a581-d8ca9342329b · outbound

This paper cites The fisher corpus: A resource for the next generations of speech-to-text.

Granite-speech: open-source speech-aware LLMs with strong English ASR capabilities The fisher corpus: A resource for the next generations of speech-to-text

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-15T21:52:53.433226Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:52:53.433226Z digest=sha256:3019e0b17d41a8b0cf03c010d92bea0155cd6c9c6c4f19bc0e55f8d60951ea18

Observation 09139005-8832-4ad4-87ab-d9c7b3017d56 · outbound

This paper cites Automatic speech recognition performance on a voicemail transcription task,.

Granite-speech: open-source speech-aware LLMs with strong English ASR capabilities Automatic speech recognition performance on a voicemail transcription task,

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:52:53.968927Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T21:52:53.438206Z digest=sha256:1f4139a693e0111c2c807037247a7b962a93fdcc26dbc91a08e697aafdaa0230

Observation 33654b2f-1065-47b5-a3aa-2c0e7450496b · outbound

This paper cites Ted-lium: an automatic speech recognition dedicated corpus.

Granite-speech: open-source speech-aware LLMs with strong English ASR capabilities Ted-lium: an automatic speech recognition dedicated corpus

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:52:53.952292Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T21:52:53.441980Z digest=sha256:801ed85a1d4e4f7bbc72506c40c69ac2bdd508ec95b249a0822eb7a1d48da724

Observation 12f3d3f5-9e33-42c4-83c2-806a408c3274 · outbound

This paper cites Conformer: Convolution-augmented transformer for speech recognition,.

Granite-speech: open-source speech-aware LLMs with strong English ASR capabilities Conformer: Convolution-augmented transformer for speech recognition,

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-15T21:52:53.445669Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:52:53.445669Z digest=sha256:1d0d0134214daead308c8eee97648238b98498be5ec4ddd9d6a7f4f72c11f443

Observation d72047bd-102f-44c4-9ab6-0fa34fdccb1c · outbound

This paper cites Google USM: Scaling Automatic Speech Recognition Beyond 100 Languages.

Granite-speech: open-source speech-aware LLMs with strong English ASR capabilities Google USM: Scaling Automatic Speech Recognition Beyond 100 Languages

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-15T21:52:53.449723Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:52:53.449723Z digest=sha256:cc86dc1869ebfce16dd2d0b6b61ba1cfd3b421935a8a5c3bfebdab48ddc12873

Observation b0d139f5-92ff-43f9-be9c-05e5b2970ce8 · outbound

This paper cites Relaxing the conditional independence assumption of ctc-based asr by conditioning on intermediate predictions,.

Granite-speech: open-source speech-aware LLMs with strong English ASR capabilities Relaxing the conditional independence assumption of ctc-based asr by conditioning on intermediate predictions,

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:52:53.937016Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T21:52:53.453651Z digest=sha256:8448538942c67be602dd3c2c0afc80b81db1d04cb740a96c8de38ac19c922bfc

Observation a36927d8-4819-476d-bcec-02b4b3e33403 · outbound

This paper cites Specaugment: A simple data augmentation method for automatic speech recognition,.

Granite-speech: open-source speech-aware LLMs with strong English ASR capabilities Specaugment: A simple data augmentation method for automatic speech recognition,

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-15T21:52:53.457463Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:52:53.457463Z digest=sha256:835c60b56bf8ae19fd2a74464f6690f95c336db1f9bafa60884c81d8877affef

Observation 94aaac3d-c411-413d-af7c-8a0019520b4b · outbound

This paper cites Blip-2: Bootstrapping language- image pre-training with frozen image encoders and large language models,.

Granite-speech: open-source speech-aware LLMs with strong English ASR capabilities Blip-2: Bootstrapping language- image pre-training with frozen image encoders and large language models,

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-15T21:52:53.461600Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:52:53.461600Z digest=sha256:787f4c4325d107f90da424540b5f4d42bda4a8d970657641352fd69f579c9511

Observation 8e2c80bd-c28e-4abb-ad6b-c60c48a28863 · outbound

This paper cites CTRL: A Conditional Transformer Language Model for Controllable Generation.

Granite-speech: open-source speech-aware LLMs with strong English ASR capabilities CTRL: A Conditional Transformer Language Model for Controllable Generation

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-15T21:52:53.465845Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:52:53.465845Z digest=sha256:ac1e7b1dd86305296a02d68af5de055c11edde48e7e392bd88fbb71ddc8eeef8

Observation e60dadd5-2439-4d47-9ba5-499fe5cce6ab · outbound

This paper cites Covost 2 and massively multilingual speech translation.

Granite-speech: open-source speech-aware LLMs with strong English ASR capabilities Covost 2 and massively multilingual speech translation

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:52:53.910207Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T21:52:53.469827Z digest=sha256:d448299a98120877def30ba0053ef85360432acefa51c472853650c3daaf5954

Observation 93146828-f972-49d4-bbea-4aab46f489e0 · outbound

This paper cites Phi-4 Technical Report.

Granite-speech: open-source speech-aware LLMs with strong English ASR capabilities Phi-4 Technical Report

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-15T21:52:53.473880Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:52:53.473880Z digest=sha256:3078305cf874cb392d493e36a87d14b319b1261e3b2f4dd0df3a4cc83af86ff5

Observation 6005560c-182a-46d4-9536-0a01c83892b6 · outbound

This paper cites Making Monolingual Sentence Embeddings Multilingual using Knowledge Distillation.

Granite-speech: open-source speech-aware LLMs with strong English ASR capabilities Making Monolingual Sentence Embeddings Multilingual using Knowledge Distillation

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-15T21:52:53.477932Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:52:53.477932Z digest=sha256:957ebea0259f0ae7a2595eb1b67ad70b76eeb4668e9a9299bbe7b7d7498717e9

Observation 502cc7a6-425c-4255-b28e-df989ad13d68 · outbound

This paper cites Madlad-400: A multilingual and document-level large audited dataset,.

Granite-speech: open-source speech-aware LLMs with strong English ASR capabilities Madlad-400: A multilingual and document-level large audited dataset,

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-15T21:52:53.482350Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:52:53.482350Z digest=sha256:d9b7429c3efaecafa5e7cc7b8e0ec2ebcef89301e87e313b5b0acd4e36746259

Observation a63f6ee0-374b-49b7-b375-85d08fea060d · outbound

This paper cites Fleurs: Few-shot learning evaluation of universal representations of speech,.

Granite-speech: open-source speech-aware LLMs with strong English ASR capabilities Fleurs: Few-shot learning evaluation of universal representations of speech,

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-15T21:52:53.486162Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:52:53.486162Z digest=sha256:68f736b40b1ce3ab91dc57090d3eb8fd277677c28f3bfb5d78e99cc98874d6e8

Observation 5c6622fe-f227-491e-a667-268cc1f97c06 · outbound

This paper cites Bold: Dataset and metrics for measuring biases in open-ended language generation,.

Granite-speech: open-source speech-aware LLMs with strong English ASR capabilities Bold: Dataset and metrics for measuring biases in open-ended language generation,

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-15T21:52:53.490661Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:52:53.490661Z digest=sha256:9cd12f4945efd783ab714d8d17622a36678917b90ef223885387835db02fe627

Observation 2393d253-c8ed-4797-b656-8061a375544a · outbound

This paper cites Unveiling Safety Vulnerabilities of Large Language Models.

Granite-speech: open-source speech-aware LLMs with strong English ASR capabilities Unveiling Safety Vulnerabilities of Large Language Models

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-15T21:52:53.494772Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:52:53.494772Z digest=sha256:2696a43c90b68abb9545d264ef795c3572583189380feccf3ccb7d5a7faceb6d

Observation cefda020-9789-4074-a462-be030d2b2df0 · outbound

This paper cites Toxigen: A large-scale machine-generated dataset for implicit and adversarial hate speech detection,.

Granite-speech: open-source speech-aware LLMs with strong English ASR capabilities Toxigen: A large-scale machine-generated dataset for implicit and adversarial hate speech detection,

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:52:53.871963Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T21:52:53.499480Z digest=sha256:b5f9ee9d92a48cce27b88fa58ac6a1c7c97aca81f9d3d823b6a2bb6d987c1a41

Pith citing papers

Observation 22ffc734-21bc-448e-a66a-fcf96a976f1b · inbound

A Text-To-Text Alignment Algorithm for Better Evaluation of Modern Speech Recognition Systems cites this paper.

A Text-To-Text Alignment Algorithm for Better Evaluation of Modern Speech Recognition Systems Granite-speech: open-source speech-aware LLMs with strong English ASR capabilities

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-05-18T13:12:37.206996Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-18T13:12:34.387094Z digest=sha256:747fe3fb4a211b43631bee0ba221569be6bc42f0e49ed6a4eff3fc5f0feb402e

Observation e6790552-d8e3-47ce-ac4c-714f6a30f0cd · inbound

SARA: Stress Test Reasoning in Audio Deepfake Detection cites this paper.

SARA: Stress Test Reasoning in Audio Deepfake Detection Granite-speech: open-source speech-aware LLMs with strong English ASR capabilities

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-03T12:18:37.983799Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T12:18:37.983799Z digest=sha256:bf78723c640e8d19f39be7c95717361a02bbeb1c6ee751dff5f602314534856e

Observation 18061900-3098-4b7d-9cc9-a622bba4be20 · inbound

LLMs and Speech: Integration vs. Combination cites this paper.

LLMs and Speech: Integration vs. Combination Granite-speech: open-source speech-aware LLMs with strong English ASR capabilities

Reference 20

Resolution
verified exact
arxiv_id, observed 2026-05-15T10:45:28.325752Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-15T10:41:22.138517Z digest=sha256:8215077ab996aa1a0cf5a6b20f6e58e59f94698697a4f249a944a9607f9d4a8e

Observation 16985378-7e67-4464-912a-37f983a69116 · inbound

LLMs and Speech: Integration vs. Combination cites this paper.

LLMs and Speech: Integration vs. Combination Granite-speech: open-source speech-aware LLMs with strong English ASR capabilities

Reference 20

Resolution
unresolved
no resolver link, observed 2026-07-14T20:46:17.286576Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T20:46:17.286576Z digest=sha256:3c5fac8e8d4275da248c7566e6f764a17b939c0bf544ac3e2e964f9eaa450a6a

Observation 4df159a1-f3ed-4d07-8975-23a9ebc674ea · inbound

Contextual Biasing for ASR in Speech LLM with Common Word Cues and Bias Word Position Prediction cites this paper.

Contextual Biasing for ASR in Speech LLM with Common Word Cues and Bias Word Position Prediction Granite-speech: open-source speech-aware LLMs with strong English ASR capabilities

Reference 35

Resolution
verified exact
arxiv_id, observed 2026-05-10T14:35:34.019668Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-10T14:35:20.515352Z digest=sha256:0c2d1160a1e65d196431fce9d17449494eeaa4472716d95ca4ed2fa39ffae2a5

Observation a31a449a-6343-4493-8af2-a46a04d1d513 · inbound

Leveraging LLM-GNN Integration for Open-World Question Answering over Knowledge Graphs cites this paper.

Leveraging LLM-GNN Integration for Open-World Question Answering over Knowledge Graphs Granite-speech: open-source speech-aware LLMs with strong English ASR capabilities

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-05-10T13:25:26.699284Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-10T13:20:48.652042Z digest=sha256:28fe50238351cd40607b5ef6e836af68c1cf442d812bd89e19308cf0df5bffec

Observation 7bf99927-7cc4-45e0-85a0-5e3dfd5f5c54 · inbound

Do LLM Decoders Listen Fairly? Benchmarking How Language Model Priors Shape Bias in Speech Recognition cites this paper.

Do LLM Decoders Listen Fairly? Benchmarking How Language Model Priors Shape Bias in Speech Recognition Granite-speech: open-source speech-aware LLMs with strong English ASR capabilities

Reference 42

Resolution
verified exact
arxiv_id, observed 2026-05-11T14:31:04.605261Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-05-09T21:45:45.422465Z digest=sha256:2cb4a443c4ecc2d1ad0e83454dd4bee4e3978e536902760264787131845ae5b6

Observation e6a51871-0ab5-4b1f-bc3d-348f12742567 · inbound

In-Sync: Adaptation of Speech Aware Large Language Models for ASR with Word Level Timestamp Predictions cites this paper.

In-Sync: Adaptation of Speech Aware Large Language Models for ASR with Word Level Timestamp Predictions Granite-speech: open-source speech-aware LLMs with strong English ASR capabilities

Reference 26

Resolution
verified exact
arxiv_id, observed 2026-05-10T13:35:26.712969Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-10T13:25:50.524448Z digest=sha256:f923c0c5a74e6a9b76e7f6e3a61b267e4e51be9a1fa3d50a4aa872b3a41d2734

Observation b2a77d1b-79d6-4a98-a185-6d978f28ec75 · inbound

AppTek Call-Center Dialogues: A Multi-Accent Long-Form Benchmark for English ASR cites this paper.

AppTek Call-Center Dialogues: A Multi-Accent Long-Form Benchmark for English ASR Granite-speech: open-source speech-aware LLMs with strong English ASR capabilities

Reference 19

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T09:46:27.868915Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-07T09:12:14.094145Z digest=sha256:12a72012ac9ff76080e82568b73c0e76fcfe3ea5041dbaaad64b01c58d6c4096

Observation 874f631d-0f08-44f5-8d27-7950535caca7 · inbound

SALSA: Speech Aware LLM Adaptation via Learned Steering Activation Vectors cites this paper.

SALSA: Speech Aware LLM Adaptation via Learned Steering Activation Vectors Granite-speech: open-source speech-aware LLMs with strong English ASR capabilities

Reference 82

Resolution
verified exact
arxiv_id, observed 2026-06-28T19:32:35.154982Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-06-28T19:23:19.422735Z digest=sha256:eb361df95170dda5c3d84fdf00e2bca9bf19ae604f0ee92e51c43de44afae90d

Observation 5ee582ce-42d8-4a83-9417-3dada4a240f3 · inbound

Afrispeech Semantics: Evaluating Audio Semantic Reasoning in Spoken Language Models Across Domains and Accents cites this paper.

Afrispeech Semantics: Evaluating Audio Semantic Reasoning in Spoken Language Models Across Domains and Accents Granite-speech: open-source speech-aware LLMs with strong English ASR capabilities

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-06-30T22:15:05.353310Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-06-30T22:11:44.891731Z digest=sha256:58290afcec5c93bd398e2a91c421d05505b38bed36f3ddb4e6b00fa39fbb354c

Observation b356711d-5b63-473c-8af0-7f025d922e29 · inbound

ESPnet3: Infrastructure for Scalable Speech and Audio Research in the Foundation Model Era cites this paper.

ESPnet3: Infrastructure for Scalable Speech and Audio Research in the Foundation Model Era Granite-speech: open-source speech-aware LLMs with strong English ASR capabilities

Reference 19

Resolution
verified exact
arxiv_id, observed 2026-07-04T08:19:44.203712Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-06-26T11:54:08.457573Z digest=sha256:b67a15ee6957eb95bc2d8b07f0113555620ba31378db47ad5c1114f7e331015e

Observation ba8605d6-31bf-4caf-b7d9-7e994a65b86e · inbound

From Text Metrics to Model Internals: A Study of Whisper ASR Hallucination Detection cites this paper.

From Text Metrics to Model Internals: A Study of Whisper ASR Hallucination Detection Granite-speech: open-source speech-aware LLMs with strong English ASR capabilities

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-07-04T12:29:52.175868Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-06-26T06:40:59.791646Z digest=sha256:84e0fbc3ae788050d7673d846ae949549b1e390f543de9b615d11ceada356de1

Observation 98c3d0c3-1d63-4b9d-a12f-d15df855b221 · inbound

Rethinking Speech-LLM Integration for ASR: Effective Joint Speech-Text Training by Interleaving cites this paper.

Rethinking Speech-LLM Integration for ASR: Effective Joint Speech-Text Training by Interleaving Granite-speech: open-source speech-aware LLMs with strong English ASR capabilities

Reference 23

Resolution
verified exact
arxiv_id, observed 2026-07-03T15:18:32.627821Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-07-03T15:17:52.966144Z digest=sha256:c25cea416584063058fa602fb7787822724fe82dbe63d701347cff09a1dce2fb

Observation df653451-76cf-4596-bad0-1ab5861d3ae5 · inbound

CHILDES-Aligned: A Curated Children's Speech Dataset via Multi-Model Timestamp Ensembling cites this paper.

CHILDES-Aligned: A Curated Children's Speech Dataset via Multi-Model Timestamp Ensembling Granite-speech: open-source speech-aware LLMs with strong English ASR capabilities

Reference 54

Resolution
unresolved
no resolver link, observed 2026-07-12T00:48:21.716770Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T00:48:21.716770Z digest=sha256:7d99171ff238facb37e1d07faed13be0797a630be4fb21ae295bc7f454d9136d