Pith. sign in

Paper Citation Record · LEDGER

Speech-IFEval: Evaluating Instruction-Following and Quantifying Catastrophic Forgetting in Speech-Aware Language Models

As of 23 August 2026, this Paper Citation Record lists 39 of 39 outbound references and 7 inbound Pith citation observations for arXiv:2505.19037.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.19037 v1

Coverage vector

measured 39 of 39 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T14:24:35.057623Z

measured 46 of 46 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-23T06:30:58.430688+00:00

measured 7 of 7 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T14:24:32.537583Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-07-07T14:53:55.827006Z

Reference resolution

39 of 39 outbound references displayed

  • verified exact1
  • verified fuzzy17
  • unresolved19
  • parse uncertain0
  • malformed identifier2
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 2a6b7b26-a6d3-4ab5-961c-069dd69f4cf5 · outbound

This paper cites Speech-IFEval: Evaluating Instruction-Following and Quantifying Catastrophic Forgetting in Speech-Aware Language Models.

Speech-IFEval: Evaluating Instruction-Following and Quantifying Catastrophic Forgetting in Speech-Aware Language Models Speech-IFEval: Evaluating Instruction-Following and Quantifying Catastrophic Forgetting in Speech-Aware Language Models

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T14:24:32.537583Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:24:32.537583Z digest=sha256:d1b42f70a3429c8933db1f6284efbdf184bb863c65744f4fb9e17ef8f649a8b9

Observation 645f1457-e873-4d0c-abd9-065f813b9006 · outbound

This paper cites To assess their instruction-following capabilities, several bench- marks have been developed [27–29].

Speech-IFEval: Evaluating Instruction-Following and Quantifying Catastrophic Forgetting in Speech-Aware Language Models To assess their instruction-following capabilities, several bench- marks have been developed [27–29]

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:24:37.568904Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-07T14:24:32.606279Z digest=sha256:27bd786cb798fb427aed9a9ac876c765c2b585dce8f62346a82353a9e29a9b8b

Observation b79fd799-4ded-4b67-9ec4-f50950a2b1f6 · outbound

This paper cites an unresolved cited work.

Speech-IFEval: Evaluating Instruction-Following and Quantifying Catastrophic Forgetting in Speech-Aware Language Models Unresolved cited work

Reference 3

Resolution
malformed identifier
raw_fallback, observed 2026-08-07T14:24:37.551140Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-07T14:24:32.692638Z digest=sha256:f49102df9b55feb2223b15198d4462d69f270099aae3dc6edea4313fa0471c67

Observation b0caaaed-1d0c-46bf-be43-0a9af3964cb2 · outbound

This paper cites Model details As shown in Table 3, we evaluate publicly available SLMs [9–11, 13–16] alongside their corresponding LLM components [2, 6, 7], which serve as reference systems.

Speech-IFEval: Evaluating Instruction-Following and Quantifying Catastrophic Forgetting in Speech-Aware Language Models Model details As shown in Table 3, we evaluate publicly available SLMs [9–11, 13–16] alongside their corresponding LLM components [2, 6, 7], which serve as reference systems

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:24:37.531817Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-07T14:24:32.722816Z digest=sha256:9e7f0af4eb384801bb4cc58cd24641196cb59b16ec5b953751f3585c6a78b4ea

Observation 20672c36-746a-4d22-abc8-4c6153f0fb13 · outbound

This paper cites Results on Speech-IFEval Table 2 presents a comprehensive evaluation of Speech-IFEval, focusing exclusively on instruction-following ability while dis- regarding speech perception.

Speech-IFEval: Evaluating Instruction-Following and Quantifying Catastrophic Forgetting in Speech-Aware Language Models Results on Speech-IFEval Table 2 presents a comprehensive evaluation of Speech-IFEval, focusing exclusively on instruction-following ability while dis- regarding speech perception

Reference 5

Resolution
malformed identifier
raw_fallback, observed 2026-08-07T14:24:37.511838Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-07T14:24:32.757494Z digest=sha256:389e8ac55b075363a0b62265e9e6d51b126f03a0e9fad40048402a7d10aa2c10

Observation 886d49cc-d10f-45f7-aafa-9bf14b787058 · outbound

This paper cites This separation enables a more precise evaluation of whether and how SLMs retain their textual competencies af- ter speech-text training.

Speech-IFEval: Evaluating Instruction-Following and Quantifying Catastrophic Forgetting in Speech-Aware Language Models This separation enables a more precise evaluation of whether and how SLMs retain their textual competencies af- ter speech-text training

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:24:37.494175Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-07T14:24:32.841787Z digest=sha256:afdde5210048c3872a1df2e0c7cd7710e73b89e1f3ca007e81071f2f878f2902

Observation ed31a8db-fcc6-4ab7-8375-94b7b68b3dfe · outbound

This paper cites GPT-4 Technical Report.

Speech-IFEval: Evaluating Instruction-Following and Quantifying Catastrophic Forgetting in Speech-Aware Language Models GPT-4 Technical Report

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T14:24:32.959694Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:24:32.959694Z digest=sha256:0393ce324600ceb991db90052ea76e184a58047a34c4d9a278ee90d9b3dc9395

Observation b3177d3d-c58e-4962-8df1-29832954970a · outbound

This paper cites Qwen Technical Report.

Speech-IFEval: Evaluating Instruction-Following and Quantifying Catastrophic Forgetting in Speech-Aware Language Models Qwen Technical Report

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T14:24:33.034169Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:24:33.034169Z digest=sha256:d2d421fef73724d0ca4c10cbb33cc5b5898f4a41b7f8a9cc98030a64ba3e9df2

Observation 7f97b3a4-e44b-484d-90ae-53fa2119c4e0 · outbound

This paper cites Qwen2.5 Technical Report.

Speech-IFEval: Evaluating Instruction-Following and Quantifying Catastrophic Forgetting in Speech-Aware Language Models Qwen2.5 Technical Report

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T14:24:33.197355Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:24:33.197355Z digest=sha256:39214e60d7c1513cdab5005471dd47b0bc481b63a3575d56192859b47f9ca655

Observation e0960d5c-1757-4510-807f-811fa14e015a · outbound

This paper cites Llama 2: Open Foundation and Fine-Tuned Chat Models.

Speech-IFEval: Evaluating Instruction-Following and Quantifying Catastrophic Forgetting in Speech-Aware Language Models Llama 2: Open Foundation and Fine-Tuned Chat Models

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T14:24:33.272243Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:24:33.272243Z digest=sha256:b1f9f22d7c34aa966e557feeb7f0fb4f4dab1dfa4e303283c6b2b248c98b6f34

Observation b1b10236-a63c-468b-a77f-539b5803095f · outbound

This paper cites The Llama 3 Herd of Models.

Speech-IFEval: Evaluating Instruction-Following and Quantifying Catastrophic Forgetting in Speech-Aware Language Models The Llama 3 Herd of Models

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T14:24:33.346212Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:24:33.346212Z digest=sha256:0432d24026468fca1b60ddb8776ab10fae26909d7f1d617510a2700590a5c5b4

Observation 3b2ad39f-88a2-4679-a9fe-e9404ce3d280 · outbound

This paper cites Vicuna: An open-source chatbot impressing gpt-4 with 90%* chatgpt quality,.

Speech-IFEval: Evaluating Instruction-Following and Quantifying Catastrophic Forgetting in Speech-Aware Language Models Vicuna: An open-source chatbot impressing gpt-4 with 90%* chatgpt quality,

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T14:24:33.413497Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:24:33.413497Z digest=sha256:830b1be0807023295fb76680bacecb0296180e2e00799629f632a9b15cfc7be2

Observation 39f79dbb-a3d9-48b2-bac6-14fe91ccb93e · outbound

This paper cites Gemini: A Family of Highly Capable Multimodal Models.

Speech-IFEval: Evaluating Instruction-Following and Quantifying Catastrophic Forgetting in Speech-Aware Language Models Gemini: A Family of Highly Capable Multimodal Models

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T14:24:33.470714Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:24:33.470714Z digest=sha256:979df736bc4eaddcc58b4a0e3048f45c8c8a28312f1668715781911825259e74

Observation bf30c3c8-2ba5-426a-9aa2-36a01b5e3467 · outbound

This paper cites Qwen-Audio: Advancing Universal Audio Understanding via Unified Large-Scale Audio-Language Models.

Speech-IFEval: Evaluating Instruction-Following and Quantifying Catastrophic Forgetting in Speech-Aware Language Models Qwen-Audio: Advancing Universal Audio Understanding via Unified Large-Scale Audio-Language Models

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T14:24:33.559481Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:24:33.559481Z digest=sha256:ae8df781236617b6e42bf924d5d7334e01be46cd14c566b75757edf0a0ec0037

Observation 2a20b2b8-01bb-47b4-a429-7979f2e0a8a3 · outbound

This paper cites Qwen2-Audio Technical Report.

Speech-IFEval: Evaluating Instruction-Following and Quantifying Catastrophic Forgetting in Speech-Aware Language Models Qwen2-Audio Technical Report

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T14:24:33.631428Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:24:33.631428Z digest=sha256:7d8309f33c5abb30bbe571744fa2143d455c48797331f656aa7acd64bfa2547c

Observation 0593cf21-cea0-43b3-aead-54ce8edbed39 · outbound

This paper cites DeSTA2: Developing Instruction-Following Speech Language Model Without Speech Instruction-Tuning Data.

Speech-IFEval: Evaluating Instruction-Following and Quantifying Catastrophic Forgetting in Speech-Aware Language Models DeSTA2: Developing Instruction-Following Speech Language Model Without Speech Instruction-Tuning Data

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T14:24:33.707065Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:24:33.707065Z digest=sha256:66138487c54520393b5aca3af6f133182e1d4e04e58a0332d6ce5eb6e557f0b8

Observation fe847266-4cc8-46a1-b98d-39bbdb5c14e8 · outbound

This paper cites Desta: Enhancing speech language models through descriptive speech-text alignment,.

Speech-IFEval: Evaluating Instruction-Following and Quantifying Catastrophic Forgetting in Speech-Aware Language Models Desta: Enhancing speech language models through descriptive speech-text alignment,

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:24:37.465571Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-07T14:24:33.750426Z digest=sha256:8952ad8bd798dc81c323ac934752c3f6147d62bac808c84c859975a9620692da

Observation 49741d8f-c750-49ea-94fd-1277e830530b · outbound

This paper cites SALMONN: Towards generic hearing abilities for large language models,.

Speech-IFEval: Evaluating Instruction-Following and Quantifying Catastrophic Forgetting in Speech-Aware Language Models SALMONN: Towards generic hearing abilities for large language models,

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T14:24:33.808780Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:24:33.808780Z digest=sha256:abe85776f4c2665ac9873e789b648b014a848bf42790ac232b57a2007d9fc74f

Observation d81abbce-7f91-47df-b675-06c6dfdce61f · outbound

This paper cites Joint audio and speech understanding,.

Speech-IFEval: Evaluating Instruction-Following and Quantifying Catastrophic Forgetting in Speech-Aware Language Models Joint audio and speech understanding,

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:24:37.438341Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-07T14:24:33.870828Z digest=sha256:6b920c6af6a46eebe40f753ab671bd59303d81b0dd18de5a0353f0723253eed7

Observation 9b55b015-bacf-4433-979d-63512062722d · outbound

This paper cites Distilling an End-to-End Voice Assistant Without Instruction Training Data.

Speech-IFEval: Evaluating Instruction-Following and Quantifying Catastrophic Forgetting in Speech-Aware Language Models Distilling an End-to-End Voice Assistant Without Instruction Training Data

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T14:24:33.963171Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:24:33.963171Z digest=sha256:62c412cabac04cb029834eb281c80bb86f5570814cfd6c92bfefa1f29ec05aa6

Observation 71b93392-4140-458c-b4ae-827014a5d997 · outbound

This paper cites BLSP-Emo: Towards Empathetic Large Speech-Language Models.

Speech-IFEval: Evaluating Instruction-Following and Quantifying Catastrophic Forgetting in Speech-Aware Language Models BLSP-Emo: Towards Empathetic Large Speech-Language Models

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T14:24:34.022294Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:24:34.022294Z digest=sha256:1376082fc13715a47eb95947596c07283600f07b40ea1d387c370dfc858d2c8c

Observation 39dc1cdd-a81c-44c2-89b0-8999099d6d37 · outbound

This paper cites WavLLM: Towards Robust and Adaptive Speech Large Language Model.

Speech-IFEval: Evaluating Instruction-Following and Quantifying Catastrophic Forgetting in Speech-Aware Language Models WavLLM: Towards Robust and Adaptive Speech Large Language Model

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T14:24:34.067293Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:24:34.067293Z digest=sha256:8e14b08123636b247269870ef688d61ae8a3e864176b169606f7acdf0280ffb0

Observation 388bbf6e-702d-4033-a8b4-9560f4d0f11b · outbound

This paper cites Speech-copilot: Leveraging large language models for speech processing via task decomposition, modularization, and program generation,.

Speech-IFEval: Evaluating Instruction-Following and Quantifying Catastrophic Forgetting in Speech-Aware Language Models Speech-copilot: Leveraging large language models for speech processing via task decomposition, modularization, and program generation,

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T14:24:34.101553Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:24:34.101553Z digest=sha256:152698e87352aedbfc8f6f38ff465ccc7c27ac3fff0605ffedc0565734ef15a9

Observation b2e998b2-2c2f-400a-9e0b-91183918f54f · outbound

This paper cites Dynamic-superb: Towards a dynamic, col- laborative, and comprehensive instruction-tuning benchmark for speech,.

Speech-IFEval: Evaluating Instruction-Following and Quantifying Catastrophic Forgetting in Speech-Aware Language Models Dynamic-superb: Towards a dynamic, col- laborative, and comprehensive instruction-tuning benchmark for speech,

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:24:37.412066Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-07T14:24:34.176459Z digest=sha256:e811db4570267225b6697f6df5211440ab9f388818217e79940fb32d590d4b45

Observation 1b177bc9-34fe-4da1-b978-29237c697713 · outbound

This paper cites AIR-bench: Benchmarking large audio-language models via generative comprehension,.

Speech-IFEval: Evaluating Instruction-Following and Quantifying Catastrophic Forgetting in Speech-Aware Language Models AIR-bench: Benchmarking large audio-language models via generative comprehension,

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:24:37.396637Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-07T14:24:34.249708Z digest=sha256:a0a56f43583f6498ef7e2f44a7f8369e8dcc8c17944894ed7136e5583962e0a2

Observation 731379bd-e6a0-4014-bfab-9508deb50827 · outbound

This paper cites Dynamic-SUPERB phase-2: A collabora- tively expanding benchmark for measuring the capabilities of spo- ken language models with 180 tasks,.

Speech-IFEval: Evaluating Instruction-Following and Quantifying Catastrophic Forgetting in Speech-Aware Language Models Dynamic-SUPERB phase-2: A collabora- tively expanding benchmark for measuring the capabilities of spo- ken language models with 180 tasks,

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:24:37.380361Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-07T14:24:34.318244Z digest=sha256:fb11150149eb0020c1d3fd0b828df9682b144f9c88c2a51f1b35b99f253a8761

Observation 09c4d480-1e0c-4b01-b5e6-24f7f22a6f04 · outbound

This paper cites Beyond single-audio: Advancing multi-audio processing in audio large language models,.

Speech-IFEval: Evaluating Instruction-Following and Quantifying Catastrophic Forgetting in Speech-Aware Language Models Beyond single-audio: Advancing multi-audio processing in audio large language models,

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:24:37.297798Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-07T14:24:34.366953Z digest=sha256:c99683d4d7f1e213a995e6d884f8cf3981846be4b5bc6d87cc44f5dd25a241e5

Observation de478c3f-6be0-4767-bfbd-9d0a3fde9c88 · outbound

This paper cites Benchmarking Open-ended Audio Dialogue Understanding for Large Audio-Language Models.

Speech-IFEval: Evaluating Instruction-Following and Quantifying Catastrophic Forgetting in Speech-Aware Language Models Benchmarking Open-ended Audio Dialogue Understanding for Large Audio-Language Models

Reference 29

Resolution
verified exact
local_arxiv, observed 2026-08-07T14:24:35.328214Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-07T14:24:34.431110Z digest=sha256:35bb10bb0e1b003084ae392c904209154066baa582b82ad86130b23e02abacd7

Observation 3cce925a-6056-447c-98bd-937c24735c65 · outbound

This paper cites MMAU: A mas- sive multi-task audio understanding and reasoning benchmark,.

Speech-IFEval: Evaluating Instruction-Following and Quantifying Catastrophic Forgetting in Speech-Aware Language Models MMAU: A mas- sive multi-task audio understanding and reasoning benchmark,

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:24:37.106691Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-07T14:24:34.486517Z digest=sha256:9970c20304d67e767aa24deb31e21c6ca158990f8d7505600519efcabe2eae41

Observation 91d20702-f918-456a-839c-3d382f463706 · outbound

This paper cites An Empirical Investigation of Catastrophic Forgetting in Gradient-Based Neural Networks.

Speech-IFEval: Evaluating Instruction-Following and Quantifying Catastrophic Forgetting in Speech-Aware Language Models An Empirical Investigation of Catastrophic Forgetting in Gradient-Based Neural Networks

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T14:24:34.571005Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:24:34.571005Z digest=sha256:3a5fcd999fa244d947a7f7433c17f96aaf67b655d760a84b3c1374835ba30afe

Observation 8c2aa1b4-b780-47e3-a703-5d15781dc6ba · outbound

This paper cites Self- powered LLM modality expansion for large speech-text models,.

Speech-IFEval: Evaluating Instruction-Following and Quantifying Catastrophic Forgetting in Speech-Aware Language Models Self- powered LLM modality expansion for large speech-text models,

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:24:36.840917Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-07T14:24:34.631664Z digest=sha256:14dbbcc12a6522915ce2b2a6c5e10e8f6534241a38c51c708a75f3dc1a1516d2

Observation 4658673b-028a-4740-b015-d0fcb7c79845 · outbound

This paper cites Instruction-Following Evaluation for Large Language Models.

Speech-IFEval: Evaluating Instruction-Following and Quantifying Catastrophic Forgetting in Speech-Aware Language Models Instruction-Following Evaluation for Large Language Models

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T14:24:34.667540Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:24:34.667540Z digest=sha256:517d3e03e865dc029c2c146569bc4aa38cbbb13ead408fb9650a00b9d8daa1d0

Observation 500d6a1e-db70-4441-bb02-518cb789caa2 · outbound

This paper cites InFoBench: Evaluating instruction following ability in large language models,.

Speech-IFEval: Evaluating Instruction-Following and Quantifying Catastrophic Forgetting in Speech-Aware Language Models InFoBench: Evaluating instruction following ability in large language models,

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:24:36.691726Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-07T14:24:34.709219Z digest=sha256:41590de6586cb06fe3d1724ce4b7141bb4bcbc1a00925a43ea7de08f83452b74

Observation 4c43b894-71b9-40e9-8346-da0bbdde1cff · outbound

This paper cites FOFO: A benchmark to evaluate LLMs’ format- following capability,.

Speech-IFEval: Evaluating Instruction-Following and Quantifying Catastrophic Forgetting in Speech-Aware Language Models FOFO: A benchmark to evaluate LLMs’ format- following capability,

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:24:36.515256Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-07T14:24:34.769578Z digest=sha256:aa2474856a35f0a3475fdc634a6d48b847d3a84fe3f54dd6ae1bf3dc588a1f53

Observation 4978150e-c95e-4ebd-a420-3180ce5c811c · outbound

This paper cites Audiobench: A universal benchmark for audio large language models,.

Speech-IFEval: Evaluating Instruction-Following and Quantifying Catastrophic Forgetting in Speech-Aware Language Models Audiobench: A universal benchmark for audio large language models,

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T14:24:34.865350Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:24:34.865350Z digest=sha256:2f716b0d5ef3a62e1587aecf8286c6b20859ca387abc86bf86d51988f7daee06

Observation d78f16a3-78e7-42bd-a8b8-854def2df714 · outbound

This paper cites Can large language models be an alternative to human evaluations?.

Speech-IFEval: Evaluating Instruction-Following and Quantifying Catastrophic Forgetting in Speech-Aware Language Models Can large language models be an alternative to human evaluations?

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:24:36.357419Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-07T14:24:34.930392Z digest=sha256:89ee1d9498d87249c5366f4e7d76f618665c9d7fcbdede1551014bb389aed60e

Observation f273a5f7-3a38-41ae-87b4-05f016f38ec7 · outbound

This paper cites Chain of thought prompting elicits reasoning in large language models,.

Speech-IFEval: Evaluating Instruction-Following and Quantifying Catastrophic Forgetting in Speech-Aware Language Models Chain of thought prompting elicits reasoning in large language models,

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:24:36.205126Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-07T14:24:34.967697Z digest=sha256:132832902986945239ba516c87ff6eecb405ef8f910e281c0d14412b9ac65cd0

Observation 1be62cdc-ea16-423e-997e-5f9a5c1f4106 · outbound

This paper cites Robust speech recognition via large-scale weak supervision,.

Speech-IFEval: Evaluating Instruction-Following and Quantifying Catastrophic Forgetting in Speech-Aware Language Models Robust speech recognition via large-scale weak supervision,

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:24:36.045834Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-07T14:24:35.003064Z digest=sha256:fc1e5b4afc0442ebedf7b56d1fd46fe711f5871d913c06e89fb62ef4bee6ad21

Observation 6576f5f4-121c-4680-8a4c-92503297dc82 · outbound

This paper cites emotion2vec: Self-supervised pre-training for speech emotion representation,.

Speech-IFEval: Evaluating Instruction-Following and Quantifying Catastrophic Forgetting in Speech-Aware Language Models emotion2vec: Self-supervised pre-training for speech emotion representation,

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:24:35.834786Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-07T14:24:35.057623Z digest=sha256:96653060b09fb2ae75ad3273c81c438ec6009c53ae80a45dcb62685485401e55

Pith citing papers

Observation 2a6b7b26-a6d3-4ab5-961c-069dd69f4cf5 · inbound

Speech-IFEval: Evaluating Instruction-Following and Quantifying Catastrophic Forgetting in Speech-Aware Language Models cites this paper.

Speech-IFEval: Evaluating Instruction-Following and Quantifying Catastrophic Forgetting in Speech-Aware Language Models Speech-IFEval: Evaluating Instruction-Following and Quantifying Catastrophic Forgetting in Speech-Aware Language Models

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T14:24:32.537583Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:24:32.537583Z digest=sha256:d1b42f70a3429c8933db1f6284efbdf184bb863c65744f4fb9e17ef8f649a8b9

Observation b1951685-0dca-4ca2-a233-c7ecd738c91f · inbound

Reducing Object Hallucination in Large Audio-Language Models via Audio-Aware Decoding cites this paper.

Reducing Object Hallucination in Large Audio-Language Models via Audio-Aware Decoding Speech-IFEval: Evaluating Instruction-Following and Quantifying Catastrophic Forgetting in Speech-Aware Language Models

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T05:44:08.447266Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:44:08.447266Z digest=sha256:ce65501247a03ff801738b710db35afb2cd791843b4db237e4a6129b8081876a

Observation 73550c0b-7382-42e0-916b-836d52fb3f4c · inbound

A Survey of Automatic Evaluation Methods on Text, Visual and Speech Generations cites this paper.

A Survey of Automatic Evaluation Methods on Text, Visual and Speech Generations Speech-IFEval: Evaluating Instruction-Following and Quantifying Catastrophic Forgetting in Speech-Aware Language Models

Reference 234

Resolution
unresolved
no resolver link, observed 2026-08-07T10:17:47.081584Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:17:47.081584Z digest=sha256:9cc1d0d216bd9ebe47740cf8477c46dc36bdafdf4522c889a89188162dfa30df

Observation 02887523-90a9-4bb8-bad3-a48d717f4ffe · inbound

A Survey of Audio Reasoning in Multimodal Foundation Models cites this paper.

A Survey of Audio Reasoning in Multimodal Foundation Models Speech-IFEval: Evaluating Instruction-Following and Quantifying Catastrophic Forgetting in Speech-Aware Language Models

Reference 117

Resolution
verified exact
arxiv_id, observed 2026-05-21T02:09:24.086329Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-21T02:08:06.976461Z digest=sha256:5980eda1c5b4ad9c81440c52ba02f01d1d6481a8f8efbfe7588b84b79960636e

Observation 4b5dee97-94de-4cce-8e2c-254715369ddf · inbound

Preference-ASR: A Preference-Aware Test Set for Benchmarking ASR in the Era of Speech LLMs cites this paper.

Preference-ASR: A Preference-Aware Test Set for Benchmarking ASR in the Era of Speech LLMs Speech-IFEval: Evaluating Instruction-Following and Quantifying Catastrophic Forgetting in Speech-Aware Language Models

Reference 23

Resolution
verified exact
arxiv_id, observed 2026-06-30T07:34:21.796509Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-06-30T07:26:38.380118Z digest=sha256:c25b0667bee6042ed43834e401c776b60f8556d188977a8598c2c50598d45da3

Observation 963c8688-4e22-42df-96b9-f4e92dd0db1e · inbound

REDDIT: Correcting Model-Generated Timestamp Drift in ASR without Forgetting via Replay-Based Distribution Editing cites this paper.

REDDIT: Correcting Model-Generated Timestamp Drift in ASR without Forgetting via Replay-Based Distribution Editing Speech-IFEval: Evaluating Instruction-Following and Quantifying Catastrophic Forgetting in Speech-Aware Language Models

Reference 28

Resolution
verified exact
local_arxiv, observed 2026-07-07T14:53:55.828890Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-07-07T14:52:35.127533Z digest=sha256:bdd17188f927366d2ab29c5c5af6f68df8359f7a8a00108da04425173b58c344

Observation f2bd38d6-64e9-4c8f-9144-a718b81b1678 · inbound

REDDIT: Correcting Model-Generated Timestamp Drift in ASR without Forgetting via Replay-Based Distribution Editing cites this paper.

REDDIT: Correcting Model-Generated Timestamp Drift in ASR without Forgetting via Replay-Based Distribution Editing Speech-IFEval: Evaluating Instruction-Following and Quantifying Catastrophic Forgetting in Speech-Aware Language Models

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-02T08:34:04.723035Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T08:34:04.723035Z digest=sha256:dfb51d039e3d91b3f4158db048a8472084373c7d8618cc38aea05942af42a996