Pith. sign in

hub Canonical reference

Reinforcement learning outperforms supervised fine-tuning: A case study on audio question an- swering

Canonical reference. 83% of citing Pith papers cite this work as background.

15 Pith papers citing it
Background 83% of classified citations

hub tools

citation-role summary

background 6

citation-polarity summary

years

2026 10 2025 5

roles

background 6

polarities

background 5 support 1

representative citing papers

LaSR: Context-Aware Speech Recognition via Latent Reasoning

cs.CL · 2026-05-30 · unverdicted · novelty 6.0

LaSR improves context-aware terminology recognition in speech LLMs by aligning latent CoT supervision on acoustic regions and introducing latent reasoning periods, shown on a new academic corpus to outperform standard fine-tuning without added latency.

Learning When to Think While Listening in Large Audio-Language Models

cs.CL · 2026-05-26 · unverdicted · novelty 6.0

A wait-think-answer controller for LALMs is trained via SFT followed by six-reward DAPO, raising row-weighted accuracy from 67.6% to 70.3% and cutting post-endpoint thinking length by 14% on synthetic spoken QA while remaining functional on real recorded audio.

TinyMU: A Compact Audio-Language Model for Music Understanding

cs.SD · 2026-04-17 · unverdicted · novelty 5.0

TinyMU is a 229M-parameter compact music understanding model that achieves 82% of state-of-the-art large audio-language model performance on the MuChoMusic benchmark while being 35 times smaller.

FSA-GRPO: Teaching Auditory LLMs to Use Few-shot Demonstrations

eess.AS · 2026-05-26 · unverdicted · novelty 4.0

FSA-GRPO applies reinforcement learning with a few-shot-aware reward to auditory LLMs, improving few-shot performance on children's ASR, speech translation, and audio tasks when trained only on adult data.

A Survey of Audio Reasoning in Multimodal Foundation Models

eess.AS · 2026-05-20 · unverdicted · novelty 2.0

A survey that provides a unified formulation of audio reasoning and reviews advances across Audio-to-Text, Audio-to-Speech, Audio-Visual, and Agentic paradigms while discussing challenges and future directions.

citing papers explorer

Showing 15 of 15 citing papers.