Pith. sign in

Omni-r1: Do you really need audio to fine-tune your audio llm?

13 Pith papers cite this work. Polarity classification is still indexing.

13 Pith papers citing it

citation-role summary

background 3

citation-polarity summary

years

2026 12 2025 1

roles

background 3

polarities

background 2 support 1

representative citing papers

Benign Fine-Tuning Breaks Safety Alignment in Audio LLMs

cs.CR · 2026-04-17 · conditional · novelty 8.0

Benign fine-tuning on audio data breaks safety alignment in Audio LLMs by raising jailbreak success rates up to 87%, with the dominant risk axis depending on model architecture and embedding proximity to harmful content.

SHALA-LLM: Smartly Handling Ambiguous Labels in Aligning LLMs

cs.LG · 2026-06-03 · unverdicted · novelty 6.0

SHALA-LLM is a new RL framework for LLM alignment that learns directly from annotator label distributions and prioritizes ambiguous samples, reducing Jensen-Shannon Distance by up to 62.1% and raising F1 by up to 16.7% on NLI and ER benchmarks.

Learning When to Think While Listening in Large Audio-Language Models

cs.CL · 2026-05-26 · unverdicted · novelty 6.0

A wait-think-answer controller for LALMs is trained via SFT followed by six-reward DAPO, raising row-weighted accuracy from 67.6% to 70.3% and cutting post-endpoint thinking length by 14% on synthetic spoken QA while remaining functional on real recorded audio.

Step-Audio 2 Technical Report

cs.CL · 2025-07-22 · unverdicted · novelty 6.0

Step-Audio 2 integrates a latent audio encoder, reasoning-centric reinforcement learning, and discrete audio token generation into language modeling to deliver state-of-the-art performance on audio understanding and conversational benchmarks.

FSA-GRPO: Teaching Auditory LLMs to Use Few-shot Demonstrations

eess.AS · 2026-05-26 · unverdicted · novelty 4.0

FSA-GRPO applies reinforcement learning with a few-shot-aware reward to auditory LLMs, improving few-shot performance on children's ASR, speech translation, and audio tasks when trained only on adult data.

A Survey of Audio Reasoning in Multimodal Foundation Models

eess.AS · 2026-05-20 · unverdicted · novelty 2.0

A survey that provides a unified formulation of audio reasoning and reviews advances across Audio-to-Text, Audio-to-Speech, Audio-Visual, and Agentic paradigms while discussing challenges and future directions.

citing papers explorer

Showing 13 of 13 citing papers.