pith. sign in

hub Mixed citations

Funaudiollm: V oice understanding and generation foundation models for natural interaction between humans and llms

Mixed citation behavior. Most common role is background (57%).

20 Pith papers citing it
Background 57% of classified citations

hub tools

citation-role summary

background 4 method 3

citation-polarity summary

representative citing papers

HumanOmni-Speaker: Identifying Who said What and When

cs.CV · 2026-03-23 · unverdicted · novelty 6.0

HumanOmni-Speaker introduces a Visual Delta Encoder and VR-SDR benchmark that enable end-to-end speaker diarization and recognition by sampling video at 25 fps and compressing inter-frame motion residuals into 6 tokens per frame.

Logics-Parsing-Omni Technical Report

cs.AI · 2026-03-10 · unverdicted · novelty 6.0

Omni Parsing framework converts complex multimodal signals into locatable, enumerable, and traceable structured knowledge via hierarchical detection, recognition, and interpreting with strict evidence alignment.

Two-Dimensional Quantization for Geometry-Aware Audio Coding

cs.SD · 2025-12-01 · unverdicted · novelty 6.0

Q2D2 uses 2D geometric grid projections to quantize feature pairs in neural audio codecs, yielding implicit codebooks that improve efficiency and utilization over RVQ, VQ, and FSQ while maintaining reconstruction quality.

GLM-4-Voice: Towards Intelligent and Human-Like End-to-End Spoken Chatbot

cs.CL · 2024-12-03 · conditional · novelty 6.0 · 2 refs

GLM-4-Voice builds an end-to-end spoken chatbot by deriving a 175bps single-codebook tokenizer from ASR, synthesizing interleaved speech-text data, and continuing pre-training of GLM-4-9B on up to 1 trillion tokens before fine-tuning on conversational speech.

FormalASR: End-to-End Spoken Chinese to Formal Text

cs.CL · 2026-05-19 · unverdicted · novelty 5.0

FormalASR fine-tunes small Qwen3-ASR models on new spoken-to-formal Chinese datasets to achieve direct transcription with up to 37.4% relative CER reduction over verbatim baselines.

PASK: Toward Intent-Aware Proactive Agents with Long-Term Memory

cs.AI · 2026-04-09 · unverdicted · novelty 4.0

PASK introduces the DD-MM-PAS paradigm for streaming proactive agents with intent-aware detection, hybrid memory modeling, and a new real-world benchmark where the IntentFlow model matches top LLMs on latency while finding deeper intents.

citing papers explorer

Showing 20 of 20 citing papers.