Research ·

Audio LLMs Know When They Can't Hear You

75Developing1 reportarXiv cs.AI

AI brief

AI-written

Why it mattersHelps developers improve stability of voice-interactive AI products.

Study finds transcription reliability can be extracted from audio large model encoders, with a lightweight predictor outperforming existing baselines

What happened

To address the problem where audio large models often make transcription errors when input recording quality is poor, and struggle to accurately judge their own transcription reliability, the research team first validated the limitations of existing detection methods. They then discovered that transcription reliability signals are concentrated in the representations of frozen audio encoders, and designed a lightweight predictor based on this finding. The predictor can assess transcription reliability before content generation, trigger user clarification prompts, and operate without modifying the underlying audio large model. Its detection accuracy far exceeds existing methods, and its reliability labels can be transferred across different families of audio large models.

Key facts

Research problem
Audio large models cannot independently judge speech transcription reliability, and are prone to generating error responses due to poor input audio quality
Core finding
Transcription reliability signals are strongly represented in frozen audio encoder representations, and reliability labels can be transferred across different families of audio large models
Proposed solution
A lightweight reliability predictor built on frozen audio encoder representations, which judges transcription reliability before content generation and requires no modifications to the underlying model
Detection performance
In-domain macro F1 reaches 81.10%, cross-domain macro F1 reaches 78.09%, outperforming the strongest baseline by 10.33 and 11.93 percentage points respectively

Background

Audio large models currently deliver capable voice interaction functionality, but they often transcribe content incorrectly and generate off-target responses when input recordings lack sufficient clarity or have very low signal-to-noise ratios. Existing detection approaches include speech quality assessment, generation uncertainty estimation, and transcription word error rate estimation, but these methods capture limited signals of transcription failure, and the models themselves have extremely low accuracy when judging their own transcription reliability.

Why it matters

For the audio large model industry, this solution solves the invalid input identification challenge in voice interaction scenarios and can drastically reduce error response rates. For developers, the lightweight predictor’s ability to operate without modifying underlying models significantly cuts integration and adaptation costs. For end users, the system can proactively request clarification when voice input is invalid, preventing off-topic model responses and improving the overall voice interaction experience.

Written by AI from the original article. It may contain mistakes; the original is the source of truth.

Source

  1. arXiv cs.AI ↗Audio LLMs Know When They Can't Hear YouStudies audio LLM's ability to recognize unreliable self-transcription of degraded audio.
Back to AI News