Research ·
Target Speaker Unlearning for LLM-Based ASR at Inference Time
AI brief
AI-writtenWhy it mattersMeets voice privacy needs and improves ASR system compliance.
Lightweight mountable module enables dynamic forgetting of specified speakers in ASR, meeting meeting transcription privacy needs
What happened
Addressing privacy pain points in multi-speaker automatic speech recognition (ASR) for transcription, a research team has introduced the first targeted speaker unlearning ASR (TSU-ASR) task. They developed a lightweight, mountable registration-conditioned gating module that works with frozen two-stream speech LLMs, enabling dynamic addition of unheard opted-out speakers at inference time: the system will not transcribe speech from these users, but still flags their active meeting status. Tests on the English AMI and Chinese AliMeeting datasets show transcription accuracy for opted-out speakers dropped from 72.3% to 48.2% (AMI) and from 73.6% to 27.3% (AliMeeting), while transcription error rates for all other speakers remained largely unchanged.
Key facts
- Output Type
- Newly released privacy-preserving multi-speaker ASR solution published on arXiv
- Core Module
- Lightweight Registration-Conditioned Gating (ECG) module mountable on frozen two-stream speech LLMs
- Test Datasets
- English AMI dataset, Chinese AliMeeting dataset
- Benchmark Performance
- Transcription accuracy for opted-out speakers dropped by 24.1 and 46.3 percentage points from baselines on the two test sets, with no significant change in transcription error rates for other speakers
- Target Use Cases
- Dynamic privacy transcription requirements for online video conferencing platforms
Background
Mainstream multi-speaker ASR and speaker diarization solutions lack flexible privacy protection capabilities: to avoid being automatically transcribed, meeting participants currently often have to leave the call entirely, which fails to match the granular privacy interaction needs of the billions of daily online meetings.
Why it matters
For the industry, this solution delivers an implementable transcription privacy protection pathway for large-scale online meeting platforms, closing a key privacy gap in existing smart meeting products. For developers, the lightweight mountable module can be integrated quickly without modifying the underlying frozen speech LLM. For end users, this allows people to opt out of AI transcription without leaving a meeting, balancing collaborative efficiency and privacy safety.
What to watch
Stakeholders can track the module's real-world deployment performance in commercial conference systems, as well as future optimizations to strengthen privacy protection for opted-out speakers.
Written by AI from the original article. It may contain mistakes; the original is the source of truth.