Research ·

Target Speaker Unlearning for LLM-Based ASR at Inference Time

77Developing1 reportarXiv cs.CL

AI brief

AI-written

Why it mattersMeets voice privacy needs and improves ASR system compliance.

Lightweight mountable module enables dynamic forgetting of specified speakers in ASR, meeting meeting transcription privacy needs

What happened

Addressing privacy pain points in multi-speaker automatic speech recognition (ASR) for transcription, a research team has introduced the first targeted speaker unlearning ASR (TSU-ASR) task. They developed a lightweight, mountable registration-conditioned gating module that works with frozen two-stream speech LLMs, enabling dynamic addition of unheard opted-out speakers at inference time: the system will not transcribe speech from these users, but still flags their active meeting status. Tests on the English AMI and Chinese AliMeeting datasets show transcription accuracy for opted-out speakers dropped from 72.3% to 48.2% (AMI) and from 73.6% to 27.3% (AliMeeting), while transcription error rates for all other speakers remained largely unchanged.

Key facts

Output Type
Newly released privacy-preserving multi-speaker ASR solution published on arXiv
Core Module
Lightweight Registration-Conditioned Gating (ECG) module mountable on frozen two-stream speech LLMs
Test Datasets
English AMI dataset, Chinese AliMeeting dataset
Benchmark Performance
Transcription accuracy for opted-out speakers dropped by 24.1 and 46.3 percentage points from baselines on the two test sets, with no significant change in transcription error rates for other speakers
Target Use Cases
Dynamic privacy transcription requirements for online video conferencing platforms

Background

Mainstream multi-speaker ASR and speaker diarization solutions lack flexible privacy protection capabilities: to avoid being automatically transcribed, meeting participants currently often have to leave the call entirely, which fails to match the granular privacy interaction needs of the billions of daily online meetings.

Why it matters

For the industry, this solution delivers an implementable transcription privacy protection pathway for large-scale online meeting platforms, closing a key privacy gap in existing smart meeting products. For developers, the lightweight mountable module can be integrated quickly without modifying the underlying frozen speech LLM. For end users, this allows people to opt out of AI transcription without leaving a meeting, balancing collaborative efficiency and privacy safety.

What to watch

Stakeholders can track the module's real-world deployment performance in commercial conference systems, as well as future optimizations to strengthen privacy protection for opted-out speakers.

Written by AI from the original article. It may contain mistakes; the original is the source of truth.

Source

  1. arXiv cs.CL ↗Target Speaker Unlearning for LLM-Based ASR at Inference TimeIt proposes TSU-ASR to skip transcription of opt-out speakers during inference.
Back to AI News