, 2021
Awesome Speaker Diarization
A curated list of awesome Speaker Diarization papers, libraries, datasets, and other resources.
This page lists names, links and short descriptions. The original list on GitHub is the source and belongs to its authors.
Publications >Special topics
Publications >Other
Software >Framework
FunASR
Industrial-grade speech recognition toolkit with built-in speaker diarization (cam++), VAD, ASR (SenseVoice/Paraformer), and punctuation. 50+ languages, 170x realtime.
MiniVox
MiniVox is an open-source evaluation system for the online speaker diarization task.
SpeechBrain
SpeechBrain is an open-source and all-in-one speech toolkit based on PyTorch.
SIDEKIT for diarization (s4d)
An open source package extension of SIDEKIT for Speaker diarization.
pyAudioAnalysis
Python Audio Analysis Library: Feature Extraction, Classification, Segmentation and Applications.
AaltoASR
Speaker diarization scripts, based on AaltoASR.
LIUM SpkDiarization
LIUM_SpkDiarization is a software dedicated to speaker diarization (i.e. speaker segmentation and clustering). It is written in Java, and includes the most recent developments in the domain (as of 2013).
kaldi-asr
Example scripts for speaker diarization on a portion of CALLHOME used in the 2000 NIST speaker recognition evaluation.
kaldi-speaker-diarization
Icelandic speaker diarization scripts using kaldi.
Alize LIA_SpkSeg
ALIZÉ is an opensource platform for speaker recognition. LIA_SpkSeg is the tools for speaker diarization.
pyannote-audio
Neural building blocks for speaker diarization: speech activity detection, speaker change detection, speaker embedding.
pyBK
Speaker diarization using binary key speaker modelling. Computationally light solution that does not require external training data.
Speaker-Diarization
Speaker diarization using uis-rnn and GhostVLAD. An easier way to support openset speakers.
EEND
End-to-End Neural Diarization.
VBx
Variational Bayes HMM over x-vectors diarization. x-vector extractor recipe
RE-VERB
RE: VERB is speaker diarization system, it allows the user to send/record audio of a conversation and receive timestamps of who spoke when.
StreamingSpeakerDiarization
Streaming speaker diarization, extends pyannote.audio to online processing
simple_diarizer
Simplified diarization pipeline using some pretrained models. Made to be a simple as possible to go from an input audio file to diarized segments.
Picovoice Falcon
A lightweight, accurate, and fast speaker diarization engine written in C and available in Python, running on CPU with minimal overhead.
DiaPer
Pytorch implementation for DiaPer: End-to-End Neural Diarization with Perceiver-Based Attractors including models pre-trained on free and public data.
sherpa-onnx
Support speaker diarization, speech recognition, and text-to speech on various platforms with various language bindings.
FluidAudio
A native Swift speaker diarization library for Apple platforms, using CoreML for efficient, real-time audio processing with high accuracy.
Software >Evaluation
pyannote-metrics
A toolkit for reproducible evaluation, diagnostic, and error analysis of speaker diarization systems.
SimpleDER
A lightweight library to compute Diarization Error Rate (DER).
DiarizationLM
Implements Word Error Rate (WER), Word Diarization Error Rate (WDER), and concatenated minimum-permutation Word Error Rate (cpWER).
dscore
Diarization scoring tools.
Sequence Match Accuracy
Match the accuracy of two sequences with Hungarian algorithm.
spyder
Simple Python package for fast DER computation.
CDER
Conversational DER from The Conversational Short-phrase Speaker Diarization (CSSD) Task: Dataset, Evaluation Metric and Baselines
Selective Collar
Selective collar for mitigating evaluation bias in speaker diarization from Beyond Uniform Forgiveness: Introducing the Selective Collar to Mitigate Evaluation Bias in Speaker Diarization
Software >Clustering
Sequence Match Accuracy
Match the accuracy of two sequences with Hungarian algorithm.
uis-rnn-sml
A variant of UIS-RNN, for the paper Supervised Online Diarization with Sample Mean Loss for Multi-Domain Data.
DNC
Transformer-based Discriminative Neural Clustering (DNC) for Speaker Diarisation. Like UIS-RNN, it is supervised.
SpectralCluster
Spectral clustering with affinity matrix refinement operations, auto-tune, and speaker turn constraints.
sklearn.cluster
scikit-learn clustering algorithms.
PLDA
Probabilistic Linear Discriminant Analysis & classification, written in Python.
PLDA
Open-source implementation of simplified PLDA (Probabilistic Linear Discriminant Analysis).
Auto-Tuning Spectral Clustering
Auto-tuning Spectral Clustering method that does not need development set or supervised tuning.
Software >Speaker embedding
resemble-ai/Resemblyzer
PyTorch implementation of generalized end-to-end loss for speaker verification, which can be used for voice cloning and diarization.
Speaker_Verification
Tensorflow implementation of generalized end-to-end loss for speaker verification.
PyTorch_Speaker_Verification
PyTorch implementation of "Generalized End-to-End Loss for Speaker Verification" by Wan, Li et al. With UIS-RNN integration.
Real-Time Voice Cloning
Implementation of "Transfer Learning from Speaker Verification to Multispeaker Text-To-Speech Synthesis" (SV2TTS) with a vocoder that works in real-time.
conformer-speaker-encoder
Massively multilingual conformer-based speaker recognition models in TFLite format.
deep-speaker
Third party implementation of the Baidu paper Deep Speaker: an End-to-End Neural Speaker Embedding System.
x-vector-kaldi-tf
Tensorflow implementation of x-vector topology on top of Kaldi recipe.
kaldi-ivector
Extension to Kaldi implementing the standard i-vector hyperparameter estimation and i-vector extraction procedure.
voxceleb-ivector
Voxceleb1 i-vector based speaker recognition system.
pytorch_xvectors
PyTorch implementation of Voxceleb x-vectors. Additionaly, includes meta-learning architectures for embedding training. Evaluated with speaker diarization and speaker verification.
ASVtorch
ASVtorch is a toolkit for automatic speaker recognition.
asv-subtools
ASV-Subtools is developed based on Pytorch and Kaldi for the task of speaker recognition, language identification, etc. The 'sub' of 'subtools' means that there are many modular tools and the parts constitute the whole.
WeSpeaker
WeSpeaker is a research and production oriented speaker verification, recognition and diarization toolkit, which supports very strong recipes with on-the-fly data preparation, model training and evaluation, as well as runtime C++ codes.
ReDimNet
Neural network architecture presented in the paper Reshape Dimensions Network for Speaker Recognition
Software >Speaker change detection
change_detection
Code for Speaker Change Detection in Broadcast TV using Bidirectional Long Short-Term Memory Networks.
tidydiarize
Diarization inside OpenAI Whisper decoder
Software >Audio feature extraction
python_speech_features
This library provides common speech features for ASR including MFCCs and filterbank energies. https://python-speech-features.readthedocs.io/en/latest/
pyAudioAnalysis
Python Audio Analysis Library: Feature Extraction, Classification, Segmentation and Applications.
Software >Audio data augmentation
pyroomacoustics
Pyroomacoustics is a package for audio signal processing for indoor applications. It was developed as a fast prototyping platform for beamforming algorithms in indoor scenarios. https://pyroomacoustics.readthedocs.io
gpuRIR
Python library for Room Impulse Response (RIR) simulation with GPU acceleration
rir_simulator_python
Room impulse response simulator using python
WavAugment
WavAugment performs data augmentation on audio data. The audio data is represented as pytorch tensors
EEND_dataprep
Recipes for generating simulated conversations used to train end-to-end diarization models.
Software >Other software
VB Diarization
VB Diarization with Eigenvoice and HMM Priors.
DOVER-Lap
Python package for combining diarization system outputs
Diar-az
Data formatting tool to support the ruv-di dataset. Kaldi to Gecko to Kaldi and corpus and back
voxsolo
Verbatim per-speaker audio from diarization: keeps one speaker bit-exact, silences other speakers and overlapped speech; exports EDL/SRT/VTT/Audacity labels for editors.
Datasets >Diarization datasets
2000 NIST Speaker Recognition Evaluation
Evaluation Plan
2003 NIST Rich Transcription Evaluation Data
telephone speech, broadcast news
CALLHOME American English Speech
CH109 whitelist
The ICSI Meeting Corpus
License
The AMI Meeting Corpus
License
Fisher English Training Speech Part 1 Speech
Fisher English Training Speech Part 1 Transcripts
Fisher English Training Part 2, Speech
Fisher English Training Part 2, Transcripts
VoxConverse
VoxConverse is an audio-visual diarisation dataset consisting of over 50 hours of multispeaker clips of human speech, extracted from YouTube videos
MiniVox
MiniVox is an open-source evaluation system for the online speaker diarization task.
The AliMeeting Corpus
Together with audios
Datasets >Speaker embedding training sets
TIMIT
Published in 1993, the TIMIT corpus of read speech is one of the earliest speaker recognition datasets.
VCTK
Most were selected from a newspaper plus the Rainbow Passage and an elicitation paragraph intended to identify the speaker's accent.
LibriSpeech
Large-scale (1000 hours) corpus of read English speech.
Multilingual LibriSpeech (MLS)
Multilingual LibriSpeech (MLS) dataset is a large multilingual corpus suitable for speech research. The dataset is derived from read audiobooks from LibriVox and consists of 8 languages - English, German, Dutch, Spanish, French, Italian, Portuguese, Polish.
LibriVox
Free public domain audiobooks. LibriSpeech is a processed subset of LibriVox. Each original unsegmented utterance could be very long.
VoxCeleb 1&2
VoxCeleb is an audio-visual dataset consisting of short clips of human speech, extracted from interview videos uploaded to YouTube.
The Spoken Wikipedia Corpora
Volunteer readers reading Wikipedia articles.
CN-Celeb
A Free Chinese Speaker Recognition Corpus Released by CSLT@Tsinghua University.
BookTubeSpeech
Audio samples extracted from BookTube videos - videos where people share their opinions on books - from YouTube. The dataset can be downloaded using BookTubeSpeech-download.
DeepMine
A speech database in Persian and English designed to build and evaluate speaker verification, as well as Persian ASR systems.
NISP-Dataset
This dataset contains speech recordings along with speaker physical parameters (height, weight, ... ) as well as regional information and linguistic information.
VoxBlink2
Multilingual dataset from VoxBlink2: A 100K+ Speaker Recognition Corpus and the Open-Set Speaker-Identification Benchmark
Datasets >Augmentation noise sources
Other learning materials >Books
Other learning materials >Tech blogs
Literature Review For Speaker Change Detection
by Halil Erdoğan
Speaker Diarization with Kaldi
by Yoav Ramon
Other learning materials >Video tutorials
Speaker Diarization: Optimal Clustering and Learning Speaker Embeddings
by Microsoft Research
Robust Speaker Diarization for Meetings: the ICSI system
by Microsoft Research
【机器之心&博文视点】入门声纹技术|第二讲:声纹分割聚类与其他应用
by Quan Wang
Related lists in Computer Science
See categoryTable of Contents
hesreallyhim/awesome-claude-code
A hand-picked collection of the finest of resources for the most awesome of agents, Claude Code, the undisputed champion of coding companions, from the unstoppable team…
Awesome Agent Skills
VoltAgent/awesome-agent-skills
A curated collection of 1000+ agent skills from official dev teams and the community, compatible with Claude Code, Codex, Gemini CLI, Cursor, and more.
Awesome Machine Learning
josephmisiti/awesome-machine-learning
A curated list of awesome Machine Learning frameworks, libraries and software.
Awesome Production Machine Learning
EthicalML/awesome-production-machine-learning
A curated list of awesome open source libraries to deploy, monitor, version and scale your machine learning
AWESOME DATA SCIENCE
academic/awesome-datascience
:memo: An awesome Data Science repository to learn and apply for real world problems.
Static Analysis
analysis-tools-dev/static-analysis
⚙️ A curated list of static analysis (SAST) tools and linters for all programming languages, config files, build tools, and more. The focus is on tools which improve…