Skip to content

Entry

FluidAudio

Appears in 4 awesome lists

SDK for real-time on-device audio intelligence on iOS/macOS (diarization, identification, VAD, separation, embeddings, ASR), with CoreML models converted directly from PyTorch to leverage Apple Neural Engine performance.

Open github.comfluidinference/fluidaudio

Found in these lists

Awesome Speaker Diarization

Section: Framework · A native Swift speaker diarization library for Apple platforms, using CoreML for efficient, real-time audio processing with high accuracy.

FreshScore 84

Awesome Swift

Section: Audio · SDK for real-time on-device audio intelligence on iOS/macOS (diarization, identification, VAD, separation, embeddings, ASR), with CoreML models converted directly from PyTorch to leverage Apple Neural Engine performance.

FreshScore 90

Awesome Ios

Section: Audio · Swift framework for local speech recognition, speaker diarization, voice activity detection, and text-to-speech using Core ML.

FreshScore 90

Awesome Swift

Section: Audio · On-device speech processing for iOS and macOS: ASR, TTS, VAD, and speaker diarization

ActiveScore 69

FunASR

industrial-grade ASR toolkit; 170× realtime on GPU, 50+ languages, built-in VAD, punctuation, speaker diarization, and emotion detection. Includes non-autoregressive SenseVoice and LLM-based Fun-ASR-Nano models.

In 9 listsDetails

Kaldi

Kaldi is a toolkit for speech recognition written in C++ and licensed under the Apache License v2.0. Kaldi is intended for use by speech recognition researchers.

In 5 listsDetails

SpeechBrain

PyTorch speech toolkit with recipes for ASR, TTS, speaker recognition, and speech enhancement.

In 5 listsDetails

sherpa-onnx

Speech-to-text, text-to-speech, speaker diarization, speech enhancement, source separation, and VAD using next-gen Kaldi with onnxruntime without Internet connection. Support embedded systems, Android, iOS, HarmonyOS, Raspberry Pi, RISC-V, RK NPU, Axera NPU, Ascend NPU, x86_64 servers, websocket…

In 4 listsDetails

pyAudioAnalysis

Python library for audio segmentation, silence removal via dynamic thresholding, spectral features, and classification.

In 4 lists

ModernAVPlayer

Persistence player to resume playback after bad network connection even in background mode, manage headphone interactions, system interruptions, now playing informations and remote commands.

In 3 lists

AudioKit

Powerful audio synthesis, processing and analysis, without the steep learning curve.

In 3 lists

pyannote-audio

Neural building blocks for speaker diarization: speech activity detection, speaker change detection, speaker embedding.

In 3 lists