#speech (3)

Whisper OSS

OpenAIの音声認識モデル。多言語の文字起こし・翻訳に対応し、日本語の精度も高い定番OSS。

stars: 107.8k price: open_source license: MIT ja: full level: beginner
faster-whisper OSS

Whisperを最大4倍高速化した再実装。同じ精度でメモリ消費も少なく、文字起こしの実運用での標準になっている。

stars: 25.1k price: open_source license: MIT ja: full level: intermediate
Kokoro TTS OSS

わずか82Mパラメータで高品質な音声を生成する軽量TTSモデル。速度・コスト・品質のバランスで注目を集める。

stars: 8.5k price: open_source license: Apache-2.0 ja: partial level: intermediate