Open source speech to text and text to speech, tested

Speech to text and text to speech tools you can run yourself. The list is ordered by whether each project came up when Argusic installed it fresh, with the recording of that attempt one click away.

Tested between and . Each row shows its own test date; a project can change after that day.

55 of 59 tested projects run. 57 more waiting for a test.

In short: 42 of the 59 tested projects started as-is on a fresh machine: GPT-SoVITS, whisper.cpp, whisperX, leon, sherpa-onnx, RealtimeSTT, openwhispr, and speech_recognition, and 34 more. 13 more started once a stand-in replaced a service they expect, such as a database: Amphion, mlx-audio, ODS, MOSS-TTS, elevenlabs-python, parlor, vllm-mlx, and TalkingHead, and 5 more. 4 could not be verified: ai-video-editor, dia, dsnote, and orbit; the log shows where each one stopped.

Measured by Argusic on a fresh machine every time. Every number links to its evidence.

#projectverdictArgusic Scorelanguagestarstested on
1GPT-SoVITSRuns100 / 100Python62,522

1 min voice data can also be used to train a good TTS model! (few shot voice cloning)

What the test found: GPT-SoVITS API server starts and responds to TTS inference requests: HTTP 200 with a valid synthesized WAV audio file returned for English text-to-speech using a default reference audio on CPU. 23 minutes.

2whisper.cppRuns100 / 100C++54,220

Port of OpenAI's Whisper model in C/C++

What the test found: whisper.cpp builds from source via cmake, all 6 unit tests pass, whisper-cli transcribes speech to text with a real tiny.en model, and whisper-server starts and binds to a TCP port. 5 minutes.

3whisperXRuns100 / 100Python24,418

WhisperX: Automatic Speech Recognition with Word-level Timestamps (& Diarization)

What the test found: WhisperX 3.8.7rc1 installs, all 12 unit tests pass, the CLI responds with version and help text, and the full transcription pipeline (ASR + alignment) runs end-to-end on CPU with no errors. 10 minutes.

4leonRuns100 / 100TypeScript17,561

🧠 Leon is your open-source personal assistant.

What the test found: Leon 1.0.0-beta.10+dev is installed with managed Node.js v24.21.0, pnpm 12.4.1, Python 3.11.9, uv 0.10.12, cmake 4.3.2, and ninja 1.13.2. The HTTP server boots on port 5366 and responds 200. All 371 unit tests pass. The web app, LLM providers, voice services, and the PaddleOCR model are not configured and were not... 6 minutes.

5sherpa-onnxRuns100 / 100C++15,165

Speech-to-text, text-to-speech, speaker diarization, speech enhancement, source separation, and VAD using next-gen Kaldi with onnxruntime without Internet...

What the test found: sherpa-onnx 1.13.8 installed in a Python 3.12 venv; streaming ASR, offline Whisper ASR, and Piper TTS all produce correct output with downloaded ONNX models; the CLI binary responds and lists available commands. 5 minutes.

6RealtimeSTTRuns100 / 100Python10,166

A robust, efficient, low-latency speech-to-text library with advanced voice activity detection, wake word activation and instant transcription.

What the test found: Installed RealtimeSTT v1.1.2 in a Python 3.12 venv with PyAudio 0.2.13 (from Debian package), torch 2.14.0+cpu, torchaudio 2.11.0+cpu, faster-whisper 1.2.1, silero-vad 6.2.3, openwakeword 0.4.0, fastapi 0.141.1 and uvicorn 0.54.0; 591 unit tests pass and the real-model golden transcription test produces correct output. 8 minutes.

7openwhisprRuns100 / 100JavaScript9,150

Voice-to-text dictation app with local (Nvidia Parakeet/Whisper) and cloud models (BYOK). Privacy-first and available cross-platform.

What the test found: npm install succeeds, all 944 tests pass with Node 24 and the custom loader bootstraps, test script runs all test suites without failures. 71 minutes.

8speech_recognitionRuns100 / 100Python8,994

Speech recognition module for Python, supporting several engines and APIs, online and offline.

What the test found: SpeechRecognition 3.17.0 installed in a venv; 41 of 56 tests pass with 15 expected skips; Google Speech API recognition returns transcriptions from local audio files; CLI sprc binary prints help. 4 minutes.

9annyangRuns100 / 100TypeScript6,817

💬 Speech recognition for your site

What the test found: All 165 tests pass, build produces ESM/CJS/IIFE bundles in dist/, and the library is usable via npm import or script tag. 11 minutes.

10FunClipRuns100 / 100Python6,379

FunASR-powered video transcription, subtitle generation, and LLM-assisted clipping tool with a local Gradio UI.

What the test found: FunClip v2.2.1 is fully installed in a Python 3.12 venv with all dependencies, all 90 tests pass, the CLI recognition+clipping pipeline processes a real video with Paraformer-Large ASR and produces valid clipped MP4 output, and the Gradio web interface launches and serves HTTP 200 on localhost. 49 minutes.

11abogenRuns100 / 100Python6,094

Generate audiobooks from EPUBs, PDFs and text with synchronized captions.

What the test found: All 1592 tests pass, Flask web UI serves HTTP 200 on GET /, and the real Kokoro-82M TTS engine with af_heart voice generates non-silent, ffmpeg-decodable WAV audio from text input. 22 minutes.

12dograhRuns100 / 100Python5,828

Open source voice AI platform. Self-hosted alternative to Vapi and Retell. On Prem, BYOK across Speech to Speech or LLM/STT/TTS, with a visual workflow...

What the test found: Backend Python venv installed with all core deps, PostgreSQL 16 running on port 5432, Redis 7.4.1 running on port 6379, Node.js v22 deployed, pipecat submodule installed editable, FastAPI app imports with speechmatics as the only blocking provider, 493/514 collectable tests pass against real services. 57 minutes.

13YouDub-webuiRuns100 / 100Python5,607

Open-source AI video localization and dubbing for YouTube/Bilibili: speech recognition, subtitle translation, voice cloning, audio mixing and rendering. 开源 AI...

What the test found: Backend API is fully operational: uvicorn serves the FastAPI app on port 9998, all 421 backend tests pass, and every major endpoint (health, login, tasks, settings, cookies) responds correctly with expected status codes. 20 minutes.

14fastrtcRuns100 / 100JavaScript4,631

The python library for real-time communication

What the test found: The fastrtc library installs in a venv, all 29 tests pass, the Kokoro TTS model generates audio, and a demo app starts under uvicorn and serves its OpenAPI schema. 7 minutes.

15MOSS-TTS-NanoRuns100 / 100Python4,457

A 100M-parameter multilingual TTS model for real-time CPU inference, voice cloning, and 48 kHz stereo generation

What the test found: Dependencies installed in venv; PyTorch CPU inference via infer.py produces valid WAV audio; FastAPI web server responds to health checks and generates 48kHz stereo speech via POST /api/generate with demo presets. 16 minutes.

16WhisperLiveRuns100 / 100Python4,308

A nearly-live implementation of OpenAI's Whisper.

What the test found: WhisperLive is installed in a Python 3.12 venv with all dependencies, 318 of 332 unit tests pass (14 metrics tests skipped as they require a dedicated metrics server port), and the server launches with the REST API health endpoint responding HTTP 200 on port 8000 and the WebSocket server listening on port 9090. 11 minutes.

17RealtimeTTSRuns100 / 100Python4,040

Converts text to speech in realtime

What the test found: realtimetts 0.8.10 installed and built wheel+sdist; 369 of 380 tests pass covering BaseEngine, alignment, language router, sentence splitting, streaming, stream player, release guard (except ssh-keygen), and Qwen studio; Qwen server starts and fails only on CUDA GPU requirement; pyttsx3 SystemEngine requires espeak. 13 minutes.

18whisper-asr-webserviceRuns100 / 100Python3,354

OpenAI Whisper ASR Webservice API

What the test found: The Whisper ASR Box webservice is installed and running: the uvicorn server serves Swagger UI at /docs, returns JSON language detection at /detect-language, and returns transcription text at /asr all with HTTP 200. 14 minutes.

19TTS-WebUIRuns100 / 100TypeScript3,282

A single Gradio + React WebUI with extensions for ACE-Step, OmniVoice, Kimi Audio, Piper TTS, GPT-SoVITS, CosyVoice, XTTSv2, DIA, Kokoro, OpenVoice, ParlerTTS...

What the test found: TTS-WebUI server starts and responds HTTP 200 on both Gradio proxy (7767) and main (7770) interfaces; 102 of 105 tests pass (3 skipped require real API credentials); all core extensions import cleanly; CLI troubleshoot passes. 6 minutes.

20GPARuns100 / 100Python3,109

[AutoArk] GPA (General Purpose Audio) can do ASR, TTS and voice conversion with one tiny model!

What the test found: GPA v1.5 native PyTorch ASR, TTS, and training all run on CPU with weights downloaded from HuggingFace. ASR transcribes Chinese and English correctly; TTS produces valid 16 kHz mono audio from text plus a reference clip; training executes a forward/backward step. 10 minutes.

21vexaRuns100 / 100Python2,865

Open-source meeting transcription API for Google Meet, Microsoft Teams & Zoom. Auto-join bots, real-time WebSocket transcripts, MCP server for AI agents...

What the test found: 11 Python service test suites and the JS terminal client test suite run successfully in a container with Node 22.13 and Python 3.12.3; the compose/Docker deployment path cannot execute without Docker engine being available. 17 minutes.

22gTTSRuns100 / 100Python2,636

Python library and CLI tool to interface with Google Translate's text-to-speech API

What the test found: gTTS 2.5.4 installed, all 108 pytest tests pass, and the CLI generates valid MP3 audio via the Google Translate TTS API. 1 minute.

23pyttsx3Runs100 / 100Python2,534

Offline Text To Speech synthesis for python

What the test found: pyttsx3 v2.99 with eSpeak-NG v1.53.0 backend on Python 3.12.3 (Linux): import, speech synthesis, WAV file saving, property get/set, event loop, and all 9 applicable tests pass. 4 minutes.

24transcribe.cppRuns100 / 100C++1,996

ggml speech-to-text inference for 16+ model families

What the test found: transcribe.cpp builds from source with cmake and GCC 13, produces a working transcribe-cli binary that understands all CLI options and reads 16 kHz mono WAV files, and passes all 39 unit/smoke tests on CPU without a model file. 4 minutes.

25Genie-TTSRuns100 / 100Python1,794

GPT-SoVITS ONNX Inference Engine & Model Converter

What the test found: GENIE TTS package installs and builds in a Python 3.12 venv; all 4 language G2P modules produce correct phone IDs; FastAPI server serves all endpoints (/stop, /clear_reference_audio_cache, /load_character, /set_reference_audio, /tts, /unload_character) returning HTTP 200; full TTS inference with predefined character... 9 minutes.

26video-podcast-makerRuns100 / 100Python1,664

Topic → 4K narrated video for coding agents. v5.3.0: local TTS (edge free + azure, no external engine), manifest-based Asset Engine, Remotion composition...

What the test found: All 255 Python tests pass, TypeScript type-checks pass, and Remotion 4.0.438 renders a 2-frame smoke test to a valid H.264 MP4 verified with ffmpeg decode check; the only source change needed was bumping zod from ^3.23.0 to ^4.3.6 in package.json. 7 minutes.

27voxtypeRuns100 / 100Rust1,620

Voice-to-text with push-to-talk for Wayland compositors

What the test found: Voxtype v1.1.0 builds and passes its full test suite (1159 passed, 0 failed) on Ubuntu 24.04 amd64 with Rust 1.99.0, after manually providing alsa and clang/LLVM development libraries extracted from apt packages without root privileges. The CLI binary responds correctly to --version and --help. 27 minutes.

28sopranoRuns100 / 100Python1,607

Soprano: Instant, Ultra-Realistic Text-to-Speech

What the test found: Soprano TTS 0.2.0 is installed from source in a venv at /tmp/venv. The CLI produces valid 32 kHz WAV audio on CPU. The FastAPI server serves the /v1/audio/speech endpoint (HTTP 200, valid WAV). The Gradio WebUI module loads successfully. Streaming playback requires the system PortAudio library (not available in this... 15 minutes.

29amicalRuns100 / 100TypeScript1,551

🎙️ AI Dictation App - Open Source and Local-first ⚡ Type 3x faster, no keyboard needed. 🆓 Powered by open source models, works offline, fast and accurate.

What the test found: Installation completed: Node 24.10.0, pnpm 10.34.5, all 1566 packages installed, whisper native addon compiled (.node binary at packages/whisper-wrapper/native/linux-x64/whisper.node), electron 44.3.0 available, workspace builds pass, and the test suite runs (149/150 test files pass, 1505/1508 tests pass, 3... 23 minutes.

30ComfyUI_Custom_Nodes_AlekPetRuns100 / 100JavaScript1,531

Custom nodes that extend the capabilities of Comfyui

What the test found: The repository installs as a ComfyUI custom_nodes extension: pip dependencies install cleanly, ComfyUI starts on CPU, the server answers HTTP 200, and all 8 AlekPet custom node groups load successfully with zero startup errors. 34 minutes.

31SoniTranslateRuns100 / 100Python1,417

Synchronized Translation for Videos. Video dubbing

What the test found: SoniTranslate web app built with Gradio 4.19.2 launches via python app_rvc.py --cpu_mode and serves a fully constructed UI at http://127.0.0.1:7860 returning HTTP 200 on / and /config endpoints. 22 minutes.

32Irodori-TTSRuns100 / 100Python1,410

A Flow Matching-based Text-to-Speech Model with Emoji-driven Style Control

What the test found: Irodori-TTS installs and all core modules import cleanly on CPU. The TextToLatentRFDiT model constructs and runs forward passes. The Gradio web UI builds and serves HTTP 200. Inference requires downloading a HuggingFace checkpoint. 8 minutes.

33transcribe-anythingRuns100 / 100Python1,404

Multi-backend whisper app. Blazing fast. Mac-arm optimized. Easy install. Input a local file or url and this service will transcribe it using Whisper AI...

What the test found: transcribe-anything 4.1.2 installed in a uv-managed Python 3.11 venv with Python 3.10 available for iso-envs; 265/278 tests pass, CLI help works, all backends ready for first-use env creation. 21 minutes.

34MOSS-TTSDRuns100 / 100Python1,403

A multilingual model for long-form, multi-speaker dialogue synthesis with flexible speaker control and zero-shot voice cloning

What the test found: MOSS-TTSD v1.0 model loads from HuggingFace, processor encodes conversations, forward pass returns logits, and generate() produces audio codes and text tokens on CPU with bfloat16 precision. 25 minutes.

35AI-Powered-Video-Tutorial-GeneratorRuns98 / 100Python314

Create and edit AI video tutorials with illustrated lessons, expressive presenters, distinct voices, and a native timeline. Windows desktop app with local...

What the test found: All 6 workspace packages build successfully. JS unit tests pass 156/156 (packages/contracts/scenes/themes). Renderer tests pass 100/117 (17 pre-existing version-mismatch failures). Python pipeline tests pass 942/956 (2 pre-existing failures: Linux path normalization vs Windows test, and stale manifest byte count). The... 22 minutes.

36qwen-audio-agentRuns96 / 100JavaScript2,860

A realtime voice runtime that keeps Agents talking, working, and present. Real-time Voice Runtime for AI Agents

What the test found: The qwen-audio-agent v1.11.0 project installs, builds, launches a Gateway serving HTTP 200 on port 3101, and passes all 1789 automated tests with zero failures. 9 minutes.

37whisper-ctranslate2Runs93.3 / 100Python1,359

Whisper command line client compatible with original OpenAI client based on CTranslate2.

What the test found: Installed in a venv; the CLI answers version 0.5.7, all 17 unit tests pass, and 6 of 7 e2e tests pass including real small/medium-model transcription of the bundled samples matched against reference outputs, with only diarization failing because it needs a real HuggingFace token and pyannote/torch. 22 minutes.

38WhisperJAVRuns92 / 100Python2,312

ASR/STT subtitle generator. Uses Qwen3-ASR, local LLM, Whisper, TEN-VAD. Noise-robust for JAV

What the test found: Project installed from source at /work/repo with 79 packages in a Python 3.12 venv. All 4 verification commands pass: import, --version (1.9.3), --help (full CLI), --dump-params (safe config probe). 617/617 clean tests pass; 43 failures and 11 errors are pre-existing (version drift, restructured config, missing... 11 minutes.

39claude-code-video-toolkitRuns90 / 100Python2,179

AI-native video production toolkit for Claude Code

What the test found: Hello-world, sprint-review template, render-baseline test, and concept-explainer-short template all render valid MP4 videos; migration script fixed and installs all 27 skills/commands; verify_setup.py confirms all prerequisites met. 16 minutes.

40read-aloudRuns90 / 100JavaScript1,757

An awesome browser extension that reads aloud webpage content with one click

What the test found: Read Aloud extension package builds cleanly (474KB ZIP), all 49 JS files pass node syntax check, the extension loads into Chromium 127 via Playwright without errors, manifest declares all TTS permissions and a service worker, and 14 distinct TTS engine implementations are wired in tts-engines.js. 5 minutes.

41minutesRuns85 / 100Rust1,539

Open-source, local-first Granola/Otter alternative that Claude Code, Codex, Cursor, and any MCP client can query. Meetings, calls, and voice memos transcribed...

What the test found: The minutes CLI v0.27.0 binary builds and runs: 59 commands available, demo generates 5 sample meetings with correct markdown frontmatter and YAML metadata, capabilities endpoint reports 50+ features, whisper-guard passes 14/14 tests, 1827/1864 minutes-core tests pass with 37 container-limited failures, 3/4... 30 minutes.

42HandyRuns60 / 100Rust33,185

A free, open source, and extensible speech-to-text application that works completely offline.

What the test found: The Handy app builds from source (debug binary + frontend), passes its full Rust unit test suite (271/271) and Playwright frontend tests (2/2), and launches and stays up on the Xvfb virtual display, enumerating its CPU compute device and 69-model catalog via CLI. 67 minutes.

43AmphionRuns with mocks92 / 100Python10,305

Amphion (/æmˈfaɪən/) is a toolkit for Audio, Music, and Speech Generation. Its purpose is to support reproducible research and help junior researchers and...

What the test found: Core dependencies installed, evaluation metrics compute correctly on audio pairs, model classes import cleanly. pyworld-dependent f0 metrics unavailable without python3-dev headers. Cython monotonic_align extension uncompiled for same reason. 25 minutes.

44mlx-audioRuns with mocks92 / 100Python8,013

A text-to-speech (TTS), speech-to-text (STT) and speech-to-speech (STS) library built on Apple's MLX framework, providing efficient speech analysis on Apple...

What the test found: mlx-audio v0.5.6 installed in a Python venv with mlx[cpu] backend, 354 of 376 tests pass (22 skipped due to missing torch/diffusers for parity), FastAPI server starts and responds 200 on /v1/models, all five CLI tools print help. 8 minutes.

45ODSRuns with mocks92 / 100Python7,147

ODS V3 Pre-Release: Public testing and refinement ahead of the official V3 launch. Turn your PC, Mac, or Linux box into a private AI server.

What the test found: Project builds correctly, all 418 bats tests, 14/14 smoke tests (after ai_err fix), ~300 shell/Python contract tests, and the vite frontend build pass. The Docker runtime install cannot complete in this environment, but the dry-run installer, CLI (-v2.6.0), all standalone test suites, and the frontend are functional. 15 minutes.

46MOSS-TTSRuns with mocks92 / 100Python4,183

An open-source model family for long-form speech, dialogue synthesis, voice design, sound effects, and real-time streaming TTS

What the test found: MOSS-TTS Python package installs and builds; FastAPI server starts and serves /health endpoint on port 8768; CLI entry point prints usage; all core Python modules import; the project is ready for model weight download and GPU inference but cannot run inference in this container because no CUDA GPU is available and the... 31 minutes.

47elevenlabs-pythonRuns with mocks92 / 100Python3,134

The official Python SDK for the ElevenLabs API.

What the test found: The elevenlabs Python SDK v2.69.0 is installed in a venv and all 219 tests pass using httpx.MockTransport to mock ElevenLabs API endpoints. 15 minutes.

48parlorRuns with mocks92 / 100Python2,074

On-device, real-time multimodal AI with features similar to GPT-Live

What the test found: Project installs, builds, and the server starts with Gemma 4 E2B loaded and serving HTTP+WebSocket; the end-to-end AI pipeline (turn detection, llama.cpp inference, transcript parsing, TTS) all initialize and run correctly when tested manually, but the full test suite cannot complete because CPU-only inference exceeds... 77 minutes.

49vllm-mlxRuns with mocks92 / 100Python1,615

High-performance OpenAI and Anthropic compatible LLM inference server for Apple Silicon. Native MLX, continuous batching, multimodal models, MCP tool calling...

What the test found: vllm-mlx 0.5.0 installs into a venv on x86 Linux, its CLI help command prints usage for 8 subcommands, and 1559 of its Linux-compatible unit tests pass. 23 minutes.

50TalkingHeadRuns with mocks92 / 100JavaScript1,596

Talking Head (3D): A JavaScript class for real-time lip-sync using full-body 3D avatars.

What the test found: TalkingHead 3D avatar library installed with npm, local brunette.glb avatar loads successfully in headless Chrome via WebGL, and the full test suite of 32 streaming API tests passes with 0 failures. 14 minutes.

51MiniMax-MCPRuns with mocks92 / 100Python1,583

Official MiniMax Model Context Protocol (MCP) server that enables interaction with powerful Text to Speech, image generation and video generation APIs.

What the test found: Package installed in a venv, all 9 pre-existing tests pass (3 server + 6 utils), and the MCP server binary launches and waits for stdio input. 2 minutes.

52VibeVoice-ComfyUIRuns with mocks92 / 100Python1,566

A comprehensive ComfyUI integration for Microsoft's VibeVoice text-to-speech model, enabling high-quality single and multi-speaker voice synthesis directly...

What the test found: 5/5 nodes registered. Torch 2.14.1+cu130, Transformers 4.57.6, Diffusers 0.39.0. Needs real model files + Qwen2.5-1.5B tokenizer for TTS. 22 minutes.

53Portable-Local-StudioRuns with mocks92 / 100JavaScript1,477

Portable local AI studio for Windows, Linux, and macOS. Zero-setup GUI for Image Generation, GGUF LLMs, Text to Speech & Speech to Text

What the test found: Frontend deps installed, React UI builds and is served on localhost:1420, all 13 regression tests pass, server API endpoints respond correctly. 7 minutes.

54OuteTTSRuns with mocks92 / 100Python1,437

Interface for OuteTTS models.

What the test found: outetts 0.4.4 installed with all dependencies including llama-cpp-python 0.3.9 (CPU), Python modules import cleanly, default speakers load, JS test suite 12/12 passes. 15 minutes.

55meetilyRuns with mocks47.3 / 100Rust31,532

Privacy first, AI meeting assistant with 4x faster Parakeet/Whisper live transcription, speaker diarization, and Ollama summarization built on Rust. 100% local...

What the test found: Meetily Tauri app builds from source, all Rust (188) and frontend (8) tests pass, and the desktop app launches successfully on Xvfb, initializing all subsystems including Tauri GUI, notification system, database, tray icon, and model managers. 61 minutes.

56ai-video-editorCould not verify80 / 100JavaScript903

Open-source, local-first video editor where creators and AI agents edit the same real timeline.

What the test found: npm install, TypeScript typecheck, and Vite build all pass; the dev server starts and serves the application at localhost:5173; the headless timeline CLI command engine parses commands and returns structured JSON errors. 9 minutes.

57diaCould not verify50 / 100Python19,396

A TTS model capable of generating ultra-realistic dialogue in one pass.

What the test found: Package nari-tts installs and all modules import, config loads from HF, CLI help parses, and audio utility functions work correctly, but the 1.6B-parameter model cannot load on this container's 755 MB of RAM (needs at least 3.2 GB). 10 minutes.

58dsnoteCould not verify20 / 100C++1,684

Speech Note Linux app. Note taking, reading and translating with offline Speech to Text, Text to Speech and Machine translation.

What the test found: CMake configured successfully in build_final2 with Qt6 6.4.2, Catch2, fmt, ffmpeg; full build fails at dsnote_lib because 18+ external C/C++ API symbols from webrtc_vad, rnnoise, espeak-ng, piper, april-asr, sam, rdrview, cld2, google_pinyinim, xdo, qhotkey, libarchive, taglib, ssplitcpp, maddy, html2md, cpp-pinyin... 48 minutes.

59orbitCould not verify10 / 100Python353

Self-hosted AI gateway for private RAG, natural-language data access, and tool-calling agents.

What the test found: could not be verified; the log shows where it stopped. 28 minutes.

-hyprwhsprNot yet tested-Python1,234-
-quillmanNot yet tested-Python1,211-
-LiveStream-Agent-StudioNot yet tested-Python1,182-
-dia2Not yet tested-Python1,179-
-PatterNot yet tested-Python1,067-
-whisper-flowNot yet tested-Python983-
-NaturalVoiceSAPIAdapterNot yet tested-C++965-
-VoiceStreamAINot yet tested-Python961-
-botium-speech-processingNot yet tested-JavaScript943-
-jiwerNot yet tested-Python933-
-vonage-php-sdk-coreNot yet tested-PHP929-
-voicyNot yet tested-JavaScript914-
-TheWhisperNot yet tested-Python898-
-lobe-ttsNot yet tested-TypeScript805-
-chaplinNot yet tested-Python759-
-CrispASRNot yet tested-C++739-
-podcast-makerNot yet tested-TypeScript709-
-expo-speech-recognitionNot yet tested-TypeScript689-
-openlrcNot yet tested-Python681-
-sonusNot yet tested-JavaScript639-
-flutter_edge_aiNot yet tested-JavaScript636-
-examplesNot yet tested-TypeScript628-
-mimiNot yet tested-Rust618-
-MimiNot yet tested-Rust618-
-ComfyUI-VibeVoiceNot yet tested-Python600-
-opentypelessNot yet tested-Rust589-
-Playwright-reCAPTCHANot yet tested-Python586-
-openreaderNot yet tested-TypeScript537-
-ai-avatar-systemNot yet tested-Python526-
-ComfyUI-VoxCPMNot yet tested-Python515-
-Audar-ASR-V1Not yet tested-Python508-
-aspeakNot yet tested-Rust497-
-explainrooNot yet tested-JavaScript492-
-deepgram-python-sdkNot yet tested-Python469-
-watch-skillNot yet tested-Python463-
-elevenlabs-jsNot yet tested-TypeScript447-
-VRCTNot yet tested-Python445-
-ai-skillsNot yet tested-Python431-
-splicrNot yet tested-TypeScript431-
-VoiceFlowNot yet tested-Python423-
-project-ravenNot yet tested-TypeScript421-
-parakeet-rsNot yet tested-Rust406-
-vonage-node-sdkNot yet tested-TypeScript398-
-izwiNot yet tested-Rust390-
-onnx-asrNot yet tested-Python382-
-LocalText2VoiceNot yet tested-Python380-
-whisper-obsidian-pluginNot yet tested-TypeScript380-
-yovoiceNot yet tested-TypeScript373-
-VocelloNot yet tested-Swift372-
-audiotextNot yet tested-Python351-
-Whisper-TikTokNot yet tested-Python342-
-sdkNot yet tested-TypeScript341-
-MsEdgeTTSNot yet tested-TypeScript340-
-openai-chat-api-workflowNot yet tested-Ruby319-
-manim-voiceoverNot yet tested-Python317-
-xcodecNot yet tested-Python315-
-saynaNot yet tested-Rust314-

runs installed and started with its real dependencies. runs with mocks started after stand-ins replaced external services such as a database or a third-party API. could not verify neither the standard agent nor the stronger one got it running within the time limit; the log shows where it stopped.

How we tested

On this list as of the latest test: 42 projects ran as-is, 13 with mocks, 4 could not be verified, 57 still waiting. Languages tested: C++, JavaScript, Python, Rust, TypeScript. Every attempt used a clean single-use machine, the subject at a pinned version, and a 45-minute limit; the complete procedure is on the methodology page.

Frequently asked questions (FAQs)

How is this list ranked?

By measurement, not opinion: projects Argusic installed and launched on a fresh machine come first, then those that ran with mocks in place of external services, then those it could not verify. Ties go to the Argusic Score, then how popular it is on its own source.

Why are some projects unranked?

57 projects are still waiting for a test or for a finished attempt. They are listed without a rank until Argusic has measured them.

Where is the evidence?

Every row links to the project's Argusic page, where each run has a full log and a terminal recording stored with a sha256 fingerprint. The same pages exist for every one of the tested projects, on this list or not.

More lists in this category

All lists: Best. All tested projects: subjects.