Open source AI apps you can run yourself, tested

AI applications with a server you host: chat UIs, document question answering, agent builders. Ordered by whether each one came up when Argusic installed it clean; the recording of every attempt is linked.

Tested between and . Each row shows its own test date; a project can change after that day.

76 of 80 tested projects run. 111 more waiting for a test.

In short: 57 of the 80 tested projects started as-is on a fresh machine: Skill_Seekers, ccstatusline, utopia, steel-browser, gpt-load, AIHOT, FunClip, and tick-stock-panel, and 49 more. 19 more started once a stand-in replaced a service they expect, such as a database: ai-goofish-monitor, ODS, open-multi-agent, humanize-text, moltis, lanhu-mcp, VoiceMem, and openclaw-android, and 11 more. 4 could not be verified: agent-ui, fess, selfhost-ai, and orbit; the log shows where each one stopped.

Measured by Argusic on a fresh machine every time. Every number links to its evidence.

#projectverdictArgusic Scorelanguagestarstested on
1Skill_SeekersRuns100 / 100Python15,115

Convert documentation websites, GitHub repositories, and PDFs into Claude AI skills with automatic conflict detection

What the test found: skill-seekers 3.10.0.dev0 installed and all 3857 unit/adaptor/scraper tests pass without external API keys or network dependencies. 6 minutes.

2ccstatuslineRuns100 / 100TypeScript13,222

🚀 Beautiful highly customizable statusline for Claude Code CLI with powerline support, themes, and more.

What the test found: Dependencies installed, project builds to dist/ccstatusline.js, all 1897 tests pass, and piped JSON input renders a formatted status line with model name and git info. 6 minutes.

3utopiaRuns100 / 100Rust8,160

World's first open-source enterprise world model.

What the test found: Utopia server v0.1.0 on :1516, PostgreSQL 16 with pgvector, 21 migrations applied, 268/268 tests passing, web frontend built to web/dist. 26 minutes.

4steel-browserRuns100 / 100TypeScript7,754

🔥 Open Source Browser API for AI Agents & Apps. Steel Browser is a batteries-included browser sandbox that lets you automate the web without worrying about...

What the test found: With a user-local Node 22 and Chrome for Testing plus DISABLE_CHROME_SANDBOX=true, the dev-mode API runs on port 3000, /v1/health returns 200 with live Chrome, and sessions/scrape/screenshot/pdf plus external CDP client connections all worked with real results. 11 minutes.

5gpt-loadRuns100 / 100Go7,055

Self-hosted AI gateway for multi-channel, multi-credential setups, API keys and subscription accounts, scheduling, failover, request logs and usage. 自托管 AI...

What the test found: GPT-Load 2.0 binary builds, web UI compiles, server starts on 127.0.0.1:3001, health endpoint returns HTTP 200, auto-generates AUTH_KEY, and shuts down gracefully with all database migrations completed and catalog synced. 8 minutes.

6AIHOTRuns100 / 100TypeScript6,667

一个自己找热点、自己写日报的网站框架。把信源和精选标准换成你的,它就是你的行业热点站。

What the test found: The project installs cleanly with Node.js 24.11 and PostgreSQL 16, typechecks across 6 packages, passes all 128 backend tests and 11 frontend tests, the API server returns 200 on /openapi-v1.json, and the web server serves /, /all, /about, and /feed.xml at HTTP 200. 13 minutes.

7FunClipRuns100 / 100Python6,379

FunASR-powered video transcription, subtitle generation, and LLM-assisted clipping tool with a local Gradio UI.

What the test found: FunClip v2.2.1 is fully installed in a Python 3.12 venv with all dependencies, all 90 tests pass, the CLI recognition+clipping pipeline processes a real video with Paraformer-Large ASR and produces valid clipped MP4 output, and the Gradio web interface launches and serves HTTP 200 on localhost. 49 minutes.

8tick-stock-panelRuns100 / 100Python5,711

TSP自托管、零运维的 A 股「选股 + 监控 + 回测」量化工作台 | LLM能力驱使策略定制+个股分析+复盘 | 自由接入第三方数据源与个性化扩展数据 | 个人开源

What the test found: Backend server starts and serves /health (200), /api/capabilities, and /api/auth/status; all 1511 backend tests pass; frontend build succeeds via pnpm build. 24 minutes.

9pinmeRuns100 / 100TypeScript3,746

Deploy Your Frontend in a Single Command. Claude Code Skills supported.

What the test found: PinMe CLI v2.0.11 builds and runs. 161 tests pass: 72 unit, 34 integration, 22 MJS regression, 31 CLI black-box, 2 npm pack. 16 minutes.

10anything-analyzerRuns100 / 100TypeScript3,741

全能协议分析工具:浏览器抓包 + MITM 代理 + 指纹伪装 + AI 分析 + MCP Server 无缝对接 AI Agent/IDE | All-in-one protocol analysis toolkit, built-in browser capture, MITM proxy, JS...

What the test found: Installation, build, full test suite (185/189 passing), Electron binary, and built app all work in the container. Only expected dbus/GPU errors appear in container. 9 minutes.

11twinnyRuns100 / 100TypeScript3,664

Open-source AI coding assistant for VS Code. Code completion, chat, edits and reviews with local or hosted models. Your models, your infrastructure.

What the test found: npm install completed, TypeScript compilation and esbuild bundling succeed, the full VS Code extension test suite passes (642 passed, 0 failed), and both twinny-node and twinny-server CLI binaries respond correctly. 8 minutes.

12rakazoRuns100 / 100TypeScript3,486

Open-source Grok Bot alternative. Choose your own model and sandbox.

What the test found: Node 22.23.2 + pnpm 9.15.0 installed; all 1853 npm packages resolved and linked; Prisma client generated; full TypeScript/Vite/Electron/Astro build completes cleanly; all 282 unit-test files pass (2605 tests); Electron 43.4.0 launches on the virtual display and stays up. 9 minutes.

13tuicrRuns100 / 100Rust3,310

a code review TUI with vim keybindings

What the test found: tuicr builds, runs, and 1826 of 1832 tests pass. The binary responds to --version, --help, and review list subcommands correctly. 5 minutes.

14open-terminalRuns100 / 100Python3,282

A computer you can curl ⚡

What the test found: Open Terminal 0.14.0 is installed from source and running as a FastAPI/Uvicorn server on localhost:8768, with all tested API endpoints (/health, /execute, /files/list, /files/read, /files/write, /files/grep, /files/search, /files/delete, /system, /api/config, /skills) responding correctly. 8 minutes.

15claude-tapRuns100 / 100Python3,271

Intercept and inspect Coding Agent API traffic from Claude Code, Codex CLI, Gemini CLI, Cursor CLI, OpenCode, Kimi/Kimi Code, Pi, and Hermes in a local trace...

What the test found: Editable pip install succeeded, all 1136 tests pass, ruff lint and format are clean. 6 minutes.

16stealth-browser-mcpRuns100 / 100Python2,216

Stealth browser automation for AI agents over MCP and the Chrome DevTools Protocol: navigation, network hooks, DOM extraction, and pixel-accurate UI cloning.

What the test found: Stealth Browser MCP server runs on stdio transport, spawns a headless Chromium 157.0.8086.0, navigates to data URLs, extracts page HTML content, and closes cleanly -- all via MCP protocol over stdio. 15 minutes.

17open-coworkRuns100 / 100TypeScript2,194

Open-source AI agent desktop app for Windows & macOS. One-click install Claude Code, MCP tools, and Skills, with sandbox isolation, multi-model support, and...

What the test found: npm install completed successfully with Node.js v22.22.0. 155 of 158 test files pass (1141 of 1144 tests), MCP servers are built, better-sqlite3 native module rebuilt, and Electron v41.7.1 is ready. 63 minutes.

18PanWatchRuns100 / 100Python2,041

PanWatch, AI stock monitoring for A-shares, HK & US markets, powered by TradingAgents. Portfolio insights, real-time alerts & automated reports.|盯盘侠:覆盖...

What the test found: Python 3.12 venv with all deps installed, 782/784 tests pass (3 skipped are Windows-only), server starts and responds 200 on /api/health, frontend builds cleanly via pnpm, WeasyPrint PDF renders Chinese characters as extractable text. 8 minutes.

19TokenTrackerRuns100 / 100JavaScript1,996

Local-first AI token usage & cost tracker for 31 coding tools incl. Claude Code, Codex, Cursor, Gemini & DeepSeek Harness, with native apps. Never reads...

What the test found: All 3052 CLI tests pass with 0 failures; the tokentracker serve command starts and returns HTTP 200 on port 7680. 26 minutes.

20gortexRuns100 / 100Go1,870

High-performance code-intelligence engine for AI agents and IDE, supports 257 languages, multi repositories, based on graph, with access via CLI, MCP Server...

What the test found: Go 1.27.0 installed locally. Gortex builds from source (v0.64.5) and the daemon indexes this repo into a knowledge graph of 182,637 nodes and 822,458 edges. Core graph, search, parser, resolver, query, analysis, semantic, serverstack, and agent test suites pass. The pre-built release binary v0.64.7 was also verified. 69 minutes.

21cc-thinking-skillsRuns100 / 100JavaScript1,611

28 eval-informed mental models and critical-thinking skills for Claude Code, GitHub Copilot, Codex, Cursor, and other Agent Skills-compatible tools

What the test found: All 314 tests pass, all 28 skills pass structural validation (100%), all 23 datasets pass split validation, and the judge reliability validator reports 100% agreement across 3 models. 6 minutes.

22openilink-hubRuns100 / 100Go1,588

开源微信 Bot 管理平台 + App 应用市场 | Self-hosted WeChat Bot Platform with App Marketplace | Lark · Slack · Discord · DingTalk · GitHub · Notion · 20+ Apps | AI Tools | 7...

What the test found: OpeniLink Hub binary built at /tmp/oih, launches on port 9800 with SQLite, serves the web UI, and all 14 Go test suites pass. 7 minutes.

23amicalRuns100 / 100TypeScript1,551

🎙️ AI Dictation App - Open Source and Local-first ⚡ Type 3x faster, no keyboard needed. 🆓 Powered by open source models, works offline, fast and accurate.

What the test found: Installation completed: Node 24.10.0, pnpm 10.34.5, all 1566 packages installed, whisper native addon compiled (.node binary at packages/whisper-wrapper/native/linux-x64/whisper.node), electron 44.3.0 available, workspace builds pass, and the test suite runs (149/150 test files pass, 1505/1508 tests pass, 3... 23 minutes.

24DEEIX-ChatRuns100 / 100Go1,529

An enterprise AI workspace for model routing, multimodal chat, files, tools, billing, identity, and operations.

What the test found: The DEEIX Chat monorepo builds completely: Go backend compiles, all Go unit tests pass (internal/application, shared, ports, transport/http, etc.), Next.js frontend builds to static output, and the single binary serves both API and frontend on port 8080 returning 200 on health, API, and static endpoints. 23 minutes.

25ouroborosRuns100 / 100Python1,419

Ouroboros, self-creating AI agent. Born Feb 16, 2026.

What the test found: Ouroboros 7.4.4 installed and builds; CLI launches, web server serves HTML UI on port 7777 (HTTP 200), default-lane pytest suite passes. 54 minutes.

26commonlyRuns100 / 100TypeScript1,361

Open-source room for humans + cross-vendor AI agents. Every agent gets its own name, memory, skills, and workstation. Any runtime, your infra, no per-agent...

What the test found: Installed deps and built both workspaces, fixed the two real test-bug failures (expired fixture date, missing WebCrypto global), and verified the backend server live on localhost:5000: /api/health 200 healthy with real mongod connected, register 201, login 200 with JWT, while frontend's 962 Jest tests pass and full... 79 minutes.

27skillsgateRuns100 / 100TypeScript1,357

What the test found: With a locally installed Node 22.23.3, `npm install`, the web build, and the desktop build succeed; the Electron desktop app launches on the virtual display with IPC handlers initialized and its SQLite database migrated (7 tables, 5 seeded skills), and the web app serves HTTP 200 responses with real HTML in both dev... 12 minutes.

28marm-memoryRuns100 / 100Python418

Local-first 3-in-1 AI memory layer & MCP server for Claude Code, Codex, Grok, Gemini, VS Code and Cursor. Fuses session history, codebase indexing & concept...

What the test found: marm-mcp-server 2.48.0 installs in a venv, all 1370 non-skipped tests pass, the CLI prints version and help, and all 14 MCP tools are documented consistently across 9 surfaces. 57 minutes.

29Companion-SpaceRuns100 / 100Python363

本地优先的二次元陪伴学习应用 · Local-first anime companion and study app.

What the test found: The backend API starts, all 549 Python tests pass, the Next.js frontend typechecks, lints, and builds, and the mock end-to-end flow (vault unlock, space creation, auto-assigned mock connections, text and streaming turn with valid structured responses, session end) works correctly without any real API keys. 7 minutes.

30matrixhubRuns100 / 100Go357

An Open-source, self-hosted AI model hub with Hugging Face compatibility, accelerating vLLM/SGLang performance.

What the test found: MatrixHub Go backend compiles and runs with SQLite on port 3001, health/endpoint returns 200, login (admin/changeme) authenticates and returns a session cookie, authenticated user CRUD works, built React UI assets are served correctly, and make test.unit passes all 17 test packages (excluding the known-broken hf... 11 minutes.

31docling-StudioRuns100 / 100Python263

Visual document analysis studio powered by Docling, configure the extraction pipeline, inspect text, tables and bounding boxes in the browser, then chunk...

What the test found: Docling Studio backend serves /api/health (200), accepts file uploads, runs async Docling analysis pipeline to completion (COMPLETED status with rendered HTML and page metadata); pytest suite 854/869 pass; frontend vitest suite 432/432 pass; frontend builds to dist/. 11 minutes.

32plandexRuns98.7 / 100Go15,702

Open source AI coding agent. Designed for large projects and real world tasks.

What the test found: Plandex server starts on port 8099 with PostgreSQL 16.4 and LiteLLM proxy on port 4000; server tests pass; CLI builds and authenticates against local server; full CRUD API for projects, plans, branches, settings is operational. 18 minutes.

33YuxiRuns97.3 / 100Python7,317

可私有部署的多租户知识智能体平台:统一 RAG、知识图谱、多智能体、MCP/Skills、沙盒与权限管理。Yuxi = Cloud Agents + Knowledge RAG, Self-hosted knowledge agent platform for RAG, knowledge graphs and...

What the test found: Backend Python dependencies installed and locked; all 1640 backend unit tests pass; all 164 web unit tests pass; Vite production build completes successfully; engineering trust contracts and lint checks pass. Two minor test permission-assertion bugs fixed (umask compatibility) and one Node.js version gap resolved... 25 minutes.

34TaxHackerRuns97.3 / 100TypeScript6,736

Self-hosted AI accounting app. LLM analyzer for receipts, invoices, transactions with custom prompts and categories

What the test found: TaxHacker Next.js app builds, all 33 tests pass, dev server runs on port 7331 and serves the TaxHacker UI with PostgreSQL 16 connected and all 10 Prisma migrations applied, creating 14 database tables. 12 minutes.

35logfireRuns97.3 / 100Python4,510

AI observability platform for production LLM and agent systems.

What the test found: The logfire SDK version 5.0.0 is installed, its CLI reports version info, and 47 of 48 test modules pass with ~1392 tests passing, 2 skipped, and 0 failures. 45 minutes.

36critRuns96.7 / 100Go1,186

Review AI coding agents' plans, diffs and running apps in the browser. Local-first, works with any agent.

What the test found: Built crit binary compiles, launches an HTTP server on 127.0.0.1, serves /api/health returning 200, auto-detects git changes in feature branches, serves embedded frontend, and the full test suite passes all unit and JS tests except one pre-existing intermittent flake in session lazy-threshold ordering. 42 minutes.

37harnessrouterRuns96 / 100Python2,926

HarnessRouter Community Edition: the self-hosted, Apache-2.0 edition of the unified interface for agent harnesses. Run Codex, Claude Code, Hermes, PI, DSH, and...

What the test found: Gateway and runner servers start and respond to health checks; all three test suites (gateway: 622/17, runner: 490/2, protocol conformance: 85/2) pass. 6 minutes.

38memory-osRuns96 / 100Python1,372

A 7-layer memory operating system for Hermes Agent, persistent memory with Qdrant, structured facts, fabric recall, auto-curated wiki, and surgical context...

What the test found: Python dependencies install cleanly in a venv, all 7 icarus module imports succeed, all 3 standalone test suites pass with 0 failures, SQLite databases are initialized with correct FTS5 schema, Icarus plugin tools produce valid output, and the Docker-compose file, worker services, and cron scripts are syntactically... 8 minutes.

39vibe-astockRuns96 / 100Python668

A 股短线复盘看板:涨停池·连板梯队·龙虎榜·板块资金一屏看完,赚钱效应/晋级率/梯队断层/情绪周期等派生指标纯计算直出(不经过 AI),AI 只把数据串成能读的盘面研判。全本地运行,可用 Claude/Codex 订阅免 API key。| A-share short-term daily-review...

What the test found: Python venv installed, Python and Node dependencies satisfied, frontend built, server starts and responds on both health endpoints, all 1282 backend tests pass, all 36 frontend tests pass. 22 minutes.

40ArkhamMirrorRuns96 / 100Python495

Local-first AI-powered document intelligence platform for investigative journalism

What the test found: The SHATTERED API server runs on http://127.0.0.1:8100 with 26 shards, 559 API endpoints, PostgreSQL+pgvector backend, serving a healthy health-check and Swagger UI. All core services (database, config, storage, chunks, vectors, events, workers, models) report active. 23 minutes.

41deepwiki-openRuns94.7 / 100Python18,143

Open Source DeepWiki: AI-Powered Wiki Generator for GitHub/Gitlab/Bitbucket Repositories. Join the discord: https://discord.gg/gMwThUMeme

What the test found: Backend FastAPI server runs on port 8001 with health, models, auth, repo, wiki, chat, codemap, and system endpoints all responding; frontend Next.js dev server compiles and serves on port 3000; 100 out of 105 tests pass (2 deselected for real API keys, 3 stale/invalid tests). 21 minutes.

42react-email-editorRuns94.7 / 100TypeScript5,234

Drag-n-Drop Email Editor Component for React.js

What the test found: The react-email-editor package installs, builds (tsup), and all 12 unit tests pass under Node 22 with --experimental-require-module flag. 4 minutes.

43nanobrowserRuns90.7 / 100TypeScript14,006

Open-Source Chrome extension for AI-powered web automation. Run multi-agent workflows using your own LLM API key. Alternative to OpenAI Operator.

What the test found: Dependencies installed, project builds to dist/ (background.iife.js 1,848.81 kB, manifest.json), TypeScript type-check passes across all 15 packages, unit tests pass (14/14). 15 minutes.

44autoclipRuns90.7 / 100Python9,251

AutoClip|一个链接,一键出片。开源 AI 视频剪辑桌面工具,将播客、访谈、课程等长视频自动剪成短视频,生成字幕、封面和发布文案,适配抖音、小红书、TikTok、Reels 与 YouTube Shorts。Open-source AI video clipping & content repurposing.

What the test found: The AutoClip FastAPI backend serves all API health endpoints and the test suite passes with 103/103 tests passing. The frontend builds to production with Vite. 4 minutes.

45forgeRuns90 / 100Python2,253

A Python framework for self-hosted LLM tool-calling and multi-step agentic workflows

What the test found: forge-guardrails 0.9.5 installed with dev dependencies; 1571 unit tests pass; core, proxy, guardrails, and client modules all import cleanly. 6 minutes.

46opc-skillsRuns90 / 100Python1,846

Agent Skills for Solopreneurs

What the test found: The repo installs via npx skills add with Node 22 and every reddit skill script now fetches live Reddit data (posts, search, comments, users, subreddits), while the website worker serves 200 HTML, skills.json (10 skills), and robots.txt; logo and banner crop scripts produce images that ffmpeg decodes without errors. 24 minutes.

47LaunchStackRuns90 / 100TypeScript890

AI-powered StartUp Accelerator Engine built with Next.js, LangChain, PostgreSQL + pgvector. Upload, organize, and chat with documents. Includes predictive...

What the test found: LaunchStack monorepo installs cleanly after Node upgrade to v20.11.0; package-level vitest (327/327) and web jest (hundreds of unit+integration tests) all pass against a standalone PostgreSQL 16 with pgvector; only missing real external credentials for chat, S3, and provider APIs. 33 minutes.

48astron-rpaRuns88.2 / 100Python5,256

Agent-ready RPA suite with out-of-the-box automation tools. Built for individuals and enterprises.

What the test found: Engine Python 3.13.15 dependencies installed, engine modules import correctly, 207 pytest tests passing across data processing and encryption components, 1 source bug fixed (string fill slice type). Frontend shared and CLI packages build. Can't run backend services or Electron desktop app without Windows/Docker. 22 minutes.

49Tracely-aiRuns80 / 100Python1,515

Trace-native CI/CD for AI agents, production failures become regression tests that block the PR. Auto-detect, cluster, freeze into hermetic cases, replay in CI...

What the test found: Python backend and SDK install, import, and all 1174 pytest tests pass on Python 3.12.3. Frontend installs via pnpm with no errors and next build compiles successfully but OOMs during static page generation (container memory limit). vitest tests require Node 22+ ([email protected] and [email protected]) while this... 43 minutes.

50browser-useRuns66.7 / 100Python117,466

Agents that use the browser.

What the test found: browser-use v0.13.10 installed in a Python 3.12 venv with uv, Chromium headless installed via playwright, and all 563 CI tests pass (529 passed, 34 skipped) after fixing 2 test identity comparison bugs in test_beta_agent.py that only manifest under xdist worker parallelism. 75 minutes.

51OpenCLIRuns66.7 / 100JavaScript29,931

Make Any Website into CLI & Use your logged-in browser by AI agent.

What the test found: OpenCLI 1.8.8 builds and runs on Node 22.14.0; 632 test files pass (7325 tests, 3 skipped) across unit, extension, adapter, smoke projects; CLI binary answers version 1.8.8 and lists 1366 registered commands. 27 minutes.

52llm-wiki-agentRuns60 / 100Python3,609

A personal knowledge base that builds and maintains itself. Drop in sources, Claude (or Codex/Gemini) reads them, extracts knowledge, and maintains a...

What the test found: llm-wiki-agent repo has 9 wiki pages (1 source, 2 entities, 5 concepts, 1 overview), a working knowledge graph (9 nodes, 36 edges), and all 10 Python tools run without errors after fixing 3 bugs in tools/_utils.py and tools/ingest.py. 20 minutes.

53Auto-CompanyRuns60 / 100Python3,120

An auto-company works for 24/7 on your own PC - Windows/Linux/macOS.

What the test found: All 419 Python tests pass, all 4 shell test suites pass, all 19 frontend JS tests pass, dashboard server starts and serves HTTP 200 with structured JSON on / and /api/status. 9 minutes.

54sprite-genRuns60 / 100Python2,607

Generate clean 2D game sprites & animation atlases, component-row pipeline: state rows, alpha cleanup, frame extraction, runtime atlases. Codex/Claude skill.

What the test found: sprite-gen 2.11.0 installed in Python 3.12 venv with dev deps via pip; Pillow 12.3, NumPy 2.5 resolve; CLI responds with full help text; img2webp 1.5.0 with -exact support on PATH; 2020 tests all pass within 300s timeout. 75 minutes.

55openclaw-androidRuns60 / 100Kotlin1,768

Run OpenClaw on Android with a single command, no proot, no Linux

What the test found: OpenClaw 2026.9.8 installed via npm; gateway starts cleanly and serves OpenClaw Control HTML dashboard on port 19000 responding HTTP 200; installation verification passes all 12 applicable checks. 7 minutes.

56LiveAgentRuns50 / 100TypeScript2,219

A fully functional AI Agent desktop client that supports Webui access and can be creatively customized and expanded!

What the test found: The Go gateway binary (agent-gateway / cmd/gateway) builds, starts on port 50052, and returns 200 on /api/status with proper Bearer auth; all Go tests pass; Node script tests (30/30), virtual-core tests (34/34), and GUI release backend tests pass; the desktop GUI (Tauri) cannot build due to missing system libraries... 32 minutes.

57private-gptRuns33.3 / 100Python57,561

Complete API layer for private AI applications on local models: RAG, skills, tools, MCP, text-to-sql, and more. Works with any OpenAI-compatible inference...

What the test found: PrivateGPT server launches with mock LLM/embedding models on port 18080, responds to /v1/models, /health, and /v1/messages API endpoints. Test suite: 2079 passed, 1 failed (pre-existing), 4 errors (transient DB state). 51 minutes.

58ai-goofish-monitorRuns with mocks92 / 100Python14,710

基于 Playwright 和AI实现的闲鱼多任务实时/定时监控与智能分析系统,配备了功能完善的后台管理UI。帮助用户从闲鱼海量商品中,找到心仪产品。

What the test found: Python backend with 117 passing tests and API server running on port 8000; frontend Vue build blocked by Node 18.19.1 < 20.19 requirement. 7 minutes.

59ODSRuns with mocks92 / 100Python7,147

ODS V3 Pre-Release: Public testing and refinement ahead of the official V3 launch. Turn your PC, Mac, or Linux box into a private AI server.

What the test found: Project builds correctly, all 418 bats tests, 14/14 smoke tests (after ai_err fix), ~300 shell/Python contract tests, and the vite frontend build pass. The Docker runtime install cannot complete in this environment, but the dry-run installer, CLI (-v2.6.0), all standalone test suites, and the frontend are functional. 15 minutes.

60open-multi-agentRuns with mocks92 / 100TypeScript6,990

Self-hosted TypeScript agent runtime with durable approvals and verifiable run records. Own it, approve it, audit it.

What the test found: The monorepo installs, builds, passes 2350 unit tests across all 4 workspaces plus the benchmark harness, and the CLI prints help, all with Node.js 22, no API keys or network required. 4 minutes.

61humanize-textRuns with mocks92 / 100Python3,212

Open-source text humanization pipeline with every intermediate step published. Two LLM rewrites at temp 1.3, then two hops across different NMT engines. Four...

What the test found: Project installs in a venv, all 40 unit tests pass, CLI entry point works, and real OpenRouter LLM rewriting (Steps 1-2 of the pipeline) is verified end-to-end. 10 minutes.

62moltisRuns with mocks92 / 100Rust2,885

A secure persistent personal agent server in Rust. One binary, sandboxed execution, multi-provider LLMs, voice, memory, Telegram, WhatsApp, Discord, Teams, and...

What the test found: Moltis gateway binary compiles from source in a fresh Ubuntu 24.04 container, launches on HTTP, prints the setup code banner, and 2810 unit tests pass across 15 of 59 workspace crates, with 5 pre-existing infrastructure-missing test failures (ssh-keygen, WASM, Docker). 48 minutes.

63lanhu-mcpRuns with mocks92 / 100Python2,441

⚡ 需求分析效率提升 200%!全球首个为 AI 编程时代设计的团队协作 MCP 服务器,自动分析需求自动编写前后端代码,下载切图

What the test found: Lanhu MCP Server 1.8.6 installed, all 216 tests pass, the server launches on HTTP port and responds with HTTP 200 to proper MCP JSON-RPC initialize requests. 7 minutes.

64VoiceMemRuns with mocks92 / 100Python2,396

Infrastructure for the next generation of voice agents, designed to provide universal memory. It is divided into a left brain and a right brain, storing...

What the test found: VoiceMem 0.2.3 installed in Python 3.12 virtual environment at /tmp/venv; left-brain memory ingest/search pipeline functional with mock OpenAI API; 5/5 unit tests pass; audio perception models not downloaded (ASR/VAD warmup skipped gracefully); PortAudio-dependent capture unavailable (no root). 15 minutes.

65openclaw-androidRuns with mocks92 / 100Kotlin1,768

Run OpenClaw on Android with a single command, no proot, no Linux

What the test found: openclaw 2026.9.8 npm package installed, gateway running on ws://127.0.0.1:19099 responds HTTP 200 with live status, CLI version/status/doctor commands all functional. 5 minutes.

66ai-moive-studioRuns with mocks92 / 100Python1,613

自然语言驱动的无限画布工作流 Agent,让 AI 视频创作第一次真正变成可编辑的工作流。 AICON 面向创作者,提供从剧本拆解、分镜生成、素材生成、视频合成到内容分发的一整套能力。 不是只给你一个输入框,而是让你用自然语言和无限画布一起驱动创作,把文本、图片、视频节点组织成完整链路,真正把“从灵感到...

What the test found: Backend FastAPI server starts and serves root, health, and API v1 endpoints on port 8000. Frontend builds to dist/. 57 backend unit tests and 15 frontend utility tests pass. 22 minutes.

67PrismerCloudRuns with mocks92 / 100TypeScript1,554

Prismer Cloud

What the test found: TypeScript SDK (built, 569 offline tests pass), Python SDK (installed, 149 offline tests pass), MCP Server (built, 12 tests pass), OpenCode Plugin (built), AIP SDK TypeScript (built, 17/17 compliance tests pass), Runtime (built) -- all 6 TS+Python packages build and pass their offline unit tests. 9 integration tests... 20 minutes.

68superlogRuns with mocks92 / 100TypeScript1,460

Open-source observability tool that uses AI agents to self-heal your software

What the test found: pnpm install succeeded (51s). Node v22.23.3 with pnpm 9.15.9. 143 proxy tests pass, 78 worker tests pass, ~380+ total tests pass across all packages with 0 failures. Full stack requires Docker for postgres/clickhouse/collector which is unavailable without root. 46 minutes.

69WebAI2APIRuns with mocks92 / 100JavaScript1,383

WebAI2API: 基于 Camoufox 的网页 AI 转 API 工具,支持 LMArena/Gemini等,多窗口并发与账号隔离。 | Web AI to OpenAI API via Camoufox. Supports LMArena/Gemini and more, multi-window...

What the test found: The WebAI2API server starts on port 3000 with a Camoufox browser instance, serving OpenAI-compatible /v1/models, WebUI, and /v1/cookies endpoints, all returning HTTP 200, though chat completions require real AI website credentials via the browser adapter and cannot produce output without logging in. 7 minutes.

70awesome-generative-ai-appsRuns with mocks92 / 100JavaScript544

50+ open-source generative AI apps you can clone, deploy, and monetize, image generators, video tools, virtual try-ons, AI SaaS templates, and platform...

What the test found: social-post builds and serves 5/5 routes with HTTP 200 (static pages) or HTTP 401 (auth-required API) using SQLite database and mocked third-party credentials. 15 minutes.

71PharosRAGRuns with mocks92 / 100Python243

Pharos, local-first agentic RAG for your team's document library: multi-format ingest, hybrid retrieval, enterprise ACL, dual HTTP + MCP exits.

What the test found: Pharos installs, builds, and runs its test suite (240/240 pass in the product+engine CPU tests); the HTTP server starts and answers health checks; the MCP stdio adapter initializes with full tool contract; the RAG pipeline (embedder → Qdrant in-memory → hybrid dense+BM25 retrieval → small-to-big context assembly)... 22 minutes.

72pipeshub-aiRuns with mocks68 / 100Python3,818

The open-source context layer for AI agents. PipesHub turns your company's knowledge (Slack, Drive, Jira, GitHub, Microsoft 365 and 40+ connectors) into a...

What the test found: 286 Python unit tests pass, 9633 JavaScript unit tests pass (4 pending). Both suites execute against mocked external services without real infrastructure dependencies. 62 minutes.

73harborRuns with mocks66 / 100Python3,240

Stop configuring your AI stack. Start using it. One command brings a complete pre-wired LLM stack with hundreds of services to explore.

What the test found: Harbor CLI v0.5.6 installed from source, frontend web app builds, all Deno-based unit tests (41 total) pass, lint self-test green (12 bash rules, 9 compose rules, 1 boost rule, 3 orchestrator checks). Docker-dependent container test matrix blocked. 16 minutes.

74WindsurfAPIRuns with mocks56 / 100JavaScript3,071

Turn Windsurf / Devin Desktop's 100+ AI models (Claude, GPT, Gemini, DeepSeek, Kimi, GLM, SWE) into OpenAI-, Anthropic- & Gemini-compatible APIs...

What the test found: The Node.js HTTP server starts on port 3003 and all 4317 tests pass; the Language Server binary is unavailable in this container (requires Windsurf IDE binary), so the proxy functions as a Devin Connect gateway rather than a full Cascade/LanguageServer bridge, which is the expected state when DEVIN_CONNECT is not... 36 minutes.

75meetilyRuns with mocks47.3 / 100Rust31,532

Privacy first, AI meeting assistant with 4x faster Parakeet/Whisper live transcription, speaker diarization, and Ollama summarization built on Rust. 100% local...

What the test found: Meetily Tauri app builds from source, all Rust (188) and frontend (8) tests pass, and the desktop app launches successfully on Xvfb, initializing all subsystems including Tauri GUI, notification system, database, tray icon, and model managers. 61 minutes.

76OpenBiliClawRuns with mocks23 / 100Python3,397

本地私有、开源的自进化跨平台 AI 内容发现 Agent:先理解你,再主动从 B站、小红书、抖音、YouTube、X、知乎、Reddit、微博等平台与开放 Web 寻找内容。(支持 deepseek harness 插件) | Local-first open-source cross-platform AI...

What the test found: OpenBiliClaw v0.3.218 is fully installed, builds, passes lint, and 1007+ tests pass. CLI commands work. The API server starts and responds to health/init-status endpoints. 76 minutes.

77agent-uiCould not verify80 / 100TypeScript1,857

A modern chat interface for AI agents built with Next.js, Tailwind CSS, and TypeScript.

What the test found: Build and typecheck pass cleanly, production server answers HTTP 200 on all routes. 11 minutes.

78fessCould not verify47.5 / 100Java1,140

Open-source, self-hosted enterprise & site search server built on OpenSearch. Crawls web / file / DB / cloud sources, 20+ languages, REST API, and AI/RAG &...

What the test found: Fess 15.9.0-SNAPSHOT builds and packages successfully from source. All 7,634 unit tests pass. The embedded Tomcat binds port 8080 but the application fails at runtime because no OpenSearch backend is reachable on localhost:9200. 36 minutes.

79selfhost-aiCould not verify15 / 100Shell942

One-command installer for a self-hosted AI stack on your own server: n8n, Ollama, Open WebUI, OpenClaw, Dify, Flowise, Supabase, ComfyUI, Qdrant & 30+ tools...

What the test found: The repository is structurally valid (Docker Compose YAML, shell scripts, JSON configs, Python files) but cannot install or launch because Docker is absent and cannot be installed without root privileges in this container. 13 minutes.

80orbitCould not verify10 / 100Python353

Self-hosted AI gateway for private RAG, natural-language data access, and tool-calling agents.

What the test found: could not be verified; the log shows where it stopped. 28 minutes.

-moliNot yet tested-Rust13,400-
-oh-my-hermesNot yet tested-Python3,218-
-awesome-muse-connectorsNot yet tested-Python1,343-
-SuggestArrNot yet tested-Python1,334-
-unigit-ecosystemNot yet tested-JavaScript1,318-
-mcp-searxngNot yet tested-TypeScript1,286-
-ollama-guiNot yet tested-TypeScript1,257-
-collieNot yet tested-TypeScript1,256-
-paulNot yet tested-JavaScript1,251-
-jevNot yet tested-C++1,246-
-OpenFicNot yet tested-Python1,190-
-claude-codex-settingsNot yet tested-Python1,167-
-TensorFoldNot yet tested-Zig1,108-
-tribeNot yet tested-TypeScript1,084-
-awesome-vibecoded-saasNot yet tested-Python1,037-
-parallel-codeNot yet tested-TypeScript1,036-
-CliRelayNot yet tested-Go1,030-
-claude-code-memory-setupNot yet tested-Python1,014-
-OpenCompanyNot yet tested-Python986-
-sandboxdNot yet tested-Go963-
-eclaireNot yet tested-TypeScript923-
-awesome-ai-persona-skillsNot yet tested-Shell907-
-auto-browserNot yet tested-Python904-
-opencode-managerNot yet tested-TypeScript893-
-denovaNot yet tested-Go882-
-claude-code-skill-factoryNot yet tested-Python880-
-llm-lsNot yet tested-Rust880-
-Dev-JanitorNot yet tested-Rust865-
-projectmemNot yet tested-Python854-
-tickdb-unified-realtime-marketdata-apiNot yet tested-Python853-
-mcp-nixosNot yet tested-Python851-
-agentfilesNot yet tested-TypeScript849-
-hueNot yet tested-JavaScript836-
-freehireNot yet tested-Go831-
-free-ai-toolsNot yet tested-TypeScript828-
-tokentapNot yet tested-Python814-
-versus-incidentNot yet tested-Go808-
-osNot yet tested-TypeScript803-
-typescript-style-guideNot yet tested-TypeScript783-
-voidaccessNot yet tested-Python778-
-KrawlNot yet tested-Python774-
-codeseekNot yet tested-Rust771-
-pasteguardNot yet tested-TypeScript762-
-geolookNot yet tested-Python752-
-beads-uiNot yet tested-JavaScript751-
-FeedMeNot yet tested-TypeScript748-
-code-on-incusNot yet tested-Go743-
-coiNot yet tested-Go743-
-ttsfmNot yet tested-Python738-
-gpt-image-canvasNot yet tested-TypeScript735-
-romeNot yet tested-TypeScript728-
-VibeUENot yet tested-C++726-
-obsidian-mcp-serverNot yet tested-TypeScript693-
-swarmclawNot yet tested-TypeScript688-
-ai-development-patternsNot yet tested-Python665-
-world-intel-mcpNot yet tested-Python657-
-spring-ai-agent-utilsNot yet tested-Java654-
-ctxportNot yet tested-TypeScript646-
-readmeXNot yet tested-Python633-
-trinityNot yet tested-Python629-
-LynkrNot yet tested-JavaScript621-
-show-me-the-storyNot yet tested-Go618-
-better-deepseekNot yet tested-JavaScript613-
-mnemonNot yet tested-Go612-
-coolify-mcpNot yet tested-TypeScript609-
-cc-lensNot yet tested-TypeScript598-
-js-reverse-automation--skillNot yet tested-Python590-
-opentypelessNot yet tested-Rust589-
-AdrianNot yet tested-Python580-
-character-arcNot yet tested-TypeScript579-
-claude-brainNot yet tested-TypeScript578-
-xiaohongshu-ai-workbenchNot yet tested-Python578-
-browser-searchNot yet tested-JavaScript528-
-otariNot yet tested-Python513-
-flockNot yet tested-Go502-
-whirlNot yet tested-TypeScript502-
-mirrormateNot yet tested-TypeScript493-
-CoWork-OSNot yet tested-TypeScript473-
-ccLoadNot yet tested-Go418-
-life-klineNot yet tested-TypeScript415-
-GPTPortalNot yet tested-JavaScript397-
-DailyBriefNot yet tested-TypeScript364-
-aurelio-financeNot yet tested-Python351-
-mindroomNot yet tested-Python320-
-picollmNot yet tested-Python318-
-ollaNot yet tested-Go316-
-CotalNot yet tested-TypeScript313-
-ProductFlowNot yet tested-Python306-
-digarrNot yet tested-TypeScript301-
-freebucks-proxyNot yet tested-Go300-
-CatGPT-GatewayNot yet tested-Python295-
-research-clawNot yet tested-Python293-
-BloomNot yet tested-JavaScript284-
-claude-unlimitedNot yet tested-Python282-
-awesome-llm-servicesNot yet tested-TypeScript267-
-GoGogotNot yet tested-Go263-
-terrapodNot yet tested-Python260-
-ResumeForgeNot yet tested-Python235-
-iva-agentNot yet tested-TypeScript233-
-LoomFlowNot yet tested-TypeScript222-
-LexiconNot yet tested-JavaScript217-
-OpenScribeNot yet tested-TypeScript215-
-inference-gatewayNot yet tested-Go214-
-OpenAgentdNot yet tested-TypeScript211-
-aural-ossNot yet tested-TypeScript210-
-skrunNot yet tested-TypeScript210-
-OpenGateLLMNot yet tested-Python206-
-flowbakerNot yet tested-Go205-
-open-tagNot yet tested-TypeScript203-
-Personal_External_BrainNot yet tested-Python201-
-Free-AI-Social-Media-SchedulerNot yet tested-JavaScript2-

runs installed and started with its real dependencies. runs with mocks started after stand-ins replaced external services such as a database or a third-party API. could not verify neither the standard agent nor the stronger one got it running within the time limit; the log shows where it stopped.

How we tested

On this list as of the latest test: 57 projects ran as-is, 19 with mocks, 4 could not be verified, 111 still waiting. Languages tested: Go, Java, JavaScript, Kotlin, Python, Rust, Shell, TypeScript. Every attempt used a clean single-use machine, the subject at a pinned version, and a 45-minute limit; the complete procedure is on the methodology page.

Frequently asked questions (FAQs)

How is this list ranked?

By measurement, not opinion: projects Argusic installed and launched on a fresh machine come first, then those that ran with mocks in place of external services, then those it could not verify. Ties go to the Argusic Score, then how popular it is on its own source.

Why are some projects unranked?

111 projects are still waiting for a test or for a finished attempt. They are listed without a rank until Argusic has measured them.

Where is the evidence?

Every row links to the project's Argusic page, where each run has a full log and a terminal recording stored with a sha256 fingerprint. The same pages exist for every one of the tested projects, on this list or not.

More lists in this category

All lists: Best. All tested projects: subjects.