Open source coding agents that ran on a fresh machine

Agents that write and change code from a prompt, ordered by whether each one installed and started when Argusic tried it fresh, with the recorded session one click away.

Tested between and . Each row shows its own test date; a project can change after that day.

77 of 81 tested projects run. 163 more waiting for a test.

In short: 64 of the 81 tested projects started as-is on a fresh machine: herdr, oh-my-claudecode, ANUS, vibesdk, tuios, twinny, omnara, and git-ai, and 56 more. 13 more started once a stand-in replaced a service they expect, such as a database: rea, claude-context, grok-cli, jevgrep, tdd-guard, DemoGPT, 10x, and opentag, and 5 more. 4 could not be verified: nWave, router, Codewhale, and spec-kitty; the log shows where each one stopped.

Measured by Argusic on a fresh machine every time. Every number links to its evidence.

#projectverdictArgusic Scorelanguagestarstested on
1herdrRuns100 / 100Rust42,896

the runtime your coding agents live on

What the test found: herdr 0.8.2 builds, runs its server, passes all Rust unit tests, Python maintenance tests, Bun integration tests, and plugin marketplace tests with no failures. 19 minutes.

2oh-my-claudecodeRuns100 / 100TypeScript39,679

Teams-first Multi-agent orchestration for Claude Code

What the test found: Oh-My-ClaudeCode v5.2.0 CLI tool builds and responds to version/help/info commands; 1,094+ tests pass with 3 env-constrained failures (Node 18 engine gap + shallow git). 31 minutes.

3ANUSRuns100 / 100JavaScript6,550

A free coding agent in your terminal. It runs on the smartest free model that is up today.

What the test found: The ANUS CLI installs cleanly, all 28 node --test tests pass, and the --version flag returns 0.2.2. 2 minutes.

4vibesdkRuns100 / 100TypeScript5,400

An open-source vibe coding platform that helps you build your own vibe-coding platform, built entirely on Cloudflare stack

What the test found: The VibeSDK project installs, builds, typechecks, and passes all 514 unit tests (root vitest + SDK bun tests) using Bun 1.4.2 and Node.js 20.18.0 inside the container. 11 minutes.

5tuiosRuns100 / 100Go5,030

A terminal window manager that knows what your agents are doing. Tiling panes, workspaces, sessions that survive restarts, and one Inbox for every coding agent.

What the test found: tuios and tuios-web binaries build from source, the full Go test suite (49 packages) passes, and the daemon creates sessions, accepts typed input via send-text/send-keys, and captures pane output showing shell command execution. 22 minutes.

6twinnyRuns100 / 100TypeScript3,664

Open-source AI coding assistant for VS Code. Code completion, chat, edits and reviews with local or hosted models. Your models, your infrastructure.

What the test found: npm install completed, TypeScript compilation and esbuild bundling succeed, the full VS Code extension test suite passes (642 passed, 0 failed), and both twinny-node and twinny-server CLI binaries respond correctly. 8 minutes.

7omnaraRuns100 / 100Go2,897

The open-source managed agent platform. A self-hostable alternative to Claude Managed Agents and OpenAI's Agents API.

What the test found: Omnara builds, all 124 unit test packages pass, all 34 integration test packages pass against live PostgreSQL 18 + Redis 7 + MinIO S3, all 4 server binaries build and the API binary starts and responds on port 8080, and the frontend web app builds successfully. 43 minutes.

8git-aiRuns100 / 100Rust2,820

A Git extension for tracking the AI-generated code in your repos

What the test found: git-ai builds from source with cargo, all 2481 unit tests pass, 1310+ sampled integration tests across checkpointing, blame, diff, rewrite ops, agent presets, worktrees, and async daemon mode all pass with RUSTUP_TOOLCHAIN=stable and GIT_AI_TEST_BINARY_PATH set. 45 minutes.

9teaql-agent-kitRuns100 / 100Python2,814

A model-mediated harness for reliable agentic software development.

What the test found: Python-based teaql_workspace.py tool verifies and applies runtime source configurations for all 7 supported languages; the build-teaql-app skill, golden examples, and reference documentation are complete and well-formed; the npx skills add install path is gated by a Node.js 22 requirement not met in this container. 3 minutes.

10agent-toolkitRuns100 / 100Python2,539

A curated collection of skills for AI coding agents. Skills are packaged instructions and scripts that extend agent capabilities across development...

What the test found: The build script assembles all 56 plugins into dist/ (232 files, 1.7MB), the marketplace.json passes CI validation, dist is in sync with source, version bumping works, and every SKILL.md, Python script, and bash script in the repository is syntactically valid. 4 minutes.

11fable-methodRuns100 / 100Python2,297

The Fable Workflow: how Claude Fable 5 worked, distilled into skills any model can run, with the eval that keeps it honest. Think / act / prove.

What the test found: The Fable Method repository installs cleanly, its 4 skills (fable-method, fable-loop, fable-judge, fable-domain) are at ~/.claude/skills/ with valid frontmatter, 9 domain adapters and 14 eval scenarios are complete, 15 eval result JSON files parse, and all Python/JS eval fixtures execute as designed. 3 minutes.

12open-coworkRuns100 / 100TypeScript2,194

Open-source AI agent desktop app for Windows & macOS. One-click install Claude Code, MCP tools, and Skills, with sandbox isolation, multi-model support, and...

What the test found: npm install completed successfully with Node.js v22.22.0. 155 of 158 test files pass (1141 of 1144 tests), MCP servers are built, better-sqlite3 native module rebuilt, and Electron v41.7.1 is ready. 63 minutes.

13minimax-codeRuns100 / 100TypeScript1,998

An open-source coding agent for your terminal, powered by MiniMax.

What the test found: MiniMax Code 0.5.5 builds from source, passes 6003 of 6007 tests across smoke, BYOK, release-tools, status-contract, policy, sandbox, artifact, and capability suites, and answers to mcode --version. 27 minutes.

14pr-lensRuns100 / 100TypeScript1,874

Review code 100X faster. Lens draws every PR as animated architecture and data-flow walkthroughs, inside the pull request itself. Use it as a GitHub App...

What the test found: 872 tests pass across all 7 workspace packages, build and typecheck are clean, CLI binary runs and prints version 0.11.0. 3 minutes.

15nimbalystRuns100 / 100TypeScript1,852

Nimbalyst - The open-source visual workspace for Claude Code, Codex, and OpenCode. Run multiple coding agents in parallel, edit their work visually in...

What the test found: Nimbalyst monorepo installs, builds (runtime, extension-sdk, collab-bundle), and passes its full vitest unit test suite with 1546 passing test files and 12983 passing tests. 41 minutes.

16CoreCoderRuns100 / 100Python1,796

Minimal AI coding agent (~1,000 lines of Python) inspired by Claude Code. Works with any LLM. Think NanoGPT for coding agents. Formerly NanoCoder.

What the test found: CoreCoder installed in a virtual environment; 185 tests all green; CLI --help and --demo mode both work; ruff reports one pre-existing EXE001 lint on examples/plan_hooks_demo.py. 1 minute.

17zeroRuns100 / 100Go1,697

The coding agent that answers to you, your model, your machine, your rules.

What the test found: Zero v0.9.0 builds from source, passes all 87 test packages, passes go vet, format check, vulncheck, smoke test, and answers all CLI commands (help, models list, providers list, doctor) on linux/amd64. 9 minutes.

18zeroRuns100 / 100Go1,697

The coding agent that answers to you, your model, your machine, your rules.

What the test found: Zero builds from source, all 87 test targets pass, release build produces version 0.9.0 and passes smoke test, and the CLI launches and responds to commands. 9 minutes.

19spec-kittyRuns100 / 100Python1,678

Spec-Driven Development with organizational governance. Specs tell AI agents what to build; Charter governs how they build it. Git-native missions, enforceable...

What the test found: Spec Kitty CLI 3.2.7rc1 installed in venv at /work/repo/.venv, all major test suites pass with 1903+ tests passing and 0 failures, spec-kitty init creates a working project with Codex agent integration. 82 minutes.

20pstack-claudeRuns100 / 100JavaScript1,621

Claude Code, Codex, Copilot, Pi, OpenCode, Gemini, and Prime Agent versions of Poteto's pstack. Rigorous agent workflows with Cursor primitives translated for...

What the test found: pstack v0.9.45 installs cleanly: bun tools/generate.mjs passes all 16 validation checks, bun test passes 227/227 tests, markdownlint passes on 165 files with 0 errors, the skills CLI installs all 54 skills and diff matches the source tree byte-for-byte, the vendored scripts compile cleanly, and the working tree is... 3 minutes.

21pi-mcp-adapterRuns100 / 100TypeScript1,579

Token-efficient MCP adapter for Pi coding agent

What the test found: Installed and running on Node 22.23.3. Tests pass 1997/2000, typecheck passes, public exports verified, CLI works. 9 minutes.

22ouroborosRuns100 / 100Python1,419

Ouroboros, self-creating AI agent. Born Feb 16, 2026.

What the test found: Ouroboros 7.4.4 installed and builds; CLI launches, web server serves HTML UI on port 7777 (HTTP 200), default-lane pytest suite passes. 54 minutes.

23oh-my-agentRuns100 / 100TypeScript1,337

Mechanical verification for AI coding agents, skills pack or full harness (stop-hook gates, artifact checks, independent judges).

What the test found: oh-my-agent v13.2.1 builds successfully, the oma CLI binary prints version 13.2.1, and all 4264 tests in 323 test files pass with no failures. 53 minutes.

24vibecode-pro-max-kitRuns100 / 100JavaScript1,145

Your AI forgets. This remembers. Spec-driven coding harness for vibecoders, product owners, CEOs and real builders, self-improving context memory, 15 agents...

What the test found: Node 22 installed, resolve-manifest.js fixed (exclude bug), e2e assertions widened: all 4 e2e tests pass, install succeeds end-to-end with discover-skills OK, compute-sync-plan reports full sync (0 stale/0 missing). 28 minutes.

25EnterpriseAgentFrameworkRuns100 / 100Java845

ReachAI企业级智能体开发平台:快速、安全完成已有业务系统智能化改造,让 AI 在 OA、ERP、CRM 等原系统中查数据、填表单、办业务。ReachAI: Quickly and securely bring AI to existing enterprise systems, enabling AI to...

What the test found: Backend builds cleanly and all 2246 unit/integration tests pass across 9 Maven modules. Frontend builds with npm run build and 453 of 454 vitest tests pass across 10 test suites (1 pre-existing timezone-dependent dashboard test fails in UTC). 34 minutes.

26wcgwRuns100 / 100Python678

Shell and coding agent on mcp clients

What the test found: wcgw 5.6.5 is fully installed, all 63 pytest tests pass, the MCP server entry point (wcgw) and CLI (wcgw_local) both respond to --help, and all package imports resolve cleanly. 2 minutes.

27building-a-coding-agent-from-scratch-courseRuns100 / 100Python615

Learn harness engineering by building Claude Code from scratch. Free, open-source course: 8 articles, 6 videos, 1 codebase.

What the test found: Installed and built with uv, 2978 unit tests and 80 integration tests pass, format and lint clean, decode CLI reports version 0.1.0. 5 minutes.

28zhikuncodeRuns100 / 100Java513

Codex/Claude Code/Cursor的开源增强版,专注一句话实现复杂长程任务(教育、编程、办公、生活、娱乐、游戏)。部署在你自己的服务器上,团队用浏览器打开就能编程, 包括手机。CLI & Web UI 双入口,Multi-Agent 协作,原生直连千问/DeepSeek...

What the test found: Python FastAPI service runs on port 8000 with health endpoint OK; Java Spring Boot backend runs on port 8080 with actuator health status UP and SQLite global+project databases operational; React frontend builds cleanly with 207/223 tests passing; all three tiers (Python 105/107, Java 2392/2454, React 207/223) pass... 18 minutes.

29AsyncRuns100 / 100TypeScript474

IDE, A native-feeling AI coding workspace that blends chat, planning, agent execution, and project navigation into a unified desktop experience.

What the test found: npm install completes, full build (typecheck + esbuild + vite) succeeds, all 117 test files (709 tests) pass, and the Electron app starts with the dist bundle rendering the agent window via did-finish-load. 8 minutes.

30agent-specRuns100 / 100Rust457

`agent-spec` is an AI-native BDD/spec verification tool for task execution.

What the test found: agent-spec 1.4.0 builds and all 1111 tests pass on Rust 1.98.1 with no modifications. 3 minutes.

31GitoRuns100 / 100Python444

An AI-powered GitHub code review tool that uses LLMs to detect high-confidence, high-impact issues, such as security vulnerabilities, bugs, and maintainability...

What the test found: All 157 tests pass and gito CLI version command runs from a Python 3.12 venv, with only the version-shell test requiring a PATH fix. 3 minutes.

32agentsRuns98.7 / 100Python40,296

Multi-harness agentic plugin marketplace for Claude Code, Codex, Cursor, OpenCode, GitHub Copilot, Google Antigravity, and Pi

What the test found: uv sync installs all deps for both projects; pytest suite passes 619/621; all 6 harness generators emit artifacts; codex doctor reports 20/20 checks ok; structural validation passes across 6 harnesses with strict mode. 11 minutes.

33deepseek-harness-desktopRuns98.3 / 100JavaScript783

Open-source Windows desktop client and GUI for DeepSeek Harness, zero-setup installer with Codex, plugins, skills, SSH, mobile remote access, and 11 skins.

What the test found: Dependencies installed, all 40 workspace projects build and typecheck. Script tests 200/200 pass. Desktop test suite 1229/1229 pass (0 fail, 26 skipped). DSH web runtime starts and serves HTTP (401 with auth active). One pre-existing issue: dsh-ssh engine.test.ts needs ssh-keygen (openssh-client) unavailable in this... 22 minutes.

34DeepSeek-ReasonixRuns97.3 / 100Go35,748

A reliable coding agent for complex software engineering tasks.

What the test found: Reasonix builds as a single Go binary (bin/reasonix), answers --help with full usage including CLI/TUI, web, serve, ACP, and bot subcommands, passes all 147+ test packages with 0 failures, and produces valid JSON diagnostics via doctor --json. 19 minutes.

35omnigentRuns97.3 / 100Python10,676

Omnigent is an open-source AI agent framework and meta-harness: orchestrate Claude Code, Codex, Cursor, Pi, and custom agents, swap harnesses without...

What the test found: The omnigent CLI reports version 0.13.0.dev0, the server listens on TCP and returns HTTP 200 on root/health/OpenAPI endpoints, and 929+ unit tests across 10 module directories pass without failure. 28 minutes.

36cli-continuesRuns96.7 / 100TypeScript1,563

resume any AI coding session in another tool, Claude Code, Copilot, Gemini, Codex, Cursor

What the test found: Node.js upgraded to v22.14.0 (via /tmp/), pnpm installed and builds approved, project compiles with tsc with zero errors, all 28 test files pass (917/919), CLI binary functions (help, version, scan, list all work), and biome passes with warnings only. 7 minutes.

37critRuns96.7 / 100Go1,186

Review AI coding agents' plans, diffs and running apps in the browser. Local-first, works with any agent.

What the test found: Built crit binary compiles, launches an HTTP server on 127.0.0.1, serves /api/health returning 200, auto-detects git changes in feature branches, serves embedded frontend, and the full test suite passes all unit and JS tests except one pre-existing intermittent flake in session lazy-threshold ordering. 42 minutes.

38harnessrouterRuns96 / 100Python2,926

HarnessRouter Community Edition: the self-hosted, Apache-2.0 edition of the unified interface for agent harnesses. Run Codex, Claude Code, Hermes, PI, DSH, and...

What the test found: Gateway and runner servers start and respond to health checks; all three test suites (gateway: 622/17, runner: 490/2, protocol conformance: 85/2) pass. 6 minutes.

39qwen-audio-agentRuns96 / 100JavaScript2,860

A realtime voice runtime that keeps Agents talking, working, and present. Real-time Voice Runtime for AI Agents

What the test found: The qwen-audio-agent v1.11.0 project installs, builds, launches a Gateway serving HTTP 200 on port 3101, and passes all 1789 automated tests with zero failures. 9 minutes.

40little-coderRuns96 / 100TypeScript2,655

A harness optimized to smaller LLMs

What the test found: little-coder v1.19.0 installs and runs: npm install succeeds (182 packages), all 715 vitest tests pass, 4/4 pytest RPC client tests pass, TypeScript typechecks cleanly, the CLI launcher starts pi and lists available models correctly. 4 minutes.

41paritok-4b-v1Runs96 / 100Python1,453

Non-destructive compression gateway for AI coding agents. Cuts token bills 25% on turn 1 to past 85% in long or saturated sessions, and fits ~3× more turns in...

What the test found: Python package installed in a venv, the proxy serves on port 8088, answers /health (200) and /stats (200) and /v1/messages/count_tokens (200), and all 316 tests in the test suite pass. 8 minutes.

42MystiRuns96 / 100TypeScript1,141

AI coding dream team of agents for VS Code. Claude Code + openai Codex collaborate in brainstorm mode, debate solutions, and synthesize the best approach for...

What the test found: Mysti v0.4.0 installs, builds via webpack, and passes all 367 automated tests with vitest v2 on Node 18, with the vscode API mocked. 5 minutes.

43agent-deckRuns96 / 100Go1,037

Terminal session manager for AI coding agents. One TUI for Claude, Gemini, OpenCode, Codex, and more.

What the test found: Agent-deck vdev builds and runs in the container. The binary responds to version, help, and ls commands. 34 of 34 directly testable internal packages pass. The WebSocket terminal bridge functions correctly. Only pre-existing environment-specific test failures: keepalive client-count assertions (container PTY behavior)... 62 minutes.

44deer-workflowRuns96 / 100TypeScript548

An open-source graph engineering runtime that keeps orchestration in TypeScript and delegates semantic work to replaceable Agent runtimes.

What the test found: The deer-workflow repository installs and builds cleanly with Bun 1.4.2, all 89 tests pass, and the CLI runs showing help for create/skill/run commands. 2 minutes.

45CodewhaleRuns94.7 / 100Rust41,084

Open-source Rust agent engine and terminal client for Codewhale, with provider choice, tools, approvals and receipts.

What the test found: Codewhale 0.9.11 builds, passes all tests, launches CLI help/doctor/exec against a mock API. Debug profile is memory-safe; release build hits OOM on constrained hardware. D-Bus dev headers were missing and were mocked via vendored crate sources. 42 minutes.

46Vibe-SkillsRuns93.5 / 100Python3,606

Intelligent Skill routing and workflow orchestration for AI agents, +21.12 pp reward, −29.6% tokens on SkillsBench with DeepSeekV4Flash-VE.

What the test found: VibeSkills v4.1.0 installs successfully, check verifies all 292 receipt-owned files, the vgo_cli Python launcher works end-to-end, and 129+ core unit tests pass with 3 isolated test failures due to missing pwsh in this container. 37 minutes.

47gentle-aiRuns93.3 / 100Go7,599

Gentle-AI configures the AI coding agents you already use: Claude Code, Cursor, OpenCode, Codex, Pi, and more. Choose persistent memory, Organic-Driven...

What the test found: gentle-ai 2.0.0 builds from source, prints version (2.0.0-20260906031619-2c25e878eea4) and help text, all 53 Go test packages pass. 17 minutes.

48agentsRuns93.3 / 100TypeScript1,599

A durable process per agent, with memory that survives restarts

What the test found: Node 22.19.0 with pnpm 9.15.0 installed, all dependencies resolved and built, type checks pass, boundary checks pass, 68 of 71 tests pass across packages/pi and packages/sandbox-adapter. 39 minutes.

49DeepCodeRuns90.7 / 100Python16,691

"DeepCode: Open Agentic Coding (Agent Harness & Loop Engineering & Multi-Agent Orchestration)"

What the test found: The Python package is installed and builds, the CLI reports version 2.1.0, the app-server runtime probe passes, and the full test suite (1586 tests) passes. 3 minutes.

50ripwireRuns90 / 100C++2,423

The ripgrep of AI context: a zero-dependency C++23 CLI + MCP server for coding agents. Find what you want without reading the repo, then check you built what...

What the test found: Build succeeds (plain dev build/ and Release build-install/), binary parses the repo source tree and emits deterministic minified XML with ranked symbols via Personalized PageRank, installed to /home/runner/.local/bin/ripwire with skills and hooks, --version/--help work, Python indexing confirms correct function... 21 minutes.

51no_humanRuns89.3 / 100Python330

From ticket to reviewed pull request. Free and open-source, on your machine.

What the test found: The no-human project (v0.2.0) is fully installed: uv sync completes, web board builds from source, the nh CLI responds with version 0.2.0, nh start binds and serves HTTP 200 in setup mode, the full node test suite (1584/1584) and the Python repoguard suite (142/142) pass. Task dispatch requires a Claude subscription... 71 minutes.

52serenaRuns88 / 100Python30,098

A powerful MCP toolkit for coding, providing semantic retrieval and editing capabilities - the IDE for your agent

What the test found: Serena is installed, initialized, and the MCP server starts successfully with all 708 tests passing. 12 minutes.

53cc-safety-netRuns86.7 / 100TypeScript1,582

A pre-execution guard for AI coding agents. It blocks destructive Git and file system commands, plus common attempts to access sensitive files, before a tool...

What the test found: bun install and bun run build succeed. bun run check passes lint, formatting, typecheck, knip, and duplication checks, with 3673 tests passing out of 3677 total across 213 files. Coverage is 98.91%. All E2E tests pass (packed-runtime, hermes-openclaw, and protection contracts). 15 minutes.

54claudexorRuns86 / 100TypeScript496

Multi-harness control plane for Claude Code, Codex, Cursor, and OpenCode: quota-aware rotation across multiple Claude/Codex subscriptions, shared thread...

What the test found: pnpm install completed, pnpm build succeeded (31 packages in 18s), the CLI runs and reports version 3.9.8, the full test suite passes (350 test files, 4692 tests), canary stories pass (6 files, 50 tests), and typecheck passes (61 tasks). 20 minutes.

55runjamRuns85 / 100Rust228

One desktop for all your AI coding Agent, Claude Code, Codex CLI & Gemini CLI. Auto-detect, one-click install, unified chat, file explorer, terminal & editor...

What the test found: Frontend builds with Vite, backend compiles with Cargo, all 98 tests (13 JS + 85 Rust) pass, and the desktop binary launches on Xvfb and starts its proxy server. 16 minutes.

56rtkRuns80 / 100Rust82,684

CLI proxy that reduces LLM token consumption by 60-90% on common dev commands. Single Rust binary, zero dependencies

What the test found: RTK (Rust Token Killer) v0.42.4 builds from source, passes all 2936 tests with 0 failures, and runs as a working CLI proxy producing token-compressed filtered output for ls, git, cargo, and other commands. 8 minutes.

57thinkrailRuns60 / 100TypeScript513

Vibe code with pi in a lightweight, real IDE that customises itself around the way you work, The Vibe You Need

What the test found: 890 deps installed, 14 packages typecheck, 14/14 test suites pass (987 in server alone), web app builds, E2E 183/183 passes, server /health returns 200. 18 minutes.

58jcodeRuns53.3 / 100Rust20,351

High performance coding agent harness written in rust

What the test found: jcode v0.81.7-dev is built from source, installed via symlinks in ~/.jcode/builds/current and ~/.local/bin, and works correctly with OpenRouter provider. jcode run and jcode repl both complete successfully with real AI responses. 20 minutes.

59delegate-skillsRuns50 / 100JavaScript2,331

Delegate a coding task to a separate coding agent CLI, review the diff, land the commit yourself, one per implementer.

What the test found: The delegate-skills package installs all 18 skills, the test suite passes all checks against fake CLI shims, and the POSIX shim no longer depends on the external dirname command. 30 minutes.

60moai-adkRuns50 / 100Go1,232

Agentic development harness for Claude Code, SPEC-driven plan/run/sync, TRUST 5 quality gates, model+effort routing, and Claude×GLM multi-LLM cost control...

What the test found: moai-adk v3.1.3 builds and runs. Binary at ./bin/moai responds to 'moai version' with v3.1.3. Test suite runs with 3 pre-existing failures (2 test-isolation, 1 codex-binary not available). templ code generation works. All ast-grep rule tests pass with version 0.40.5. 57 minutes.

61clawcodexRuns50 / 100TypeScript905

Token efficient Claude Code full Python rebuild. AI Coding Agent in 310K LoC Python.

What the test found: Python virtual environment installs all dependencies cleanly, the 'clawcodex' CLI starts and reports version 1.6.0, and the full test suite of 10,435 tests passes with 0 failures. 51 minutes.

62agnixRuns50 / 100Rust445

The missing linter and lsp for AI coding assistants. Validate CLAUDE.md, AGENTS.md, SKILL.md, hooks, MCP. Plugin for all major IDEs included, with autofixes.

What the test found: agnix 0.52.2 installed and working from both npm (user dir) and cargo source build; 4700+ tests passing across all 6 workspace crates; binary validates agent configs and produces text, JSON, SARIF, and GitHub annotations output formats with i18n support. 20 minutes.

63cli-agent-orchestratorRuns48 / 100Python1,397

Multi-agent orchestration for AI coding CLIs, Claude Code, Kiro, Codex, and more, coordinated in isolated tmux sessions

What the test found: CAO server starts, responds 200 on /health and /docs, and passes all unit tests (2986+ passed) except for 4 pre-existing test-isolation wiki_lint failures and Node version-gated web UI build. 43 minutes.

64lean-ctxRuns40 / 100Rust3,870

LeanCTX, Context Gateway for AI Systems. Control what your AI can see. Open-source Engine for context selection, supported controls, and evidence.

What the test found: lean-ctx 3.10.1 release binary is built, installed at ~/.cargo/bin/lean-ctx, and fully functional with all CLI commands working. 65 minutes.

65reaRuns with mocks92 / 100TypeScript19,434

Reverse engineer anything with agents, from app behavior down to native binaries.

What the test found: REA compiles and runs on Node.js 22.19.0 with npm 11.16.0. The full TypeScript build succeeds, all 353 Vitest test files (1756 individual tests) pass covering domain logic, adapters, composition, MCP boundary, acceptance, conformance, and evaluation layers, and the CLI lists all 60+ analysis commands without error. 8 minutes.

66claude-contextRuns with mocks92 / 100TypeScript12,593

Code search MCP for Claude Code. Make entire codebase the context for any coding agent.

What the test found: Install, build, typecheck, lint config, and all 35 tests pass across the monorepo. The core indexing pipeline runs end-to-end with mock providers. The MCP server starts but requires real OpenAI API key and Milvus credentials for production use. 15 minutes.

67grok-cliRuns with mocks92 / 100TypeScript3,487

An open-source coding agent for the Grok API

What the test found: All 48 test files pass (248 tests), typecheck passes via tsc --noEmit, and CLI help displays correctly via bun run dist/index.js --help. 9 minutes.

68jevgrepRuns with mocks92 / 100TypeScript2,436

Find code by asking what it does. A CLI for coding agents that uses Jev to discover relevant files and source context.

What the test found: Repository builds, types check, lints clean. All 4 test suites pass (70 pass, 0 fail). End-to-end CLI search works against a mock provider: auth saves credentials, doctor verifies access, search returns file locations and source excerpts. 6 minutes.

69tdd-guardRuns with mocks92 / 100TypeScript2,360

Automated TDD enforcement for Claude Code

What the test found: TDD Guard installs, builds, and runs. All unit tests (770) and linter tests (24) pass. Core validation logic works with mock AI client. Jest, Vitest, Go, and pytest reporter integrations work. Validator integration and 6 language-specific reporter tests need real API credentials or additional language runtimes not... 53 minutes.

70DemoGPTRuns with mocks92 / 100Python1,906

🤖 Create agentic apps in a second with your prompts. Everything you need to create an LLM Agent - tools, prompts, frameworks, and models - all in one place.

What the test found: Install built a working venv with all dependencies. All 4 tests pass against a mock OpenAI server (2 LLM: OpenAIModel, OpenAIChatModel; 2 RAG: add_files, add_text). Streamlit app launches on port 8592 and returns HTTP 200. The demogpt CLI entry point uses subprocess('streamlit') which resolves only from the venv bin... 55 minutes.

7110xRuns with mocks92 / 100TypeScript1,371

⚡️ 10x - Up to 20x faster AI coding with multi-step Superpowers. Open-source agent with smart model routing, BYOK, fully self-hosted.

What the test found: Build of core, shared, and CLI packages succeeded. All 234 tests pass (184 core, 50 shared). CLI binary at apps/cli/dist/index.js answers --version (0.1.0) and --help. Web build failed due to out-of-memory (SIGKILL) during Next.js static page generation. 9 minutes.

72opentagRuns with mocks92 / 100TypeScript1,363

Mention any ACP coding agent from Slack, GitHub, GitLab, Linear, or Lark. OpenTag runs Claude Code, Codex, Cursor and more on your own machine, then replies...

What the test found: The repository installs, typechecks, lints, and builds cleanly; the full vitest suite passes (1502 passed) against a user-space PostgreSQL 17.10; the credential-free paired-relay contract smoke passes (113 tests); and the built CLI answers --version 0.11.0 and --help. 27 minutes.

73HolyClaudeRuns with mocks86 / 100JavaScript2,580

AI coding workstation: Claude Code + web UI + 8 AI CLIs + headless browser + 50+ tools

What the test found: All 285 Node.js unit tests, 5 Python notification tests, and the CLI persistence shell test pass. Repository structure is validated end to end with product facts contracts, security policy, documentation, and patch scripts all verified. 2 minutes.

74ospecRuns with mocks86 / 100JavaScript452

Spec-driven, agentic workflow framework for AI coding agents. Turn a request into a verifiable goal loop, plan, act, verify, with durable specs and evidence in...

What the test found: npm install completed without errors (Node 18 engine warnings only); ospec CLI v2.1.0 responds to --version, --help, init, status, change, session commands; release smoke test and content scan both pass; index rebuild tool works; global install via --prefix succeeds. 5 minutes.

75poirotRuns with mocks56 / 100Python249

Poirot is a deep research agent kernel built for those who care about how agents are architected.

What the test found: poirot installs via pip into a venv, all 2716 unit/integration tests pass, and the CLI binary reports help for run and cli subcommands. 30 minutes.

76ouroborosRuns with mocks44.3 / 100Python6,194

Agent OS: the agent gets smarter on its own. We just hold the line: Interview-gated, staged evaluation, budgeted evolution loop. MCP server, 14 runtimes...

What the test found: All unit tests pass (0 failures in ~21000 tests across 34 unit directories), ruff lint and format clean, CLI boots with correct version, MCP server help displays correctly, 2 Windows-specific bugs were caught and fixed. 73 minutes.

77openinterpreterRuns with mocks30.7 / 100Rust68,523

A coding agent for open models like Kimi K3 and GLM 5.3

What the test found: The codex-rs workspace builds successfully, the binary launches, and the core protocols/config/otel/api tests all pass. CLI tests pass at 99.4% with only 2 pre-existing failures unrelated to the installation: features list doesn't fetch cloud config, and sandbox enforcement needs bubblewrap. 71 minutes.

78nWaveCould not verify80 / 100Python617

AI agents that guide you from idea to working code, with you in control at every step.

What the test found: nwave-ai v3.22.2 installed in a venv, deploys to ~/.claude/ for Codex CLI, all 10 doctor checks pass, 34 agents and 151 skills install successfully, and all key Python modules (nwave_ai.cli, des.domain, nwave_ai.sync, nwave_ai.outcomes) import without errors. 7 minutes.

79routerCould not verify40 / 100Go5,580

Model router for agentic systems. Routes every prompt to the right model in <50ms. Cut costs 40-70% with just an endpoint change.

What the test found: could not be verified; the log shows where it stopped. 14 minutes.

80CodewhaleCould not verify20 / 100Rust41,083

Open-source Rust agent engine and terminal client for Codewhale, with provider choice, tools, approvals and receipts.

What the test found: could not be verified; the log shows where it stopped. 19 minutes.

81spec-kittyCould not verify10 / 100Python1,678

Spec-Driven Development with organizational governance. Specs tell AI agents what to build; Charter governs how they build it. Git-native missions, enforceable...

What the test found: could not be verified; the log shows where it stopped. 87 minutes.

-zeronNot yet tested-Rust3,121-
-AegisNot yet tested-Python1,328-
-weaveNot yet tested-Rust1,316-
-helmorNot yet tested-TypeScript1,309-
-autoprompt-skillNot yet tested-JavaScript1,298-
-splashNot yet tested-C++1,285-
-pi-from-scratchNot yet tested-TypeScript1,273-
-openpetsNot yet tested-TypeScript1,269-
-repository-harnessNot yet tested-Rust1,239-
-opencode-telegram-botNot yet tested-TypeScript1,233-
-open-stepsNot yet tested-Shell1,225-
-gentle-shellNot yet tested-TypeScript1,221-
-alookNot yet tested-TypeScript1,193-
-ref-tools-mcpNot yet tested-TypeScript1,179-
-aws-agent-skillsNot yet tested-Python1,162-
-openharnessNot yet tested-Dart1,158-
-ctxNot yet tested-Rust1,151-
-deja-vuNot yet tested-Go1,150-
-MiniCodeNot yet tested-TypeScript1,137-
-deeptideNot yet tested-Rust1,103-
-dao-codeNot yet tested-TypeScript1,082-
-pydantic-deepagentsNot yet tested-Python1,077-
-Tianshu-harnessNot yet tested-TypeScript1,073-
-CORALNot yet tested-Python1,052-
-parallel-codeNot yet tested-TypeScript1,036-
-PiDeckNot yet tested-TypeScript1,034-
-SWE-AFNot yet tested-Go1,032-
-dscodeNot yet tested-JavaScript1,021-
-claude-code-memory-setupNot yet tested-Python1,014-
-zoetropeNot yet tested-Rust1,014-
-easy-agentNot yet tested-TypeScript1,008-
-codealmanacNot yet tested-TypeScript997-
-fenceNot yet tested-Go988-
-sandboxdNot yet tested-Go963-
-luvusNot yet tested-Rust958-
-DSHANot yet tested-Java952-
-gooey-piNot yet tested-TypeScript941-
-WhaleNot yet tested-Go932-
-codex-slidesNot yet tested-TypeScript926-
-agent-pluginsNot yet tested-Python915-
-opencode-managerNot yet tested-TypeScript893-
-haxNot yet tested-C885-
-pi-webNot yet tested-TypeScript866-
-herdr-reviewrNot yet tested-Rust851-
-agentfilesNot yet tested-TypeScript849-
-memorixNot yet tested-TypeScript836-
-dsh-browserNot yet tested-TypeScript785-
-hippo-memoryNot yet tested-TypeScript773-
-agentacctNot yet tested-Python766-
-okf-agent-memoryNot yet tested-Go756-
-codannaNot yet tested-Rust755-
-birdviewNot yet tested-TypeScript724-
-agenttrailNot yet tested-JavaScript717-
-MegaMemoryNot yet tested-TypeScript716-
-soloNot yet tested-Go696-
-pi-extensionsNot yet tested-TypeScript667-
-herdr-web-uiNot yet tested-TypeScript648-
-casbin-gatewayNot yet tested-Go642-
-superdesign-skillNot yet tested-JavaScript633-
-SWE-ReXNot yet tested-Python617-
-agenttyNot yet tested-C++615-
-svg-diagramNot yet tested-JavaScript614-
-mcp-pointerNot yet tested-TypeScript599-
-claw-orchestratorNot yet tested-TypeScript586-
-Clean-Coder-AINot yet tested-Python585-
-brain.mdNot yet tested-JavaScript564-
-smart-ralphNot yet tested-Shell557-
-pi-dynamic-workflowsNot yet tested-TypeScript555-
-pinloop-cliNot yet tested-TypeScript554-
-tmux-ideNot yet tested-TypeScript550-
-foremergeNot yet tested-Rust538-
-appworldNot yet tested-Python529-
-phiNot yet tested-Go526-
-agentboxNot yet tested-TypeScript523-
-bug-hunterNot yet tested-JavaScript519-
-modelsNot yet tested-Rust514-
-Orkas-VideoStudioNot yet tested-TypeScript498-
-machinistNot yet tested-Go492-
-talkcodyNot yet tested-TypeScript479-
-10xProductivityNot yet tested-Python478-
-pigoNot yet tested-Go476-
-LinghunNot yet tested-TypeScript474-
-gwqNot yet tested-Go473-
-sno-stationNot yet tested-TypeScript469-
-cerseiNot yet tested-Rust459-
-opencode-goal-pluginNot yet tested-TypeScript454-
-ProjectAtlasNot yet tested-Rust440-
-apm-studioNot yet tested-TypeScript436-
-HGMNot yet tested-Python436-
-diriNot yet tested-Rust425-
-ownmemNot yet tested-JavaScript423-
-concord-mcpNot yet tested-TypeScript418-
-wmuxNot yet tested-TypeScript415-
-dryforgeNot yet tested-Python411-
-martin-loopNot yet tested-TypeScript409-
-FrameNot yet tested-JavaScript403-
-harness-remoteNot yet tested-JavaScript403-
-youtube-tutorialsNot yet tested-Python402-
-opencode-primerNot yet tested-JavaScript397-
-cc-sessions-viewerNot yet tested-Rust395-
-deepx-codeNot yet tested-Go390-
-perchoNot yet tested-TypeScript390-
-robotics-agent-skillsNot yet tested-Python369-
-openqodexNot yet tested-TypeScript368-
-pi-appNot yet tested-TypeScript365-
-dscodeNot yet tested-TypeScript364-
-nvxNot yet tested-Python359-
-crew44Not yet tested-Go356-
-developer-kitNot yet tested-Python356-
-zotNot yet tested-Go349-
-genieNot yet tested-TypeScript346-
-vibe-log-cliNot yet tested-TypeScript339-
-openliveNot yet tested-TypeScript337-
-kelosNot yet tested-Go336-
-stashNot yet tested-Python335-
-armoryNot yet tested-Python328-
-agentfs-claudeNot yet tested-TypeScript323-
-devoNot yet tested-Rust323-
-golembotNot yet tested-TypeScript323-
-brain0Not yet tested-Rust321-
-OpenLoreNot yet tested-TypeScript319-
-packmindNot yet tested-TypeScript317-
-JoySafeterNot yet tested-Python313-
-metisNot yet tested-TypeScript313-
-CodeAFNot yet tested-Go312-
-compose-skillNot yet tested-Shell301-
-ai-prd-workflowNot yet tested-Python298-
-nopusNot yet tested-TypeScript298-
-cupcakeNot yet tested-Rust296-
-funzzyNot yet tested-Rust296-
-pi-desktopNot yet tested-TypeScript293-
-sudocodeNot yet tested-TypeScript293-
-codevNot yet tested-TypeScript288-
-spellbookNot yet tested-Python287-
-broccoliNot yet tested-Python286-
-harmonyos-ai-skillNot yet tested-Shell279-
-cloudroom-coreNot yet tested-Rust272-
-runnerNot yet tested-Rust272-
-lazyskillsNot yet tested-Go270-
-umadevNot yet tested-Rust261-
-pi-desktopNot yet tested-TypeScript259-
-acrylNot yet tested-TypeScript256-
-skill-validatorNot yet tested-Go256-
-better-designNot yet tested-TypeScript254-
-meta-llm-charterNot yet tested-TypeScript249-
-llmtrimNot yet tested-Rust244-
-herdr-auto-titleNot yet tested-Go238-
-jev-reviewNot yet tested-TypeScript238-
-mulmoterminalNot yet tested-TypeScript236-
-AlbatrossNot yet tested-Rust235-
-alphacodeNot yet tested-Rust233-
-Vibe_coding_guideNot yet tested-JavaScript231-
-gographNot yet tested-Go228-
-agentsNot yet tested-Python225-
-CodeKanbanNot yet tested-Go225-
-habit-hooksNot yet tested-Python221-
-probityNot yet tested-TypeScript221-
-openusageNot yet tested-Go218-
-procoderNot yet tested-Go211-
-agent-sandboxNot yet tested-Python208-
-agentic-osNot yet tested-Python206-
-code-cliNot yet tested-TypeScript202-
-codex-command-centerNot yet tested-TypeScript201-

runs installed and started with its real dependencies. runs with mocks started after stand-ins replaced external services such as a database or a third-party API. could not verify neither the standard agent nor the stronger one got it running within the time limit; the log shows where it stopped.

How we tested

On this list as of the latest test: 64 projects ran as-is, 13 with mocks, 4 could not be verified, 163 still waiting. Languages tested: C++, Go, Java, JavaScript, Python, Rust, TypeScript. Every attempt used a clean single-use machine, the subject at a pinned version, and a 45-minute limit; the complete procedure is on the methodology page.

Frequently asked questions (FAQs)

How is this list ranked?

By measurement, not opinion: projects Argusic installed and launched on a fresh machine come first, then those that ran with mocks in place of external services, then those it could not verify. Ties go to the Argusic Score, then how popular it is on its own source.

Why are some projects unranked?

163 projects are still waiting for a test or for a finished attempt. They are listed without a rank until Argusic has measured them.

Where is the evidence?

Every row links to the project's Argusic page, where each run has a full log and a terminal recording stored with a sha256 fingerprint. The same pages exist for every one of the tested projects, on this list or not.

More lists in this category

All lists: Best. All tested projects: subjects.