Open source AI agent frameworks, installed and run

AI agent frameworks you can run yourself. The list is ordered by whether each project came up when Argusic installed it fresh, with the recording of that attempt one click away.

Tested between and . Each row shows its own test date; a project can change after that day.

114 of 116 tested projects run. 52 more waiting for a test.

In short: 85 of the 116 tested projects started as-is on a fresh machine: ragflow, deer-flow, anything-llm, agentic-awesome-skills, CowAgent, ai-job-search, herdr, and oh-my-claudecode, and 77 more. 29 more started once a stand-in replaced a service they expect, such as a database: mem0, marketingskills, nanobot, agno, DeepTutor, Claude-Code-Game-Studios, distilly, and learn-harness-engineering, and 21 more. 2 could not be verified: astrid and skillhub; the log shows where each one stopped.

Measured by Argusic on a fresh machine every time. Every number links to its evidence.

#projectverdictArgusic Scorelanguagestarstested on
1ragflowRuns100 / 100Go91,830

RAGFlow is a leading open-source Retrieval-Augmented Generation (RAG) engine that fuses cutting-edge RAG with Agent capabilities to create a superior context...

What the test found: Python 3.13 venv with 697 deps installed, Go toolchain auto-resolved to go1.27, C++ librag_tokenizer_c_api.a built with clang++-20, Go binaries ragflow_server and ragflow-cli compiled successfully, Go unit tests pass (all packages), Python unit tests pass (5145/5145), frontend builds successfully, native static libs... 40 minutes.

2deer-flowRuns100 / 100Python83,506

An open-source long-horizon SuperAgent harness that researches, codes, and creates. With the help of sandboxes, memories, tools, skill, subagents and message...

What the test found: Backend Gateway boots and serves health checks on port 8003 (HTTP 200), 1642 backend tests pass, 160 blocking-IO tests pass, 24811 frontend unit tests pass, all dependencies installed. 62 minutes.

3anything-llmRuns100 / 100JavaScript66,820

Stop renting your intelligence. Own it with AnythingLLM. Everything you need for a powerful local-first agent experience

What the test found: AnythingLLM installs and runs in development mode on port 3001; the test suite passes all 1119 tests across 62 suites covering server and collector modules. 17 minutes.

4agentic-awesome-skillsRuns100 / 100Python47,355

AAS Core is the local, agent-first control plane for complete catalog discovery, agent-owned selection, stack validation, and planning, backed by 2,400+...

What the test found: Dependencies installed, full test suite (130 tests) passes, validation passes for 2459 skills, build chain generates catalog, and AAS CLI responds to commands. The web-app portion requires Node >=22 and cannot launch in this container. 8 minutes.

5CowAgentRuns100 / 100Python47,274

Open-source personal AI assistant & Agent Harness. Plans tasks, runs tools and skills, self-evolves with memory and knowledge. Multi-agent, multi-model...

What the test found: CowAgent installed in a Python 3.12 venv with all deps, CLI responds at version 2.1.9, web console serves successfully on port 9899, and the test suite passes 1948 tests. 5 minutes.

6ai-job-searchRuns100 / 100Python45,265

The job search that runs on your machine. AI job application framework built on Claude Code: evaluate postings, tailor CVs, write cover letters, prep...

What the test found: Python test suite (321 unittest tests) and all 6 Bun portal CLI test suites (231 tests) pass. LinkedIn search CLI makes live queries and returns results. 4 Python analysis tools (security_guards, framework_version, lint_skills, verify_pdf) run successfully. 8 minutes.

7herdrRuns100 / 100Rust42,896

the runtime your coding agents live on

What the test found: herdr 0.8.2 builds, runs its server, passes all Rust unit tests, Python maintenance tests, Bun integration tests, and plugin marketplace tests with no failures. 19 minutes.

8oh-my-claudecodeRuns100 / 100TypeScript39,679

Teams-first Multi-agent orchestration for Claude Code

What the test found: Oh-My-ClaudeCode v5.2.0 CLI tool builds and responds to version/help/info commands; 1,094+ tests pass with 3 env-constrained failures (Node 18 engine gap + shallow git). 31 minutes.

9PageIndexRuns100 / 100Python38,976

📑 PageIndex: Document Index for Vectorless, Reasoning-based RAG

What the test found: PageIndex SDK v0.2.10 installed in a venv, all 796 unit tests pass, the CLI parses correctly, and the package imports without error. 3 minutes.

10500-AI-Agents-ProjectsRuns100 / 100Python38,378

The 500 AI Agents Projects is a curated collection of AI agent use cases across various industries. It showcases practical applications and provides links to...

What the test found: Web app builds with vite and serves HTTP 200 on port 5173. All 3 CrewAI MCP course lessons execute end-to-end with OpenRouter. SQL query agent creates demo database and answers natural language questions. Unit test generator agent produces passing 21/21 pytest suites. News summarizer falls back to mock data and... 12 minutes.

11ai-website-cloner-templateRuns100 / 100TypeScript36,241

Clone any website with one command using AI coding agents

What the test found: The next.js app installs, builds, typechecks, lints, and runs with a dev server responding HTTP 200 on localhost:3000. 6 minutes.

12book-to-skillRuns100 / 100Python34,167

Turn any technical book PDF into a Claude Code skill, ready to study, reference, and use while you work.

What the test found: The book-to-skill package is installed as an editable pip package, passes ruff lint, all 627 tests pass, skill validation passes, and the CLI extracts text from documents producing structured JSON metadata and clean text output. 3 minutes.

13invisible_dotsRuns100 / 100TypeScript31,843

Open-source, self-hosted alternative to OpenAI Dots, Grok Bot. Built to be undetectable by anti-bot systems.

What the test found: npm ci installs 428 packages without error; npm run build produces the CLI binary (apps/cli/dist/invisible-dots.mjs) and the server binary (apps/web/.next/standalone/apps/web/server.js). The test suite passes 2103 tests across 150 files (0 failures, 3 skipped). The server starts, creates its PGlite database, applies... 49 minutes.

14cogneeRuns100 / 100Python31,642

Cognee is the open-source AI memory platform for agents. Give your AI agents persistent long-term memory with small models for free

What the test found: cognee 1.6.1 installs, builds, starts, and runs: the local LLM-free remember/recall pipeline produces search results from ingested text, the CLI demo loads a bundled knowledge graph, the FastAPI server answers HTTP 200 on /health, and 388 unit tests pass across 5 test suites with zero failures. 33 minutes.

15nanoclawRuns100 / 100TypeScript30,894

A lightweight alternative to OpenClaw that runs in containers for security. Connects to WhatsApp, Telegram, Slack, Discord, Gmail and other messaging apps, has...

What the test found: The NanoClaw project installs, compiles with tsc, passes typecheck with tsc --noEmit, and all 2945 host tests pass via vitest. 10 minutes.

169routerRuns100 / 100JavaScript30,455

Unlimited FREE AI coding. Connect Claude Code, Codex, Cursor, Cline, Copilot, Antigravity to FREE Claude/GPT/Gemini via 40+ providers. Auto-fallback, RTK -40%...

What the test found: Root npm install completed, npm run build succeeded, the standalone Next.js app serves the login page (200), models API (200 with JSON model list), and version endpoint (200), and the test suite runs 2562 passing tests. 42 minutes.

17page-agentRuns100 / 100TypeScript29,359

JavaScript in-page GUI agent. Control web interfaces with natural language.

What the test found: Install, sequential build, typecheck, lint, and all 69 unit tests pass. The project is fully functional with Node v22. via nvm. 27 minutes.

18haystackRuns100 / 100Python26,695

Open-source AI orchestration framework for building context-engineered, production-ready LLM applications. Design modular pipelines and agent workflows with...

What the test found: Haystack 3.3.0-rc0 installed and buildable, 6569 unit tests pass, Pipeline/InMemoryDocumentStore/retrievers/ChatMessage/serialization all work correctly. 9 minutes.

19awesome-claude-code-subagentsRuns100 / 100Shell25,578

A collection of 100+ specialized Claude Code subagents covering a wide range of development use cases

What the test found: The repository clone is complete: 161 agent definitions with valid YAML frontmatter across 10 categories, a syntax-valid interactive installer (needs TTY), a fully operational subagent-catalog CLI skill with 12-hour TTL cache and fetch/search/list/invalidate commands, and all documentation (README, CONTRIBUTING... 5 minutes.

20doltRuns100 / 100Go24,600

Dolt, Git for Data

What the test found: Dolt v2.3.5 binary runs as a CLI tool and MySQL-compatible SQL server; version control operations (init, add, commit, branch, merge, log) work via both CLI and SQL procedures; bats integration tests pass on core functionality. 28 minutes.

21khazix-skillsRuns100 / 100Python21,250

数字生命卡兹克开源的 AI Skills 合集 | Agent Skills: leader(帮你定义目标), neat-freak 洁癖, hv-analysis, khazix-writer & more, Claude Code, Codex & 40+ agents

What the test found: aihot skill installed and its live API returns real AI news (HTTP 200 with real items); neat-freak all 11 structural regressions and 21 trigger evals pass; hv-analysis md_to_pdf.py generates valid PDFs via WeasyPrint; storage-analyzer build_report.py and server.py produce and serve disk-usage reports; scan.py... 5 minutes.

22openfangRuns100 / 100Rust18,215

Open-source Agent Operating System

What the test found: OpenFang 0.6.9 builds, passes all 2699 tests, and the API server boots on port 4200 serving agents, health endpoint, and dashboard. 23 minutes.

23img2threejsRuns100 / 100Python17,705

Rebuild the object in a reference image as a code-only, procedural, quality-gated, animation-ready Three.js model. Token-efficient image-to-3D.

What the test found: 1416 Python stdlib-only test suite tests pass, all forge CLI tools report help text, the --experimental-strip-types Node flag was patched to npx tsx for Node 18 compatibility, and the project runs without external dependencies. 29 minutes.

24vibe-coding-cnRuns100 / 100Python17,268

Vibe Coding 从入门到精通教程|AI 结对编程工作流|Prompt、Skill、Workflow、上下文管理、codex实战指南

What the test found: All repository quality gates pass except check-research-raw which requires gh CLI auth. Node v22 and Python venv with all dependencies are installed from /tmp. Submodules initialized. Markdown lint and all Python-based checks (links, details, structure, directory coverage, metadata, AI citations, external resources... 6 minutes.

25agent-frameworkRuns100 / 100Python14,007

A framework for building, orchestrating and deploying AI agents and multi-agent workflows with support for Python and .NET.

What the test found: Python Agent Framework installs and its entire test suite passes across core, tools, providers (OpenAI, Anthropic, Gemini, Ollama, Mistral, Bedrock, Foundry), orchestrations, declarative, A2A, hosting, Redis, ChatKit, GitHub Copilot, and AG-UI packages. 28 minutes.

26claude-skillsRuns100 / 100Python11,776

67 Specialized Skills for Full-Stack Developers. Transform Claude Code into your expert pair programmer.

What the test found: python3 scripts/validate-skills.py passes with 0 errors, make test runs 12/12 passing, and the site builds 99 static pages successfully via Node 20. 6 minutes.

27cocoindexRuns100 / 100Rust11,651

Incremental engine for long horizon agents 🌟 Star if you like it!

What the test found: Python package imports, Rust extension builds, all core tests pass (611 Rust + 857 Python), mypy type-checks 297 files clean, CLI runs all subcommands, and the files_transform example pipeline successfully converts markdown files to HTML with correct incremental change detection. 13 minutes.

28loop-engineeringRuns100 / 100TypeScript11,442

Practical patterns, starters & CLI tools for loop engineering with AI coding agents. Design systems that prompt and orchestrate agents (inspired by Addy Osmani...

What the test found: All 15 TypeScript tools in tools/ build and pass their test suites (298 passed), unified loop CLI (tools/loop) runs doctor/status/audit with real audit output scoring 100/100 L3, before-after-demo.sh exercises the full empty→L2 scaffold pipeline end to end. 28 minutes.

29voltagentRuns100 / 100TypeScript10,753

AI Agent Engineering Platform built on an Open Source TypeScript AI Agent Framework

What the test found: Monorepo builds (29/29 packages), lints (0 errors, 5 warnings), and all 25 packages with runnable unit tests pass (~2776 tests). The only test failure is @voltagent/e2e which requires an unreachable PostgreSQL database. Two test files with a pre-existing ESM dependency conflict are excluded. 26 minutes.

30GraftRuns100 / 100TypeScript9,771

Turbocharge Claude Code, Cursor, Codex, Gemini & every coding agent: faster, cheaper, with contextual understanding specific to your codebase.

What the test found: npm install completes, TypeScript compiles, the CLI builds/checks/queries a structural code graph on a real repo without any keys, and all 1220 tests run with 0 failures. 10 minutes.

31mobilerunRuns100 / 100Python9,588

Automate your mobile devices with natural language commands - an LLM agnostic mobile Agent 🤖

What the test found: The mobilerun framework installs, builds, CLI reports v0.6.19, and all 979 pytest tests pass (967 core + 12 langfuse after installing optional dep) in 23s. 5 minutes.

32openbrowserRuns100 / 100TypeScript9,559

Let AI agents browse the web. An autonomous toolkit for browser-based AI agents.

What the test found: Bun 1.4.2 installed, all 42 npm dependencies resolved, TypeScript builds pass for all 3 packages, 364/364 core tests pass, CLI prints help/version correctly, and a headless Playwright browser opens and navigates to https://example.com/. 8 minutes.

33ccpmRuns100 / 100Shell8,408

Project management skill system for Agents that uses GitHub Issues and Git worktrees for parallel agent execution.

What the test found: CCPM skill installed at /home/runner/.codex/skills/ccpm, .claude/ directory initialized with all subdirectories, gh 2.64.0 available via ~/.local/bin/gh, and all 14 local workflow scripts execute with correct output against real data. GitHub sync commands blocked by missing interactive auth. 4 minutes.

34Vision-AgentsRuns100 / 100Python8,154

Open Vision Agents by Stream. Build voice and vision agents quickly with any model or video provider. Uses Stream's edge network for ultra-low latency.

What the test found: The vision-agents core framework plus anthropic, gemini, openai, deepgram, elevenlabs, getstream, and smart_turn plugins are installed; 733 unit tests pass, the CLI binary responds, and all selected plugin imports resolve correctly. 42 minutes.

35OctopRuns100 / 100Python7,931

A smarter, self-hosted AI assistant, multi-user, multi-agent.

What the test found: Octop v1.0.2b5 dependencies installed, CLI operational (bootstrap, version), server starts on user-specified port and responds to /api/health with HTTP 200, unit test suite passes across all subgroups. 45 minutes.

36steel-browserRuns100 / 100TypeScript7,754

🔥 Open Source Browser API for AI Agents & Apps. Steel Browser is a batteries-included browser sandbox that lets you automate the web without worrying about...

What the test found: With a user-local Node 22 and Chrome for Testing plus DISABLE_CHROME_SANDBOX=true, the dev-mode API runs on port 3000, /v1/health returns 200 with live Chrome, and sessions/scrape/screenshot/pdf plus external CDP client connections all worked with real results. 11 minutes.

37big-AGIRuns100 / 100TypeScript7,138

AI suite powered by state-of-the-art models and providing advanced AI/AGI functions. Includes AI personas, AGI functions, world-class Beam multi-model chats...

What the test found: All npm commands (install, test, build, tscheck, lint) pass cleanly on Node.js v22.14.0. The test suite runs 25 subtests: 4 pass (including a live OpenRouter model listing returning 388 models), 21 skip due to absent API keys, 0 fail. 9 minutes.

38agent-governance-toolkitRuns100 / 100Python6,408

AI Agent Governance Toolkit, Policy enforcement, zero-trust identity, execution sandboxing, and reliability engineering for autonomous AI agents. Covers 10/10...

What the test found: All core Python packages install from source: agent-governance-toolkit-core, agent-governance-toolkit-cli, agent-control-specification (Rust wheel), agent-governance-toolkit-compliance. The agent-os test suite passes 2993/2993 selected tests, the ACS SDK test suite passes 306/306 selected tests, and the quickstart... 18 minutes.

39agentscope-javaRuns100 / 100Java5,889

Build distributed, production-grade, long-running agents.

What the test found: AgentScope Java 2.0.4-SNAPSHOT builds and all 3082 test cases across core, harness, and extension modules pass with zero failures. 79 minutes.

40opengapRuns100 / 100TypeScript2,972

A framework-agnostic, git-native standard for defining AI agents

What the test found: OpenGAP v0.5.0 builds, all 51 unit tests pass, CLI works for init/validate/info/export commands, and validates example agents successfully. 5 minutes.

41harmonistRuns100 / 100Python2,252

Portable AI agent orchestration with mechanical protocol enforcement. 186 agents, zero runtime dependencies.

What the test found: Harmonist pack integrates and enforces the agent protocol pipeline on Python 3.12, all 373 test assertions passed across 22 shell/Python test suites covering integration, hooks, upgrade, memory, telemetry, supply-chain, regression, and cross-platform paths; the sole gap is lsof(1) missing from the container, blocking... 6 minutes.

42agentic_securityRuns100 / 100Python2,019

Agentic LLM Vulnerability Scanner / AI red teaming kit 🧪

What the test found: Agentic Security LLM vulnerability scanner is fully installed, all 390 unit and integration tests pass, the FastAPI server starts and responds on /health (HTTP 200) and /v1/self-probe (returns mock OpenAI-style completion). 19 minutes.

43waku-agentRuns100 / 100Python1,962

Waku Waku! Waku Agent is a local-first AI agent harness you actually own, including loop, memory, eval, all in code built to stay legible as it grows.

What the test found: Waku 0.1.8 installed and verified: all 1714 offline deterministic tests pass, the dashboard serves HTTP 200 on localhost:7777, the CLI responds to waku --help, the loop runs end-to-end with a real LLM provider (OpenRouter free tier), and lint passes cleanly. 7 minutes.

44jidoRuns100 / 100Elixir1,877

🤖 Autonomous agent framework for Elixir. Built for distributed, autonomous behavior and dynamic workflows.

What the test found: Elixir 1.18.5 + Erlang/OTP 25 installed under ~/local_erlang and ~/elixir; mix deps.get, mix compile, and mix test all succeed with 2470 passing tests and 0 failures. 16 minutes.

45trpc-agent-goRuns100 / 100Go1,850

A Go framework for building production agent systems with graph workflows, tools, memory, A2A, AG-UI, MCP, evaluation, and observability.

What the test found: The tRPC-Agent-Go root module (trpc.group/trpc-go/trpc-agent-go) builds and all 235 test suites pass, the test/ module passes with Go 1.24.4 auto-downloaded, and the examples/ module builds successfully, all from a fresh container with Go installed from the upstream tarball. 14 minutes.

46bubRuns100 / 100Python1,679

Bub it. Build it. A tiny agent runtime, composable with plugins.

What the test found: Bub installed from source, builds with uv, all 579 tests pass with zero failures, the CLI accepts all commands, and the framework makes live agent calls through the default OpenRouter free model returning correct text responses. 3 minutes.

47AgentlyRuns100 / 100Python1,657

[GenAI Application Development Framework] 🚀 Build GenAI application quick and easy 💬 Easy to interact with GenAI agent in code using structure data and...

What the test found: The agently 4.1.4.8 package installs, imports, and creates agents via its fluent API. 3,776 of 3,826 tests pass. 9 failures are pre-existing characterization snapshot mismatches (expected on different package versions), and 1 requires a Docker daemon. 45 minutes.

48MasterAgentRuns100 / 100C++1,646

Build AI agents that run 100% on-device. Sub-100ms latency on Qualcomm NPU. Zero cloud dependency.

What the test found: OAK v0.3.0 OSS build compiles and passes 6/6 tests + 5/5 eval suites on gcc 13.3, no errors. 4 minutes.

49better-agentsRuns100 / 100TypeScript1,563

Standards for building agents, better

What the test found: The better-agents CLI builds, typechecks, and passes 15 unit tests; the init command successfully generates a complete agent project scaffold including src/, prompts/, tests/scenarios/, AGENTS.md, .env, and .mcp.json. 5 minutes.

50chidoriRuns100 / 100Rust1,367

The agent framework where every run is durable, replayable, and resumable by default.

What the test found: Chidori v3.8.1 binary installed at ~/.chidori/bin/chidori; runs TypeScript agents directly via its embedded JS engine, serves agents over HTTP on localhost, replays runs deterministically from journaled call logs with zero LLM calls, scaffolds agent projects from templates, and both the TypeScript and Python SDKs... 8 minutes.

51commonlyRuns100 / 100TypeScript1,361

Open-source room for humans + cross-vendor AI agents. Every agent gets its own name, memory, skills, and workstation. Any runtime, your infra, no per-agent...

What the test found: Installed deps and built both workspaces, fixed the two real test-bug failures (expired fixture date, missing WebCrypto global), and verified the backend server live on localhost:5000: /api/health 200 healthy with real mongod connected, register 201, login 200 with JWT, while frontend's 962 Jest tests pass and full... 79 minutes.

52EnterpriseAgentFrameworkRuns100 / 100Java845

ReachAI企业级智能体开发平台:快速、安全完成已有业务系统智能化改造,让 AI 在 OA、ERP、CRM 等原系统中查数据、填表单、办业务。ReachAI: Quickly and securely bring AI to existing enterprise systems, enabling AI to...

What the test found: Backend builds cleanly and all 2246 unit/integration tests pass across 9 Maven modules. Frontend builds with npm run build and 453 of 454 vitest tests pass across 10 test suites (1 pre-existing timezone-dependent dashboard test fails in UTC). 34 minutes.

53agentsRuns98.7 / 100Python40,296

Multi-harness agentic plugin marketplace for Claude Code, Codex, Cursor, OpenCode, GitHub Copilot, Google Antigravity, and Pi

What the test found: uv sync installs all deps for both projects; pytest suite passes 619/621; all 6 harness generators emit artifacts; codex doctor reports 20/20 checks ok; structural validation passes across 6 harnesses with strict mode. 11 minutes.

54plandexRuns98.7 / 100Go15,702

Open source AI coding agent. Designed for large projects and real world tasks.

What the test found: Plandex server starts on port 8099 with PostgreSQL 16.4 and LiteLLM proxy on port 4000; server tests pass; CLI builds and authenticates against local server; full CRUD API for projects, plans, branches, settings is operational. 18 minutes.

55ponytailRuns98 / 100JavaScript158,133

Makes your AI agent think like the laziest senior dev in the room. The best code is the code you never wrote.

What the test found: All 110 tests pass: the root project's behavioral, plugin, correctness, and hardware-awareness tests (84), the Pi extension's mode/injection tests (23), and the MCP server's instruction-building tests (3) all succeed. 3 minutes.

56DeepSeek-ReasonixRuns97.3 / 100Go35,748

A reliable coding agent for complex software engineering tasks.

What the test found: Reasonix builds as a single Go binary (bin/reasonix), answers --help with full usage including CLI/TUI, web, serve, ACP, and bot subcommands, passes all 147+ test packages with 0 failures, and produces valid JSON diagnostics via doctor --json. 19 minutes.

57omnigentRuns97.3 / 100Python10,676

Omnigent is an open-source AI agent framework and meta-harness: orchestrate Claude Code, Codex, Cursor, Pi, and custom agents, swap harnesses without...

What the test found: The omnigent CLI reports version 0.13.0.dev0, the server listens on TCP and returns HTTP 200 on root/health/OpenAPI endpoints, and 929+ unit tests across 10 module directories pass without failure. 28 minutes.

58YuxiRuns97.3 / 100Python7,317

可私有部署的多租户知识智能体平台:统一 RAG、知识图谱、多智能体、MCP/Skills、沙盒与权限管理。Yuxi = Cloud Agents + Knowledge RAG, Self-hosted knowledge agent platform for RAG, knowledge graphs and...

What the test found: Backend Python dependencies installed and locked; all 1640 backend unit tests pass; all 164 web unit tests pass; Vite production build completes successfully; engineering trust contracts and lint checks pass. Two minor test permission-assertion bugs fixed (umask compatibility) and one Node.js version gap resolved... 25 minutes.

59harnessrouterRuns96 / 100Python2,926

HarnessRouter Community Edition: the self-hosted, Apache-2.0 edition of the unified interface for agent harnesses. Run Codex, Claude Code, Hermes, PI, DSH, and...

What the test found: Gateway and runner servers start and respond to health checks; all three test suites (gateway: 622/17, runner: 490/2, protocol conformance: 85/2) pass. 6 minutes.

60agent-scriptsRuns95 / 100Shell7,283

Scripts for agents, shared between my repositories.

What the test found: All 12 test suites pass; npm install completes with 0 vulnerabilities; Python, Node.js, and Ruby scripts compile and run correctly; browser-tools CLI loads and reports help; docs-list enumerates all docs with metadata checks; no Chrome binary available for browser launch tests. 26 minutes.

61txtaiRuns94.7 / 100Python12,998

💡 All-in-one AI framework for semantic search, LLM orchestration and language model workflows

What the test found: txtai 9.14.0 installed in venv, all core embeddings/graph/workflow/agent/cloud/vector/scoring tests pass, API responds to search/index/count endpoints with real model inference, ~470 automated tests pass across 30 test suites. 83 minutes.

62Vibe-SkillsRuns93.5 / 100Python3,606

Intelligent Skill routing and workflow orchestration for AI agents, +21.12 pp reward, −29.6% tokens on SkillsBench with DeepSeekV4Flash-VE.

What the test found: VibeSkills v4.1.0 installs successfully, check verifies all 292 receipt-owned files, the vgo_cli Python launcher works end-to-end, and 129+ core unit tests pass with 3 isolated test failures due to missing pwsh in this container. 37 minutes.

63gentle-aiRuns93.3 / 100Go7,599

Gentle-AI configures the AI coding agents you already use: Claude Code, Cursor, OpenCode, Codex, Pi, and more. Choose persistent memory, Organic-Driven...

What the test found: gentle-ai 2.0.0 builds from source, prints version (2.0.0-20260906031619-2c25e878eea4) and help text, all 53 Go test packages pass. 17 minutes.

64crawl4aiRuns91.4 / 100Python84,979

Open-source web crawler and scraper for LLMs and AI agents: any website into clean, LLM-ready Markdown. Run it yourself, or use Crawl4AI Cloud with one key.

What the test found: Crawl4AI 0.9.4 installed in venv, Playwright chromium browser downloaded, DB initialized, AsyncWebCrawler can crawl real URLs with markdown output, cache operations work, and 360 core tests pass with 11 skipped. 14 minutes.

65nanobrowserRuns90.7 / 100TypeScript14,006

Open-Source Chrome extension for AI-powered web automation. Run multi-agent workflows using your own LLM API key. Alternative to OpenAI Operator.

What the test found: Dependencies installed, project builds to dist/ (background.iife.js 1,848.81 kB, manifest.json), TypeScript type-check passes across all 15 packages, unit tests pass (14/14). 15 minutes.

66autoclipRuns90.7 / 100Python9,251

AutoClip|一个链接,一键出片。开源 AI 视频剪辑桌面工具,将播客、访谈、课程等长视频自动剪成短视频,生成字幕、封面和发布文案,适配抖音、小红书、TikTok、Reels 与 YouTube Shorts。Open-source AI video clipping & content repurposing.

What the test found: The AutoClip FastAPI backend serves all API health endpoints and the test suite passes with 103/103 tests passing. The frontend builds to production with Vite. 4 minutes.

67agency-agents-zhRuns90 / 100Shell21,121

🎭 277 个即插即用的 AI 专家角色, 支持 Claude Code/Cursor/Copilot 等 20 种工具,覆盖工程/设计/营销/金融等 20 个部门。含 64 个中国市场原创智能体(小红书/抖音/微信/飞书/钉钉/Qt 上位机/机械设计)。搭配编排器...

What the test found: All 277 AI agent definition files are valid, the npm package builds, the conversion pipeline produces output for all 18 supported tools, and the install script copies agents to 6 detected tool directories without errors. 5 minutes.

68hiveRuns90 / 100Python11,087

Multi-Agent Harness for Production AI

What the test found: All Python dependencies installed and both framework and tools packages build successfully. Core test suite (2500 tests) passes cleanly. Tools test suite (6840 of 7261 tests) passes with one pre-existing endpoint mismatch failure. The hive CLI launches and displays help text. 20 minutes.

69ralph-claude-codeRuns90 / 100Shell9,669

Autonomous AI development loop for Claude Code with intelligent exit detection

What the test found: Ralph is installed globally with 7 CLI commands, all 1319 tests pass, project setup works, and dry-run mode simulates loops without API calls. 13 minutes.

70ECCRuns80 / 100JavaScript275,235

The agent harness performance optimization system. Skills, instincts, memory, security, and research-first development for Claude Code, Codex, Opencode, Cursor...

What the test found: ECC v2.2.2 installed and verified: npm dependencies installed (212 packages), all 10 CI validation scripts pass, all 31 CI test files pass, 173 of 173 Python tests pass, ECC welcome CLI displays v2.2.2 and help system works. 24 minutes.

71go-microRuns80 / 100Go23,090

A framework for building agents and services

What the test found: Go Micro builds and tests run with the CLI binary working; README updated to satisfy first-agent wayfinding doc integrity tests that enforce install-troubleshooting, micro agent demo/quickcheck/examples/zero-to-hero, examples/INDEX.md, first-agent, support, no-secret, your-first-agent, debugging-agents, inspect, and... 31 minutes.

72hermes-webuiRuns80 / 100Python18,808

Hermes WebUI: The best way to use Hermes Agent from the web or from your phone!

What the test found: Hermes WebUI server launches on port 8789, responds with HTTP 200 on /health and /, and 11,844+ automated tests pass with 269 skipped (31 due to missing hermes-agent); 7 pre-existing failures involve Node.js ES module import from smd.min.js and test ordering in sharded runs, none affecting production runtime. 33 minutes.

73ag-uiRuns80 / 100TypeScript16,387

AG-UI: the Agent-User Interaction Protocol. Bring Agents into Frontend Applications.

What the test found: The AG-UI monorepo installs, builds (35/35 projects), and tests (5576 passed, 1 dotnet-only failure due to dotnet not being available in the container) under Node.js v22.23.3. The dojo demo viewer serves HTTP 200. The core SDK (@ag-ui/core, @ag-ui/client, @ag-ui/encoder, etc.) exports a working protocol implementation. 16 minutes.

74graphifyRuns66.7 / 100Python124,825

Turn any codebase, with its docs, SQL schemas, configs, and PDFs, into a queryable knowledge graph. A /graphify skill for Claude Code, Cursor, Codex, and...

What the test found: graphifyy CLI installs, builds knowledge graphs from source code locally with tree-sitter AST, and passes 5234/5288 tests (11 pre-existing skillgen failures from shallow clone, no fix available). 11 minutes.

75browser-useRuns66.7 / 100Python117,466

Agents that use the browser.

What the test found: browser-use v0.13.10 installed in a Python 3.12 venv with uv, Chromium headless installed via playwright, and all 563 CI tests pass (529 passed, 34 skipped) after fixing 2 test identity comparison bugs in test_beta_agent.py that only manifest under xdist worker parallelism. 75 minutes.

76OpenCLIRuns66.7 / 100JavaScript29,931

Make Any Website into CLI & Use your logged-in browser by AI agent.

What the test found: OpenCLI 1.8.8 builds and runs on Node 22.14.0; 632 test files pass (7325 tests, 3 skipped) across unit, extension, adapter, smoke projects; CLI binary answers version 1.8.8 and lists 1366 registered commands. 27 minutes.

77crewAIRuns60 / 100Python59,454

Framework for orchestrating role-playing, autonomous AI agents. By fostering collaborative intelligence, CrewAI empowers agents to work together seamlessly...

What the test found: CrewAI v1.15.22 is installed with all optional extras, the CLI reports version 1.15.22, and 5424 of 5427 crewai tests pass (3 pre-existing unrelated failures), 87 crewai-core tests pass, 489 crewai-tools tests pass. 59 minutes.

78adk-pythonRuns60 / 100Python21,743

An open-source, code-first Python toolkit for building, evaluating, and deploying sophisticated AI agents with flexibility and control.

What the test found: uv virtual environment installed with all extras, google-adk 2.9.2 builds and imports, full unit test suite passes at 16167/16167 (0 failures, 83 skipped, 26 expected failures, 2 unexpected passes), CLICK-based CLI entry point is functional. 14 minutes.

79text-to-cadRuns60 / 100Python18,363

Give your agent CAD superpowers.

What the test found: cadgen 0.6.6 installed from source: OCP kernel loads, CLI responds, Python test suites for cadgen CLI/store/memo and all eight tested skill packages pass, cross-language JS runtime bundle and Playwright snapshot browser present. 32 minutes.

80edictRuns60 / 100Python16,976

🏛️ 三省六部制 · OpenClaw Multi-Agent Orchestration System, 9 specialized AI agents with real-time dashboard, model config, and full audit trails

What the test found: All 68 tests pass on Python 3.12.3; the dashboard HTTP server serves the React frontend, health check, and all API endpoints correctly. 6 minutes.

81MiMo-CodeRuns60 / 100TypeScript13,613

MiMo Code: Where Models and Agents Co-Evolve

What the test found: MiMoCode CLI runs, typechecks, builds SDK, and passes ~2000+ tests across all core packages (agent, config, auth, git, tool, session, cli, server, workflow, effect, actor, history) with only ~15 environmentally-specific failures (umask, Bun version, test isolation). 37 minutes.

82fast-agentRuns60 / 100Python3,928

Code, Build and Evaluate agents - excellent Model and Skills/MCP/ACP/A2A Support

What the test found: All 8730 unit tests pass, CLI boots and responds, playback LLM agent runs end-to-end, format/lint/typecheck all pass. 30 minutes.

83Agentlas-OSRuns60 / 100Python1,577

Agent OS: keep specialist agents in a hub, spin up a temporary orchestrator per task. Local-first, works with any model.

What the test found: Agentlas OS runtime v1.2.57 installed to ~/.agentlas/runtime/current with hephaestus doctor ok; ontology runtime operational with SQLite storage, Model2Vec int8 embeddings, FTS adapters, and auto-ingest working from ~/.agentlas/inbox directories. 6 minutes.

84harness-sdkRuns55 / 100Python8,736

Build an agent harness and control it end-to-end. Open-source SDK for production AI agents in Python & TypeScript - any model, any cloud.

What the test found: Python SDK strands-agents v0.1.dev1 and strands-harness v0.0.1.dev1 installed from source, both packages import correctly, and over 2,000 unit tests pass across both packages. 68 minutes.

85UpsonicRuns33.3 / 100Python7,956

Build autonomous AI agents in Python.

What the test found: Upsonic v0.77.3 installed from source in editable mode; all 1490+ tests pass; CLI is functional. 23 minutes.

86mem0Runs with mocks92 / 100Python66,814

The Memory Layer for AI Agents - Drop-in memory infrastructure for AI agents and apps. Context that persists. Built for production.

What the test found: Mem0 SDK 2.2.0 installs cleanly in a Python 3.12 venv with all optional dependencies; 1845/1922 pytest test cases pass (77 skipped due to missing external services), 0 failures. 49 minutes.

87marketingskillsRuns with mocks92 / 100JavaScript53,679

Marketing skills for Claude Code and AI agents. CRO, copywriting, SEO, analytics, and growth engineering.

What the test found: All 150 CLI contract tests pass, 50 skills pass both shell and official skills-ref validation, and version consistency checks pass. 2 minutes.

88nanobotRuns with mocks92 / 100Python48,864

Ultra-lightweight, open-source, self-hosted personal AI agent framework in Python with WebUI, tools, memory, MCP, multi-agent workflows, automation, and chat...

What the test found: nanobot v0.3.0 installed from source in a venv, all 6014 automated tests pass, the gateway launches successfully, and ruff linting passes cleanly. 28 minutes.

89agnoRuns with mocks92 / 100Python42,612

Build, run, and manage agent platforms.

What the test found: Agno v3.0.11 installs and builds successfully in the dev venv. The framework core (Agent, Team, Workflow, tools) and 1595 unit tests pass. 109 unrelated test files show collection errors due to missing optional SDK packages. The format and ruff validation pipelines pass cleanly. 43 minutes.

90DeepTutorRuns with mocks92 / 100Python40,928

DeepTutor: Lifelong Personalized Tutoring. https://deeptutor.info/.

What the test found: DeepTutor 1.6.11 installs successfully in a venv, all 6409 unit tests pass, the CLI and API server (health endpoints 200) launch without LLM credentials, and the app is ready for LLM configuration. 56 minutes.

91Claude-Code-Game-StudiosRuns with mocks92 / 100Shell25,892

Turn Claude Code into a full game dev studio, 49 AI agents, 72 workflow skills, and a complete coordination system mirroring real studio hierarchy.

What the test found: All 14 hooks pass syntax and runtime checks, all 8 scripts pass syntax and runtime checks, all 74 skills have valid frontmatter and sections, all 49 agents exist in catalog, all 148 spec files exist at declared paths, the config resolution pipeline works correctly, and Claude Code CLI v2.1.282 is available but... 7 minutes.

92distillyRuns with mocks92 / 100Python25,395

Distilly, Distill how they think into reusable Skills for any Agent or Bot. Formerly Colleague Skill(原同事 Skill).

What the test found: Python dependencies installed in venv, all 73 pytest tests pass, the Node.js distilly CLI prints version 1.0.0 and installs the Skill payload correctly. 4 minutes.

93learn-harness-engineeringRuns with mocks92 / 100TypeScript19,512

Harness engineering beginner tutorial, from 0 to 1

What the test found: Root npm install succeeded, VitePress docs dev server serves all pages with 200, all 15 project Electron apps build successfully, and project-01 starter Electron app launches and stays running in the container. 15 minutes.

94reaRuns with mocks92 / 100TypeScript19,434

Reverse engineer anything with agents, from app behavior down to native binaries.

What the test found: REA compiles and runs on Node.js 22.19.0 with npm 11.16.0. The full TypeScript build succeeds, all 353 Vitest test files (1756 individual tests) pass covering domain logic, adapters, composition, MCP boundary, acceptance, conformance, and evaluation layers, and the CLI lists all 60+ analysis commands without error. 8 minutes.

95E2BRuns with mocks92 / 100Python14,238

Open-source, secure environment with real-world tools for enterprise-grade agents.

What the test found: Dependencies installed, JS SDK builds (47 exports), Python SDK imports, CLI executable produces help and version output, and 323 JS tests plus ~450 Python tests pass using mocked HTTP endpoints. 33 minutes.

96Agent-SRuns with mocks92 / 100Python12,558

Agent S: an open agentic framework that uses computers like a human

What the test found: Package installs, all core library modules import, 5/5 provider tests pass, CLI argument parser works (requires mouseinfo/tkinter workaround for native launch on Linux without python3-tk). 5 minutes.

97MemOSRuns with mocks92 / 100TypeScript11,760

Self-evolving memory OS for LLM & AI Agents: ultra-persistent memory, hybrid-retrieval, and cross-task skill reuse, with 35.24% token savings and DeepSeek...

What the test found: MemoryOS 2.0.33 installs, builds, and passes 875/875 pytest cases across all modules (LLM, embedder, vector DB, graph DB, chunker, parser, reranker, memory, scheduler, API lifecycle, CLI, context, multi-cube, dream) with no failures, 9 skipped only where external backends are required. 5 minutes.

98artemisRuns with mocks92 / 100Python11,141

ARTEMIS turns natural-language instructions into reliable Android automation. It automates end-to-end workflows, captures logs, and integrates seamlessly with...

What the test found: uv sync --dev installs cleanly; artemis doctor runs; 2185/2185 pytest suite passes with a mock API key; 3 pre-existing code and fixture bugs fixed. 26 minutes.

99aichatRuns with mocks92 / 100Rust10,491

All-in-one LLM CLI tool featuring Shell Assistant, Chat-REPL, RAG, AI Tools & Agents, with access to OpenAI, Claude, Gemini, Ollama, Groq, and more.

What the test found: aichat 0.30.0 builds from source, passes all 21 unit tests, and successfully proxies requests to an OpenAI-compatible API backend including chat completions (200), model listing, and web playground (200). 7 minutes.

100PraisonAIRuns with mocks92 / 100Python9,202

PraisonAI 🦞, Hire a 24/7 AI Workforce. Stop writing boilerplate and start shipping autonomous self-improving agents that research, plan, code, and execute...

What the test found: Both praisonaiagents and PraisonAI packages install and import successfully. Agent creation works. All core agent unit tests pass (480/480). LLM token tracking race condition fixed with ContextVar. LLM class supports __deepcopy__ for agent cloning. Knowledge tests pass with chonkie installed. 193/194 LLM tests pass... 32 minutes.

101mcp-agentRuns with mocks92 / 100Python8,570

Build effective agents using Model Context Protocol and simple workflow patterns

What the test found: mcp-agent 0.2.6 installs, builds, CLI boots, and all 1501 tests pass after fixing the abstract generate_stream method in the base AugmentedLLM class and patching the Bedrock streaming test mock. 20 minutes.

102lamdaRuns with mocks92 / 100Python8,553

Android Full-Stack Device Control Platform: WebRTC/H.264 remote desktop, UI/OCR/image-matching automation, one-click MITM, built-in Frida, proxy/VPN/frp/P2P...

What the test found: lamda Python client v10.8 installs, all modules import, and the gRPC protocol stack works end-to-end via a mocked server. 7 minutes.

103MiroThinkerRuns with mocks92 / 100Python8,420

MiroThinker is a deep research agent optimized for complex research and prediction tasks. Our latest models, MiroThinker-1.7, achieves 74.0 and 75.3 on the...

What the test found: All 5 packages install and all source modules import correctly. The agent pipeline executes end-to-end with a mock LLM server: connects to MCP tool servers, calls LLM, processes response, writes task log with status=success. Flask trace visualizer answers HTTP 200. 12 minutes.

104awesome-agentic-ai-zhRuns with mocks92 / 100Python7,426

A trilingual (繁中 / English / 简中) learning roadmap for agentic AI: from LLM basics to multi-agent systems, with 240+ curated resources and hands-on examples. 中文...

What the test found: The mkdocs site builds to _build/site/, the mdBook builds to book/dist/, and all 25 example test suites (stages 1-7) pass with mock/offline mocks, no real API keys required. 21 minutes.

105CyberStrikeAIRuns with mocks92 / 100Go7,200

The system of action for AI-native cybersecurity, where intent becomes governed execution, evidence becomes operational memory, and every operation improves...

What the test found: The CyberStrikeAI binary builds cleanly with Go 1.25.0, serves its web UI on HTTP port 8080 with a 200 status, accepts admin login via POST /api/auth/login returning a full permissions token, and all 28 internal Go test packages (including database, handler, multiagent, security, mcp, and workflow) pass without... 17 minutes.

106open-multi-agentRuns with mocks92 / 100TypeScript6,990

Self-hosted TypeScript agent runtime with durable approvals and verifiable run records. Own it, approve it, audit it.

What the test found: The monorepo installs, builds, passes 2350 unit tests across all 4 workspaces plus the benchmark harness, and the CLI prints help, all with Node.js 22, no API keys or network required. 4 minutes.

107intentkitRuns with mocks92 / 100Python6,513

IntentKit is an open-source, self-hosted cloud agent cluster that manages a collaborative team of AI agents for you.

What the test found: Project installs with uv, the intentkit Python package and FastAPI server import correctly, and 1760 of 1774 non-BDD/non-migration tests pass using SQLite in-memory engine as a PostgreSQL replacement inside the container. 56 minutes.

108MiroFlowRuns with mocks92 / 100Python3,119

🏆 Top-1 on 5+ benchmarks | Web UI | Supports MiroThinker, Claude, Kimi, OpenAI

What the test found: MiroFlow agent pipeline runs end-to-end with a mock OpenRouter API: uv sync installs all deps, the agent loads config, initializes the LLM client, calls the API, processes tool-using context, and extracts a boxed final answer. 10 minutes.

109full-stack-ai-agent-templateRuns with mocks92 / 100Python1,934

Full-stack AI app generator, FastAPI + Next.js with AI Agents, RAG, streaming, auth, and 20+ integrations out of the box.

What the test found: All 681 tests pass; the CLI generates projects and runs type-checking on all generated matrix configurations including pydantic_deep_pg, pydantic_ai_code_execution, and pydantic_ai_skills without errors. 37 minutes.

110DirectorRuns with mocks92 / 100Python1,550

AI video agents framework for next-gen video interactions and workflows.

What the test found: Backend serves all 25 agent listings, config check, session CRUD, and collection endpoints via a mock VideoDB server. Frontend dev server builds and serves the Vue app. Both run simultaneously on :8000 and :8080. 16 minutes.

111kestraRuns with mocks87 / 100Java29,395

Event Driven Orchestration & Scheduling Platform for Mission Critical Applications

What the test found: Kestra repository builds and compiles successfully with JDK 25 / Gradle 9.7. All 5903 backend tests pass except HttpClientTest (Testcontainers/Docker unavailable). All 2743 frontend unit tests pass, type-checking is clean, lint reports 0 errors. A one-line production fix was applied so DNS resolution failures wrap in... 55 minutes.

112parlantRuns with mocks85.3 / 100Python18,301

Build reliable customer-facing AI agents with Parlant: an interaction control harness optimized for controlled, consistent, and predictable LLM interactions.

What the test found: Parlant v3.3.1 installs, imports, its server launches and responds on port 8800, and 113 of 115 non-engine tests pass when backed by the mock Emcie NLP server. 24 minutes.

113pydantic-aiRuns with mocks72 / 100Python20,487

How Python does AI. Agents, realtime voice, image generation, embeddings. Every model, every interface, typed end to end.

What the test found: Full install succeeds, agent runs end-to-end with the test model, CLI boots and displays help, and the vast majority of tests pass without any API keys. 51 minutes.

114vibe-kanbanRuns with mocks46 / 100Rust28,290

Get 10X more out of Claude Code, Codex or any coding agent

What the test found: Rust workspace (excluding tauri-app) compiles and passes cargo check. Server binary builds and starts, binding to random available port and responding to HTTP requests. 3 crate test suites pass (10/10 unit tests). Frontend TypeScript type checks pass for web-core and local-web. The npx CLI, Tauri desktop app, and web... 82 minutes.

115astridCould not verify20 / 100Rust10,278

Astrid is a portable, capability-secure operating system for composable software.

What the test found: could not be verified; the log shows where it stopped. 72 minutes.

116skillhubCould not verify10 / 100Java5,159

Self-hosted, open-source agent skill registry for enterprises. Publish & version skill packages, govern with RBAC and audit logs, deploy on-premise with...

What the test found: All 4 backend Maven modules compile and 1483 of 1489 tests pass (6 fail only due to Docker unavailability). Frontend dependencies install and individual vitest tests pass. The full app cannot start because Docker is not available to launch Postgres, Redis, and MinIO. 88 minutes.

-moliNot yet tested-Rust13,400-
-sreNot yet tested-TypeScript1,300-
-AWorldNot yet tested-Python1,240-
-LightAgentNot yet tested-Python1,231-
-openclackyNot yet tested-Ruby1,204-
-agent-landing-zoneNot yet tested-Python1,187-
-GPT-RAGNot yet tested-Python1,187-
-openharnessNot yet tested-Dart1,158-
-pydantic-deepagentsNot yet tested-Python1,077-
-CORALNot yet tested-Python1,052-
-FoundryNot yet tested-Python873-
-semantixNot yet tested-Go821-
-bladesNot yet tested-Go816-
-aeonNot yet tested-Shell767-
-crabtalkNot yet tested-Rust745-
-zero2AgentNot yet tested-Python699-
-swarmclawNot yet tested-TypeScript688-
-Deuz-SDKNot yet tested-TypeScript687-
-agentosNot yet tested-TypeScript677-
-daydreamsNot yet tested-TypeScript619-
-mnemonNot yet tested-Go612-
-forallNot yet tested-Rust578-
-deerflow-bookNot yet tested-JavaScript544-
-agents-universeNot yet tested-Python498-
-agentsilexNot yet tested-Python456-
-HyphaNot yet tested-TypeScript447-
-AutoHarnessNot yet tested-Python431-
-flowcraftNot yet tested-Go416-
-pydantic-ai-skillsNot yet tested-Python377-
-AI-companyNot yet tested-Python371-
-agenticaNot yet tested-Python352-
-alphoraNot yet tested-Python350-
-yoagentNot yet tested-Rust340-
-CyberClawNot yet tested-Python336-
-aftNot yet tested-Rust319-
-agentaoNot yet tested-Python307-
-AgenticXNot yet tested-Python301-
-vizra-adkNot yet tested-PHP295-
-SicoNot yet tested-Python278-
-animaworksNot yet tested-Python266-
-datNot yet tested-Java259-
-KADATHNot yet tested-Python258-
-agentdescentNot yet tested-Python252-
-designpowersNot yet tested-Shell251-
-awesome-ai-agents-2026Not yet tested-Python245-
-agents-aeaNot yet tested-Python241-
-LeAgentNot yet tested-Python232-
-skills-constitutionNot yet tested-Python232-
-streamcore-serverNot yet tested-Go232-
-LibraNot yet tested-Python222-
-Trend2Video-ProNot yet tested-Python214-
-agentic-osNot yet tested-Python206-

runs installed and started with its real dependencies. runs with mocks started after stand-ins replaced external services such as a database or a third-party API. could not verify neither the standard agent nor the stronger one got it running within the time limit; the log shows where it stopped.

How we tested

On this list as of the latest test: 85 projects ran as-is, 29 with mocks, 2 could not be verified, 52 still waiting. Languages tested: C++, Elixir, Go, Java, JavaScript, Python, Rust, Shell, TypeScript. Every attempt used a clean single-use machine, the subject at a pinned version, and a 45-minute limit; the complete procedure is on the methodology page.

Frequently asked questions (FAQs)

How is this list ranked?

By measurement, not opinion: projects Argusic installed and launched on a fresh machine come first, then those that ran with mocks in place of external services, then those it could not verify. Ties go to the Argusic Score, then how popular it is on its own source.

Why are some projects unranked?

52 projects are still waiting for a test or for a finished attempt. They are listed without a rank until Argusic has measured them.

Where is the evidence?

Every row links to the project's Argusic page, where each run has a full log and a terminal recording stored with a sha256 fingerprint. The same pages exist for every one of the tested projects, on this list or not.

More lists in this category

All lists: Best. All tested projects: subjects.