Open source web scraping tools, installed and run
Web scraping and crawling tools you can run yourself. The list is ordered by whether each project came up when Argusic installed it fresh, with the recording of that attempt one click away.
Tested between and . Each row shows its own test date; a project can change after that day.
61 of 64 tested projects run. 74 more waiting for a test.
In short: 47 of the 64 tested projects started as-is on a fresh machine: ai-website-cloner-template, colly, proxy_pool, katana, newspaper, Skill_Seekers, SeleniumBase, and crawlab, and 39 more. 14 more started once a stand-in replaced a service they expect, such as a database: Scrapling, pinchtab, node-crawler, twscrape, brightdata-mcp, diskover-community, finvizfinance, and Scweet, and 6 more. 3 could not be verified: WechatSogou, fess, and agentql; the log shows where each one stopped.
Measured by Argusic on a fresh machine every time. Every number links to its evidence.
| # | project | verdict | Argusic Score | language | stars | tested on |
|---|---|---|---|---|---|---|
| 1 | ai-website-cloner-template | Runs | 100 / 100 | TypeScript | 36,241 | |
Clone any website with one command using AI coding agents What the test found: The next.js app installs, builds, typechecks, lints, and runs with a dev server responding HTTP 200 on localhost:3000. 6 minutes. | ||||||
| 2 | colly | Runs | 100 / 100 | Go | 25,548 | |
Elegant Scraper and Crawler Framework for Golang What the test found: installed Go 1.24.9, built colly v2.1.0, all tests passed, CLI binary functional. 1 minute. | ||||||
| 3 | proxy_pool | Runs | 100 / 100 | Python | 23,753 | |
Python ProxyPool for web spider What the test found: The ProxyPool Flask API server runs on port 5010 backed by real Redis, responds to all endpoints (/get, /pop, /all, /delete, /count), and all 248 automated tests pass. 6 minutes. | ||||||
| 4 | katana | Runs | 100 / 100 | Go | 17,631 | |
A next-generation crawling and spidering framework. What the test found: Katana v1.7.0 builds and runs: go test ./... passes all unit tests, the 5 functional tests (standard crawl, depth 3, headless, two authenticated recorded-flow variants) all pass, and the binary crawls a live HTTP target discovering URLs. 8 minutes. | ||||||
| 5 | newspaper | Runs | 100 / 100 | Python | 15,170 | |
newspaper3k is a news, full-text, and article metadata extraction in Python 3. Advanced docs: What the test found: newspaper3k 0.3.0 installs in a virtualenv on Python 3.12, all 40 unit tests pass, and core Article download/parse/nlp pipeline produces correct output from static HTML. 4 minutes. | ||||||
| 6 | Skill_Seekers | Runs | 100 / 100 | Python | 15,115 | |
Convert documentation websites, GitHub repositories, and PDFs into Claude AI skills with automatic conflict detection What the test found: skill-seekers 3.10.0.dev0 installed and all 3857 unit/adaptor/scraper tests pass without external API keys or network dependencies. 6 minutes. | ||||||
| 7 | SeleniumBase | Runs | 100 / 100 | Python | 13,053 | |
📊 Browser automation framework for scraping, testing, and completing tasks with Python. Supports pytest. Stealth options. Over 100 examples. What the test found: SeleniumBase 4.54.7 is installed in a Python 3.12 venv; 8 framework unit tests, 7 offline HTML tests, and 2 real web tests passed using headless Chrome 154. 5 minutes. | ||||||
| 8 | crawlab | Runs | 100 / 100 | Go | 12,277 | |
Distributed web crawler admin platform for spiders management regardless of languages and frameworks. 分布式爬虫管理平台,支持任何语言和框架 What the test found: Crawlab Go backend compiles, runs as master with MongoDB, serves the /system-info API endpoint on port 8080 returning 200, and its test suite runs clean for 4 core packages; the Vue frontend builds with Vite. 17 minutes. | ||||||
| 9 | webmagic | Runs | 100 / 100 | Java | 11,671 | |
A scalable web crawler framework for Java. What the test found: All 8 Maven modules (webmagic-core, extension, scripts, selenium, saxon, samples, coverage) compile and their tests pass with 0 failures. The library is fully functional as a Java 11 crawler framework. 41 minutes. | ||||||
| 10 | camofox-browser | Runs | 100 / 100 | JavaScript | 11,480 | |
Stealth headless browser for AI agents, bypass Cloudflare, bot detection, and anti-scraping. Drop-in Puppeteer/Playwright replacement. What the test found: Server starts with Camoufox browser under Xvfb, responds on port 9377, creates browser tabs that navigate to real URLs, returns accessibility snapshots with element refs, and processes click interactions correctly. All API endpoints function end-to-end. 4 minutes. | ||||||
| 11 | crawlee-python | Runs | 100 / 100 | Python | 9,582 | |
Crawlee, A web scraping and browser automation library for Python to build reliable crawlers. Extract data for AI, LLMs, RAG, or GPTs. Download HTML, PDF, JPG... What the test found: Dependencies installed with uv sync --all-extras, Playwright browsers (chromium + firefox) downloaded, all 2480 unit tests pass, lint and type-check pass, and the package builds and imports successfully at version 1.10.3. 8 minutes. | ||||||
| 12 | helium | Runs | 100 / 100 | Python | 8,328 | |
Lighter web automation with Python What the test found: Helium 7.0.3 installed from source via pip install -e . Selenium 4.49.0 with Chrome 154 headless via Selenium Manager; all 332 tests pass with 2 regressions fixed (date input JS fallback, subprocess module discovery). 33 minutes. | ||||||
| 13 | JMComic-Crawler-Python | Runs | 100 / 100 | Python | 7,521 | |
Python API for JMComic | 提供Python API访问禁漫天堂,同时支持网页端和移动端 | 禁漫天堂GitHub Actions下载器🚀 What the test found: Package jmcomic 2.7.7 installed in a Python 3.12.3 venv with all dependencies; 238/240 tests pass including real network calls against JMComic API servers; CLI commands jmcomic and jmv produce correct help output. 10 minutes. | ||||||
| 14 | trafilatura | Runs | 100 / 100 | Python | 6,934 | |
Python & Command-line tool to gather text and metadata on the Web: Crawling, scraping, extraction, output as CSV, JSON, HTML, MD, TXT, XML What the test found: Trafilatura 2.2.0 is installed in editable mode in a .venv, all 428 tests pass, and both CLI and Python extraction work correctly on sample HTML input. 6 minutes. | ||||||
| 15 | curl_cffi | Runs | 100 / 100 | Python | 6,682 | |
Python binding for curl-impersonate fork via cffi. A http client that can impersonate browser tls/ja3/http2 fingerprints. What the test found: After installing the pre-built curl-cffi wheel (v0.16.3) into a fresh venv, the library imports, the CLI runs (curl-cffi get, doctor, --help), real HTTP requests with browser impersonation return 200, and the unit test suite passes 600+ tests across all 20 test modules including sync/async HTTP, websockets... 38 minutes. | ||||||
| 16 | ferret | Runs | 100 / 100 | Go | 6,012 | |
Declarative data automation language and Go runtime for structured extraction workflows. What the test found: Go 1.27.1 installed, all dependencies resolved, all 68 test packages pass with race detection, CLI binary compiles and runs FQL queries end-to-end. 7 minutes. | ||||||
| 17 | scrapy-redis | Runs | 100 / 100 | Python | 5,642 | |
Redis-based components for Scrapy. What the test found: scrapy-redis 0.9.1 is installed in a venv at /work/repo/.venv with real Redis 6.2 running on localhost:6379; all 102 tests pass covering connection pooling, dupefilter, queue (FIFO/LIFO/priority), scheduler, spider setup, idle backoff, multi-key consumption, and utility modules. 9 minutes. | ||||||
| 18 | scrape-it | Runs | 100 / 100 | JavaScript | 4,068 | |
🔮 A Node.js scraper for humans. What the test found: scrape-it installs, builds, and passes its full test suite on Node 18.19.1 with cheerio pinned to 1.0.0. 3 minutes. | ||||||
| 19 | toapi | Runs | 100 / 100 | Python | 3,556 | |
Every web site provides APIs. What the test found: toapi 2.2.4 installed from source, all 7 tests pass, CLI help works, Flask server answers 200 with JSON, version 2.2.4 confirmed. 2 minutes. | ||||||
| 20 | php-curl-class | Runs | 100 / 100 | PHP | 3,300 | |
PHP Curl Class makes it easy to send HTTP requests and integrate with web APIs What the test found: php-curl-class repository is fully installed with PHP 8.5.8 static binary and all Composer dependencies; the full test suite passes with 264 tests and 23390 assertions and zero failures. 43 minutes. | ||||||
| 21 | google-play-scraper | Runs | 100 / 100 | JavaScript | 2,977 | |
Node.js scraper to get data from Google Play What the test found: The google-play-scraper npm package installs, all 84 tests pass against live Google Play, and the library's app/list/search/suggest/developer/reviews/similar/permissions/datasafety/categories methods all fetch and parse real data correctly. 8 minutes. | ||||||
| 22 | crawler | Runs | 100 / 100 | PHP | 2,830 | |
https://spatie.be/docs/crawler What the test found: PHP 8.4.23 and Composer 2.7.1 installed from static binaries, all dependencies resolved via `composer update --prefer-stable`, and the full test suite passes (229/229) including both fake-HTTP unit tests and real-HTTP integration tests against a local PHP server. 6 minutes. | ||||||
| 23 | geziyor | Runs | 100 / 100 | Go | 2,781 | |
Geziyor, blazing fast web crawling & scraping framework for Go. Supports JS rendering. What the test found: The Geziyor web scraping framework builds and passes every test across all 7 packages, and a compiled demo successfully scraped a live website. 16 minutes. | ||||||
| 24 | spider | Runs | 100 / 100 | Rust | 2,762 | |
Foundational low latency web data collecting in Rust What the test found: Rust 1.99.0 installed, cargo build succeeds, spider CLI binary crawls and scrapes example.com returning status 200 and expected page content, all 608 spider core tests pass, workspace-wide crate tests all pass. 56 minutes. | ||||||
| 25 | gecco | Runs | 100 / 100 | Java | 2,512 | |
Easy to use lightweight web crawler(易用的轻量化网络爬虫) What the test found: Build succeeded (mvn compile, package, test-compile all exit 0); demo crawler successfully fetched and parsed https://www.baidu.com/, producing JSON output. 5 minutes. | ||||||
| 26 | Crawler-Detect | Runs | 100 / 100 | PHP | 2,408 | |
🕷 CrawlerDetect is a PHP class for detecting bots/crawlers/spiders via the user agent What the test found: The CrawlerDetect PHP library installs, builds, and passes its full test suite of 28 tests with 170,886 assertions on PHP 8.0.30. 5 minutes. | ||||||
| 27 | goclone | Runs | 100 / 100 | Go | 2,267 | |
Website Cloner - Utilizes powerful Go routines to clone websites to your computer within seconds. What the test found: goclone binary builds from source, all 9 test suites across 3 packages pass, the built binary's help and version commands print correct output. 11 minutes. | ||||||
| 28 | stealth-browser-mcp | Runs | 100 / 100 | Python | 2,216 | |
Stealth browser automation for AI agents over MCP and the Chrome DevTools Protocol: navigation, network hooks, DOM extraction, and pixel-accurate UI cloning. What the test found: Stealth Browser MCP server runs on stdio transport, spawns a headless Chromium 157.0.8086.0, navigates to data URLs, extracts page HTML content, and closes cleanly -- all via MCP protocol over stdio. 15 minutes. | ||||||
| 29 | gain | Runs | 100 / 100 | Python | 2,019 | |
Web crawling framework based on asyncio. What the test found: All 7 tests pass, ruff lint passes, library imports correctly, and item attribute extraction from HTML produces expected results. 1 minute. | ||||||
| 30 | x-crawl | Runs | 100 / 100 | TypeScript | 1,889 | |
Flexible Node.js crawler library What the test found: x-crawl library builds successfully; its crawlData, crawlHTML, async/sync modes, baseUrl, timeout, errorCollect, proxy rotation, and Puppeteer browser APIs work correctly against a real HTTP server and CONNECT proxy. 34 minutes. | ||||||
| 31 | ruia | Runs | 100 / 100 | Python | 1,738 | |
Async Python 3.6+ web scraping micro-framework based on asyncio What the test found: Ruia 0.8.5 is installed in a Python 3.12 venv with uvloop; pip install -e .[uvloop] succeeds; 66 of 67 tests pass including all field extraction, item parsing, middleware, request fetching, and spider orchestration tests against live httpbin.org and news.ycombinator.com. 22 minutes. | ||||||
| 32 | selectolax | Runs | 100 / 100 | Python | 1,688 | |
Python binding to Modest and Lexbor engines. Fast HTML5 parser with CSS selectors for Python. What the test found: selectolax C extension successfully built, installed in a venv, all 279 tests passing, and basic HTML parsing + CSS selector usage verified. 4 minutes. | ||||||
| 33 | patchright-python | Runs | 100 / 100 | Python | 1,552 | |
Undetected Python version of the Playwright testing and automation library. What the test found: Patchright 1.63.0 is installed, Chromium browser is downloaded and launchable, the CLI reports its version, and 2237 of 2239 Chromium tests pass (sync + async). 37 minutes. | ||||||
| 34 | wreq-python | Runs | 100 / 100 | Python | 1,484 | |
An ergonomic, privacy-aware Python HTTP Client What the test found: wreq Python package builds, installs, and passes all 57 unit tests against a local httpbin server, supporting HTTP/1.1 and HTTP/2 requests, cookie jars, redirect policies, TLS emulation (JA3/JA4), multipart uploads, streaming, and compression. 15 minutes. | ||||||
| 35 | wombat | Runs | 100 / 100 | Ruby | 1,359 | |
Lightweight Ruby web crawler/scraper with an elegant DSL which extracts structured data from pages. What the test found: Wombat 3.3.0 builds and installs as a gem under Ruby 3.2.3, all 56 RSpec tests pass, and a live Wombat.crawl scrape of a local WEBrick page returns correctly parsed structured data. 12 minutes. | ||||||
| 36 | Douyin_TikTok_Download_API | Runs | 97.3 / 100 | Python | 20,511 | |
🚀 抖音、TikTok 数据采集与无水印视频下载 API,自托管,支持 MCP 调用与 Docker 一键部署。| Self-hosted TikTok & Douyin scraper and no-watermark video downloader, async REST API, MCP server... What the test found: All 2960 tests pass (2450 unit + 510 integration), the CLI outputs version 5.0.1 and lists 10 commands, the API server starts and serves healthz/readyz/swagger/console root with 200, the web console SPA is built, and both the PostgreSQL+TimescaleDB and Redis backends and the Go downloader sidecar run. 48 minutes. | ||||||
| 37 | rod | Runs | 96 / 100 | Go | 7,124 | |
A Chrome DevTools Protocol driver for web automation and scraping. What the test found: Rod compiles with go build ./... passes all vet checks, and launches Chromium 128 to control a browser page via DevTools Protocol. The main rod test suite and lib/cdp, lib/launcher (except known flaky TestManaged), and lib/proto sub-packages all pass. 59 minutes. | ||||||
| 38 | browserless | Runs | 93.3 / 100 | JavaScript | 1,839 | |
The headless Chrome/Chromium driver on top of Puppeteer. Take screenshots, generate PDFs, extract text and HTML with a production-ready API. What the test found: The browserless monorepo installs under Node.js 24 with puppeteer 25, all packages build, Chrome 154 launches headlessly, and 390+ tests pass across 10 packages (browserless, goto, screenshot, capture, function, devices, errors, screencast, lighthouse, ai). 32 minutes. | ||||||
| 39 | crawl4ai | Runs | 91.4 / 100 | Python | 84,979 | |
Open-source web crawler and scraper for LLMs and AI agents: any website into clean, LLM-ready Markdown. Run it yourself, or use Crawl4AI Cloud with one key. What the test found: Crawl4AI 0.9.4 installed in venv, Playwright chromium browser downloaded, DB initialized, AsyncWebCrawler can crawl real URLs with markdown output, cache operations work, and 360 core tests pass with 11 skipped. 14 minutes. | ||||||
| 40 | firecrawl-mcp-server | Runs | 90.2 / 100 | JavaScript | 7,569 | |
🔥 Official Firecrawl MCP Server - Adds powerful web scraping and search to Cursor, Claude and any other LLM clients. What the test found: Firecrawl MCP Server v3.24.1 builds, installs, and starts successfully with 25 tools in stdio mode; 81 of 84 tests pass. 24 minutes. | ||||||
| 41 | lux | Runs | 90 / 100 | Go | 31,752 | |
👾 Fast and simple video download library and CLI tool written in Go What the test found: Lux builds and runs: the binary outputs version info and help, extracts real YouTube video metadata, and 14 of 48 test packages pass (all failures are network-dependent extractor tests or DNS-timeout content-type lookups; no code defects remain). 11 minutes. | ||||||
| 42 | news-please | Runs | 90 / 100 | Python | 2,492 | |
news-please - an integrated web crawler and information extractor for news that just works What the test found: news-please 1.6.13 installs and its library API extracts structured article data (title, authors, publication date, description, maintext, language) from HTML via both HTTP and raw HTML input; CLI entry points respond with help output. 5 minutes. | ||||||
| 43 | skills | Runs | 85 / 100 | Python | 6,116 | |
Browser automation CLI built for AI agents. Break through anti-bot walls, hand off to humans across platforms when stuck. Parallel multi-task execution... What the test found: browser-act-cli 1.4.2 installed, skill handshake completed, and the chrome browser type (bundled Chromium) works end-to-end, browser create, open, navigate, state inspection, content extraction via get markdown, and session close all succeed; stealth features require a server-side-validated BrowserAct API key that... 5 minutes. | ||||||
| 44 | scrapy | Runs | 80 / 100 | Python | 64,658 | |
Scrapy, a fast high-level web crawling & scraping framework for Python. What the test found: Scrapy 2.19.0 is installed and fully functional. The CLI, 4539/4540 tests, and an end-to-end crawl all work. The sole failing test is a container IPv6 limitation, not a code defect. 11 minutes. | ||||||
| 45 | changedetection.io | Runs | 60 / 100 | Python | 34,834 | |
Best and simplest tool for website change detection, web page monitoring, and website change alerts. Perfect for tracking content changes, price drops, restock... What the test found: changedetection.io 0.60.7 installs into a venv, launches and serves HTTP 200, its full basic test suite passes (680 passed in the main group, exit 0, plus 508 unit tests), and a real watch added via the API detected a content change and exposed it in the UI. 76 minutes. | ||||||
| 46 | jsoup | Runs | 60 / 100 | Java | 11,405 | |
jsoup: the Java HTML parser, built for HTML editing, cleaning, scraping, and XSS safety. What the test found: jsoup compiles from source with JDK 21 and Apache Maven, all 2,341 unit tests pass, and the built JAR correctly parses HTML and extracts DOM content. 2 minutes. | ||||||
| 47 | fscrawler | Runs | 60 / 100 | Java | 1,453 | |
Elasticsearch File System Crawler (FS Crawler) What the test found: FSCrawler 3.1-SNAPSHOT builds successfully via Maven, distributable ZIP is packaged, binary launches and shows help/version, and 687 unit tests pass when run with the `-Dtest="*Test"` workaround. 82 minutes. | ||||||
| 48 | Scrapling | Runs with mocks | 92 / 100 | Python | 86,312 | |
🕷️ An adaptive Web Scraping framework that handles everything from a single request to a full-scale crawl! Don't be shy, join here... What the test found: Scrapling v0.4.15 installed and all non-browser tests pass; the core parser, CLI, spider framework, request-based fetchers, and integration with Scrapy all work correctly. 11 minutes. | ||||||
| 49 | pinchtab | Runs with mocks | 92 / 100 | Go | 10,349 | |
High-performance browser automation bridge and multi-instance orchestrator with advanced stealth injection and real-time dashboard. What the test found: Pinchtab builds from source, 53 of 54 Go test packages pass (1 pre-existing handlers failure), and the bridge serves HTTP 200 health checks with Chrome 131 headless. 41 minutes. | ||||||
| 50 | node-crawler | Runs with mocks | 92 / 100 | TypeScript | 6,794 | |
Web Crawler/Spider for NodeJS + server-side jQuery ;-) What the test found: crawler v2.1.1 builds from TypeScript, all 51 tests pass under Node 22 with nock-mocked HTTP, and the module loads correctly to process crawl tasks. 4 minutes. | ||||||
| 51 | twscrape | Runs with mocks | 92 / 100 | Python | 2,844 | |
Python library and CLI for X/Twitter scraping with multi-account rotation and built-in rate-limit handling. What the test found: All 217 tests pass with 81% coverage; twscrape CLI reports version 0.20.1 and SQLite 3.45.1. 1 minute. | ||||||
| 52 | brightdata-mcp | Runs with mocks | 92 / 100 | JavaScript | 2,661 | |
A powerful Model Context Protocol (MCP) server that provides an all-in-one solution for public web access. What the test found: The Bright Data MCP server installs and all 21 unit tests pass under Node.js v22.12.0; the server starts listening when API_TOKEN is set. 4 minutes. | ||||||
| 53 | diskover-community | Runs with mocks | 92 / 100 | PHP | 1,823 | |
Diskover Community Edition - Open source file indexer, file search engine and data management and analytics powered by Elasticsearch What the test found: Python scanner installs successfully, compiles cleanly, and completes a directory crawl indexing 22 file/directory documents into a mock Elasticsearch instance. 18 minutes. | ||||||
| 54 | finvizfinance | Runs with mocks | 92 / 100 | Python | 1,716 | |
Finviz analysis python library. What the test found: All 128 offline tests pass and every package module (quote, news, insider, calendar, screener, group, earnings, forex, crypto, future, util) returns correct data from committed HTML fixtures. 2 minutes. | ||||||
| 55 | Scweet | Runs with mocks | 92 / 100 | Python | 1,640 | |
Scrape tweets, profiles, followers and following from Twitter/X, no API key needed. Python library with smart multi-account pooling, proxy support and async. What the test found: Scweet 5.8.0 installs cleanly in a Python 3.12 venv; 495 unit tests pass; the CLI builds and shows help; the library imports without network; material in tests/fixtures/ provides captured X responses for offline test coverage. 9 minutes. | ||||||
| 56 | fredy | Runs with mocks | 92 / 100 | JavaScript | 1,563 | |
❤️ Fredy - [F]ind [R]eal [E]state [D]amn Eas[y] - Fredy keeps searching for new apartments, houses, and flats in Europe on platforms like ImmoScout24... What the test found: Fredy backend serves HTTP 200 on port 9998 after DB migration and all 4260 offline tests pass with Node v22.23.3. 6 minutes. | ||||||
| 57 | boss-zhipin-scraper | Runs with mocks | 92 / 100 | Python | 1,525 | |
Boss直聘爬虫 / BOSS直聘职位数据抓取工具,基于 Chrome CDP 协议复用真实登录态,绕过字体反爬,输出明文薪资 JSON/CSV + 薪资技能分析。A Chrome-CDP-based BOSS Zhipin job scraper/crawler. What the test found: Virtual environment created with all dependencies, both Python scripts compile clean, and the full mock-based test suite (109 tests) passes. The CLI prints correct version (2.2.0), help output, and city listings. Real Chrome CDP scraping cannot be verified in this container, --check correctly reports no Chrome... 2 minutes. | ||||||
| 58 | openserp | Runs with mocks | 92 / 100 | Go | 1,455 | |
Self-hosted SERP API for AI, SEO & automation. Browser-rendered Google, Bing, Yandex, Baidu, DuckDuckGo and Ecosia search with page extraction 🎉 What the test found: OpenSERP v0.8.13 builds, all 12 unit test suites pass, and the HTTP server starts, serving /health and /ready endpoints on 127.0.0.1:7000. 8 minutes. | ||||||
| 59 | WebAI2API | Runs with mocks | 92 / 100 | JavaScript | 1,383 | |
WebAI2API: 基于 Camoufox 的网页 AI 转 API 工具,支持 LMArena/Gemini等,多窗口并发与账号隔离。 | Web AI to OpenAI API via Camoufox. Supports LMArena/Gemini and more, multi-window... What the test found: The WebAI2API server starts on port 3000 with a Camoufox browser instance, serving OpenAI-compatible /v1/models, WebUI, and /v1/cookies endpoints, all returning HTTP 200, though chat completions require real AI website credentials via the browser adapter and cannot produce output without logging in. 7 minutes. | ||||||
| 60 | google-maps-scraper | Runs with mocks | 86 / 100 | Go | 6,326 | |
scrape data from Google Maps. Extracts data such as the name, address, phone number, website URL, rating, reviews number, latitude and longitude... What the test found: All 17 Go packages with tests passed with race detection (no failures), the binary compiled and prints version/help correctly, and all 10 JS skill tests plus shell helper tests passed. 6 minutes. | ||||||
| 61 | Scrapegraph-ai | Runs with mocks | 56 / 100 | Python | 31,618 | |
Python scraper based on AI What the test found: 191 unit tests pass after fixing 13 categories of test failures covering models_tokens assertions, cleanup_html behaviors, logger capture, file paths, missing dependencies, mock patterns, and mock class completeness. 46 minutes. | ||||||
| 62 | WechatSogou | Could not verify | 80 / 100 | Python | 6,398 | |
基于搜狗微信搜索的微信公众号爬虫接口 What the test found: wechatsogou 4.5.4 is installed in a Python 3.12 venv; pip install -e . succeeds; 24 offline tests pass (constants, tools, URL generation, HTML parsing); real-network API tests fail only because DNS is unavailable in the container. 13 minutes. | ||||||
| 63 | fess | Could not verify | 47.5 / 100 | Java | 1,140 | |
Open-source, self-hosted enterprise & site search server built on OpenSearch. Crawls web / file / DB / cloud sources, 20+ languages, REST API, and AI/RAG &... What the test found: Fess 15.9.0-SNAPSHOT builds and packages successfully from source. All 7,634 unit tests pass. The embedded Tomcat binds port 8080 but the application fails at runtime because no OpenSearch backend is reachable on localhost:9200. 36 minutes. | ||||||
| 64 | agentql | Could not verify | 20 / 100 | Python | 1,479 | |
AgentQL is a suite of tools for connecting your AI to the web. Featuring a query language and Playwright integrations for interacting with elements and... What the test found: could not be verified; the log shows where it stopped. | ||||||
| - | moli | Not yet tested | - | Rust | 13,400 | - |
| - | httpcloak | Not yet tested | - | Go | 1,332 | - |
| - | google-maps-scraper-kit | Not yet tested | - | Python | 1,312 | - |
| - | Beanbun | Not yet tested | - | PHP | 1,261 | - |
| - | Decodo | Not yet tested | - | Java | 1,247 | - |
| - | reverse-api-engineer | Not yet tested | - | Python | 1,226 | - |
| - | newspaper4k | Not yet tested | - | Python | 1,150 | - |
| - | wreq | Not yet tested | - | Rust | 1,060 | - |
| - | parse-video | Not yet tested | - | Go | 1,057 | - |
| - | stormcrawler | Not yet tested | - | Java | 1,005 | - |
| - | x-reader | Not yet tested | - | Python | 964 | - |
| - | x-kit | Not yet tested | - | Python | 943 | - |
| - | siteone-crawler | Not yet tested | - | Rust | 937 | - |
| - | icrawler | Not yet tested | - | Python | 935 | - |
| - | chatWeb | Not yet tested | - | Python | 918 | - |
| - | TikHub-API-Python-SDK | Not yet tested | - | Python | 903 | - |
| - | scrapyrt | Not yet tested | - | Python | 885 | - |
| - | agent-skills | Not yet tested | - | JavaScript | 875 | - |
| - | Scavenger | Not yet tested | - | Python | 842 | - |
| - | spidr | Not yet tested | - | Ruby | 836 | - |
| - | jvppeteer | Not yet tested | - | Java | 805 | - |
| - | seonaut | Not yet tested | - | Go | 804 | - |
| - | patchright-nodejs | Not yet tested | - | TypeScript | 794 | - |
| - | Krawl | Not yet tested | - | Python | 774 | - |
| - | gpt-promo-scanner | Not yet tested | - | Python | 760 | - |
| - | xxl-crawler | Not yet tested | - | Java | 756 | - |
| - | google-search-results-python | Not yet tested | - | Python | 755 | - |
| - | scrapecraft | Not yet tested | - | Python | 713 | - |
| - | ketch | Not yet tested | - | Go | 697 | - |
| - | deepcrawl | Not yet tested | - | TypeScript | 684 | - |
| - | NetDiscovery | Not yet tested | - | Java | 644 | - |
| - | primp | Not yet tested | - | Rust | 615 | - |
| - | social-media-profile-scrapers | Not yet tested | - | Python | 580 | - |
| - | Stealth-Requests | Not yet tested | - | Python | 567 | - |
| - | freeproxy | Not yet tested | - | Python | 558 | - |
| - | reader | Not yet tested | - | TypeScript | 558 | - |
| - | oc | Not yet tested | - | JavaScript | 533 | - |
| - | browser-search | Not yet tested | - | JavaScript | 528 | - |
| - | scrapple | Not yet tested | - | Python | 505 | - |
| - | fundus | Not yet tested | - | Python | 479 | - |
| - | quantumproxies.io | Not yet tested | - | Java | 459 | - |
| - | email-sleuth | Not yet tested | - | Rust | 426 | - |
| - | sparkler | Not yet tested | - | Python | 419 | - |
| - | camofox-browser | Not yet tested | - | JavaScript | 411 | - |
| - | wreq-js | Not yet tested | - | TypeScript | 408 | - |
| - | Hydra | Not yet tested | - | Java | 396 | - |
| - | trends-checker | Not yet tested | - | Python | 396 | - |
| - | ir-search | Not yet tested | - | Python | 392 | - |
| - | nebula | Not yet tested | - | Go | 379 | - |
| - | news-crawl | Not yet tested | - | Java | 379 | - |
| - | graphlit-mcp-server | Not yet tested | - | TypeScript | 378 | - |
| - | top-user-agents | Not yet tested | - | JavaScript | 378 | - |
| - | crawler | Not yet tested | - | PHP | 371 | - |
| - | scrapy-zyte-smartproxy | Not yet tested | - | Python | 365 | - |
| - | hQuery.php | Not yet tested | - | PHP | 360 | - |
| - | twitter-scraper-selenium | Not yet tested | - | Python | 346 | - |
| - | Laravel-Crawler-Detect | Not yet tested | - | PHP | 326 | - |
| - | aliexpress-product-scraper | Not yet tested | - | JavaScript | 321 | - |
| - | extractor | Not yet tested | - | TypeScript | 321 | - |
| - | SpotifyScraper | Not yet tested | - | Python | 316 | - |
| - | agent-fetch | Not yet tested | - | TypeScript | 314 | - |
| - | 4chan-downloader | Not yet tested | - | Python | 312 | - |
| - | gplay-scraper | Not yet tested | - | Python | 310 | - |
| - | xvideos | Not yet tested | - | TypeScript | 304 | - |
| - | MinerU-HTML | Not yet tested | - | Python | 292 | - |
| - | teracrawl | Not yet tested | - | TypeScript | 285 | - |
| - | geo-aeo-tracker | Not yet tested | - | TypeScript | 284 | - |
| - | go-movies | Not yet tested | - | Go | 275 | - |
| - | laravel-seo-scanner | Not yet tested | - | PHP | 274 | - |
| - | Uni-CLI | Not yet tested | - | TypeScript | 274 | - |
| - | facebook-ads-library-mcp | Not yet tested | - | Python | 270 | - |
| - | OddsHarvester | Not yet tested | - | Python | 257 | - |
| - | Project-Eyes-On | Not yet tested | - | Python | 245 | - |
| - | examples | Not yet tested | - | TypeScript | 244 | - |
runs installed and started with its real dependencies. runs with mocks started after stand-ins replaced external services such as a database or a third-party API. could not verify neither the standard agent nor the stronger one got it running within the time limit; the log shows where it stopped.
How we tested
On this list as of the latest test: 47 projects ran as-is, 14 with mocks, 3 could not be verified, 74 still waiting. Languages tested: Go, Java, JavaScript, PHP, Python, Ruby, Rust, TypeScript. Every attempt used a clean single-use machine, the subject at a pinned version, and a 45-minute limit; the complete procedure is on the methodology page.
Frequently asked questions (FAQs)
How is this list ranked?
By measurement, not opinion: projects Argusic installed and launched on a fresh machine come first, then those that ran with mocks in place of external services, then those it could not verify. Ties go to the Argusic Score, then how popular it is on its own source.
Why are some projects unranked?
74 projects are still waiting for a test or for a finished attempt. They are listed without a rank until Argusic has measured them.
Where is the evidence?
Every row links to the project's Argusic page, where each run has a full log and a terminal recording stored with a sha256 fingerprint. The same pages exist for every one of the tested projects, on this list or not.
More lists in this category
- browser automation tools (shares Uni-CLI, WebAI2API, brightdata-mcp, browser-search, camofox-browser, camofox-browser, ferret, moli, oc, pinchtab, reader, stealth-browser-mcp, teracrawl with this list)
- API clients (shares Stealth-Requests, curl_cffi, httpcloak, php-curl-class, primp, wreq, wreq-js, wreq-python with this list)
- self-hosted AI apps (shares Krawl, Skill_Seekers, WebAI2API, browser-search, fess, moli, stealth-browser-mcp with this list)
- MCP servers (shares Douyin_TikTok_Download_API, SeleniumBase, Skill_Seekers, brightdata-mcp, firecrawl-mcp-server, stealth-browser-mcp with this list)
- AI agent frameworks (shares ai-website-cloner-template, crawl4ai, moli with this list)
- search engines (shares fess, openserp with this list)
- low-code platforms (shares skills with this list)
- static site generators (shares siteone-crawler with this list)
- vector databases (shares chatWeb with this list)
- API gateways
- CI/CD tools
- CMS platforms
- LLM gateways
- VPN tools
- backup tools
- code editors
- open source coding agents
- Open source databases, installed and queried
- developer CLI tools
- e-commerce platforms
- ebook readers
- game engines
- open source games
- home automation tools
- map tools
- message queues
- Open source music servers, installed and played
- observability tools
- Open source office suites, installed and launched
- self-hosted password managers
- project management tools
- screen recorders
- self-hosted analytics
- speech tools
- uptime monitors
- open source video editors
- video players
- whiteboard tools
- wikis
- workflow automation tools
- self-hosted Notion alternatives
- self-hosted dashboards
- self-hosted git servers