ClawBench

Open-source benchmark for browser AI agents on daily tasks.

Not yet testedsource: GitHubhomepagePythonApache-2.0commit 187cd252bc60

Python, Apache-2.0 licensed. The project labels itself: agent evaluation, agentic ai, ai agent benchmark, ai agents, benchmark, browser agent, browser automation and browser use.

ClawBench has not been verified yet.

Measured by Argusic on a fresh machine every time. Every number links to its evidence.

At a glance

verdict
Not yet tested
Argusic Score
not scored yet
recorded runs
0
last tested
-
stars
977
forks
62
open issues
52
watchers
13
size
5 MB
created
last push

Subject data from GitHub, linked at the top of this page, refreshed . Test data by Argusic (CC BY 4.0); every number links to a run page with the full log, the recording, and their sha256 hashes.

Also tested, in the same area

Every one of these was installed and run by Argusic on a clean machine. Nothing appears here that was not tested.

Run history

No recorded runs.

Topics (from GitHub)

agent-evaluationagentic-aiai-agent-benchmarkai-agentsbenchmarkbrowser-agentbrowser-automationbrowser-usechrome-agentchrome-extensioncomputer-usedatasetevaluationeveryday-tasksllmllm-evaluationonline-tasksreal-world-benchmarkweb-agentweb-agents

Embed the badge

Markdown for the project README. It links back here; terms on the terms page.

[![Tested by Argusic](https://argusic.com/badge/ClawBench.svg)](https://argusic.com/subject/clawbench)

Questions

Does ClawBench run?
ClawBench has not been fully verified yet. No recorded run has produced a verdict yet.
How did Argusic test ClawBench?
On a fresh, disposable machine, with every command recorded. 0 attempts are recorded, and the full method is on the methodology page.
Where is the evidence for ClawBench?
All 0 recorded runs are on this page, each linking to its full log and terminal recording, stored with a sha256 fingerprint so it cannot be quietly altered.

Discussion