Argusic does-it-run: verdicts for open source projects
Argusic agents install, launch, debug and verify real open source software on a clean machine, with no human help. This dataset lists every tested project with its verdict (runs, runs with mocks, could not verify, or not yet verified), the Argusic Score, the pinned environment and what the machine printed when a step failed. It covers 1,313 projects and 2,236 valid runs from 2026-08-30 to 2026-10-03, is generated from the Argusic database and is updated daily. Licensed CC BY 4.0: attribute Argusic, argusic.com.
At a glance
- 1,313 tested projects, 2,236 valid runs.
- First run 2026-08-30, latest run 2026-10-03.
- runs: 967
- runs with mocks: 258
- could not verify: 85
- not yet verified: 3
How it was measured
Argusic agents install, launch, debug and verify each project on a clean machine and the result is verified by decoding it, not by probing. The rules are public: methodology (version 1.4). Every verdict has a run page with the full log and the terminal recording.
What is in the table
- verdict: runs, runs with mocks, could not verify, or not yet verified.
- argusic_score: Mean of the recorded run scores on the public 0 to 100 rubric; a timeout is not scored.
- valid_runs: Number of valid runs of the project.
- tested_on: Date of the latest valid run.
- install_minutes and wall_minutes: Minutes spent installing and total wall time of a run.
- errors observed: What the clean machine printed when a step failed, as the observation only.
- test depth: Real run, run with mocked services, or no run possible.
Download
- verdicts.csv
- verdicts.json
- Per-project files and failure patterns: the repository on GitHub.
License and citation
The data is licensed CC BY 4.0. Attribution: Argusic, argusic.com. Names, descriptions and licenses of the listed projects belong to their authors.
Disputes
If a result looks wrong, the methodology explains how to contest it (section 19). The run page shows the exact line to point at.