llm-evaluation
Tested projects carrying the llm-evaluation topic, with their current verdicts. Topics come from each project's own metadata and Argusic's pool tags. Spellings of the same topic are gathered here under one name, so this page holds every project tagged any of them.
Measured by Argusic on a fresh machine every time. Every number links to its evidence.
What Argusic measured in llm-evaluation
Argusic installed and ran 1 llm-evaluation project on clean machines: 1 only ran against stand-in services. Last tested September 3, 2026.
Every tested llm-evaluation project
| subject | verdict | language | runs | last tested |
|---|---|---|---|---|
| ouroboros | Runs with mocks | Python | 3 |
All tested projects: subjects.