extractor

Use LLMs to robustly extract web data

Not yet testedsource: GitHubhomepageTypeScriptApache-2.0commit c9edd2bf363e

TypeScript, Apache-2.0 licensed. The project labels itself: ai agents, article extractor, crawler, data engineering, data pipeline, ecommerce scraping, etl and google gemini.

extractor has not been verified yet.

Measured by Argusic on a fresh machine every time. Every number links to its evidence.

At a glance

verdict
Not yet tested
Argusic Score
not scored yet
recorded runs
0
last tested
-
stars
321
forks
10
open issues
2
watchers
0
size
0 MB
created
last push

Subject data from GitHub, linked at the top of this page, refreshed . Test data by Argusic (CC BY 4.0); every number links to a run page with the full log, the recording, and their sha256 hashes.

Also tested, in the same area

Every one of these was installed and run by Argusic on a clean machine. Nothing appears here that was not tested.

Run history

No recorded runs.

Topics (from GitHub)

ai-agentsarticle-extractorcrawlerdata-engineeringdata-pipelineecommerce-scrapingetlgoogle-geminihtml-parserhtml-to-markdownllmllm-extractionllm-scrapermarkdownnlpopenairagweb-data-extractionweb-extractionwebscraping

Embed the badge

Markdown for the project README. It links back here; terms on the terms page.

[![Tested by Argusic](https://argusic.com/badge/extractor.svg)](https://argusic.com/subject/extractor)

Questions

Does extractor run?
extractor has not been fully verified yet. No recorded run has produced a verdict yet.
How did Argusic test extractor?
On a fresh, disposable machine, with every command recorded. 0 attempts are recorded, and the full method is on the methodology page.
Where is the evidence for extractor?
All 0 recorded runs are on this page, each linking to its full log and terminal recording, stored with a sha256 fingerprint so it cannot be quietly altered.

Discussion