headless crawlers work the crypto web in real browsers: whitepapers, docs, governance, code. every page they pull is cleaned, deduped and tokenized into queen's dataset.
Crawlnet says it crawls crypto-related web content with headless browsers, processes pages into a dataset, and uses the dataset to pretrain Queen from scratch.
Project documentation describes the workflow, but no independent crawl-to-training test or inspection establishes execution.
To settle it: Inspect live crawler state and crawl reports, and compare dataset/model artifacts and training records.
Skeptic: plausible
The published repository describes a small model trained from scratch and documents a dataset containing crawler pages, which supports the basic crawler-to-model workflow. The stronger site wording that the dataset consists of crawler material alone is not accurate as written; see the data-source discrepancy below. [E6] [E10]