One link.
Web data, ready.
Readable Markdown and checked fields from public pages, ready for an AI agent, a RAG pipeline or a citation. When a page can’t be read, you get the reason.
Readable Markdown and checked fields from public pages, ready for an AI agent, a RAG pipeline or a citation. When a page can’t be read, you get the reason.
YOUR RESULTS
This visit only · cleared when you leave the page
Use it in your agent: Connect MCP · Run it yourself
SELECTED RUN
RECORDED RESULTS
Readable Markdown, product fields checked against the page, and a refusal with its reason. Three results OctoCrawl recorded, shown as they were.
936fdf0.B000NI69YA, recorded 23 Sep 2026 on a local OctoCrawl at source commit 991097f. Each value was matched against the captured page and signed off in a 100-product review, where 99 of 100 products came back complete and the one that did not withheld its fields. Amazon.sg support is in Beta.936fdf0. OctoCrawl stopped at the robots.txt check, before fetching the page, and returned no content instead of an empty page. At that commit the same reason was also given when robots.txt could not be read, and this record does not say which; OctoCrawl now reports the two apart.HOW IT WORKS
No install or sign-up. Paste any public http(s) address; you get 3 previews a day.
It respects robots.txt and reads only what anyone can open, then reports the status, final URL and time.
Copy or download readable Markdown or the result JSON. Amazon.sg product pages add checked fields.
WHY OCTOCRAWL
A crawler that only checks for a response can hand your agent a login wall, a challenge page or an empty shell as if it were the page. OctoCrawl reports what it actually read.
Fields you ask for are read from the page’s own JSON-LD, microdata, meta tags and tables, without a model, and each names where it came from. A field the page doesn’t state stays empty, with the reason.
A page OctoCrawl can’t read comes back blocked, incomplete, timed out or failed, with a reason and a diagnostic code. The preview reads robots.txt first, and never signs in or solves a CAPTCHA.
AGPL-3.0. Run it yourself with no daily limit, through your own network and proxy (HTTPS_PROXY), with a local Chromium for pages that need a browser.
Measured on our own test sets
blocked.Local API · main@6024703 · 30 Sep 2026Our own suites, not an independent audit. Method, and how other tools did on the same suite
RUN IT YOURSELF
Clone the repository, start OctoCrawl on your computer, and call it from your agent through MCP, from your code through REST or the SDK, or from a Firecrawl v1 client. Results stay on your machine.
Node.js 22.13 or later. The API listens on this computer only.
Local only: the service listens on 127.0.0.1. Verified with Codex; setups for Claude Code, Cursor and OpenCode are in the guide.
Local only. @w2l/sdk lives in the repository; it is not on npm yet.
Partial and local only: Firecrawl v1 scrape and crawl, Markdown and links. No search, map, extract or v2. A migration aid, not a full compatibility layer.
WHAT WORKS TODAY
What you can use now, on this page and on your own computer, and what the roadmap has next or on hold.
HTTPS_PROXYnpxmap and a maxAge cacheNot planned: stealth, fingerprint spoofing or proxy pools; getting past logins or CAPTCHAs; scraping sales leads. Roadmap
FAQ
We can’t give legal advice; here is what the preview does. It reads a site’s robots.txt before it fetches a page. If the page is disallowed, or robots.txt can’t be reached (a server error, no answer or a timeout), it stops and reports the page as blocked; a robots.txt that answers with a 4xx status counts as no rules, as RFC 9309 provides. It never signs in, solves a CAPTCHA or gets past a verification page, and it refuses private network addresses. What you do with a page is up to you and the site’s terms.
Public pages anyone can open without signing in. The preview reads them over HTTP without running JavaScript, so a page that only appears in a browser may come back incomplete. It reads pages up to 2 MiB and files such as PDFs up to 5 MiB, and stops after 40 seconds. Amazon.sg product pages (/dp/ASIN) are in Beta. X and Reddit posts often don’t come through: robots rules, sign-in walls or verification pages can stop the preview, and a hosted X or Reddit result hasn’t been verified yet. See Limits and result states.
The preview is a limited public trial: three previews per visitor and 100 for the whole site each UTC day. A request turned down before a preview starts, such as a malformed URL, or localhost or a private IP address typed into it, doesn’t count. Once a preview starts it counts, whatever the result, including a host name that turns out to point to a private network or a page stopped by robots.txt. OctoCrawl on your own computer has no daily limit.
Results aren’t saved: your recent runs live in this page and are gone when you leave it. Each preview logs its state and the host of the page, never its path. While you type, the page asks the service for a short hint about the address, so that address appears in our hosting provider’s request log, kept for 30 days. The details are on the Privacy page.
OctoCrawl reports what it actually read: a blocked, incomplete or timed-out page is a result with a reason, and checked fields carry their source. The benchmark notes compare the three tools on the same test suite, with the limits of that comparison. For moving off Firecrawl, OctoCrawl has a partial, local Firecrawl v1 shim.
Please don’t: the preview is for trying OctoCrawl in a browser. To automate, run OctoCrawl yourself and use REST, the SDK or MCP.
The preview is free and needs no account. OctoCrawl is open source under the AGPL-3.0, and running it yourself costs nothing but your own machine. There is no paid plan today.