Skip to content

Web Crawler ​

The web crawler maps a target by following links and analyzing pages. Every crawl request — regardless of engine — appears in the activity log the same way Nikto scan traffic does.

You choose how each crawl fetches pages:

  • Standard — fast, browser-free HTTP link extraction (HTML parsing only).
  • Headless Browser — a real Chromium session that executes JavaScript, like a user.
  • Both — two passes in one scan: Standard first for a reliable static floor, Headless for JS-heavy coverage when it finishes.

Use crawls to discover paths for the Sitemap, detect technologies and outdated JavaScript libraries, capture a seed-page screenshot (headless only), and surface recommendations for follow-up work. Crawls complement Nikto's module checks; run one alone, together with Nikto/Bustah, or in any combination from the New Scan wizard.

Headless browser is optional

Standard crawls need no Chromium sidecar. Headless Browser crawls require the headless Docker compose profile (see Docker deployment). When Chromium is down, the UI disables only the Headless Browser engine — Standard still runs, and the panel shows a notice explaining the situation.

The platform refuses a Headless crawl if Chromium's renderer sandbox is not confirmed on. That is a deployment problem, not a target problem — see Docker → Hardening.

Where to launch ​

Entry pointScope
Tools → CrawlHosts this tool has scanned; New Scan opens the wizard on the Crawl tab
New Scan wizard → Crawl tabBulk (many hosts) or from Scans / Hosts
Host Crawl tab → New ScanSingle host, Crawl options only
New Scan wizard → Crawl + Nikto tabsBoth tools in one launch

Every route reaches the same wizard. The button to press depends on where you are:

  • Tools → Crawl — New Scan
  • Hosts list — New Scan/Import, or select hosts and press Scan
  • Project Scans page — New Scan
  • A host's Crawl or Nikto tab — New Scan

See Scans → New Scan wizard.

Start Scan (N scan(s)) creates one crawl scan per target when Crawl is enabled. In bulk mode there is no shared seed — each host is crawled from its own root URL (Each target is crawled from its own root — no shared seed needed for bulk crawls.).

Single-host mode shows a Seed field. It is prefilled from the target's own scheme and port — http://host/ for a host configured without SSL, and a non-default port is kept — so an HTTP-only host is not crawled against port 443. Edit it if you want to start somewhere other than the root.

Crawl engines ​

The Engine checkboxes control which fetch strategy runs. Both engines share the same seed, scope regexes, limits, proxy, and activity logging — they differ only in what drives the HTTP client.

EngineWhat it does
StandardBrowser-free HTTP crawl: parse HTML (and similar) for <a href>, forms, and other static link sources. Fast, low resource use, no Chromium slot.
Headless BrowserHeadless Chromium navigates pages, executes JavaScript, and discovers links the static parser would miss (SPAs, client-side routers, lazy-loaded menus). Slower, needs the headless sidecar.

Select one or both. You must keep at least one checked while Enable is on.

What each engine gives you ​

CapabilityStandardHeadless Browser
Link discovery from static HTML✅✅
Links rendered only by JavaScript (React/Vue/Angular SPAs, etc.)❌✅
Seed-page screenshot on host Overview and Hosts Tile view❌✅
Live JavaScript library version probe (runtime versions, not just filenames)❌✅
Page Load Wait / Skip Images options❌ (ignored)✅
Runs when Chromium sidecar is down✅❌ (checkbox disabled)
Uses a browser slot❌✅

Both engines still run technology detection, outdated JS library checks (from fetched script bodies), secret detection, scope filtering, and sitemap contribution for every page they actually fetch. The gap is which pages and scripts get fetched — not whether analysis runs on them.

When to choose Standard ​

Pick Standard alone when:

  • You want a quick map of a mostly server-rendered site (classic HTML, PHP, Jinja, etc.) without paying for a browser.
  • Chromium is unavailable (no headless profile, sidecar restarting, air-gapped install) but you still need crawl coverage.
  • The target is sensitive to browser automation and you prefer a plain HTTP client signature (still subject to your User-Agent and headers).
  • You are crawling many hosts in bulk and need predictable, lightweight runs.

Do not pick Standard alone when the app's navigation lives in JavaScript. You see the shell page and miss routes that appear only after JavaScript runs.

When to choose Headless Browser ​

Pick Headless Browser alone when:

  • The target is a SPA or heavy JS site and static HTML barely lists real routes.
  • You need the screenshot on the host Overview tab or Hosts Tile view.
  • You care about runtime library versions (the live probe reads what's actually loaded in the page, not just a <script src="…/jquery-1.2.3.min.js"> filename).
  • You can afford slower, heavier crawls and have Chromium running.

Do not pick Headless Browser alone when you only need a fast link inventory of a server-rendered blog or admin panel. Standard finishes sooner with fewer moving parts.

When to choose both ​

Checking both launches two engine passes inside one crawl scan (~2× requests vs a single engine):

  1. Standard runs as a fast, browser-free pass — it reliably lands a broad static sitemap even if the headless pass is slow, paused, wedged, or stopped early.
  2. Headless Browser runs as a separate pass — it adds JS-discovered links, the screenshot, and live version probes when it completes.

Both passes merge into the same host sitemap and the same scan's findings and activity log. The crawl summary finding names every engine that contributed (for example standard, headless browser engine).

Choose both when Chromium is up and you want a static floor plus JavaScript coverage, without depending on one engine finishing.

Standard-only is a valid production choice

Operators sometimes run Standard only on every host and reserve Headless Browser (or both) for hosts that looked incomplete or JS-heavy after triage. The UI never auto-downgrades your selection — engine choice is always explicit.

Crawler options ​

Click the Crawler header to expand or collapse the panel (expanded by default in the New Scan wizard). When Enable is checked:

OptionDefaultWhat it does
EnableOnInclude a crawl scan in Start Scan. At least one engine must stay selected.
EngineHeadless Browser onlyChoose Standard, Headless Browser, or both. Default: Headless Browser only. See Crawl engines.
Seedthe target's own root URLStarting URL (single-host only). Must begin with http:// or https://. Prefilled from the target's scheme and port.
Max Pages150Stop after this many pages are fetched (hard ceiling 5000). Applies per engine pass.
Max Depth3Maximum link depth from the seed (0–10).
Duration (Seconds)600Time budget for the crawl (up to 3600 seconds).
In Scope(empty)Regex patterns a URL must match to be crawled — one per line. Blank = no extra restriction beyond hostname scope.
Out Of Scope(empty)Regex patterns a URL must not match — one per line. Applied after In Scope.
Skip ImagesoffHeadless pass only. Replace images with a 1×1 placeholder instead of downloading them. The first page still loads images so its screenshot remains usable.
Page Load Waitdefault (settled)Headless pass only. How long Chromium waits after each navigation (settled, load event + 500ms, DOM content loaded, or none).
Request Timeout (Seconds)blank (15 direct / 60 via proxy)Maximum time for a single request. Lower on targets that drop slow probes instead of refusing them.

The Crawl tab keeps the boolean toggles (Follow Redirects, Send Browser Headers, Secret Detection — on by default for crawls — Force IPv6) and Header Mode. Identity fields (User-Agent, Root Path, Additional Headers, and optionally Virtual Host) are on the wizard's Options tab. Proxy is on Throttling & Limits — an enforced project proxy is shown locked and always wins; both engines honor it. See Scan options for Header Mode and the rest of the shared fields.

Virtual Host on crawl ​

Virtual Host on the shared Options tab behaves differently by crawl engine:

Engine selectionVirtual Host
Standard onlyFully applied — Host header, TLS SNI, dialed IP, and scope all bind to the vhost. Use this to crawl a site on a shared server while connecting to the target's real address.
Headless Browser or bothNot applied — launch is blocked if a vhost is set. The headless browser navigates the target's real hostname; cookie jar, same-origin policy, scope matching, and SNI all stay on that host. A Host-header-only override would not produce a true vhost crawl.

To crawl by hostname instead, point the seed URL at that hostname directly.

Enforced project Host headers still reach crawl wire traffic (the connection target and SNI stay on the real host — a curl --resolve-style split). That is a compliance override of the header only, not a crawl "of a vhost".

Hostname scope ​

By default the crawler stays on the exact hostname of the seed URL — it does not automatically expand to the whole registered domain (for example www.example.com does not crawl example.com or sibling subdomains unless you follow links that stay in scope). Use In Scope / Out Of Scope regexes to tighten or widen what gets fetched.

Invalid scope regexes are flagged inline and block launch until fixed.

The console rechecks browser availability every 30 seconds. If Chromium comes up after you open the page, the Headless Browser engine re-enables without a reload (Standard was never blocked).

Host Crawl tab ​

Each host has a dedicated Crawl tab (alongside Nikto, LFIC, etc.):

  • Crawl Scans — list of crawl runs for this host with New Scan, refresh, and row actions. Click a completed scan to open its findings.
  • Detected Technologies — table (Product, URL, Version) from the latest completed crawl. Copy table exports markdown. Empty until a crawl finishes.

Discovered crawl paths also appear on the host Sitemap tab.

Parent-directory indexing probe ​

After a crawl finishes, the platform may queue follow-up probes for parent directories of pages the crawl already fetched — for example checking whether /app/docs/ lists when the crawl visited /app/docs/readme.html. Each distinct in-scope parent is probed at most once; directories the crawl already fetched are skipped. Probes respect scope and share the crawl's proxy/settings. Results appear as normal Nikto/crawl findings and in the activity log.

If a crawl ends completed but with partial coverage — it stopped before finishing, after fetching some pages — the scan list shows an orange degraded badge with that reason; hover for the full text. A crawl that fetched no pages is not reported as partial coverage: it is retried, and a persistent failure shows as failed naming the transport reason (wrong scheme, connection refused, DNS). See Scans → Degraded scans.

Seed-page screenshot ​

After a Headless Browser crawl, the host Overview tab may show a Screenshot card on the right — an above-the-fold JPEG of the seed page captured by Chromium. Hosts → Tile view shows the same shot (or an empty frame labeled No screenshot when none exists). Standard-only crawls do not produce a screenshot.

  • On Overview, the card appears only when a screenshot exists (no placeholder if the host was never crawled). Tile view always shows a frame.
  • Click to enlarge on Overview opens a full-screen lightbox; press Escape or click outside to close. Clicking a tile screenshot opens host detail.
  • Re-crawling replaces the previous screenshot (one per host).
  • Navigate away and back to Overview to pick up a fresh shot after a new crawl.
  • If the only response from the seed page is a proxy error page (for example Burp unreachable), the crawl is marked failed and no screenshot is stored — the run did not reach the real target. Error pages the platform's own proxy generates are also excluded from analysis, so they never produce findings attributed to the target.

What gets photographed ​

The screenshot is always the seed page — the URL the crawl started from — and only when that page returns HTML. Pages found during the crawl, and non-HTML responses such as a .js or .css file, are never captured. A seed page that cannot be photographed leaves the host with no screenshot; the platform does not substitute another page, because a picture of the wrong page is indistinguishable from a correct one.

Chromium waits for the page's load event and re-takes the capture while the result is still visually blank, so a single-page app that fetches its content after load is photographed rendered rather than mid-spinner. The wait is bounded by what is left of the seed page's request timeout: on a very slow page the crawl keeps the page and its discovered links, and the screenshot may be blank or absent.

Where the picture came from ​

The enlarged view carries a Captured from bar along the bottom naming the page the image is of. Use it to confirm a screenshot depicts the page you expected.

  • The value is shown as plain text, not a link. It comes from the scanned target, so the console never navigates to it — copy it if you want to open it elsewhere.
  • Screenshots captured before the platform recorded this detail show unknown. That means the page was not recorded, not that it was the home page.

What a crawl produces ​

Findings ​

Crawl scans emit content-analysis findings as pages are processed — detected technologies, outdated JavaScript libraries, and similar. When the crawl finishes, an info finding titled Crawl summary: N pages crawled (… engine) summarizes pages fetched, which engine(s) ran, technologies detected, and outdated JS counts.

Triage on project Findings, the host Findings tab, the project Scans list (click a crawl row), or the host Crawl tab scan list. See Findings.

Activity log ​

Every HTTP request the crawl makes is logged to the target's activity log, regardless of engine. Use Logs in the project sidebar to audit crawled traffic alongside Nikto requests.

Recommendations ​

Some crawl analysis produces recommendations — proposed follow-up actions derived from page content. An app is reported from same-origin links in the page, not from paths the scanner itself requested. Examples:

  • WordPress detected → WPScan or WPProbe
  • Joomla detected → joomscan; Drupal detected → droopescan; Magento detected → a Magento security scanner
  • WebSocket in use → manual review via an interactive proxy (Burp Suite or Caido); automated scanning does not exercise WebSocket message flows

Review them on the project Recommendations sidebar page or the host Recommendations tab. Run is disabled for these crawl suggestions (Running recommendations is not available yet). Dismiss, Undismiss, and Delete work the same as on other recommendation rows.

Open listings of S3 / Azure Blob containers raise a different recommendation — Export full file list — which does run from the UI. See Recommendations and Cloud Storage Listings.

Control and monitor ​

Running crawl scans support Cancel and Delete from the host Crawl tab, project Scans list, Tools → Crawl, and Active Jobs. Click a crawl row anywhere to open scan detail (progress, Speed limit, activity, tasks). Pause does not apply to crawls — a crawl is one long-lived task, so Pause All skips them and tells you so. Use Cancel to stop a running crawl now. See Scans → Control a running scan.

Hostname lookup happens at scan creation. An unresolvable seed or target is rejected with a message that names the host — the crawl is never queued. Followed links stay inside the target's address tier.

Next steps ​

Proprietary software. Licensed for use under the End User License Agreement.