Skip to content
FlowHubFluxonLab
W
Document AI & Extractionfree

Domain Specific Web Content Crawler with Depth Control & Text Extraction

by Le NguyenUpdated Aug 2026
Share Post Share
WeWebhookSILoop Links (Batches)Loop Links (Bat…IfIF Crawl Depth OK?IF Crawl Depth …HtExtract Body & LinksExtract Body & …CoAttach URL/Depth to HTMLAttach URL/Dept…HRFetch HTML PageMeSeed Root Crawl ItemSeed Root Crawl…CoCollect Pages & Emit When DoneCollect Pages &…SeStore Page DataMeMerge Web PagesCoCombine & ChunkRTRespond to WebhookRespond to Webh…CoInit GlobalsSeInit Crawl ParamsInit Crawl Para…CoRequeue Link ItemRequeue Link It…CoQueue & Dedup LinksQueue & Dedup L…1234567891011121314151617
1/5
STEPS · 17
Starts on an incoming request

On a webhook request, crawls a domain in batches with depth control, extracts page text from HTML, and returns it in the response.

Tags

webhookcomplexDocument Extractiondiscoveredpending-review
CategoryDocument AI & Extraction
Triggerwebhook
Complexitycomplex
Nodes16
AddedJun 27, 2026