Enterprise Geo-Business Discovery & Data Collection Pipeline
A modular Node.js and Playwright workflow that transforms publicly available Google Maps business results into structured CSV datasets.
The pipeline separates discovery, URL batching, duplicate removal, business-detail extraction, retries, and CSV generation into independent stages for easier maintenance and scaling.
It was validated across more than 10,000 business records from multiple US and UK regions with manual checks on extracted metadata.
What the project demonstrates.
Validated on 10,000+ business records across US and UK regions
Parallel extraction, retries, batching, and duplicate removal
~98% address extraction accuracy on manual verification
95%+ phone and region extraction accuracy on manual verification
How the work was approached.
The challenge
Collect inconsistent public business listings at scale and turn them into deduplicated, structured records.
Technical approach
- Separate discovery, batching, detail extraction, retry handling, deduplication and CSV export.
- Manually inspect data-quality fields to estimate extraction accuracy.
Evidence & outcomes
- 10,000+ records collected across US and UK regions in project testing.
- Approximately 98% address and 95%+ phone/region extraction accuracy reported from manual checks.
Scope note: Accuracy figures refer to the project’s described evaluation sample, not all businesses on the web.
Want to inspect the implementation? Explore the source repository ↗