Skip to content

About

No description, website, or topics provided.

Resources

Stars

0 stars

Watchers

0 watching

Forks

Latest commit

 

History

1 Commit

Folders and files

Repository files navigation

🕷️ Intelligent Web Scraping Engine

Node.js JavaScript Puppeteer Google Sheets API

A robust, enterprise-grade web scraping engine designed to monitor real estate competitors in real-time. Built with Node.js and Puppeteer Stealth, it automatically bypasses Web Application Firewalls (WAF, like Cloudflare) using user-agent rotation and exports parsed results directly to a Google Sheets spreadsheet using a Google Service Account.


📈 Business Case & Real-World Impact

Manually tracking real estate prices and competitor availability is a slow and error-prone process. Bypassing modern security firewalls (WAFs) programmatically is challenging, as traditional automated bots get blocked instantly by Captchas and Cloudflare checks.

The Solution:

I architected and developed a stealthy crawler:

  • WAF Bypassing: Utilized puppeteer-extra-plugin-stealth coupled with a randomized User-Agent rotator and custom HTTP headers to simulate human-browser signatures.
  • Dynamic Data Parsing: Targeted complex SPA selectors using asynchronous wait structures, extracting addresses, pricing, property size, and structural metadata.
  • Automatic Sheets Export: Integrated the Google Sheets API via service account credentials, pushing parsed data batches to a spreadsheet automatically at runtime.
  • Web Dashboard Panel: Created a simple HTML/CSS dashboard monitor to visualize the extraction tasks, showing progress, logs, and a direct Google Sheet link.

Results:

  • Speed: Automatically extracts 20+ detailed listings in under 35 seconds.
  • 100% Success Rate: Successfully bypassed antibot layers during running cycles.
  • Data-Driven Advantage: Enables immediate identification of properties listing below average market value, generating actionable purchasing leads.

📸 Screenshots

Here is a preview of the monitoring dashboard and data outputs:

Web Scraper Monitor Dashboard Data Output (Google Sheets)
Scraper Monitor Spreadsheet Output

🏗️ Technical Architecture & Key Logic

1. Stealth Browser Initialization

Enabling plugins to hide Selenium/Puppeteer footprints:

const puppeteer = require('puppeteer-extra');
const StealthPlugin = require('puppeteer-extra-plugin-stealth');

puppeteer.use(StealthPlugin());

async function runStealthBrowser() {
  const browser = await puppeteer.launch({
    headless: true,
    args: [
      '--no-sandbox',
      '--disable-setuid-sandbox',
      '--disable-web-security',
      '--window-size=1920,1080'
    ]
  });
  const page = await browser.newPage();
  
  // Set custom user-agent and language headers
  await page.setUserAgent("Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/120.0.0.0 Safari/537.36");
  await page.setExtraHTTPHeaders({ 'Accept-Language': 'pt-BR,pt;q=0.9,en-US;q=0.8,en;q=0.7' });
  
  return { browser, page };
}

2. Service Account Google Sheets Sync

Securely uploading arrays via JWT auth without requesting client login:

const { google } = require('googleapis');

async function syncToSheets(dataRows) {
  const auth = new google.auth.GoogleAuth({
    keyFile: 'credentials.json',
    scopes: ['https://www.googleapis.com/auth/spreadsheets'],
  });
  
  const sheets = google.sheets({ version: 'v4', auth });
  const spreadsheetId = process.env.SPREADSHEET_ID;
  
  await sheets.spreadsheets.values.append({
    spreadsheetId,
    range: 'RawData!A:G',
    valueInputOption: 'USER_ENTERED',
    resource: { values: dataRows }
  });
}

⚙️ Setup & Local Installation

Prerequisites

  • Node.js (v18 or v20+)
  • A Google Cloud Project Service Account with credentials.json downloaded.
  • A target spreadsheet shared with your service account email.

Installation Steps

  1. Clone the repository and install dependencies:

    npm install
  2. Create a .env file from the template:

    cp .env.example .env
  3. Populate the env keys with your spreadsheet information:

    SPREADSHEET_ID="your-google-sheets-spreadsheet-id"
    PORT=4000
  4. Place your Google Cloud service account JSON file in the root of the project and rename it to credentials.json.

  5. Start the engine and local dashboard monitor:

    npm start
  6. Open your browser to http://localhost:4000 to interact with the scraper dashboard.


🇧🇷 Resumo em Português

O Motor de Web Scraping Inteligente é um serviço em Node.js projetado para extração automatizada de dados concorrentes de alta complexidade.

  • Problema: Bloqueios de WAF (Cloudflare/Captchas) e tarefas manuais demoradas para monitorar anúncios e preços imobiliários concorrentes.
  • Solução: Crawler baseado em Puppeteer Stealth que emula comportamentos e assinaturas de navegadores humanos reais, integrado à API do Google Sheets via Service Account.
  • Resultados: Extração e upload automático de mais de 20 imóveis em 35 segundos, taxa de sucesso de 100% contra bloqueios e monitoramento ativo do mercado.

Developed with ❤️ by João Melo

About

No description, website, or topics provided.

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages