didi-lot1-ai/ai_platform/modules/cloak
2026-06-27 06:42:02 -07:00
..
deploy Livrare LOT 1 - Didi 2026-06-25 14:13:25 -07:00
src/cloak Livrare LOT 1 - Didi 2026-06-25 14:13:25 -07:00
tests Livrare LOT 1 - Didi 2026-06-25 14:13:25 -07:00
INDEX.md LOT 1 - Optimizare script build -Instalare mono comanda 2026-06-27 06:42:02 -07:00
pyproject.toml Livrare LOT 1 - Didi 2026-06-25 14:13:25 -07:00
README.md Livrare LOT 1 - Didi 2026-06-25 14:13:25 -07:00

cloak — stealth-Chromium scraping service

Standalone HTTP service that scrapes Google / Bing / DuckDuckGo SERPs through a warm pool of CloakBrowser (patched stealth Chromium) instances.

Designed as tier-3 search fallback for the AI platform web module: when SearXNG + the paid rotation return thin results (often the case for very niche or recent queries), cloak provides results scraped directly from Google, Bing and DuckDuckGo's HTML SERPs.

Why a separate service?

  • Isolated lifecycle. Browser pool restarts don't take down the rest of the AI platform.
  • Bounded footprint. A small fixed pool (default 3 instances ≈ 1.2 GB RAM) versus N pools spreading across every web worker.
  • Same deployment pattern as audio, embeddings, video-analysis etc.

API

POST /v1/search
{
  "queries": ["BNR confiscare conturi 10000 euro"],
  "engines": ["google", "bing", "ddg"],
  "max_results_per_engine": 10,
  "language": "ro"            // optional hint
}

200 OK
{
  "results": [
    {"url": "...", "title": "...", "snippet": "...",
     "engine": "google", "query": "...", "rank": 1},
    ...
  ],
  "stats": [
    {"engine": "google", "query": "...", "results_count": 10,
     "blocked": false, "captcha": false, "elapsed_ms": 2750, "error": null},
    ...
  ],
  "total_elapsed_ms": 2900
}

GET /health
{
  "status": "healthy" | "degraded" | "unhealthy",
  "pool_size": 3, "pool_available": 3, "version": "0.1.0"
}

Auth is optional via Authorization: Bearer <CLOAK_AUTH_TOKEN>; when the env var is unset (default), all requests are accepted (intra-cluster service — should never be reachable from the internet).

Configuration (env)

Variable Default Notes
CLOAK_HOST 0.0.0.0 Bind address
CLOAK_PORT 8770 HTTP port
CLOAK_POOL_SIZE 3 Number of warm browsers (~400 MB each)
CLOAK_SEARCH_TIMEOUT_SEC 20 Hard timeout for entire /v1/search call
CLOAK_PAGE_TIMEOUT_MS 18000 Per-engine page load timeout
CLOAK_MAX_ENGINES 3 Cap on engines per request
CLOAK_MAX_QUERIES 5 Cap on queries per request
CLOAK_DEFAULT_MAX_RESULTS 10 Default per-engine result cap
CLOAK_MAX_RESULTS_CAP 30 Hard cap regardless of input
CLOAK_ENGINE_MIN_INTERVAL_MS 200 Throttle between successive scrapes per engine
CLOAK_HUMANIZE false Human-like mouse/keyboard timing (slower, better for behavioral anti-bot)
CLOAK_AUTH_TOKEN "" Optional bearer token. Empty = no auth
CLOAK_LOG_LEVEL INFO DEBUG, INFO, WARNING, ERROR

Deploy

cd ai_platform/modules/cloak/deploy
docker compose up -d --build
docker logs -f didiAI-cloak
# Health
curl -s http://127.0.0.1:8770/health | jq
# Smoke
curl -sS -X POST http://127.0.0.1:8770/v1/search \
  -H 'Content-Type: application/json' \
  -d '{"queries":["NYTimes climate report 2026"],"engines":["google","bing","ddg"]}' | jq '.stats'

Where it fits

agent-v3 / claims-verifier
        │
        ▼
didiAI-web-api  ──► SearXNG (free, local)        [tier 1 — always]
                   + Brave/Tavily/etc rotation   [tier 2 — one paid per call]
                   + cloak (this service)         [tier 3 — only if tier 1+2 thin]

The web module's orchestrator decides when to invoke cloak based on the number of unique results returned from tiers 1+2. agent-v3 and didi-framework do not call cloak directly.

Operational notes

  • Selectors break. Google rotates its result-DOM classes every 612 months. The scraper has fallback selectors but the primary path will eventually need re-tuning. Monitor stats.blocked / stats.results_count over time.
  • Rate limits. No formal limit on the SERP endpoints, but bursts trigger captcha. Default CLOAK_ENGINE_MIN_INTERVAL_MS=200 paces requests; tune up if you see captcha rates rise.
  • CPU/RAM. Each browser instance uses ~400 MB RAM and is single-CPU for most of a page load. The default CLOAK_POOL_SIZE=3 is sized for didi12 ≤ 5k scrape ops/day; raise to 58 if pool starves the request queue.