didi-lot1-ai/ai_platform/modules/cloak/README.md

113 lines
4.1 KiB
Markdown
Raw Permalink Blame History

This file contains ambiguous Unicode characters

This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.

# cloak — stealth-Chromium scraping service
Standalone HTTP service that scrapes Google / Bing / DuckDuckGo SERPs through a
warm pool of CloakBrowser (patched stealth Chromium) instances.
Designed as **tier-3 search fallback** for the AI platform `web` module: when
SearXNG + the paid rotation return thin results (often the case for very niche
or recent queries), `cloak` provides results scraped directly from Google,
Bing and DuckDuckGo's HTML SERPs.
## Why a separate service?
- **Isolated lifecycle.** Browser pool restarts don't take down the rest of the
AI platform.
- **Bounded footprint.** A small fixed pool (default 3 instances ≈ 1.2 GB RAM)
versus N pools spreading across every web worker.
- **Same deployment pattern** as `audio`, `embeddings`, `video-analysis` etc.
## API
```
POST /v1/search
{
"queries": ["BNR confiscare conturi 10000 euro"],
"engines": ["google", "bing", "ddg"],
"max_results_per_engine": 10,
"language": "ro" // optional hint
}
200 OK
{
"results": [
{"url": "...", "title": "...", "snippet": "...",
"engine": "google", "query": "...", "rank": 1},
...
],
"stats": [
{"engine": "google", "query": "...", "results_count": 10,
"blocked": false, "captcha": false, "elapsed_ms": 2750, "error": null},
...
],
"total_elapsed_ms": 2900
}
GET /health
{
"status": "healthy" | "degraded" | "unhealthy",
"pool_size": 3, "pool_available": 3, "version": "0.1.0"
}
```
Auth is optional via `Authorization: Bearer <CLOAK_AUTH_TOKEN>`; when the env
var is unset (default), all requests are accepted (intra-cluster service —
should never be reachable from the internet).
## Configuration (env)
| Variable | Default | Notes |
|---|---|---|
| `CLOAK_HOST` | `0.0.0.0` | Bind address |
| `CLOAK_PORT` | `8770` | HTTP port |
| `CLOAK_POOL_SIZE` | `3` | Number of warm browsers (~400 MB each) |
| `CLOAK_SEARCH_TIMEOUT_SEC` | `20` | Hard timeout for entire `/v1/search` call |
| `CLOAK_PAGE_TIMEOUT_MS` | `18000` | Per-engine page load timeout |
| `CLOAK_MAX_ENGINES` | `3` | Cap on engines per request |
| `CLOAK_MAX_QUERIES` | `5` | Cap on queries per request |
| `CLOAK_DEFAULT_MAX_RESULTS` | `10` | Default per-engine result cap |
| `CLOAK_MAX_RESULTS_CAP` | `30` | Hard cap regardless of input |
| `CLOAK_ENGINE_MIN_INTERVAL_MS` | `200` | Throttle between successive scrapes per engine |
| `CLOAK_HUMANIZE` | `false` | Human-like mouse/keyboard timing (slower, better for behavioral anti-bot) |
| `CLOAK_AUTH_TOKEN` | `""` | Optional bearer token. Empty = no auth |
| `CLOAK_LOG_LEVEL` | `INFO` | `DEBUG`, `INFO`, `WARNING`, `ERROR` |
## Deploy
```bash
cd ai_platform/modules/cloak/deploy
docker compose up -d --build
docker logs -f didiAI-cloak
# Health
curl -s http://127.0.0.1:8770/health | jq
# Smoke
curl -sS -X POST http://127.0.0.1:8770/v1/search \
-H 'Content-Type: application/json' \
-d '{"queries":["NYTimes climate report 2026"],"engines":["google","bing","ddg"]}' | jq '.stats'
```
## Where it fits
```
agent-v3 / claims-verifier
didiAI-web-api ──► SearXNG (free, local) [tier 1 — always]
+ Brave/Tavily/etc rotation [tier 2 — one paid per call]
+ cloak (this service) [tier 3 — only if tier 1+2 thin]
```
The web module's orchestrator decides when to invoke `cloak` based on the
number of unique results returned from tiers 1+2. agent-v3 and didi-framework
do not call `cloak` directly.
## Operational notes
- **Selectors break.** Google rotates its result-DOM classes every 612 months.
The scraper has fallback selectors but the primary path will eventually need
re-tuning. Monitor `stats.blocked` / `stats.results_count` over time.
- **Rate limits.** No formal limit on the SERP endpoints, but bursts trigger
captcha. Default `CLOAK_ENGINE_MIN_INTERVAL_MS=200` paces requests; tune up
if you see captcha rates rise.
- **CPU/RAM.** Each browser instance uses ~400 MB RAM and is single-CPU for
most of a page load. The default `CLOAK_POOL_SIZE=3` is sized for didi12 ≤ 5k
scrape ops/day; raise to 58 if pool starves the request queue.