didi-lot1-ai/ai_platform/modules/web/README.md

321 lines
12 KiB
Markdown

# Web Module
Unified web module for search, fetch, browse, vision, and evidence gathering. Provides REST APIs for web content extraction with automatic fallback between methods.
## What It Does
This module is designed for **fact-checking pipelines**. Given a claim, it:
1. **Searches** the web for relevant sources (SearXNG metasearch)
2. **Fetches** page content using the best method available
3. **Extracts** relevant evidence snippets using LLM
4. Returns structured evidence for verification
## Prerequisites
**Required:**
- All global prerequisites (see main [README.md](../../README.md))
- SearXNG instance. The delivered deployment runs a **single** SearXNG container (`didiAI-web-searxng`, internal `:8080`) wired via `local_gpu_stack`. The `deploy/metasearch` cluster (multiple replicas) is an **optional** scaling variant, not the default.
- LLM Inference server (llm-inference module at port 14011)
**Optional:**
- OpenAI API key (for vision/LLM fallback)
- Anthropic API key (for vision/LLM fallback)
## Features
| Component | Description |
|-----------|-------------|
| **Search** | SearXNG metasearch client with site filtering, freshness, language |
| **Fetch** | HTTP content extraction with readability-lxml |
| **Browse** | Playwright-based JavaScript rendering for dynamic pages |
| **Vision** | Screenshot + Vision LLM for complex/protected pages |
| **Evidence** | Deduplicate + LLM snippet extraction + relevance scoring |
| **Orchestrator** | Auto-fallback chain: HTTP → Playwright → Vision LLM |
### Smart Features
- **PDF Filtering**: Automatically skips direct PDF URLs (can't extract text)
- **Thinking Mode Handling**: Strips `<think>` tags from reasoning models (Qwen3)
- **Auto-Fallback**: Escalates to more powerful methods when content is poor
- **Deduplication**: Removes duplicate content based on text similarity
## Installation
```bash
cd modules/web
# Install base dependencies
uv sync
# Install with all extras (fetch, browse, vision)
uv sync --all-extras
# Install Playwright browsers
uv run playwright install chromium
# Install with dev dependencies
uv sync --extra dev --all-extras
```
## Quick Start
### Docker Deployment (Recommended)
```bash
cd deploy/
# Configure environment
cat > .env << 'EOF'
WEB_SEARXNG_BASE_URL=http://didiAI-web-searxng:8080
WEB_LLM_BASE_URL=http://didiAI-llm-api:14011
WEB_LLM_API_KEY=your-llm-api-key
WEB_VISION_MODEL=qwen-vl
WEB_TEXT_MODEL=qwen3-235b
WEB_LOG_LEVEL=INFO
EOF
# Start the server
docker compose --profile api up -d
# Check health
curl http://localhost:51100/health
```
### Example: Fact-Check a Claim
```bash
curl -X POST http://localhost:51100/v1/gather \
-H "Content-Type: application/json" \
-d '{
"claim": "Romania had the highest economic growth in the EU in 2024",
"search_queries": ["Romania GDP growth 2024 EU"],
"max_search_results": 5,
"extract_snippets": true,
"max_evidence_items": 5
}'
```
**Response:**
```json
{
"claim": "Romania had the highest economic growth in the EU in 2024",
"evidence": [
{
"url": "https://en.wikipedia.org/wiki/Economy_of_Romania",
"title": "Economy of Romania - Wikipedia",
"snippet": "Romania's nominal GDP reached approximately $423 billion in 2024, reflecting real growth of 0.9% that year...",
"relevance_score": 0.6
}
],
"stages": [
{"stage": "search", "success": true, "items_processed": 5},
{"stage": "fetch", "success": true, "items_processed": 4},
{"stage": "evidence", "success": true, "items_processed": 3}
],
"execution_time_ms": 40643.68
}
```
## API Endpoints
| Endpoint | Method | Description |
|----------|--------|-------------|
| `/v1/gather` | POST | **Main endpoint** - unified pipeline |
| `/v1/search` | POST | Execute web search only |
| `/v1/image-search` | POST | Search for images |
| `/v1/fetch` | POST | Fetch URLs (HTTP + readability) |
| `/health` | GET | Health check |
| `/ready` | GET | Readiness probe |
### Gather Request Schema
```json
{
"claim": "The claim to verify",
"search_queries": ["optional", "custom", "queries"],
"max_search_results": 10,
"site_allowlist": ["reuters.com", "bbc.com"],
"site_blocklist": ["spam-site.com"],
"fetch_method": "auto",
"auto_fallback": true,
"extract_snippets": false,
"max_evidence_items": 15,
"dedupe": true
}
```
| Field | Type | Default | Description |
|-------|------|---------|-------------|
| `claim` | string | required | The claim to gather evidence for |
| `search_queries` | array | null | Custom search queries (auto-generated if not provided) |
| `max_search_results` | int | 10 | Max search results (5-50) |
| `site_allowlist` | array | null | Only search these domains |
| `site_blocklist` | array | null | Exclude these domains |
| `fetch_method` | string | "auto" | "auto", "http", "browse", "vision" |
| `auto_fallback` | bool | true | Escalate on fetch failure |
| `extract_snippets` | bool | False | Deprecated — LLM snippet extraction (no longer used by default) |
| `max_evidence_items` | int | 15 | Max items in final pack |
| `dedupe` | bool | true | Deduplicate evidence |
## Configuration
Configure via environment variables (prefix: `WEB_`):
| Variable | Default | Description |
|----------|---------|-------------|
| `WEB_SEARXNG_BASE_URL` | - | **Required** SearXNG instance URL |
| `WEB_LLM_BASE_URL` | - | **Required** LLM inference server URL |
| `WEB_LLM_API_KEY` | - | LLM API authentication token |
| `WEB_VISION_MODEL` | `qwen-vl` | Vision model for screenshots |
| `WEB_TEXT_MODEL` | `qwen3-235b` | Text model for snippets |
| `WEB_PORT` | `51100` | API server port |
| `WEB_OPENAI_API_KEY` | - | OpenAI fallback (optional) |
| `WEB_ANTHROPIC_API_KEY` | - | Anthropic fallback (optional) |
See `.env.example` for full configuration options.
## Auto-Fallback Logic
The module uses intelligent fallback to get the best content:
```
┌─────────────────────────────────────────────────────────┐
│ URL Input │
└─────────────────┬───────────────────────────────────────┘
┌─────────────────────────────────────────────────────────┐
│ 1. PDF Check - Skip .pdf URLs (can't extract text) │
└─────────────────┬───────────────────────────────────────┘
┌─────────────────────────────────────────────────────────┐
│ 2. HTTP Fetch (fastest, cheapest) │
│ - Uses readability-lxml for content extraction │
│ - Detects JS-heavy pages and insufficient content │
└─────────────────┬───────────────────────────────────────┘
▼ if text < min_length or JS detected
┌─────────────────────────────────────────────────────────┐
│ 3. Playwright Browse (renders JavaScript) │
│ - Waits for networkidle │
│ - Extracts rendered content │
└─────────────────┬───────────────────────────────────────┘
▼ if content still poor quality
┌─────────────────────────────────────────────────────────┐
│ 4. Vision LLM (screenshot → extract) │
│ - Takes full-page screenshot │
│ - Uses vision model to extract text │
└─────────────────────────────────────────────────────────┘
```
Enable with `fetch_method: "auto"` and `auto_fallback: true`.
## Architecture
```
src/web/
├── __init__.py # Package exports
├── config.py # Unified WebSettings
├── exceptions.py # Custom exceptions
├── logging.py # Structured logging
├── orchestrator.py # Auto-fallback pipeline orchestration
├── cli.py # CLI entry point
├── schemas/ # All request/response schemas
│ ├── common.py # PageContent, FailedUrl, PageImage, shared types
│ ├── search.py # SearchRequest/Response
│ ├── fetch.py # FetchRequest/Response
│ ├── browse.py # BrowseRequest/Response
│ ├── vision.py # VisionExtractRequest/Response
│ ├── evidence.py # EvidencePackRequest/Response
│ └── gather.py # GatherRequest/Response (unified)
├── llm/ # LLM provider abstraction
│ ├── __init__.py
│ └── provider.py # LLMProviderChain (local→OpenAI→Anthropic)
├── search/ # Search providers
│ ├── __init__.py
│ └── protocol.py # SearchProvider Protocol
├── metasearch/ # SearXNG (free tier)
│ └── client.py # SearXNGClient
├── fetch/ # HTTP + readability extraction
│ └── client.py # FetchClient
├── browse/ # Playwright browser automation
│ └── client.py # BrowseClient
├── vision/ # Screenshot + Vision LLM
│ └── client.py # VisionClient (local + OpenAI + Anthropic)
├── evidence/ # Evidence processing
│ └── packer.py # EvidencePacker (dedupe, snippets, scoring)
└── api/
├── app.py # FastAPI app factory
├── dependencies.py # DI (auth, rate limiter)
├── middleware.py # RequestId, RateLimit
└── routes/
├── search.py # POST /v1/search
├── image_search.py # POST /v1/image-search
├── fetch.py # POST /v1/fetch
├── gather.py # POST /v1/gather (main endpoint)
├── info.py # GET /v1/info (catalog metadata)
└── health.py # /health, /ready
```
## Deployment
```bash
cd deploy/
# Configure environment (REQUIRED)
# Edit .env with your API keys
# Start server
docker compose --profile api up -d
# View logs
docker compose --profile api logs -f
# Stop server
docker compose --profile api down
```
### Port Allocation
| Port | Service |
|------|---------|
| 51100 | Web API |
### Network
The container joins the `didi-network` network to communicate with:
- `didiAI-llm-api:14011` - LLM inference server
## Development
```bash
# Install dev dependencies
uv sync --extra dev --all-extras
# Run tests
uv run pytest
# Run tests with coverage
uv run pytest --cov=src/web --cov-report=term-missing
# Lint and format
uv run ruff check .
uv run ruff format .
```
## Dependencies on Other Modules
| Module | Purpose | Required |
|--------|---------|----------|
| llm-inference | LLM for snippet extraction & relevance scoring | Yes (for extract_snippets) |
## License
MIT