didi-lot1-ai/ai_platform/README.md
EVOTECH IT SRL af0ec4a150 Arhitectură (diagrame) + imagini Docker module + contract furnizare + optimizări deploy
- docs/ARCHITECTURE.md (7 diagrame Mermaid) + architecture.html + diagrame originale (offer_diagram_1..7.png)
- artefacte_lot1/imagini_docker: imagini pre-construite pt toate cele 13 module (la zi)
- documentatie: Contract de furnizare nr.19
- optimizări deploy.sh/seed (MODELS_DIR, fix download python3); .gitignore (modele/.env/PV-uri excluse)
2026-06-29 16:23:03 +03:00

433 lines
14 KiB
Markdown

# ML Projects
A monorepo for machine learning services that **expose ML models via REST APIs**. This is the core purpose of the repository - each module provides production-ready API endpoints for inference, allowing applications to consume ML capabilities over HTTP.
Each module (LLM inference, RAG, etc.) is an independently deployable service with its own API, Docker configuration, and documentation.
## Quickstart (CPU-only host)
Bootstrap the **web + dashboard** stack in one command. This is the recommended path for a host that doesn't have a GPU and points at an LLM endpoint running elsewhere.
```bash
git clone <repository-url>
cd ml-projects
./bootstrap.sh
```
The interactive script prompts for the GPU host URL, paid search provider keys, and generates a fresh DB password. See [`DEPLOYMENT.md`](./DEPLOYMENT.md) for architecture, manual steps, configuration reference, and troubleshooting.
### What gets deployed by `bootstrap.sh`
| Service | Port | Role |
|---------------|--------|-------------------------------------------------|
| Web API | 51100 | Free (SearXNG) + premium (paid) search pipeline |
| Dashboard | 51300 | Monitoring, runtime config, archive, cost |
| Dashboard DB | 15432 | PostgreSQL — history + config + audit + archive |
GPU-dependent services (`llm-inference`, `embeddings`, `rerank`, `audio`, `video-analysis`) remain on a separate GPU host and are called over HTTP. Their URLs are captured in `modules/*/deploy/.env`.
## Repository Structure
```
/
├── README.md # This file
├── CLAUDE.md # Guidelines for Claude AI assistant
├── .pre-commit-config.yaml # Pre-commit hooks
├── ruff.toml # Shared linting/formatting config
├── modules/
│ ├── <module-name>/
│ │ ├── pyproject.toml # Module dependencies and metadata
│ │ ├── README.md # Module docs (incl. prerequisites)
│ │ ├── .env.example # Required environment variables
│ │ ├── src/<module_name>/ # Source code (importable package)
│ │ │ ├── __init__.py
│ │ │ └── ...
│ │ ├── tests/ # Tests
│ │ │ ├── __init__.py
│ │ │ └── ...
│ │ └── deploy/ # REQUIRED: Deployment directory
│ │ ├── deploy.sh # Main deployment script
│ │ ├── docker-compose.yml
│ │ ├── Dockerfile
│ │ └── nginx.conf # Optional: reverse proxy config
│ └── ...
└── shared/ # Optional shared utilities
└── ...
```
## Getting Started
### Prerequisites
**Required (all modules):**
- **Linux** - Ubuntu 22.04+ or similar distribution (no Windows/macOS support)
- **Docker Engine 24.0+** - With Docker Compose V2
- **Python 3.10+**
- **[uv](https://docs.astral.sh/uv/)** - Fast Python package manager
- **Git** - With pre-commit hooks support
**Optional (for GPU workloads):**
- **NVIDIA Driver 535+** - For CUDA 12.x support
- **NVIDIA Container Toolkit** - For GPU access in Docker containers
**Note:** Individual modules may have additional prerequisites (specific hardware, external services, etc.). Check each module's README for module-specific requirements.
### Initial Setup
1. Clone the repository:
```bash
git clone <repository-url>
cd ml-projects
```
2. Install pre-commit hooks:
```bash
uv tool install pre-commit
pre-commit install
```
### Working with a Module
Each module is independent. Navigate to the module directory to work with it:
```bash
cd modules/<module-name>
# Create virtual environment and install dependencies
uv sync
# Run tests
uv run pytest
# Run linting
uv run ruff check .
uv run ruff format --check .
# Start services (if docker-compose.yml exists)
docker compose up -d
```
## Module Guidelines
### Required Structure
Every module MUST have:
| File/Directory | Purpose |
|----------------|---------|
| `pyproject.toml` | Package metadata, dependencies, and build config |
| `README.md` | Module documentation (purpose, prerequisites, usage) |
| `API.md` | API documentation (required for modules with HTTP APIs) |
| `.env.example` | Required environment variables with documentation |
| `src/<module_name>/` | Source code as importable package |
| `src/<module_name>/__init__.py` | Package init with version |
| `tests/` | Test directory with pytest tests |
| `deploy/` | Deployment directory (see below) |
### Deployment Directory (Required)
Every module MUST have a `deploy/` directory containing:
| File | Purpose |
|------|---------|
| `deploy.sh` | Main deployment script with `--help`, validation, up/down/logs |
| `docker-compose.yml` | Docker Compose with profiles, health checks, pinned versions |
| `Dockerfile` | Container image for the module |
**Optional deployment files:**
- `nginx.conf` - Reverse proxy configuration
- `*.conf` - Other service configurations
**deploy.sh requirements:**
- Must validate all required environment variables (fail-fast)
- Must support `--help` with documentation
- Must handle `up`, `down`, `logs` actions
- Must source `.env` file if present
- Must NOT use fallback defaults
### pyproject.toml Requirements
Every module's `pyproject.toml` must include:
```toml
[project]
name = "<module-name>"
version = "0.1.0"
description = "Brief description"
requires-python = ">=3.10"
dependencies = [
# production dependencies
]
[project.optional-dependencies]
dev = [
"pytest>=8.0",
"pytest-cov>=4.0",
"ruff>=0.8",
]
[build-system]
requires = ["hatchling"]
build-backend = "hatchling.build"
[tool.hatch.build.targets.wheel]
packages = ["src/<module_name>"]
# Use root ruff.toml - only override if necessary
[tool.ruff]
extend = "../../ruff.toml"
```
### Docker/Docker Compose Standards
- Use `docker-compose.yml` (not `docker-compose.yaml`)
- Pin image versions (avoid `latest` tag)
- Use environment files (`.env`) for configuration
- Document all exposed ports in module README
- Include health checks for services
- **No fallback defaults** - use `${VAR}` not `${VAR:-default}` (fail-fast on missing config)
### Configuration Best Practices
**Never use fallback defaults for configuration variables.** If a required variable isn't set, the application should fail immediately with a clear error message.
```yaml
# Good - fails if OPENAI_API_KEY not set
environment:
- OPENAI_API_KEY=${OPENAI_API_KEY}
# Bad - silently uses empty string
environment:
- OPENAI_API_KEY=${OPENAI_API_KEY:-}
```
This applies to:
- Docker Compose environment variables
- Python `os.getenv()` calls
- Pydantic Settings defaults
**Why?** Silent misconfiguration leads to hard-to-debug production issues. Fail-fast behavior catches problems at startup.
### API Documentation (Required for HTTP APIs)
Modules that expose HTTP endpoints **MUST** have an `API.md` file documenting the API.
**Required sections:**
| Section | Description |
|---------|-------------|
| Base URL | Use `{BASE_URL}` variable with common configurations |
| Authentication | Auth requirements (Bearer token, API key, or "none") |
| Endpoints | All endpoints with method, path, request/response, examples |
| Error Responses | HTTP status codes and error format |
| Request Headers | Required and optional headers |
**Optional sections:** Rate Limiting, SDK Examples, Supported Backends
**Reference:** See `modules/llm-inference/API.md` for a complete example.
### Port Convention
Port allocation follows datacenter schema (5-digit ports):
| Prefix | Environment | Description |
|--------|-------------|-------------|
| 1xxxx | Production | Production services |
| 5xxxx | Development | Development/testing services |
| Second Digit | Category | Example |
|--------------|----------|---------|
| x0xxx | Web / Frontend | 10000, 50000 |
| x1xxx | API / Gateway | 11000, 51100 |
| x4xxx | LLM / AI | 14001 |
**Current Port Allocation:**
| Port | Service | Environment |
|------|---------|-------------|
| 11000 | Gateway (nginx, external entry) | Production |
| 11000 | Catalog API (internal, via gateway) | Production |
| 14001 | vLLM Qwen3.5-35B-A3B | Production |
| 14011 | LLM Inference API (router) | Production |
| 14100 | Embeddings API | Production |
| 14200 | Rerank API | Production |
| 51100 | Web API | Development |
| 51300 | Dashboard | Development |
| 54300 | Audio API | Development |
| 54400 | Extractors API | Development |
| 54500 | BusterX vLLM | Development |
| 54600 | Video Analysis API | Development |
| 8080 (intern) | SearXNG | Development |
| 8085 | Forensic Features API | Other |
| 8090 | didi_brain API | Other |
| 11000 (intern, alias `domain-check-api`) | Domain Check API (T4 — credibilitate sursă) | Other |
**Guidelines:**
- Use 1xxxx for production services
- Use 5xxxx for development/testing services
- Document your port in ENDPOINTS.md and docker-compose.yml
## Code Quality Standards
### Linting and Formatting
This repo uses **[Ruff](https://docs.astral.sh/ruff/)** for linting and formatting. Configuration is in the root `ruff.toml`.
```bash
# Check for issues
uv run ruff check .
# Auto-fix issues
uv run ruff check --fix .
# Format code
uv run ruff format .
```
### Type Hints
Type hints are **required** for all public functions and methods:
```python
# Good
def process_document(text: str, max_length: int = 512) -> list[str]:
...
# Bad - missing type hints
def process_document(text, max_length=512):
...
```
### Testing
Tests are **required** for all modules. Use pytest:
```bash
# Run all tests
uv run pytest
# Run with coverage
uv run pytest --cov=src/<module_name> --cov-report=term-missing
# Run specific test file
uv run pytest tests/test_specific.py
```
**Minimum requirements:**
- All public functions must have tests
- Aim for >80% code coverage
- Include both unit tests and integration tests where applicable
### Pre-commit Hooks
Pre-commit hooks run automatically on `git commit`. They enforce:
- Ruff linting and formatting
- YAML/TOML/JSON validation
- No trailing whitespace
- No large files (>1MB)
- No private keys committed
To run manually:
```bash
pre-commit run --all-files
```
## Contributing
### Adding a New Module
1. Create the module directory structure:
```bash
mkdir -p modules/<module-name>/src/<module_name>
mkdir -p modules/<module-name>/tests
mkdir -p modules/<module-name>/deploy
touch modules/<module-name>/src/<module_name>/__init__.py
touch modules/<module-name>/tests/__init__.py
```
2. Create `pyproject.toml` following the template above
3. Create `README.md` documenting:
- What the module does
- Installation instructions
- Usage examples
4. Create `API.md` if module has HTTP endpoints (see API Documentation section)
5. Add initial tests
6. Submit a merge request
### Merge Request Process
1. Create a feature branch: `git checkout -b feature/<description>`
2. Make changes and ensure all checks pass:
```bash
pre-commit run --all-files
uv run pytest
```
3. Push and create a merge request
4. Request review from at least one team member
5. Address review comments
6. Squash and merge when approved
### Code Review Expectations
Reviewers will check for:
- [ ] Code follows repo conventions
- [ ] Type hints present on public interfaces
- [ ] Tests cover new functionality
- [ ] Documentation updated if needed
- [ ] API.md included/updated for HTTP endpoints
- [ ] No security issues introduced
- [ ] Changes are focused (no unrelated modifications)
## Module Index
| Module | Port | Description | Status |
|--------|------|-------------|--------|
| [gateway](modules/gateway/) | 11000 | nginx reverse proxy — single external entry point | Active |
| [catalog-api](modules/catalog-api/) | 11000 | Service catalog & discovery aggregator (internal) | Active |
| [llm-inference](modules/llm-inference/) | 14011 | Unified LLM inference (Qwen3.5-35B-A3B) with multiple backends | Active |
| [embeddings](modules/embeddings/) | 14100 | OpenAI-compatible embeddings (BAAI/bge-m3) | Active |
| [rerank](modules/rerank/) | 14200 | Document reranking (BAAI/bge-reranker-v2-m3) | Active |
| [web](modules/web/) | 51100 | Web scraping & fact-checking evidence gathering | Active |
| [dashboard](modules/dashboard/) | 51300 | Admin UI + API (monitoring, config, RBAC) | Active |
| [audio](modules/audio/) | 54300 | Speech-to-text (Whisper large-v3-turbo) | Active |
| [extractors](modules/extractors/) | 54400 | Feature extraction: metadata/sentiment/OCR/NER (GLiNER)/detect (YOLOv8n) | Active |
| [video-analysis](modules/video-analysis/) | 54600 | Deepfake detection (BusterX) & semantic video analysis | Active |
| [forensic_features](modules/forensic_features/) | 8085 | Forensic detectors (rPPG/lip-sync/forgery) via MediaPipe | Active |
| [didi_brain](modules/didi_brain/) | 8090 | Result cache (Postgres/pgvector) + RAG + fact-checking | Active |
| [domain_check](modules/domain_check/) | 11000 (alias `domain-check-api`) | T4 — source credibility: WHOIS/SSL/blacklist/DNS/IP/HTTP/mail-intel → risk score. Isolated stack (own Postgres+Redis) | Active |
| [cloak](modules/cloak/) | internal | Stealth SERP scraper (tier-3 fallback for web) | Active |
## Port Allocation Summary
Port allocation follows datacenter schema (5-digit ports):
| Prefix | Environment |
|--------|-------------|
| 1xxxx | Production |
| 5xxxx | Development |
**Current Services:**
| Port | Service | Module | Environment |
|------|---------|--------|-------------|
| 11000 | Gateway (nginx, external entry) | gateway | Production |
| 11000 | Catalog API (internal) | catalog-api | Production |
| 14001 | vLLM Qwen3.5-35B-A3B | llm-inference | Production |
| 14011 | LLM API Gateway | llm-inference | Production |
| 14100 | Embeddings API | embeddings | Production |
| 14200 | Rerank API | rerank | Production |
| 51100 | Web API | web | Development |
| 51300 | Dashboard | dashboard | Development |
| 54300 | Audio API (Whisper) | audio | Development |
| 54400 | Extractors API | extractors | Development |
| 54500 | BusterX vLLM | video-analysis | Development |
| 54600 | Video Analysis API | video-analysis | Development |
| 8080 (intern) | SearXNG | web | Development |
| 8085 | Forensic Features API | forensic_features | Other |
| 8090 | didi_brain API | didi_brain | Other |
For detailed endpoint documentation, see [ENDPOINTS.md](ENDPOINTS.md).