# Implementation Plan — HIL Moderation + Brain v2 Integration **Status**: NOT STARTED **Started**: — **Last Updated**: 2026-04-30 **Owner**: tehnic@finesynergy.eu ## Companion documents - `HIL_MODERATION_DESIGN.md` — design queue + UI moderare - `BRAIN_V2_DESIGN.md` — design atomi brain + tier-uri (cu corecție: tier NU în UNIQUE, write doar premium) ## How to resume from this file If a new chat session starts, the new assistant should: 1. Read this file FIRST, completely 2. Read `HIL_MODERATION_DESIGN.md` and `BRAIN_V2_DESIGN.md` 3. Look at the "Status & Progress Tracking" section below — find first unchecked `[ ]` task in current phase 4. Verify the actual codebase state matches what's marked completed (DON'T trust this file blindly — `git log`, file contents, DB schema) 5. Continue from first unchecked task; update `[ ]` → `[x]` as work progresses If a task description seems stale (e.g., file structure changed since written), the new assistant should: - Read the current file structure with `Read` / `Bash` tools - Update the task description in this file BEFORE implementing - Note "✏️ Updated 2026-XX-XX: reason" inline --- ## ⚠️ CRITICAL — DO NOT BREAK These are existing flows that MUST keep working at every step. Any change that breaks them is a stop-the-world bug. ### Existing endpoints that MUST keep working - `POST /api/v3/pipeline/analyze` — sync analyze, returns `AnalysisSession` flat - `POST /api/v3/pipeline/analyze-async` — returns 202 + poll URL - `POST /api/v3/{techniques,ai-tampered,claims}/analyze` — individual components - `GET /api/v3/pipeline/{sessionId}/status` — polling - `GET /api/v3/pipeline/{sessionId}/result` — final result - `GET /api/v3/pipeline/history` — user history - `GET /api/v3/pipeline/history/admin` — admin history - All `/agent-v3/*` endpoints called from admin dashboard ### Existing data flows that MUST keep working - agent-v3 reads framework config from Redis keys `didi:framework:*` and `didi:config:*` - agent-v3 writes session state to Redis `didi:pipeline:*` (TTL 7d) and PG `bos_analysis.*` - didiFramework `POST /api/sync-redis` writes Redis from PG - Claims component continues to use `verification_cache` in brain (existing path, not modified) - All worker queues (techniques, ai_tampered, claims, domain, media-preprocess, aggregator) ### Feature flags as safety net EVERY new code path must check a feature flag from Redis. If flag is OFF, code path is bypassed and the OLD behavior runs. Flags are managed by `bos_parammgmt.moderation_config` table: - `triage_enabled` (default false → start OFF, activate after testing) - `brain_enabled` (default false → start OFF) - `brain_per_component.{techniques,ai_tampered,claims}` (per-component toggle) ### Rollback strategy per phase Each phase ends with a deployment checkpoint. If the next phase breaks something, we roll back BY: 1. Setting feature flag = false in `moderation_config` (no code revert needed, just SQL UPDATE) 2. If schema change breaks queries: pre-prepared rollback migration `010_rollback.sql` 3. If new code breaks startup: `git revert` last commit, redeploy previous image --- ## Architecture overview (1-paragraph recap) User → web app → agent-v3 → 4 component workers in parallel. Each worker BEFORE running LLM does brain lookup. On hit → return cached, skip LLM. On miss → run LLM + fire-and-forget write atom. After verdict aggregator persists, triage decides if session needs review. If yes, INSERT moderation_queue. Moderator opens admin dashboard `/moderation`, claims session, edits if needed, resolves. On resolve with corrections, brain atom is PATCHed to gold tier with human_validated=true. Next user analyzing same content → hit gold atom → return verified-by-analyst result instantly. ZERO config hardcoded — everything in PG `bos_parammgmt.moderation_config` + `sensitive_topic` + `moderation_role`, synced to Redis. Editable via admin dashboard tab "Moderation Settings". Separation: **didi side** (HIL, queue, triage, brain client) is implemented HERE. **Brain side** (atomi storage, lookup endpoints, brain admin UI) is implemented in AI platform repo (separate). Contracts in `BRAIN_V2_DESIGN.md`. --- ## Status & Progress Tracking Format: `[ ]` = not started, `[~]` = in progress, `[x]` = completed, `[!]` = blocked, `[s]` = skipped (not needed) Update format: when changing status, also append `→ YYYY-MM-DD note` if useful for resume context. --- ## TRACK 1 — Didi Side (independent of brain) This track delivers a working HIL moderation system without brain integration. Even with brain OFF, queue + UI work standalone. ### PHASE 1.1 — Database foundation (estimated 0.5 day) — ✅ DONE 2026-05-01 **Goal**: All schema changes applied to cluster PG. New tables seeded with defaults. Old tables get new columns. Rollback migration ready. **RESULT**: Migration `011_add_moderation.sql` (NOTE: 011 not 010 — 010 was already taken by `add_user_storage_quota`). Applied to cluster `10.11.50.167:5000` DB `DIDI`. Dry-run confirmed clean before commit. Smoke tests on existing endpoints pass. #### Task 1.1.1 — Write migration `011_add_moderation.sql` - [x] File: `didiFramework/sql/migrations/011_add_moderation.sql` (NEW, 14.5KB) - Note: lives in `didiFramework/sql/migrations/`, NOT `agent-v3/sql/migrations/` — that's where existing migrations are. - [x] Contents per design doc: - 6 columns added to `bos_analysis.analysis_session` (review_status, human_corrected, human_corrections, verified_by, verified_at, review_notes) with CHECK constraint on review_status - `bos_analysis.moderation_queue` table — FK to `analysis_session(session_id)` (UUID type, important: not TEXT) - 3 indexes on moderation_queue (status_priority, session, assigned) - `bos_parammgmt.moderation_config` single-row table with CHECK(config_id=1) — 13 fields covering triage + brain client + audit - `bos_parammgmt.sensitive_topic` with 5 seed rows (elections, health, war, covid, climate) - `bos_parammgmt.moderation_role` with 2 seed rows (moderator, senior_moderator) - Triggers `set_updated_at` on all 3 config tables - [x] Companion file: `011_rollback.sql` (NEW, 2.5KB) — drops everything in reverse order - [x] Dry-run validated (BEGIN+ROLLBACK trick) before commit - [x] Applied to cluster, verified all rows seeded correctly - [x] Smoke tested existing endpoints (`/api/admin/users`, `/api/verdicts/runtime-config`) — both still respond OK **Resume context**: - Both migration files exist on disk and have been applied to cluster. - `moderation_config.triage_enabled=false` and `brain_enabled=false` — feature flags OFF as planned. - DB schema ready for Phase 1.2 (CRUD endpoints in didiFramework). **Gotcha discovered**: `analysis_session.session_id` is `UUID` type, not `TEXT`. FK on `moderation_queue.session_id` MUST be `UUID NOT NULL REFERENCES`. Initial design assumed TEXT — corrected before apply. **Validation**: ```bash # Apply on cluster — manual, never automated docker exec didi-framework node -e " const fs=require('fs'); const {Pool}=require('pg'); const pool=new Pool({host:'10.11.50.167',port:5000,user:'bos_interface',password:'interface',database:'DIDI'}); pool.query(fs.readFileSync('/path/to/010_add_moderation.sql','utf8')).then(r=>{console.log('OK');pool.end();}).catch(e=>{console.error(e);pool.end();}); " ``` **Validation checks**: - [ ] `\d+ bos_analysis.analysis_session` shows 6 new columns - [ ] `SELECT * FROM bos_parammgmt.moderation_config` returns 1 row with defaults - [ ] `SELECT COUNT(*) FROM bos_parammgmt.sensitive_topic WHERE is_active=true` returns 5 - [ ] `SELECT * FROM bos_parammgmt.moderation_role` returns 2 rows **Rollback**: - Apply `010_rollback.sql`. All changes are additive — no data loss on rollback (review_status drops to default 'none' on old sessions). **Risk**: LOW. Pure schema additions, no destructive ops. --- ### PHASE 1.2 — didiFramework: config CRUD endpoints (estimated 0.5 day) — ✅ DONE 2026-05-01 **Goal**: Admin can read/write moderation config via REST API. Sync-redis extended to push new keys to Redis. **RESULT**: 3 routes files created + mounted + sync-redis extended. All 4 endpoints tested live (GET/PUT/POST/DELETE all working). Validation tested (rejects bad input). Sync-redis now writes 3 new keys (`config_keys` count rose from 48 to 51). Existing endpoints still respond. #### Task 1.2.1 — Create `routes/moderation-config.ts` - [x] File created. `UPDATABLE_FIELDS` whitelist (14 fields), validates body is object, returns 400 if no allowed fields, dynamic SET clause builder, audit `updated_by` from header `x-user-id`. PUT response includes "Sync to Redis to apply" message. #### Task 1.2.2 — Create `routes/sensitive-topics.ts` - [x] File created. Full CRUD with `?active=true|false|all` filter, regex validation on topic_code (`[a-z0-9_]+`), 409 conflict on duplicate (PG code 23505), 404 on missing id, soft delete (is_active=false). #### Task 1.2.3 — Create `routes/moderation-roles.ts` - [x] File created. GET list, PUT only on `:code` (immutable PK). Toggle fields whitelist + label validation. role_code IS immutable. #### Task 1.2.4 — Mount routes in `server.ts` - [x] 3 imports + 3 mounts under "Category 14: HIL Moderation Config". Existing routes untouched. #### Task 1.2.5 — Extend sync-redis - [x] Block added BEFORE `pipeline.exec()` — wrapped in try/catch (skips silently if migration 011 not applied). Writes: - `didi:config:moderation:v1:settings` (full row from moderation_config) - `didi:config:moderation:v1:sensitive_topics` (active topics only) - `didi:config:moderation:v1:roles` (all roles with permissions) **Live test results**: - `GET /api/moderation-config` → 200 with full row - `GET /api/sensitive-topics` → 200, 5 active topics - `GET /api/moderation-roles` → 200, 2 roles - `PUT /api/moderation-config` → 200, updates persist - `POST /api/sensitive-topics` → 201, validation works - `PUT /api/moderation-roles/:code` → 200, toggles work - `DELETE /api/sensitive-topics/:id` → 200, soft delete - Validation rejects: empty body, bad regex codes, non-boolean toggles - `POST /api/sync-redis` → `keys_written: 59, config_keys: 51` (was 48 before) - All existing endpoints (`/api/admin/users`, `/api/verdicts/runtime-config`) still respond OK **Resume context**: - All 3 endpoints respond at port 3005. - Redis cluster has the 3 new keys after sync. - Container `didi-framework` running healthy. - Ready for Phase 1.3 (admin dashboard UI tab "Moderation Settings"). **Validation**: ```bash # After PUT curl -s http://localhost:3005/api/moderation-config | jq '.data.confidence_low' # Should match what was PUT curl -s -X POST http://localhost:3005/api/sync-redis # Then check Redis: redis-cli -h 10.11.50.100 -p 16379 -a --no-auth-warning GET didi:config:moderation:v1:settings | jq ``` **Validation checks**: - [ ] GET returns full config including all defaults - [ ] PUT with valid body updates DB - [ ] PUT with invalid body returns 400 with error message - [ ] Sync-redis includes new keys in response manifest - [ ] Redis key contains JSON with all settings after sync **Risk**: LOW. New routes, no modification to existing behavior. --- ### PHASE 1.3 — Admin dashboard: "Moderation Settings" tab (estimated 1 day) — ✅ DONE 2026-05-01 **Goal**: User can edit all moderation config from UI. Triage rules, brain client params (CLIENT only — server config lives in AI platform dashboard), sensitive topics list, roles permissions. **RESULT**: Module `components/ModerationSettings/` (6 files) mounted as 6th toggle button in `/llm-components`. Auto-sync to Redis after every save. Dashboard rebuilt + container `didi-admin` recreated. All endpoints respond via nginx proxy. #### Task 1.3.1 — Create `components/ModerationSettings/` - [x] Folder created (NEW), 6 files: - `index.tsx` — main panel: Refresh + Sync buttons, 4 cards stacked vertically, Snackbar for success - `TriageCard.tsx` — toggle, slider for confidence_low, TextFields for risk_grey min/max + queue thresholds, auto-tune toggle - `BrainClientCard.tsx` — toggle, URL TextField, timeouts (lookup/write), sliders (confidence_min_silver, semantic_threshold), per-component checkboxes. Includes Alert: "These are CLIENT settings — brain server config lives in AI platform dashboard" - `SensitiveTopicsCard.tsx` — chip list with delete icon, add form (code+label), show inactive toggle, soft delete + reactivate - `RolesCard.tsx` — table with 4 toggle columns (can_resolve, can_escalate, can_force_gold_brain, is_active) - `api.ts` — fetch helpers + types matching backend - [x] Auto-sync after save: every save handler calls `handleSync()` automatically — user sees "Config saved. Syncing to Redis…" snackbar without manual button press - [x] Manual sync button still available top-right for explicit re-sync #### Task 1.3.2 — Mount in `/llm-components` - [x] `LLMComponentsConfig/index.tsx` modified: - Added imports: `ModerationIcon` (Shield) and `ModerationSettings` component - Extended `ComponentType` union with `'moderation'` - Added entry in `COMPONENT_CONFIGS['moderation']` (color #0288d1, ShieldIcon) - Added `'moderation': ''` in `SAMPLE_TEXTS` (not used, has own component) - Added `'moderation'` to `analysisComponents` array → 6th toggle button rendered - Added render branch: `selectedComponent === 'moderation' ? : ...` - [x] Existing 5 tabs (Techniques/AI/Claims/Source/Verdict) unchanged **Validation results** (all ✅): - [x] TypeScript clean (preexisting warnings only, not from my changes) - [x] Image built `didi-admin:latest`, container recreated, healthy - [x] `https://10.11.10.12:3000/admin/` → 200 - [x] `https://10.11.10.12:3000/framework/api/moderation-config` → 200 (proxy) - [x] `https://10.11.10.12:3000/framework/api/sensitive-topics` → 200 (proxy) - [x] `https://10.11.10.12:3000/framework/api/moderation-roles` → 200 (proxy) **Resume context**: - Open `/admin/llm-components`, click "Moderation" (last toggle). - 4 cards visible. All editable. Saves auto-sync. - Feature flags `triage_enabled` and `brain_enabled` are still **false** — agent-v3 behavior unchanged. - Ready for Phase 1.4 (triage logic + queue manager in agent-v3). **Risk**: LOW (UI only, no breaking changes). --- ### PHASE 1.4 — agent-v3: triage + queue manager (estimated 1 day) — ✅ DONE 2026-05-01 **Goal**: After every analysis finalizes (sync or async), triage decides if session enters moderation queue. All thresholds in Redis (no hardcode). Triage must NEVER block analysis on failure. **RESULT**: 2 new modules + 4 files modified. Live E2E test verified: with `triage_enabled=true, confidence_low=95`, an analysis with confidence=42 was correctly enqueued (queue_id=1, reason=low_confidence, priority=3, status=pending). With `triage_enabled=false`, no enqueue occurs. Existing endpoints unchanged. **Files created**: - `agent-v3/src/components/moderation/triage.ts` — pure decision function with Redis config + 60s TTL cache. Returns `{needsReview, priority, reason, meta}`. Disabled flag → always returns `needsReview=false`. Refresh helper `invalidateTriageCache()` for admin. - `agent-v3/src/components/moderation/queue-manager.ts` — PG ops: `enqueueForReview`, `listQueue`, `getQueueEntry`, `claimQueueEntry`, `resolveQueueEntry`, `getQueueStats`. Idempotent enqueue (checks for existing pending/in_review row before INSERT). Updates `analysis_session.review_status` in same transaction. **Files modified**: - `pipeline/executor.ts` — calls triage AFTER `persistService.persist(session)`. Wrapped in try/catch — triage failure logs warning and continues. - `queue/aggregator.ts` — same wire-up in async path after persist. - `queue/types.ts` — added `userFlagged?: boolean` to `SessionState`. - `components/pipeline/types.ts` — added `userFlagged?: boolean` to `PipelineInput`. - `api/pipeline-routes.ts` — accepts `user_flagged` in body, propagates to `PipelineInput.userFlagged`. **E2E Test results**: - Triage OFF → analyze runs → 0 queue rows ✅ - Triage ON, confidence_low=95 → analyze runs → 1 queue row with priority=3, reason=low_confidence ✅ - `analysis_session.review_status` updated to 'pending' ✅ - Triage timing: ~10ms (Redis read cached after first call) - Total pipeline time unchanged (~71s for full LLM analysis), triage adds <50ms **Resume context**: - Modules in place, wired, tested. - Feature flag `triage_enabled=false` again (default safe). - Migration 011 applied, 4 endpoints + sync-redis working, UI tab live. - Ready for Phase 1.5 (REST API moderation in agent-v3 — endpoints `/api/v3/moderation/*`). **Risk**: MEDIUM (touches pipeline hot path). Mitigation: try/catch wrapper, feature flag, Redis cache fallback to no-op. Verified live. **Goal**: After every analysis, triage logic decides if session enters moderation queue. Logic reads thresholds from Redis (no hardcode). #### Task 1.4.1 — Create `components/moderation/triage.ts` - [ ] File: `agent-v3/src/components/moderation/triage.ts` (NEW) - [ ] Function `shouldEnqueueForReview(session, userFlagged): TriageOutput` - Reads config from Redis key `didi:config:moderation:v1:settings` (cached in memory with 60s TTL) - Reads sensitive topics from `didi:config:moderation:v1:sensitive_topics` - Returns `{ needsReview, priority, reason, meta }` - [ ] Function `loadTriageConfigFromRedis()` — refresh cache - [ ] Tests: 5 unit tests covering each branch (flagged, low_confidence, sensitive_topic, none, disabled) - [ ] If `triage_enabled=false` in config, returns `{needsReview: false}` always #### Task 1.4.2 — Create `components/moderation/queue-manager.ts` - [ ] File: `agent-v3/src/components/moderation/queue-manager.ts` (NEW) - [ ] Functions: - `enqueueForReview(sessionId, priority, reason, meta)` — INSERT INTO moderation_queue + UPDATE analysis_session.review_status='pending' - `claimQueueEntry(queueId, userId)` — atomic UPDATE with WHERE status='pending' RETURNING - `resolveQueueEntry(queueId, action, corrections, notes, userId)` — UPDATE queue + analysis_session - `getQueueStats()` — counts pending/in_review by priority - [ ] All functions take a PG pool from existing `pg-pool.ts` (don't create new connection) #### Task 1.4.3 — Wire triage into pipeline executor (sync flow) - [ ] File: `agent-v3/src/components/pipeline/executor.ts` (MODIFY) - [ ] After `PersistService.persist(...)`, call: ``` const triage = await shouldEnqueueForReview({ session, userFlagged: input.userFlagged }); if (triage.needsReview) await enqueueForReview(session.session_id, ...); ``` - [ ] WRAP in try/catch — triage failure must NOT fail the analysis - [ ] If catch fires, log warning + continue (degraded mode) #### Task 1.4.4 — Wire triage into aggregator (async flow) - [ ] File: `agent-v3/src/queue/aggregator.ts` (MODIFY) - [ ] After persist, identical try/catch pattern as 1.4.3 #### Task 1.4.5 — Accept `userFlagged` in request body - [ ] File: `agent-v3/src/api/pipeline-routes.ts` (MODIFY) - [ ] Accept optional `user_flagged: boolean` in body of `/analyze` and `/analyze-async` - [ ] Default false. Pass through to PipelineInput / dispatcher. - [ ] Document in API_GUIDE.md if exists **Validation**: ```bash # Make moderation_config triage_enabled=false in DB # Run analyze → check moderation_queue is empty # Set triage_enabled=true, sync-redis # Run analyze with text designed to be low-confidence → check queue has 1 row ``` **Validation checks**: - [ ] With triage_enabled=false: zero queue entries created - [ ] With triage_enabled=true: queue gets entries for low-confidence sessions - [ ] Triage failure (e.g., Redis down) does NOT fail analysis - [ ] user_flagged=true in body → priority=1 entry in queue **Risk**: MEDIUM. New code path in critical pipeline. Mitigation: try/catch around triage, feature flag. --- ### PHASE 1.5 — agent-v3: moderation REST API (estimated 0.5 day) — ✅ DONE 2026-05-01 **Goal**: REST endpoints for queue list/detail/claim/resolve/flag/stats. Used by admin UI (Phase 1.6) and browser extension (Phase 5). **RESULT**: New `api/moderation-routes.ts` with 6 endpoints. JWT roles middleware extended. All endpoints tested live with realistic flow (insert → get → claim → resolve). Validation tested. Existing endpoints still respond. **Files created**: - `agent-v3/src/api/moderation-routes.ts` — 6 endpoints, soft role check (permits in staging without JWT, enforces in prod with JWT roles) **Files modified**: - `agent-v3/src/index.ts` — added `jwtRoles?: string[]` to Request interface, populated from `payload.realm_access.roles`. Mounted `/api/v3/moderation`. **Endpoints live**: - `GET /api/v3/moderation/queue` — paginated, filters: status, priority (CSV), assigned_to (or 'me'), limit/offset - `GET /api/v3/moderation/queue/:queueId` — entry + full session (with all component LEFT JOINs) - `POST /api/v3/moderation/queue/:queueId/claim` — atomic claim (UPDATE pending→in_review) - `PUT /api/v3/moderation/queue/:queueId/resolve` — body: `{action, corrections?, notes?, user_id?}` - `POST /api/v3/moderation/flag` — any authenticated user (extension report); validates session exists; idempotent (won't double-enqueue) - `GET /api/v3/moderation/stats` — pending, in_review, resolved_24h, by_priority, avg_time_in_queue_ms **E2E live test results**: - `GET /queue` empty list → 200 - `GET /stats` zeroed → 200 - Inserted fake queue entry, then: - `GET /queue/3` → 200 with full detail - `POST /queue/3/claim` → 200, status=in_review, assigned_to=mod-test - `PUT /queue/3/resolve` action=approved → 200, status=resolved - `POST /flag` → 200, returns queue_id=4 - Validations: - `action='bogus'` → 400 with clear message - `action='corrected'` without `corrections` → 400 - `flag` reason='spam' → 400 (not in allowed list) - Final smoke: `/api/v3/health` → 200 (no regression) **Resume context**: - 6 endpoints respond on port 24803 + via nginx proxy at `/agent-v3/api/v3/moderation/*` - Role check is permissive in staging (works without JWT roles) and strict in prod (Kong sets jwtRoles) - `corrections` JSONB diff is just stored — applied to brain only in Phase 1.8 - Ready for Phase 1.6 (admin dashboard `/moderation` page). **Risk**: LOW (new routes, no modification to existing pipeline behavior). **Goal**: Endpoints for queue list, detail, claim, resolve, stats. Used by admin UI. #### Task 1.5.1 — Create `api/moderation-routes.ts` - [ ] File: `agent-v3/src/api/moderation-routes.ts` (NEW) - [ ] Endpoints (per `HIL_MODERATION_DESIGN.md` API section): - `GET /api/v3/moderation/queue` — paginated list with filters - `GET /api/v3/moderation/queue/:queueId` — detail (queue entry + full session) - `POST /api/v3/moderation/queue/:queueId/claim` — assign to current user - `PUT /api/v3/moderation/queue/:queueId/resolve` — body: action, corrections, notes, trigger_brain_write - `POST /api/v3/moderation/flag` — for browser extension user reports - `GET /api/v3/moderation/stats` — counts + averages - [ ] Auth middleware: extract JWT roles, check includes `moderator` (or `senior_moderator` for escalate). Reject 403 otherwise. - [ ] For `/flag` endpoint: any authenticated user, NOT moderator-only - [ ] In resolve handler: - On `action='approved'` and corrections=null: nothing to brain (yet — phase 4) - On `action='corrected'`: save corrections to analysis_session.human_corrections (already wired); brain PATCH happens in phase 4 #### Task 1.5.2 — Mount in `index.ts` - [ ] File: `agent-v3/src/index.ts` (MODIFY) - [ ] Add import + `app.use('/api/v3/moderation', moderationRoutes)` - [ ] DO NOT touch existing mounts **Validation**: - [ ] curl with bad JWT → 401 - [ ] curl with user JWT (no moderator role) → 403 - [ ] curl with moderator JWT → 200, gets queue - [ ] Claim then resolve flow works end-to-end on a test session **Risk**: LOW (new routes). --- ### PHASE 1.6 — Admin dashboard: `/moderation` page (estimated 2 days) — ✅ DONE 2026-05-01 **Goal**: Moderator opens `/moderation`, sees pending queue, opens session detail, claims, resolves with action+corrections+notes. Stats dashboard at `/moderation/stats`. **RESULT**: 4 files created in new `Moderation/` folder, 2 modifications (App.tsx + AdminLayout.tsx). MVP delivered: queue list with filters/pagination, detail page with side-by-side input+verdict + Resolve dialog (action, corrections JSON, notes), stats page with cards + by-priority breakdown. JSON corrections panel is intentionally minimal (free-form JSON edit) — bogát Edit panels per component (techniques toggle, claims status changer) are deferred to follow-up MR if needed. **Files created**: - `admin-dashboard/src/components/Moderation/api.ts` — fetch helpers + types + UI helpers (priorityLabel, priorityColor, statusColor, ageMinutes) - `admin-dashboard/src/components/Moderation/ModerationQueue.tsx` — paginated table, filters (status: pending/in_review/resolved/all + priority CSV), auto-refresh 30s on pending tab, click row → detail - `admin-dashboard/src/components/Moderation/ModerationDetail.tsx` — back button, action buttons (Claim / Approve / Resolve with corrections / Reject), side-by-side cards (Input + AI Verdict), Queue metadata card, Resolve dialog with action toggle + JSON corrections + notes - `admin-dashboard/src/components/Moderation/ModerationStats.tsx` — 3 stat cards (pending, in_review, resolved_24h) + by-priority chips + avg time in queue, auto-refresh 30s **Files modified**: - `admin-dashboard/src/App.tsx` — 3 new imports + 3 routes inside ProtectedRoute/AdminLayout (`/moderation`, `/moderation/stats`, `/moderation/:queueId`) - `admin-dashboard/src/components/layout/AdminLayout.tsx` — added `Shield as ModerationIcon` import + sidebar `ListItemButton` under "Management" group **Live test results**: - `https://10.11.10.12:3000/admin/` → 200 - `https://10.11.10.12:3000/admin/moderation` → 200 - `https://10.11.10.12:3000/admin/moderation/stats` → 200 - `/agent-v3/api/v3/moderation/queue` via proxy → 200 - `/agent-v3/api/v3/moderation/stats` via proxy → 200 - E2E: inserted fake queue entry → appears in list → detail loads with full JSON - Cleanup: queue empty after test **Resume context**: - 4 routes wired, sidebar link visible always (not role-gated yet — Phase 1.7 adds Keycloak role enforcement) - All test paths working through nginx → didi-framework / agent-v3 - Resolve form supports approved + corrected (JSON) + rejected - Ready for Phase 1.7 (Keycloak roles setup) — admin/moderator user, role-gated sidebar visibility **Risk**: LOW (UI only, no breaking changes to existing routes). **Goal**: Moderator opens `/moderation`, sees pending queue, opens session, edits, resolves. #### Task 1.6.1 — Create `components/Moderation/` folder - [ ] Folder: `admin-dashboard/src/components/Moderation/` (NEW) - [ ] Files: - `ModerationQueue.tsx` — paginated table with filters - `ModerationDetail.tsx` — split view (input | verdict editable) - `ModerationStats.tsx` — cards + charts - `EditVerdictPanel.tsx` — sub-component for verdict overrides - `EditTechniquesPanel.tsx` — toggle techniques on/off - `EditAITamperedPanel.tsx` — verdict + indicators toggle - `EditClaimsPanel.tsx` — per-claim status changer - `api.ts` — fetch helpers - `types.ts` — TypeScript types matching backend #### Task 1.6.2 — Add routes in `App.tsx` - [ ] File: `admin-dashboard/src/App.tsx` (MODIFY) - [ ] Add 3 routes: - `/moderation` → ModerationQueue - `/moderation/:queueId` → ModerationDetail - `/moderation/stats` → ModerationStats - [ ] Wrap each in `` #### Task 1.6.3 — Add sidebar link - [ ] File: `admin-dashboard/src/components/dashboard/ServicesDashboard.tsx` (MODIFY) or wherever sidebar lives - [ ] Add link "Moderation" with badge showing pending count - [ ] Show only if user has role `moderator` - [ ] Auto-refresh badge count every 30s (poll `/api/v3/moderation/stats`) #### Task 1.6.4 — Detail page edit logic - [ ] When moderator changes a field (toggle technique off, change claim status, override risk_score), local state captures the diff - [ ] On "Save corrections" click, build `human_corrections` JSONB diff - [ ] PUT `/api/v3/moderation/queue/:id/resolve` with `action='corrected', corrections=diff, trigger_brain_write=true` - [ ] Show loading state, then redirect to queue list with success toast **Validation**: - [ ] Moderator login → sees Moderation in sidebar - [ ] Click Moderation → list of pending sessions - [ ] Click a session → detail view loads - [ ] Edit a technique toggle → state changes locally, dirty flag shows - [ ] Click Save → resolve API call, queue refreshes, that entry now has status=resolved **Risk**: MEDIUM (complex UI). Mitigation: ship in 2 sub-PRs (queue list first, then detail page). --- ### PHASE 1.7 — Keycloak roles setup (estimated 0.25 day) — ✅ DONE 2026-05-01 **Goal**: Realm `didi-clients` on **SSO cluster** (`` public / `` admin) has 2 new roles + 2 groups + 1 test user assigned. Sidebar link "Moderation" gated by role in admin dashboard. **RESULT**: 2 realm roles created via Admin API, 2 groups created with proper role mappings, admin@didi.local user got both moderator + senior_moderator roles. Sidebar gating wired in `AdminLayout.tsx`. In staging mode (`REACT_APP_STAGING_MODE=true`), `hasRole()` always returns true → link visible to all users for testing. In prod, only users with `moderator|senior_moderator|admin` see the link. **SSO cluster credentials discovered**: - Public URL: `https:///realms/didi-clients` - Admin URL: `https://` (needs internal DNS, resolves to 10.11.10.171) - Master credentials: `admin / admin123` (from `backend/production/.env`) - Realm: `didi-clients` **Operations performed via Admin API**: - Created realm role `moderator` (id=`6552fc4e-...`) - Created realm role `senior_moderator` (id=`e0b18224-...`) - Created group `moderators-team` (id=`8a692f87-...`) → has `[moderator]` - Created group `senior-moderators-team` (id=`0e3bae14-...`) → has `[moderator, senior_moderator]` - Assigned both roles to user `admin@didi.local` (id=`2c0f074b-97f1-4779-91a5-6b5b5bc1da8c`) — same user used in earlier tests **Files modified**: - `admin-dashboard/src/components/layout/AdminLayout.tsx`: - Added `Shield as ModerationIcon` import - `const canModerate = hasRole('moderator') || hasRole('senior_moderator') || isAdmin` - Wrapped Moderation `` in `{canModerate && (...)}` **Verification (live)**: - API: `GET /admin/realms/didi-clients/roles` → 12 roles total (was 10), includes moderator + senior_moderator - API: `GET /admin/realms/didi-clients/groups` → 6 groups total (was 4) - User admin@didi.local → realm roles: `[moderator, viewer, analyst, api_user, admin, enterprise_tier, senior_moderator]` - Admin dashboard rebuilt, container recreated and healthy - `/admin/moderation` → 200 (visible because staging mode) **NOT modified**: - `realm-import/didi-clients-realm.json` (local container artifact, no longer used since 2026-04-29 migration to SSO cluster — already deprecated) - agent-v3 routes — already use `req.jwtRoles` array populated from JWT (Phase 1.5 setup) **Resume context**: - Roles are LIVE on SSO cluster, will persist across SSO restarts - admin@didi.local already had admin role; now also has moderator + senior_moderator (additive, non-destructive) - For prod role gating to work, need to ensure JWT issued to admin user includes the new roles in `realm_access.roles` array — Keycloak does this automatically on next token refresh **Risk**: VERY LOW. Additive operations only. Did not modify any existing role/group/user. **Goal**: Roles `moderator` and `senior_moderator` exist in Keycloak realm. Test user assigned. - [ ] Realm `didi-clients` — Add 2 roles via admin console or realm-import: - `moderator` - `senior_moderator` - [ ] Add 2 groups: `moderators-team`, `senior-moderators-team` - [ ] Assign 1 test user to `moderators-team` (e.g., test@didi.local) - [ ] Verify JWT contains role on next login: decode JWT, check `realm_access.roles` - [ ] Document: who is real moderator? Add row in `moderation_role` table for each Keycloak role **Validation**: - [ ] Login as moderator → frontend gets JWT with role - [ ] Login as normal user → frontend gets JWT without role - [ ] /moderation route 200 for moderator, 403 for normal **Risk**: LOW. --- ## PHASE 1 — END-OF-TRACK CHECKPOINT After phases 1.1-1.7, the deliverable is: - ✅ Triage runs after every analysis (configurable, default OFF until tested) - ✅ Queue auto-fills based on rules in PG (editable from UI) - ✅ Moderator opens UI, claims, edits verdict, resolves - ✅ Resolved corrections saved to `analysis_session.human_corrections` (PG) - ❌ Brain integration NOT yet — that's Phase 1.8 At this checkpoint, system has full HIL workflow but doesn't propagate to brain yet. Operationally useful already (analysts can correct verdicts internally). **Deployment validation before moving to Phase 1.8:** - [ ] Run 50 test analyses through pipeline - [ ] Verify: ones matching triage rules end up in queue - [ ] Verify: moderator can resolve all of them - [ ] Verify: NO existing tests fail - [ ] Verify: latency on `/analyze` not increased (triage adds <50ms) --- ### PHASE 1.8 — Brain client integration in agent-v3 (estimated 1 day) — ✅ DONE 2026-05-01 **Goal**: Wire brain client into techniques/ai_tampered executors + moderation resolve handler. Feature flag controls activation. **Files modified**: - `agent-v3/src/shared/brain/client.ts` — added `computeContentHash`, `loadBrainConfig`, `lookupAnalysisAtom`, `writeAnalysisAtomAsync`, `patchAnalysisAtomGold`, types `AtomComponent`, `AtomCacheTier`, `AtomStaleness`, etc. Config read from Redis (60s cache). - `agent-v3/src/components/techniques/executor.ts` — brain lookup at start of `execute()`, write atom on both happy path AND early-exit path. - `agent-v3/src/components/ai-tampered/executor.ts` — same pattern. - `agent-v3/src/api/moderation-routes.ts` — `promoteAtomsToGold()` helper called from resolve handler. Iterates components, finds matching atom, PATCHes to gold. **Bug found + fixed during E2E**: `writeAnalysisAtomAsync` was called only on the happy path (after `buildFinalResult`). Early-exit branches (`buildEmptyResult` for both techniques and ai_tampered) returned BEFORE the write call → no atoms written for non-manipulation content. Fixed by duplicating write call in early-exit branches. **Resume context**: - Brain config in `moderation_config.brain_url` set to `http://10.11.10.12:8090` (local brain for testing); change to production brain when ready. - Test user `2c0f074b-...` switched to plan_type=4 (premium) for E2E testing — needs to be reverted post-test to plan_type=1. - Atoms written tier-isolated by content (read tier-agnostic, write only tier=premium). **Risk**: MEDIUM. Touches hot path (executors). Mitigation: try/catch on lookup, fire-and-forget on write, feature flag, 2s lookup timeout, fail-open to LLM. **E2E test results** (full HIL → gold cycle on local stack): | Test | Result | |---|---| | Run 1 cold (no cache) | 30.5s end-to-end, 2 atoms written silver (techniques + ai_tampered) | | Run 2 warm (same content) | 18.3s end-to-end (-40%), atom hit_count incremented to 1 | | Insert moderation queue entry | 200 | | Claim queue entry | 200 | | Resolve as 'corrected' with diff | 200 | | Brain atoms after resolve | Both promoted to gold, `human_validated=true`, `validator_user_id=mod-tester`, `human_corrections` stored as JSONB | | Cleanup (revert plan, disable brain, truncate atoms) | All systems back to safe defaults | **Final state after Track 1 + Track 2**: - All feature flags (`triage_enabled`, `brain_enabled`) → `false` (safe) - Test user plan reverted to Freemium - Brain atom DB cleaned - All 4 service containers healthy: `didi-agent-v3`, `didi-framework`, `didi-admin`, `didibrain-api` **Goal**: agent-v3 has client code ready to call brain. Feature flag OFF by default. When brain becomes available (Track 2), flip flag. #### Task 1.8.1 — Extend `shared/brain/client.ts` - [ ] File: `agent-v3/src/shared/brain/client.ts` (MODIFY — already has gatherFromBrain etc.) - [ ] Add: - `lookupAnalysisAtom(params)` — POST /v1/analysis_atom/lookup with timeout from Redis config - `writeAnalysisAtomAsync(params)` — fire-and-forget POST /v1/analysis_atom (only if tier='premium') - `patchAnalysisAtomGold(params)` — PATCH /v1/analysis_atom/:id with human_validated=true - [ ] Read `brain_enabled` flag from Redis. If false, return null/skip immediately - [ ] Read `brain_url`, `brain_lookup_timeout_ms`, `brain_per_component` from Redis - [ ] Compute `prompt_hash` (sha256 of system_prompt + user_template) and `framework_version` (sha256 of relevant Redis configs) - [ ] All errors caught silently, return null. Brain failure must NOT fail analysis. #### Task 1.8.2 — Integrate in `techniques/executor.ts` - [ ] File: `agent-v3/src/components/techniques/executor.ts` (MODIFY) - [ ] At start of execute(): ``` if (brain_enabled AND brain_per_component.techniques) { const lookup = await lookupAnalysisAtom(...); if (lookup?.hit) { if (lookup.atom.cache_tier === 'gold' || (lookup.atom.cache_tier === 'silver' AND lookup.staleness === 'fresh')) { return { ...lookup.atom.result_processed, _cache_tier: lookup.atom.cache_tier }; } } } // Else continue normal LLM run ``` - [ ] After LLM run (miss path): ``` if (tier === 'premium' AND brain_enabled AND llmConfidence >= brain_confidence_min_silver) { writeAnalysisAtomAsync(...); // fire-and-forget } ``` - [ ] DO NOT change return signature. UI just passes `_cache_tier` through verdict to extension UI. #### Task 1.8.3 — Integrate in `ai-tampered/executor.ts` - [ ] Same pattern as 1.8.2 with component='ai_tampered' #### Task 1.8.4 — Update moderation resolve handler - [ ] File: `agent-v3/src/api/moderation-routes.ts` (MODIFY) - [ ] In resolve handler, after saving corrections to PG: ``` if (action === 'corrected' AND trigger_brain_write !== false) { for (const component of ['techniques', 'ai_tampered', 'claims']) { if (corrections[component]) { const lookup = await lookupAnalysisAtom(...); if (lookup?.atom) { await patchAnalysisAtomGold({ atomId: lookup.atom.atom_id, validatorUserId: jwt.sub, humanCorrections: corrections[component], resultProcessed: applyCorrections(lookup.atom.result_processed, corrections[component]), }); } } } } if (action === 'approved') { // Same lookup, PATCH with human_corrections=null but human_validated=true (silver→gold without changes) } ``` - [ ] applyCorrections() helper: takes a result and a diff, returns corrected result. Pure function, unit-testable. - [ ] Brain failure here does NOT fail resolve. Log warning, continue. Resolve still saves to PG. **Validation** (only with brain available — see Track 2): - [ ] With brain_enabled=false in PG: brain not called, no log lines, normal flow - [ ] With brain_enabled=true but brain unreachable: 2s timeout, fallback to LLM, no failure - [ ] With brain reachable, hit path: latency drops from ~5s to ~200ms on cache hit - [ ] Resolve corrected → brain has new gold atom **Risk**: MEDIUM. Touches hot path (executors). Mitigation: feature flag, try/catch, timeout, brain optional. --- ## TRACK 2 — Brain Side — ✅ DONE 2026-05-01 (Phases 2.1-2.4) **Status**: All 4 endpoints live and tested E2E on local brain instance (10.11.10.12:8090). Production brain on 10.11.10.13 still on old version — to be deployed when integration is final. **Files created/modified in `ai_platform/modules/didi_brain/brain_api/`**: - `db.py` — extended `_SCHEMA_SQL` with `brain_analysis_atom` table + 4 indexes (idx_baa_lookup, idx_baa_gold, idx_baa_expires, idx_baa_prompt). Schema applied automatically on app startup (existing migration mechanism). - `schemas.py` — added 6 Pydantic models: `AnalysisAtomLookupRequest/Response`, `AnalysisAtomWriteRequest/Response`, `AnalysisAtomPatchRequest`, `AnalysisAtomStatsResponse`, plus `AnalysisAtomData` shared type. - `services/analysis_atom.py` (NEW) — 4 functions: `lookup`, `upsert`, `patch_to_gold`, `get_stats`. Tier decision logic (silver vs bronze based on confidence threshold). Gold-preservation in upsert ON CONFLICT (gold never downgraded). Hit count tracking. - `app.py` — added 4 endpoints + 7 imports. Existing endpoints UNCHANGED. **Endpoints live (port 8090)**: - `POST /v1/analysis_atom/lookup` — tier-agnostic match by (content_hash, component, prompt_hash). Returns hit/staleness. Bronze never served. - `POST /v1/analysis_atom` — upsert with gold preservation. Rejects tier=free (`skipped_reason="tier=free (premium-only ingest)"`). - `PATCH /v1/analysis_atom/{atom_id}` — promote to gold (sets human_validated, applies corrections, cache_tier=gold, expires_at=NULL). - `GET /v1/analysis_atom/stats` — counts by tier/component, hit rate 24h, writes 24h. **E2E test results (8/8 pass)**: | # | Test | Result | |---|---|---| | 1 | Write conf=78 → silver | ✅ | | 2 | Write conf=40 → bronze | ✅ | | 3 | Write tier=free → rejected | ✅ | | 4 | Lookup tier=free on premium silver atom → HIT | ✅ tier-agnostic | | 5 | Lookup bronze atom → MISS (filtered) | ✅ | | 6 | PATCH silver → gold + corrections | ✅ | | 7 | Lookup gold with different prompt → HIT (survives prompt change) | ✅ | | 8 | Stats reflects state | ✅ | **Bug fixed during testing**: `cache_tier` had Pydantic default "silver" which preempted the `_decide_cache_tier()` logic from llm_confidence. Changed to `cache_tier: AtomCacheTier | None = None` so server-side decision applies when caller doesn't override. **Resume context for next run**: - Brain v2 endpoints live on `http://10.11.10.12:8090` (this dev machine). - Production brain on `10.11.10.13:8090` is UNTOUCHED — still on old version. - agent-v3 `moderation_config.brain_url` defaults to `10.11.10.13` — when ready to test integration on this machine, set to `http://didibrain-api:8090` (Docker internal hostname) or `http://10.11.10.12:8090`. - DB cleaned post-test (atom table empty). **Risk**: LOW. New endpoints, additive schema. Existing `/v1/verification_cache` and `/v1/gather` flows untouched. --- ## TRACK 2 — Original detailed plan (preserved for reference) **RESOLVED 2026-05-01**: Brain code lives at `/home/admin365/didi_mono/ai_platform/modules/didi_brain/`. Same monorepo, separate Python service. Pattern to follow: existing `verification_cache.py` (claims-side cache). ### Brain stack reference | Layer | Tech | Where | |---|---|---| | FastAPI app | Python 3.11+, Uvicorn, Pydantic v2 | `brain_api/app.py` | | DB pool | asyncpg via `brain_api/db.py` (`brain_db.connect()`) | Direct PG | | Schemas (Pydantic) | `brain_api/schemas.py` | Add new models for analysis_atom | | Services (logic) | `brain_api/services/.py` | Add `analysis_atom.py` | | Storage | PG 16 + pgvector. Tables prefixed `brain_*` (alongside Atomic) | New: `brain_analysis_atom` | | Embeddings | BGE-M3 (1024-dim) via vLLM OpenAI-compat | `shared/embedding_client.py` | | Reranker | BGE-reranker-v2-m3 (only on /search & /gather, NOT on lookup) | `shared/reranker_client.py` (if exists) | ### PHASE 2.1 — Brain schema migration - [ ] File: `ai_platform/modules/didi_brain/sql/migrations/brain_analysis_atom_v1.sql` (NEW or follow brain's existing migration convention) - [ ] CREATE TABLE `brain_analysis_atom` per `BRAIN_V2_DESIGN.md` schema, with corrected UNIQUE `(content_hash, component, prompt_hash)` — tier OUT of unique key - [ ] Validate pgvector extension installed (it is — verification_cache uses it via Atomic) - [ ] Index ivfflat on content_embedding (lists=100) - [ ] TTL cleanup: cron job in `scripts/` deletes expired silver/bronze ### PHASE 2.2 — Brain Pydantic schemas - [ ] File: `ai_platform/modules/didi_brain/brain_api/schemas.py` (MODIFY — add new models) - [ ] Add: `AnalysisAtomLookupRequest`, `AnalysisAtomLookupResponse`, `AnalysisAtomWriteRequest`, `AnalysisAtomWriteResponse`, `AnalysisAtomPatchRequest`, `AnalysisAtomStatsResponse` - [ ] Match field types to `BRAIN_V2_DESIGN.md` exactly ### PHASE 2.3 — Brain service logic - [ ] File: `ai_platform/modules/didi_brain/brain_api/services/analysis_atom.py` (NEW) - [ ] Functions: - `lookup_atom(req, db, embed)` — exact match by SHA256, fallback semantic if allow_semantic_match - `write_atom(req, db, embed)` — INSERT or UPDATE if same key. Reject silently if tier='free' (per decision) - `patch_atom_gold(atom_id, req, db)` — UPDATE with human_validated=true, cache_tier='gold', expires_at=NULL - `get_stats(db)` — aggregates per tier/component ### PHASE 2.4 — Brain endpoints in `app.py` - [ ] File: `ai_platform/modules/didi_brain/brain_api/app.py` (MODIFY — add 4 routes) - [ ] `POST /v1/analysis_atom/lookup` → `lookup_atom` - [ ] `POST /v1/analysis_atom` → `write_atom` - [ ] `PATCH /v1/analysis_atom/{atom_id}` → `patch_atom_gold` - [ ] `GET /v1/analysis_atom/stats` → `get_stats` - [ ] Errors handled like existing endpoints (HTTPException with structured detail) ### PHASE 2.5 — Brain admin UI (in ai_platform/modules/dashboard) - [ ] File: `ai_platform/modules/dashboard/...` (location TBD — check this dashboard's structure) - [ ] Page: atom stats, browse atoms (paginated), force gold/demote silver/delete buttons - [ ] Existing AI platform dashboard already runs at port 51300 (per ai_platform/README.md) ### PHASE 2.6 — Bootstrap script (optional, for hit rate jump-start) - [ ] File: `ai_platform/modules/didi_brain/scripts/bootstrap_from_didi_history.py` (NEW) - [ ] Connects to DIDI cluster PG, reads `bos_analysis.analysis_session` older than 90d - [ ] For each row, computes content_hash + posts to `/v1/analysis_atom` (silver, tier='premium' only) - [ ] Skip if `confidence < 60` **Track 2 deployment**: brain Docker image rebuilt + redeployed on `10.11.10.13`. Existing `verification_cache` endpoints UNTOUCHED. --- ## TRACK 3 — Production cutover & role enforcement — ✅ Faza A+B DONE 2026-05-01 After Track 1 + Track 2 worked end-to-end on staging, the user decided to flip **brain local** (`10.11.10.12:8090`) into production and freeze the old brain on `10.11.10.13` as idle fallback. Plus enforce role-based access on admin dashboard so end-users (clients with `viewer`/`free_tier`/etc) cannot reach `/admin/*`. ### Faza A — Brain cutover 10.11.10.13 → 10.11.10.12 — ✅ DONE 2026-05-01 **Goal**: agent-v3 talks to local brain (with new analysis_atom endpoints), prod brain on `.13` runs idle as fallback. No deprovisioning yet. **Files modified**: - `agent-v3/docker-compose.yml` — `DIDI_BRAIN_URL` default changed from `http://10.11.10.13:8090` → `http://10.11.10.12:8090` (env var used by claims `verification_cache` flow) - `didiFramework/sql/migrations/011_add_moderation.sql` — `brain_url` DEFAULT changed from `http://10.11.10.13:8090` → `http://10.11.10.12:8090` (so new fresh deploys point local) - Live PG `bos_parammgmt.moderation_config.brain_url` → updated to `10.11.10.12:8090` via API + sync to Redis **Live state**: ``` brain_url: http://10.11.10.12:8090 brain_enabled: true triage_enabled: false (default safe — activate from UI when ready to use queue) ``` **Verified**: - agent-v3 recreated with new env (`docker exec didi-agent-v3 printenv DIDI_BRAIN_URL` → 10.11.10.12) - E2E analyze through pipeline → brain local hit verification_cache (claims=1) - Brain prod 10.11.10.13 still up but idle (no requests routed) - `GET http://10.11.10.13:8090/v1/analysis_atom/stats` → `{"detail":"Not Found"}` confirms old version, NEW endpoints only on local **Resume context**: - Brain prod on `.13` is left idle as 1-2 week fallback. Decommission deferred to Faza E. - All agent-v3 brain calls (gather, verification_cache, analysis_atom) → 10.11.10.12. - `triage_enabled=false` keeps queue empty until activated explicitly. ### Faza B — Role guards on admin dashboard — ✅ DONE 2026-05-01 **Goal**: end-users (clients with `viewer` role) cannot reach `/admin/*`. Get a clear 403 page with link to public app. Admin dashboard sidebar shows only relevant items per role. **Decision log update**: `viewer` (default Keycloak role on signup) = end-user / client. Should NEVER access admin dashboard. Admin dashboard is for staff only: `admin`, `moderator`, `senior_moderator`. **Files modified**: - `admin-dashboard/src/components/auth/ProtectedRoute.tsx` — extended with `requiredAnyRole?: string[]` (any-of gate). Redirects to `/unauthorized` on fail. - `admin-dashboard/src/components/auth/Unauthorized.tsx` (NEW) — 403 page with role list + "Go to Public App" + "Log out" buttons - `admin-dashboard/src/App.tsx`: - Added `/unauthorized` route - Wrapped AdminLayout in `ProtectedRoute requiredAnyRole={['admin','moderator','senior_moderator']}` - Per-route nested guards: `/framework`, `/users`, `/providers`, `/llm-components` → `requiredRole="admin"`; `/history`, `/moderation/*` → `requiredAnyRole=[admin,moderator,senior_moderator]` - `admin-dashboard/src/components/layout/AdminLayout.tsx` — sidebar restructured: - Configuration section (Framework, LLM Components, Providers) → admin-only (existing logic kept) - Management section split: Users → admin only; Analysis History → admin OR moderator; Moderation → moderator only **Permission matrix (effective)**: | Page | viewer / paid_tier / etc. | moderator | senior_moderator | admin | |------|---|---|---|---| | `/admin/*` (any) | ❌ 403 → Unauthorized | ✅ Dashboard + History + Moderation | ✅ same + force_gold_brain | ✅ everything | | `/users`, `/framework`, `/llm-components`, `/providers` | ❌ | ❌ | ❌ | ✅ | | `/history` | ❌ | ✅ | ✅ | ✅ | | `/moderation/*` | ❌ | ✅ | ✅ | ✅ | **Staging mode behavior**: `REACT_APP_STAGING_MODE=true` makes `hasRole()` always return true → all guards are no-ops locally for testing. In production with Keycloak, JWT roles enforce strictly. **Verified**: - TypeScript clean (`npx tsc --noEmit` exit 0) - `didi-admin` container rebuilt + recreated, healthy - `https://10.11.10.12:3000/admin/` → 200 - `https://10.11.10.12:3000/admin/unauthorized` → 200 (renders 403 page) - `https://10.11.10.12:3000/admin/moderation` → 200 (in staging, guards bypassed) **Risk**: LOW. New routes + UI gates only; no breaking changes to existing protected routes. ### Faza C — AI Platform dashboard reskin (React + Keycloak SSO) — ✅ DONE 2026-05-02 **RESULT**: Full reskin shipped. AI platform admin dashboard now React 19 + MUI 7 + TanStack Query + keycloak-js, side-by-side with original Jinja UI at `/v2/`. All 9 sub-phases (C.0–C.8) complete. Detailed log lives in `/home/admin365/didi_mono/ai_platform/modules/dashboard/AI_PLATFORM_RESKIN_PLAN.md` (~700 lines). **Highlights**: - 98 config keys CRUD-able via schema-driven `` across 8 modules + brain settings - 8 new brain admin endpoints (atom browse with filters, force-gold flow, taxonomy, extended stats) + matching React DataGrid pages - 5 React pages porting all 7 existing Jinja routes (Overview, History, Cost, Providers, Archive, AuditLog) with parity + drawer details - Each AI module (llm, embeddings, rerank, audio, video, catalog) gained a `RuntimeConfigClient` that polls dashboard `/api/config` every 30s and applies log-level + rate-limit changes live - Keycloak SSO: new client `ai-platform-dashboard` on `` realm `didi-clients`, distinct role `ai_platform_admin`. Hybrid backend auth (JWT first, legacy bearer fallback). DIDI admin-dashboard's existing Keycloak setup (Faza B) untouched. **Pre-requisite for production cutover** (see C.8 in plan doc): - Manual Keycloak setup: create client + role + assign to user - Set `DASHBOARD_KEYCLOAK_URL=https://`, `DASHBOARD_STAGING_MODE=false`, `VITE_STAGING_MODE=false` - Rebuild with build-args + redeploy **Pages to port** (current AI platform dashboard): - `overview.html` — services + usage stats - `archive.html` + `archive_detail.html` — knowledge graph atoms browsing - `audit.html` — audit log of actions - `config.html` — runtime config (knowledge atom config, ingest defaults) - `cost.html` — LLM cost per period - `history.html` + `history_detail.html` — ingest history - `providers.html` — search providers (Brave, Tavily, M17 SearXNG) + keys **Endpoints** (FastAPI on `dashboard/src/dashboard/api/routes/`): - `archive.py`, `config.py`, `health.py`, `history.py`, `ingest.py`, `pages.py`, `stats.py` **Auth model**: Replace `auth.py` (bearer token in PG) with Keycloak JWT middleware. Existing User table can map keycloak_id → role (still enforced in DB for audit trail, but JWT is source of truth). **Decision needed before starting**: - Where to host the React SPA? Same container or separate? Recommend same nginx serving SPA + reverse proxying API. - Brain admin UI (atom browse, force gold) — add as new section in this dashboard? Recommend yes. ### Faza D — Backend role guards (agent-v3) — ⏳ PENDING Currently agent-v3 moderation routes use **soft role check** (permits if no JWT roles array). For production, must be strict: every `/api/v3/moderation/*` endpoint MUST verify `jwtRoles.includes('moderator')` or return 403. Soft mode only when an explicit `STAGING_MODE=true` env var is present. **Single file change**: `agent-v3/src/api/moderation-routes.ts` — `requireRole()` helper tightened. ### Faza E — Decommission 10.11.10.13 brain — ⏳ PENDING (defer ~2 weeks) After local brain proves stable in production for 1-2 weeks: - Stop `didibrain-api` container on `.13` - Optionally archive its PG data (verification_cache atoms can be useful as bootstrap source) - Update DEPLOYMENT.md / ENDPOINTS.md to remove `.13` references --- ## INTEGRATION PHASE — End-to-end (estimated 1 day) After Track 1 phases 1.1-1.8 done AND Track 2 phases 2.1-2.2 done: #### Task INT.1 — Activate brain on staging - [ ] In moderation_config, set `brain_enabled=true`, `brain_url=` - [ ] Sync-redis - [ ] Run 100 analyses with same text → verify hit rate increases - [ ] Verify: cached responses come back faster (<500ms) #### Task INT.2 — Test full HIL → brain flow - [ ] Run an analysis (gets to queue via triage) - [ ] Moderator resolves with correction - [ ] Verify atom in brain became gold (`GET /v1/analysis_atom/:id` shows `cache_tier='gold', human_validated=true`) - [ ] Run same analysis again → verify response includes `_cache_tier='gold'` and frontend shows badge #### Task INT.3 — Performance test - [ ] Run 1000 concurrent analyses (mix of cached + new content) - [ ] Measure: P50/P95/P99 latency, brain lookup latency, brain write latency - [ ] Verify: P95 brain lookup <200ms, P99 <500ms - [ ] Verify: no analysis failures due to brain #### Task INT.4 — Production cutover - [ ] Apply migration on prod cluster (off-hours) - [ ] Deploy didi backend with feature flags OFF - [ ] Deploy admin dashboard - [ ] Set up moderator user on Keycloak prod - [ ] Activate triage_enabled=true - [ ] Monitor queue depth for 24h - [ ] If stable, activate brain_enabled=true - [ ] Monitor cost reduction + hit rate #### Task INT.5 — Bootstrap brain from prod history (optional) - [ ] Run script on prod historical sessions older than 90 days - [ ] Verify hit rate jump on subsequent traffic --- ## OPEN QUESTIONS (need answers before starting) These block parts of the plan. Get answers before phases that depend on them. | # | Question | Blocks | Default if unresolved | |---|---|---|---| | 1 | ~~AI platform repo location & deployment process for brain v2?~~ **RESOLVED 2026-05-01**: Brain is at `/home/admin365/didi_mono/ai_platform/modules/didi_brain/`. Stack: Python 3.11+, FastAPI, Pydantic v2, Postgres 16 + pgvector + Atomic (Rust KG). Embeddings BGE-M3 1024-dim. Service runs at `10.11.10.13:8090`. Existing services pattern in `brain_api/services/` (e.g., `verification_cache.py`). Follow same pattern for `analysis_atom`. | ~~Track 2~~ | Resolved | | 2 | First moderator user identity (existing user or new)? | Phase 1.7 | Use admin@didi.local | | 3 | Notification channel for new queue entries (email/Slack)? | Out of scope for v1 | Skip — UI polling is enough | | 4 | Senior moderator escalation flow specifics? | Phase 2 / future | Faza 2: senior=admin override only | | 5 | Soft vs hard delete on rejected sessions? | Phase 1.5 | Soft — keep for audit trail | | 6 | Privacy redaction for moderator (PII in input text)? | Phase 1.6 | Skip v1, document as known issue | | 7 | Brain embedding model exact dimensions? | Phase 2.1 | Block — query brain at start of Track 2 | | 8 | Cross-encoder reranking on atom lookup or only on /v1/search? | Phase 2.2 | Default: only on /v1/search; lookup is fast-path | | 9 | Atom TTL: 90 days silver default OK? | Phase 2.1 | Yes, configurable from UI later | --- ## DECISIONS LOG (already agreed, do not re-litigate) - 2026-04-30: SLA = instant with post-hoc correction. User sees verdict in 5s; corrections async. - 2026-04-30: Triage v1 strict — 1-2 moderators max, 10-20 reviews/day target. - 2026-04-30: Operational data (queue, status) in DIDI PG. Knowledge (atoms) in brain PG. Split. - 2026-04-30: 1 atom per (component × content) — techniques separate from ai_tampered separate from claims. - 2026-04-30: 3 tiers — gold (human_validated), silver (LLM cache), bronze (pending review, not served). - 2026-04-30: Confidence < 60 → bronze, not silver. - 2026-04-30: Brain write ONLY from premium tier. Read tier-agnostic (free benefits from premium cache). - 2026-04-30: Schema correction — UNIQUE (content_hash, component, prompt_hash) — tier OUT of unique key. - 2026-04-30: Zero hardcoded config. All in PG `bos_parammgmt.moderation_config` + sensitive_topic + moderation_role tables. Synced to Redis. - 2026-04-30: Brain admin UI lives in AI platform repo, NOT in didi admin dashboard. - 2026-04-30: Feature flags everywhere. Default OFF on first deploy. Activate progressively after validation. - 2026-05-01: `viewer` role (default Keycloak signup role) = end-user / client. NEVER reaches admin dashboard. Admin dashboard = staff-only (admin / moderator / senior_moderator). - 2026-05-01: Brain cutover — local `10.11.10.12:8090` is the production brain; old `10.11.10.13:8090` runs idle as fallback for 1-2 weeks before decommission. - 2026-05-01: AI platform dashboard reskin = full React 19 + MUI 7 + Keycloak SSO (same stack as admin dashboard). Admin-only access. Postponed to next session due to size. - 2026-05-01: Permission matrix locked: admin sees all, moderator + senior_moderator see Dashboard + History + Moderation, viewer/paid/free_tier see 403 → public app. - 2026-05-01: AI platform brain admin UI (browse atoms, force gold, stats) goes inside the reskinned AI platform dashboard, not a new app. --- ## FILE INVENTORY What gets created vs modified across the entire plan. ### New files (didi side) ``` agent-v3/sql/migrations/ ├── 010_add_moderation.sql NEW └── 010_rollback.sql NEW agent-v3/src/api/ └── moderation-routes.ts NEW agent-v3/src/components/moderation/ ├── triage.ts NEW └── queue-manager.ts NEW didiFramework/src/routes/ ├── moderation-config.ts NEW ├── sensitive-topics.ts NEW └── moderation-roles.ts NEW admin-dashboard/src/components/ModerationSettings/ ├── index.tsx NEW ├── TriageCard.tsx NEW ├── BrainClientCard.tsx NEW ├── SensitiveTopicsCard.tsx NEW ├── RolesCard.tsx NEW └── api.ts NEW admin-dashboard/src/components/Moderation/ ├── ModerationQueue.tsx NEW ├── ModerationDetail.tsx NEW ├── ModerationStats.tsx NEW ├── EditVerdictPanel.tsx NEW ├── EditTechniquesPanel.tsx NEW ├── EditAITamperedPanel.tsx NEW ├── EditClaimsPanel.tsx NEW ├── api.ts NEW └── types.ts NEW ``` ### Modified files (didi side, surgical edits only) ``` agent-v3/src/ ├── index.ts MODIFY (mount /api/v3/moderation) ├── api/pipeline-routes.ts MODIFY (accept user_flagged in body) ├── components/pipeline/executor.ts MODIFY (call triage after persist, with try/catch) ├── components/techniques/executor.ts MODIFY (brain lookup + write — phase 1.8) ├── components/ai-tampered/executor.ts MODIFY (brain lookup + write — phase 1.8) ├── queue/aggregator.ts MODIFY (call triage after persist, with try/catch) └── shared/brain/client.ts MODIFY (add lookupAnalysisAtom, writeAnalysisAtomAsync, patchAnalysisAtomGold) didiFramework/src/ ├── server.ts MODIFY (mount 3 new route modules) └── routes/sync-redis.ts MODIFY (add 3 new Redis keys) admin-dashboard/src/ ├── App.tsx MODIFY (add 3 routes for /moderation/*) └── components/ ├── LLMComponentsConfig/index.tsx MODIFY (add Moderation toggle to top-level selector) └── dashboard/ServicesDashboard.tsx MODIFY (add sidebar link for moderators) ``` ### NOT touched (existing flows preserved) ``` agent-v3/src/ ├── components/claims/executor.ts UNTOUCHED (uses verification_cache pattern, working) ├── components/source-assessment/executor.ts UNTOUCHED ├── components/pipeline/verdict-calculator.ts UNTOUCHED ├── components/pipeline/verdict-explanation.ts UNTOUCHED └── shared/persistence/* UNTOUCHED didiFramework/src/routes/ ├── verdicts.ts UNTOUCHED (already extended with runtime-config in prev session) ├── input-profiles.ts UNTOUCHED └── (everything else) UNTOUCHED admin-dashboard/src/components/ ├── LLMComponentsConfig/VerdictConfig.tsx UNTOUCHED (already extended in prev session) └── (everything else) UNTOUCHED ``` --- ## RUNBOOK FOR FRESH SESSION RESUME If you (the new assistant) are reading this in a fresh session: 1. **Read all 3 docs** (this + HIL_MODERATION_DESIGN.md + BRAIN_V2_DESIGN.md). Do not start work without context. 2. **Verify state of codebase**: ```bash ls /home/admin365/didi_mono/backend/services/orchestration-layer/agent-v3/sql/migrations/ # Look for 010_*.sql — if exists, Phase 1.1 done ls /home/admin365/didi_mono/backend/services/orchestration-layer/didiFramework/src/routes/ | grep moderation # If files exist, Phase 1.2 partly done ls /home/admin365/didi_mono/backend/admin-dashboard/src/components/Moderation* 2>/dev/null # If exists, Phase 1.3 or 1.6 partly done ``` 3. **Verify DB state**: ```bash docker exec didi-framework node -e " const {Pool} = require('pg'); const p = new Pool({host:'10.11.50.167',port:5000,user:'bos_interface',password:'interface',database:'DIDI'}); p.query(\"SELECT table_name FROM information_schema.tables WHERE table_schema='bos_parammgmt' AND table_name IN ('moderation_config','sensitive_topic','moderation_role')\").then(r => {console.log(r.rows); p.end();}); " # If 3 rows → Phase 1.1 done ``` 4. **Find first unchecked task** in the order: 1.1 → 1.2 → 1.3 → 1.4 → 1.5 → 1.6 → 1.7 → 1.8 → INT.1+ 5. **Update this file** as you go: change `[ ]` → `[x]` immediately after completing each subtask. Add `→ YYYY-MM-DD note` if useful. 6. **Critical rule**: NEVER ship code without: - Type-check passing (`npx tsc --noEmit`) - Existing endpoints still responding (curl smoke test on `/api/v3/health`, `/api/v3/pipeline/analyze` with simple text) - Feature flag default OFF for new path 7. **If stuck or in doubt**: prefer asking the user over guessing. Don't introduce hardcoded values. Don't break existing flows. 8. **Update Decisions Log** if user makes new decisions during the session. Append, don't overwrite. --- ## ESTIMATED TIMELINE Track 1 didi side: ~5-6 working days for one person, given parallel work. - Phase 1.1: 0.5 day - Phase 1.2: 0.5 day - Phase 1.3: 1 day - Phase 1.4: 1 day - Phase 1.5: 0.5 day - Phase 1.6: 2 days - Phase 1.7: 0.25 day - Phase 1.8: 1 day (mostly idle waiting for brain) Track 2 brain side: AI platform team estimate (separate) Integration: 1 day end-to-end test + cutover **Total didi side**: ~6 days. Brain side parallel. Integration adds 1 day. --- ## END OF PLAN