Sentinel Commander β Documentation (v2026.08.001)
Text Only ##
++########++
++++###+++++++
+++++++++++-++-.
++++++#++++++-+-
++######+++++--
+++ ####+ -+-
++++######++-
++++#######+-#
++++######++-+.-+
++++++###+++##+---++#
+++++++++###+#######++++++++
-+++++++++++###+#####+#+++++----..-
++-++++++++#######+######++++++++--.. .
---++++##+++#####++####++++#++++---.....
--+++++############++###++++#+++---........
--++++++##++##########++++++++++----.........
--+++++#####++#####+#+++++++++++++-.-.........
--++++#####+########++++++-+++++-++--------.-..
++++##########+####++++++-+++++++-+++---------.
S E N T I N E L C O M M A N D E R
v2026.08.001
Hybrid AI Log Monitor & Analyzer for Linux Infrastructure
Quick Navigation
What is Sentinel?
Sentinel Commander is an advanced hybrid monitoring system with a full AI layer designed for enterprise Linux infrastructure. It combines a Pull approach (inotify log tailing) with a Push approach (remote Python agents POSTing telemetry). Core integrations: Hailo AI HAT 2+ NPU (hailo-ollama), local/external Ollama LLM, ChromaDB (RAG).
As of v2026.08.001 the system has accumulated 1 050 automated tests and 80/100 AI roadmap items across 24 new AI modules covering diagnostics, verification, remediation planning, correlation, baselining, forecasting, and knowledge management.
What's New in v2026.08.001 (2026-08-03)
| Feature | Detail |
| Tests: 938 β 1050 | 112 new tests covering knowledge base, infra audit, and dependency mapping modules |
| 80/100 AI roadmap items | Up from 66/100 in v2026.07 |
| 3 critical SSH bugs fixed | SSH broken July 30βAug 1 in diagnostics + remediation; fully resolved |
knowledge.py | Runbooks, prevention hints, training pairs, KB transfer to/from other instances |
infra_audit.py | Config drift detection, zombie process auditing, cert expiry, post-reboot checklist, docs accuracy check |
dependencies.py | Inferred host dependency graph, blast-radius calc, simulated shutdown impact |
host-setup.md | New guide: preparing monitored hosts for AI diagnostics (sentinel user, sudoers, journal group) |
What's New in v2026.07.001 (2026-07-31)
AI Layer β 21 New Modules
| Module | What it does |
ai_guard.py | Prompt-injection defence for log content, hourly action cap, loop detection |
ai_verify.py | Hallucination check β validates AI claims against known infrastructure |
ai_profiles.py | Context-window profiles per task type (triage, deep analysis, chat) |
ai_runtime.py | Response cache, consistency guard, token budget, model routing |
diagnostics.py | Fixed read-only command catalog β AI picks command IDs, never writes raw shell; executes and interprets real output |
fix_verify.py | Post-fix verification ~15 min after each remediation attempt; failures fed back as anti-patterns |
remediation.py | Graduated ladder: observe β reload β restart β reboot |
remediation_plan.py | Rollback, contextual risk assessment, dry-run mode, work queue |
policy.py | Block explanation, allowlist/auto-execute proposals |
escalation.py | Escalation with context (what was already tried + why it failed) |
correlate.py | Change correlation, causal chains, cross-host patterns |
incident_analysis.py | Common denominator finder, incident timeline, cascade detection, ranked hypotheses |
trend_detect.py | Silent degradation (regression + rΒ²), missing-signal detection |
baseline.py | Per-host normal profile, seasonality detection, auth-log audit |
alert_quality.py | False-alarm mining from historical data |
playbooks.py | Procedures learned from manual fixes β builds institutional memory |
foresight.py | Capacity forecast, weekly infrastructure outlook |
unmatched.py | Random sampling of log lines that no plugin catches |
rag_utils.py | Compression, hybrid search, citations, chunk management |
knowledge.py | (preview in 07, promoted to stable in 08) |
infra_audit.py | (preview in 07, promoted to stable in 08) |
Scale
- Tests: 317 β 938 (+621 new tests covering the full AI layer)
- AI roadmap: 66/100 items complete
- Full decision audit trail for every AI action
- All AI modules fully integrated into the Scheduler maintenance loop
What's New in v2026.06.024 (since v2026.06.005)
Security Hardening
| Feature | Detail |
| 2FA / TOTP | Full stack: pyotp (RFC 6238), QR code enrollment (Google Authenticator / Authy), two-step login flow, user_totp DB table, admin can disable 2FA per user |
| bcrypt password hashing | web.password_hash: "$2b$12$..." takes priority over plaintext; "Hash password" button in Settings generates the hash; viewer password analogous |
| CSRF protection | Token in session + SameSite=Strict cookie + global fetch wrapper adds X-CSRF-Token to every POST/PUT/DELETE |
| Brute-force protection | Form login + Basic auth: IP ban after 5 failed attempts for 300 s; auto-register rate limit 10/min per IP |
| XSS hardening | All AI replies escaped via html.escape() before innerHTML β log content is attacker-controlled |
| SSH hardening | ssh_utils.py: accept-new + UserKnownHostsFile instead of StrictHostKeyChecking=no; ssh-keyscan on agent registration; known_hosts UI (view / rescan / delete keys) |
| API key scopes | Fine-grained: read:issues, write:actions, admin:users β with backwards compatibility |
| Secrets handling | {SECRET:ENV_VAR} substitution in config.yaml (vars deleted from os.environ after use); /api/config/view masks password/token/secret/api_key as *** |
| Session control | Absolute 12 h timeout (security.session_max_hours), role refresh from DB every 5 min, revoked sessions persisted in DB |
| Audit trail | config_audit table (who, when, IP, which keys), 403/401 access audit, audit trail viewer in Settings |
| SSRF + input validation | /api/admin/validate_url rejects private IP ranges; hostname regex validation on all SSH/ingest endpoints; int_param() bounds-checking on 20+ endpoints |
| Security self-check | /api/admin/security_check returns security grade A/B/C/D |
Notifications & Integrations
| Feature | Detail |
| Outbound channels | MS Teams Β· Slack Β· PagerDuty Β· Discord (embeds) Β· Telegram bot Β· Opsgenie (Events API v2) Β· ntfy.sh Β· Gotify Β· SMTP e-mail (STARTTLS 587 / SSL 465) Β· Matrix Β· Home Assistant Β· MQTT Β· generic Webhook (HMAC-SHA256 + replay protection) |
| Inbound webhooks | /api/inbound/grafana (legacy + unified alerting) Β· /api/inbound/alertmanager (Prometheus AM) Β· /api/inbound/zabbix (Media Type flat JSON) |
| Reliability | Retry queue with exponential backoff (30 s/120 s/300 s, max 3 attempts); per-severity throttling (critical/security/root 15 min, high 1 h, medium/low 4 h) |
| Lifecycle webhooks | Configurable webhooks fired on issue CREATED / ACKNOWLEDGED / RESOLVED |
| Gitea issue sync | Critical issues automatically opened in a Gitea repository |
| Prometheus | GET /metrics scrape endpoint + pushgateway export |
Analytics, Agent Fleet, Operations
- Health score per host (AβD grade), 7-day issue forecast, SLA & alert-fatigue reports
- Batch SSH (50 hosts, ThreadPoolExecutor), per-agent thresholds, CVE scanner, package inventory
/healthz probe, config backup/restore with snapshots, SIGHUP hot-reload, FIM, Ansible runner - Composite DB indexes, WAL tuning, HTTP caching (ETag + 304), SocketIO backpressure, virtual scroll
Engineering Quality
- 181 automated tests (v2026.06.024 baseline) β now 1 050 in v2026.08.001
- CI pipeline β
pytest + node --check + make build, pre-push git hook - Linting β ruff (Python), ESLint, pinned
requirements.txt
Feature Overview
| Area | Capabilities |
| AI Inference | Hailo-10H NPU (hailo-ollama) Β· CPU Ollama Β· external API Β· runtime model switch |
| RAG Knowledge Base | ChromaDB + nomic-embed-text Β· BM25 TFΓIDF fallback Β· custom file upload (.md/.txt/.pdf/.docx/.csv) Β· one-click reindex |
| AI Safety | Prompt-injection defence Β· hourly action cap Β· loop detection Β· hallucination check Β· full audit trail (ai_guard.py, ai_verify.py) |
| AI Diagnostics | Fixed read-only command catalog β AI picks IDs, never raw shell; executes and interprets real output (diagnostics.py) |
| AI Verification | Post-fix check ~15 min after each remediation; failures flagged as anti-patterns (fix_verify.py) |
| AI Remediation | Graduated ladder: observe β reload β restart β reboot Β· rollback Β· dry-run Β· contextual risk (remediation.py, remediation_plan.py) |
| AI Correlation | Causal chains Β· change correlation Β· cross-host patterns Β· cascade detection Β· incident timelines Β· ranked hypotheses (correlate.py, incident_analysis.py) |
| AI Baselining | Per-host normal profile Β· seasonality Β· silent degradation (regression + rΒ²) Β· missing signals (baseline.py, trend_detect.py) |
| AI Foresight | Capacity forecast Β· weekly outlook Β· false-alarm mining Β· unmatched log sampling (foresight.py, alert_quality.py, unmatched.py) |
| Knowledge Base | Runbooks Β· prevention hints Β· training pairs Β· KB transfer between instances (knowledge.py) |
| Infra Audit | Config drift Β· zombie processes Β· cert expiry Β· post-reboot checklist Β· docs accuracy check Β· dependency blast-radius (infra_audit.py, dependencies.py) |
| Hybrid Telemetry | Pull (inotify logs) + Push (agents via Bearer token) Β· multiple IPs per agent Β· agent version tracking (SHA) |
| Autofix | AI proposes fix β admin Approve/Reject β SSH exec on mgmt node Β· allowed-commands allowlist Β· autonomous exec |
| Predictive Analytics | TTC (Time-To-Critical) for disks Β· Mann-Kendall trend test Β· linear regression forecast Β· capacity planning |
| Security Profiler | Brute-force, sudo abuse, CVE scan, unauthorised ports, honeypot, FIM, SSL expiry |
| Notifications | 13 outbound channels Β· 3 inbound webhooks Β· retry queue Β· per-severity throttle Β· per-detector/channel toggles |
| Prometheus | GET /metrics scrape + pushgateway export; auth via scrape_token |
| Dashboard | Stat cards Β· interactive min/max/avg charts Β· trend chart Β· donut Β· health trend Β· flapping widget Β· live clock |
| Auth | viewer / admin / superadmin Β· LDAP (lldap + OpenLDAP) Β· 2FA/TOTP Β· bcrypt Β· rate-limit + IP ban Β· CSRF |
| Issue Workflow | active β acknowledged β validating β resolved Β· escalation rules Β· lifecycle webhooks |
| Auto-Remediation | One-shot SSH fix Β· allowed_commands with auto_execute Β· AUTOFAIL issues Β· SSH jump host (ProxyJump) Β· Ansible runner |
| Plugin Hot-Reload | POST /api/plugins/reload Β· SIGHUP full reload Β· Pattern Editor with regex tester + AI pattern suggestions |
| Telemetry | Anomaly detection (3Ο) Β· fixed thresholds Β· per-agent thresholds Β· InfluxDB export Β· heatmap Β· health score history |
| Topology | Agent topology map Β· plugin dependency graph Β· SNMP CDP/LLDP Β· Canvas force-directed graph |
| SSH Actions | Jump host (ProxyJump) Β· SSH modal (admin+) Β· streaming output (SSE) Β· batch SSH Β· known_hosts management |
| API Docs | GET /api/docs β Swagger UI Β· GET /api/openapi.json β OpenAPI 3.0 spec |
| Hailo TUI | hailo_models.py β Unicode TUI model manager: htop-style CPU/Mem bars, RX/TX, NPU arch+FW, TPS benchmark |
| UI i18n | Czech (default) Β· English toggle Β· localStorage persistence Β· timezone display config |
| Security | Symlink containment Β· upload limit (5 MB) Β· secure_filename Β· timing-safe token verify Β· CSP headers Β· secrets masking Β· SSRF guard |
Architecture
Text Only Log files ββinotifyβββΆ watcher.py βββΆ plugins[] βββΆ api.report_problem()
Remote agents ββPOSTβββΆ /api/v1/agent/ingest β
Grafana/AM/Zabbix ββPOSTβββΆ /api/inbound/* β
βΌ
state.py (SQLite WAL)
β
ββββββββββββββββββββββββββββββββββββ€
βΌ βΌ
ollama_service.py Flask + SocketIO :5050
(AI worker pool) (chat_service.py + routes/)
β β
βββββββββββββββββΌββββββββββββββββ ββββββββββ΄βββββββββ
βΌ βΌ βΌ βΌ βΌ
hailo-ollama Ollama CPU external scheduler.py notifier.py
(NPU :8000) (:11434) API (maintenance) (13 channels)
Core modules
| File | Responsibility |
chat_service.py | Flask/SocketIO app factory, RBAC, WebSocket |
auth.py | Authentication, LDAP, 2FA/TOTP, bcrypt, sessions, API key verify |
state.py (state_base/issues/agents) | SQLite WAL orchestration β issues, telemetry, agents |
watcher.py | inotify filesystem events, hot config reload, FIM |
plugin_manager.py | Dynamic plugin loading, pattern routing, hot-reload |
ollama_service.py | AI worker thread pool, model switching |
rag.py | ChromaDB, nomic-embed-text, BM25 fallback |
actions.py | Autofix lifecycle β create, approve, reject, SSH exec |
notifier.py | All outbound notifications + retry queue + throttling |
scheduler.py | Background maintenance (minute / hourly / nightly tiers) |
ssh_utils.py | Central SSH security β build_ssh_cmd(), host key scanning |
analytics.py | TTC, Mann-Kendall, Z-Score, health score, forecast |
topology.py | Agent topology builder, SNMP CDP/LLDP |
routes/ | Flask Blueprints: main, issues, agents, actions, system, export, integrations, chat |
AI layer modules (v2026.07+)
| File | Responsibility |
ai_guard.py | Prompt-injection defence, action cap, loop detection |
ai_verify.py | Hallucination check against known infrastructure |
ai_profiles.py | Context-window profiles per task type |
ai_runtime.py | Response cache, consistency, token budget, model routing |
diagnostics.py | Fixed read-only command catalog β AI picks IDs, not raw shell |
fix_verify.py | Post-fix verification; failures feed back as anti-patterns |
remediation.py | Graduated remediation ladder with rollback |
remediation_plan.py | Risk assessment, dry-run, work queue |
policy.py | Block explanation, allowlist/auto-execute proposals |
escalation.py | Escalation with prior-attempt context |
correlate.py | Change correlation, causal chains, cross-host patterns |
incident_analysis.py | Timeline, cascade detection, ranked hypotheses |
trend_detect.py | Silent degradation, missing signals |
baseline.py | Per-host normal profile, seasonality, auth-log audit |
alert_quality.py | False-alarm mining from historical data |
playbooks.py | Procedures learned from manual fixes |
foresight.py | Capacity forecast, weekly outlook |
unmatched.py | Sampling of uncaught log lines |
rag_utils.py | Compression, hybrid search, citations, chunking |
knowledge.py | Runbooks, prevention hints, KB transfer |
infra_audit.py | Config drift, zombies, certs, post-reboot check |
dependencies.py | Host dependency graph, blast-radius, shutdown simulation |
Data Flow
A. Ingest & Detection
- Log line written to
/var/log/sentinel/logs/ β inotify event watcher.py reads new lines via mmap PluginManager routes lines to matching detectors by file mask - Detector generates unique key (e.g.
DISK_FULL|proxmox01|/data) - State written to
problems table β task pushed to task_queue - Push path: agents POST
/api/v1/agent/ingest (or /api/v1/ingest/bulk for batches); metrics payload is checked against per-agent thresholds
B. AI Inference + RAG
- AI worker fetches task from
task_queue - Loads prompt template by channel (security, clusters, infra, root, icinga)
- RAG: ChromaDB query β relevant context injected into system prompt
- Prompt dispatched to Ollama (NPU/CPU/external)
- If response contains remediation script β
actions.py creates pending action
actions.py creates DB record with proposed SSH command - Frontend emits
new_action WebSocket event to Web UI - Admin with
admin or superadmin clicks Approve or Reject - On Approve: SSH connection to management node β command execution (allowlist pre-validated)
- STDOUT/STDERR logged β incident enters
validating state
D. Notification Pipeline
- Issue saved β
notifier.send_notification() fan-out to all enabled channels - Per-detector and per-channel toggles checked first
- Per-severity throttle applied (critical 15 min β¦ low 4 h)
- Failures enter the retry queue (30 s β 120 s β 300 s backoff)
- Lifecycle webhooks fired on CREATED / ACKNOWLEDGED / RESOLVED
Issue Workflow
Text Only detected βββΆ active βββΆ acknowledged (ββ button)
β β
β validating βββΆ resolved
β
ββββΆ (auto-resolved by detector or expiry rule)
Escalation rules: If an issue stays active or acknowledged for more than N hours without resolution, its severity is automatically raised to the next level.
Lifecycle webhooks can notify external systems on every transition, and critical issues can be mirrored into Gitea.
Installation
Bashgit clone <repo> /opt/Sentinel
sudo bash /opt/Sentinel/install.sh # Debian/Ubuntu/RHEL/Rocky/Pi OS
sudo python3 /opt/Sentinel/sentinel_init.py # interactive config wizard
sudo systemctl enable --now sentinel
The setup wizard refuses default passwords, generates the systemd unit with WatchdogSec=900 and a WAL-checkpoint ExecStartPre, and creates /var/lib/sentinel.
Key configuration (config.yaml)
YAMLweb:
port: 5050
password_hash: "$2b$12$..." # bcrypt β preferred over plaintext `password`
security:
login_max_attempts: 5
login_ban_time: 300
session_max_hours: 12
ollama:
url: "http://localhost:11434"
model: "llama3.2"
workers: 3
hailo_ollama:
enabled: false
url: "http://localhost:8000"
ldap:
enabled: false
host: "ldaps://ldap.example.com"
base_dn: "dc=example,dc=com"
prometheus:
enabled: true
scrape_token: "{SECRET:PROM_TOKEN}" # env-var substitution
telemetry_alerts:
cpu_critical: 95
disk_critical: 95
temp_critical: 85
fim:
enabled: true
paths: [/etc/passwd, /etc/shadow, /etc/ssh/sshd_config]
Requirements
- Python 3.13+ Β· Flask Β· Flask-SocketIO Β· ChromaDB Β· paho-mqtt
pyotp (2FA) Β· qrcode + pillow (QR enrollment) Β· bcrypt Β· jsonschema - Ollama with
nomic-embed-text (for embeddings) - Optional: Hailo AI HAT 2+ with hailo-ollama 5.3.0 for NPU inference
Component Ecosystem
| Component | Description | Port |
| Sentinel | Central server (this doc) | 5050 |
| sentinel-agent | Push agent on each monitored node | β |
| sentinel-overhealth | SSH pull orchestrator (cron) | β |
| sentinel-plugins | 11 detector plugins | β |
| sentinel-alert | Standalone network security dashboard (incl. MikroTik + PiHole) | 5056 |
| sentinel-hw | Physical RPi robot | 5055 |
| sentinel-app | Android mobile client | β |
| sentinel-console | TUI terminal client | β |