Multi-Agent AI Reliability Commander
The AI that doesn't just tell you something broke it tells you what will break, why, what to do about it, and proves it with numbers.
Most modern maintenance operations rely heavily on either slow reactionary metrics (fixing it when it breaks) or blind-scheduled preventive tasks (costly and inefficient). Even modern predictive maintenance platforms stop at "anomaly detected," leaving you guessing why, and what actions you should take next.
Hephaestus is an end-to-end, multi-agent artificial intelligence predictive operations brain. It closes the full operational loop autonomously by using 10 specialized intelligent agents configured in a directed graph.
Hephaestus tracks telemetry data to predict equipment failure, constructs an evidence-backed hypothesis root cause, builds a human-readable maintenance playbook, dynamically optimizes resource planning constraints (crew, parts, budgets), and mathematically simulates projected cost-savings prior to human intervention.
- High-Precision Risk Detection: Uses machine-learning models (Isolation Forests, Gradient Boosting) to compute unsupervised anomalies and failure probabilities (RUL).
- Causal Reasoning: Not just black-box predictions. Hephaestus generates hypothesis graphs with detailed certainty levels using SHAP explainability.
- Auto-Generated Intervention Options: Predicts the skills, parts, and downtime requirements for several maintenance scenarios.
- Mathematical Optimization: Balances options against your operational constraints via a multi-objective scoring function (Cost / Downtime / SLA Compliance).
- Monte Carlo Simulation Projector: Projects hypothetical 30-day trajectories and risk curves with uncertainty bounds.
- Stakeholder Reporting: Dynamically writes step-by-step playbooks for operators, financial summaries for managers, and step-by-step decision audit trails for compliance.
Hephaestus coordinates 10 autonomous software agents. Each uses different models suited to their specialty (ML models for tracking telemetry, LLMs for reasoning and synthesis, mathematical solvers for simulations):
- Intake Agent: Parses and standardizes incoming datasets.
- Quality Agent: Checks for feature drift, data missingness, or frozen sensors.
- Sentinel Agent: Analyzes the telemetry window and flags anomalous behaviour.
- Prognostics Agent: Estimates precise failure probability risk horizons.
- Causal Agent: Examines maintenance history to correlate root-cause reasons.
- Planner Agent: Creates intervention schedules (parts/skills dependencies).
- Optimizer Agent: Adjusts plans algorithmically inside hard constraints (e.g. budgets).
- Simulation Agent: Verifies expected outcomes vs. "Doing nothing".
- Reporter Agent: Converts output payload intelligence to readable summaries.
- Governance Agent: Final safety check before pushing human-in-the-loop playbooks.
This is a monorepo partitioned into three core domain layers:
hephaestus/
├── frontend/ # Next.js / TypeScript / React / Tailwind dashboard app
├── backend/ # FastAPI / Python backend hosting the primary generic REST routes
└── ml/ # The Core intelligence module (`aegis`)
└── aegis/
├── agents/ # Holds all 10 specialized Python agents + the Orchestrator
├── config/ # Decision weighting values, hyperparams and security profiles
├── data/ # Schemas, real-time validators and synthetic asset generation
├── models/ # Predictive ML models (anomaly, failure risk, survival regression)
├── planning/ # Engine housing hard/soft constraints & mathematical solvers
├── reporting/ # Final composition templates for output summaries
├── simulation/ # Scenario engine logic and Monte Carlo calculations
├── storage/ # Specialized layer interacting with DB (PostgreSQL internals)
├── telemetry/ # Tracing, scoring metrics on agent behaviour and system logs
└── tests/ # Segmented suite (Unit, Integration, E2E)
- Core & Backend API: Python 3.11+, FastAPI, Pydantic v2, Uvicorn, Celery/Redis
- Artificial Intelligence Framework: LangGraph (Orchestration), Scikit-Learn (Isolation Forest), XGBoost/LightGBM, Survival regression, SHAP (Explainability features), Ollama local inference (Primary) / Google Gemini API (Fallback).
- Solver & Simulation: Numpy / Scipy
- Frontend Dashboard: Next.js (App Router), React, TypeScript, Tailwind CSS, ECharts/Recharts
- Storage / Database: PostgreSQL, Redis, optionally TimescaleDB
- Development / DevOps: Docker + Docker Compose, GitHub Actions, Pytest, Ruff / Mypy
- Repository docs/config files should use kebab-case when multi-word naming is needed.
- API route paths use kebab-case segment style only when segments are multi-word.
- Python modules remain snake_case by design to preserve valid imports and tooling compatibility.
(Development Guide coming soon once modules are initialized)
graph TB
%% DATA SOURCES
A[Telemetry Data<br>CSV / JSON / Streams]
B[Event Logs]
C[Maintenance History]
%% INGESTION
A --> D[Intake Agent<br>FastAPI + Pydantic]
B --> D
C --> D
%% DATA VALIDATION
D --> E[Quality Agent<br>Pandas + Numpy<br>Data Validation]
%% EVENT BUS
E --> F[Event Bus<br>Redis Streams]
%% ANOMALY DETECTION
F --> G[Sentinel Agent<br>Anomaly Detection<br>Isolation Forest + Z Score]
%% FAILURE PREDICTION
G --> H[Prognostics Agent<br>Failure Prediction<br>XGBoost / LightGBM]
%% ROOT CAUSE ANALYSIS
H --> I[Causal Agent<br>Explainability<br>SHAP Analysis]
%% PLAN GENERATION
I --> J[Planner Agent<br>LLM Planning<br>Mistral / Llama via Ollama]
%% OPTIMIZATION
J --> K[Optimizer Agent<br>Objective Function Scoring]
%% SIMULATION
K --> L[Simulation Agent<br>Monte Carlo Engine]
%% PARALLEL SIMULATION
L --> M1[Simulation Worker 1]
L --> M2[Simulation Worker 2]
L --> M3[Simulation Worker N]
%% WORKER TECHNOLOGY
M1 --> N[Celery Workers + Redis Queue]
M2 --> N
M3 --> N
%% AGGREGATION
N --> O[Simulation Result Aggregator<br>Pandas / Numpy]
%% DECISION OUTPUT
O --> P[Decision Package<br>Risk + Cost + Downtime]
%% REPORT GENERATION
P --> Q[Reporter Agent<br>LLM Report Generator]
%% STORAGE
Q --> R[(PostgreSQL<br>Decision & Telemetry Storage)]
%% FRONTEND
R --> S[Dashboard<br>Next.js + React + Tailwind]