"# DevOps Monitoring & Alerting Stack
A Docker-based monitoring and alerting solution using Prometheus, Grafana, Node Exporter, and AlertManager. This stack provides comprehensive system monitoring, visualization, and alert notifications via email.
- Overview
- Prerequisites
- Quick Start
- Services
- Access Points
- Configuration
- Architecture
- Alert Rules
- Usage
- Troubleshooting
This project sets up a complete monitoring ecosystem that:
- Collects system metrics from your infrastructure
- Stores and queries metrics using Prometheus
- Visualizes data through Grafana dashboards
- Triggers alerts based on predefined rules
- Sends notifications via email
Before starting, ensure you have:
- Docker (v20.10 or higher)
- Docker Compose (v1.29 or higher)
- A valid Gmail account (for email alerts)
- At least 2GB of RAM allocated to Docker
-
Clone or navigate to the project directory:
cd d:\Devops
-
Update AlertManager configuration:
- Edit
alertmanager.ymland update Gmail credentials with an app-specific password
- Edit
-
Start all services:
docker-compose up -d
-
Verify services are running:
docker-compose ps
-
Access the services:
- Prometheus: http://localhost:9090
- Grafana: http://localhost:3000
- AlertManager: http://localhost:9093
- Port: 9090
- Purpose: Metrics collection and storage
- Configuration:
prometheus.yml - Scrape Interval: 5 seconds
- Features:
- Collects metrics from Node Exporter
- Evaluates alert rules
- Routes alerts to AlertManager
- Port: 3000
- Purpose: Metrics visualization and dashboards
- Default Credentials:
- Username:
admin - Password:
admin
- Username:
- Features:
- Create custom dashboards
- Visualize Prometheus metrics
- Alert management UI
- Port: 9100
- Purpose: Exports system and hardware metrics
- Metrics Exposed:
- CPU usage
- Memory usage
- Disk I/O
- Network statistics
- Process information
- Port: 9093
- Purpose: Handles and routes alerts
- Configuration:
alertmanager.yml - Notification Method: Email
| Service | URL | Port |
|---|---|---|
| Prometheus | http://localhost:9090 | 9090 |
| Grafana | http://localhost:3000 | 3000 |
| Node Exporter Metrics | http://localhost:9100/metrics | 9100 |
| AlertManager | http://localhost:9093 | 9093 |
Global Settings:
- Scrape interval: 5 seconds
- Target: Node Exporter on port 9100
Alert Configuration:
- Rule file: alert.rules.yml
- AlertManager target: localhost:9093Triggered alerts:
- HighCPUUsage: Triggers when CPU usage > 5% for 1 second
- HighMemoryUsage: Triggers when Memory usage > 5% for 1 second
Email notification settings:
- SMTP Server: smtp.gmail.com:587
- From Email: raghulpg2006@gmail.com
- To Email: raghulpg2006@gmail.com
- TLS: Enabled
Email Template:
- Uses custom HTML template (
email_template.html) for professional alert formatting - Features:
- Color-coded alert status (π΄ Red for active, β Green for resolved)
- Severity badges with visual indicators
- Clean, organized layout with instance details
- Direct links to AlertManager for quick access
- Professional styling with improved readability
- Grouped alerts for better organization
βββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββ
β Monitoring Stack β
βββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββ€
β β
β ββββββββββββββββββββ ββββββββββββββββββββ ββββββββββββββββ β
β β Node Exporter β β Prometheus β β Grafana β β
β β (9100) β β (9090) β β (3000) β β
β ββββββββββ¬ββββββββββ ββββββββββ¬ββββββββββ ββββββββ¬ββββββββ β
β β β β β
β β Metrics β β β
β βββββββββββββββββββββββ€ β β
β β Alert Rules β β
β ββββββΌββββββββββββββ β β
β β Alert Manager β β β
β β (9093) β β β
β ββββββ¬ββββββββββββββ β β
β β β β
β ββββββββββΌβββββββββ β β
β β Email Alert β β β
β β (Gmail SMTP) β β β
β βββββββββββββββββββ β β
β β β
β Connected to Grafana ββββββββ β
β β
βββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββ
The system monitors two critical metrics:
- Condition: CPU usage > 5%
- Duration: 1 second
- Severity: Critical
- Action: Email notification
- Condition: Memory usage > 5%
- Duration: 1 second
- Severity: Critical
- Action: Email notification
Note: Alert thresholds can be adjusted in alert.rules.yml
- Open http://localhost:9090
- Go to "Graph" tab
- Example queries:
node_cpu_seconds_total- CPU metricsnode_memory_MemTotal_bytes- Memory metrics
- Open http://localhost:3000
- Login with default credentials
- Add Prometheus as data source
- Create visualization panels
- Save as dashboard
- Open http://localhost:9093
- View active alerts
- Check alert history
# Start services
docker-compose up -d
# Stop services
docker-compose down
# View logs for all services
docker-compose logs -f
# View logs for specific service
docker-compose logs -f prometheus
docker-compose logs -f grafana
docker-compose logs -f node_exporter
docker-compose logs -f alertmanager
# Restart services
docker-compose restart
# Check service status
docker-compose ps
# Remove containers and volumes
docker-compose down -vdocker-compose.yml- Service definitions and configurationsprometheus.yml- Prometheus configurationalert.rules.yml- Alert rules and thresholdsalertmanager.yml- AlertManager configuration and notification settingsemail_template.html- Custom HTML email template for professional alert formattingREADME.md- This file
# Check logs
docker-compose logs
# Verify Docker daemon is running
docker ps- Ensure Node Exporter is running: http://localhost:9100/metrics
- Check Prometheus target status in UI: http://localhost:9090/targets
- Verify Gmail credentials in
alertmanager.yml - Ensure "Less secure app access" is enabled (or use app-specific password)
- Check AlertManager logs:
docker-compose logs alertmanager
- Ensure services are running:
docker-compose ps - Check if ports are already in use
- Try recreating containers:
docker-compose down && docker-compose up -d
- Alert thresholds (5%) are set very low for testing. Adjust in
alert.rules.ymlfor production - Gmail credentials are hardcoded. For production, use secrets management
- Default Grafana password should be changed on first login
- Regular backups recommended for Prometheus data
For issues or questions, check:
- Individual service logs:
docker-compose logs <service-name> - Service health pages at their respective ports
- Official documentation: