Sitelet https://github.com/raghulpalanikumar/devops
Skip to content

Latest commit

Β 

History

4 Commits

Folders and files

NameName
Last commit message
Last commit date
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 

Repository files navigation

"# DevOps Monitoring & Alerting Stack

A Docker-based monitoring and alerting solution using Prometheus, Grafana, Node Exporter, and AlertManager. This stack provides comprehensive system monitoring, visualization, and alert notifications via email.

πŸ“‹ Table of Contents

πŸ“– Overview

This project sets up a complete monitoring ecosystem that:

  • Collects system metrics from your infrastructure
  • Stores and queries metrics using Prometheus
  • Visualizes data through Grafana dashboards
  • Triggers alerts based on predefined rules
  • Sends notifications via email

πŸ”§ Prerequisites

Before starting, ensure you have:

  • Docker (v20.10 or higher)
  • Docker Compose (v1.29 or higher)
  • A valid Gmail account (for email alerts)
  • At least 2GB of RAM allocated to Docker

πŸš€ Quick Start

  1. Clone or navigate to the project directory:

    cd d:\Devops
  2. Update AlertManager configuration:

    • Edit alertmanager.yml and update Gmail credentials with an app-specific password
  3. Start all services:

    docker-compose up -d
  4. Verify services are running:

    docker-compose ps
  5. Access the services:

πŸ”§ Services

Prometheus

  • Port: 9090
  • Purpose: Metrics collection and storage
  • Configuration: prometheus.yml
  • Scrape Interval: 5 seconds
  • Features:
    • Collects metrics from Node Exporter
    • Evaluates alert rules
    • Routes alerts to AlertManager

Grafana

  • Port: 3000
  • Purpose: Metrics visualization and dashboards
  • Default Credentials:
    • Username: admin
    • Password: admin
  • Features:
    • Create custom dashboards
    • Visualize Prometheus metrics
    • Alert management UI

Node Exporter

  • Port: 9100
  • Purpose: Exports system and hardware metrics
  • Metrics Exposed:
    • CPU usage
    • Memory usage
    • Disk I/O
    • Network statistics
    • Process information

AlertManager

  • Port: 9093
  • Purpose: Handles and routes alerts
  • Configuration: alertmanager.yml
  • Notification Method: Email

πŸ“ Access Points

Service URL Port
Prometheus http://localhost:9090 9090
Grafana http://localhost:3000 3000
Node Exporter Metrics http://localhost:9100/metrics 9100
AlertManager http://localhost:9093 9093

βš™οΈ Configuration

Prometheus Configuration (prometheus.yml)

Global Settings:
- Scrape interval: 5 seconds
- Target: Node Exporter on port 9100

Alert Configuration:
- Rule file: alert.rules.yml
- AlertManager target: localhost:9093

Alert Rules (alert.rules.yml)

Triggered alerts:

  • HighCPUUsage: Triggers when CPU usage > 5% for 1 second
  • HighMemoryUsage: Triggers when Memory usage > 5% for 1 second

AlertManager Configuration (alertmanager.yml)

Email notification settings:

Email Template:

  • Uses custom HTML template (email_template.html) for professional alert formatting
  • Features:
    • Color-coded alert status (πŸ”΄ Red for active, βœ… Green for resolved)
    • Severity badges with visual indicators
    • Clean, organized layout with instance details
    • Direct links to AlertManager for quick access
    • Professional styling with improved readability
    • Grouped alerts for better organization

πŸ—οΈ Architecture

β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
β”‚                        Monitoring Stack                         β”‚
β”œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€
β”‚                                                                 β”‚
β”‚  β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”  β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”  β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β” β”‚
β”‚  β”‚  Node Exporter   β”‚  β”‚   Prometheus     β”‚  β”‚    Grafana   β”‚ β”‚
β”‚  β”‚   (9100)         β”‚  β”‚    (9090)        β”‚  β”‚   (3000)     β”‚ β”‚
β”‚  β””β”€β”€β”€β”€β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜  β””β”€β”€β”€β”€β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜  β””β”€β”€β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”€β”€β”˜ β”‚
β”‚           β”‚                     β”‚                    β”‚          β”‚
β”‚           β”‚    Metrics          β”‚                    β”‚          β”‚
β”‚           └──────────────────────                    β”‚          β”‚
β”‚                                 β”‚ Alert Rules        β”‚          β”‚
β”‚                            β”Œβ”€β”€β”€β”€β–Όβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”     β”‚          β”‚
β”‚                            β”‚  Alert Manager   β”‚     β”‚          β”‚
β”‚                            β”‚    (9093)        β”‚     β”‚          β”‚
β”‚                            β””β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜     β”‚          β”‚
β”‚                                 β”‚                    β”‚          β”‚
β”‚                        β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β–Όβ”€β”€β”€β”€β”€β”€β”€β”€β”           β”‚          β”‚
β”‚                        β”‚   Email Alert   β”‚           β”‚          β”‚
β”‚                        β”‚  (Gmail SMTP)   β”‚           β”‚          β”‚
β”‚                        β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜           β”‚          β”‚
β”‚                                                      β”‚          β”‚
β”‚                          Connected to Grafana β”€β”€β”€β”€β”€β”€β”€β”˜          β”‚
β”‚                                                                 β”‚
β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜

🚨 Alert Rules

The system monitors two critical metrics:

1. HighCPUUsage

  • Condition: CPU usage > 5%
  • Duration: 1 second
  • Severity: Critical
  • Action: Email notification

2. HighMemoryUsage

  • Condition: Memory usage > 5%
  • Duration: 1 second
  • Severity: Critical
  • Action: Email notification

Note: Alert thresholds can be adjusted in alert.rules.yml

πŸ“Š Usage

View Metrics in Prometheus

  1. Open http://localhost:9090
  2. Go to "Graph" tab
  3. Example queries:
    • node_cpu_seconds_total - CPU metrics
    • node_memory_MemTotal_bytes - Memory metrics

Create Grafana Dashboard

  1. Open http://localhost:3000
  2. Login with default credentials
  3. Add Prometheus as data source
  4. Create visualization panels
  5. Save as dashboard

Check Alert Status

  1. Open http://localhost:9093
  2. View active alerts
  3. Check alert history

πŸ› οΈ Commands

# Start services
docker-compose up -d

# Stop services
docker-compose down

# View logs for all services
docker-compose logs -f

# View logs for specific service
docker-compose logs -f prometheus
docker-compose logs -f grafana
docker-compose logs -f node_exporter
docker-compose logs -f alertmanager

# Restart services
docker-compose restart

# Check service status
docker-compose ps

# Remove containers and volumes
docker-compose down -v

πŸ“ Project Files

  • docker-compose.yml - Service definitions and configurations
  • prometheus.yml - Prometheus configuration
  • alert.rules.yml - Alert rules and thresholds
  • alertmanager.yml - AlertManager configuration and notification settings
  • email_template.html - Custom HTML email template for professional alert formatting
  • README.md - This file

⚠️ Troubleshooting

Services not starting

# Check logs
docker-compose logs

# Verify Docker daemon is running
docker ps

Metrics not appearing in Prometheus

Alerts not sending emails

  • Verify Gmail credentials in alertmanager.yml
  • Ensure "Less secure app access" is enabled (or use app-specific password)
  • Check AlertManager logs: docker-compose logs alertmanager

Connection refused errors

  • Ensure services are running: docker-compose ps
  • Check if ports are already in use
  • Try recreating containers: docker-compose down && docker-compose up -d

πŸ“ Notes

  • Alert thresholds (5%) are set very low for testing. Adjust in alert.rules.yml for production
  • Gmail credentials are hardcoded. For production, use secrets management
  • Default Grafana password should be changed on first login
  • Regular backups recommended for Prometheus data

πŸ“ž Support

For issues or questions, check:

  • Individual service logs: docker-compose logs <service-name>
  • Service health pages at their respective ports
  • Official documentation:

About

DevOps monitoring and observability setup using Grafana and Prometheus for real-time metrics visualization and infrastructure monitoring.

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages