This directory contains configuration for running mockd with Prometheus, Jaeger, Loki, and Grafana for full observability (metrics, tracing, and logging).
# Start the full stack
docker-compose -f docker-compose.observability-local.yml up -d
# Start mockd with tracing and log aggregation enabled
./mockd serve --no-auth \
--otlp-endpoint http://localhost:4318/v1/traces \
--loki-endpoint http://localhost:3100/loki/api/v1/push
# View logs
docker-compose -f docker-compose.observability-local.yml logs -f
# Stop everything
docker-compose -f docker-compose.observability-local.yml down| Service | URL | Description |
|---|---|---|
| mockd Mock Server | http://localhost:4280 | Mock endpoint server |
| mockd Admin API | http://localhost:4290 | Admin API & metrics |
| Prometheus | http://localhost:9090 | Metrics storage & queries |
| Jaeger UI | http://localhost:16686 | Distributed tracing |
| Loki | http://localhost:3100 | Log aggregation |
| Grafana | http://localhost:3000 | Dashboards (admin/admin) |
mockd exposes Prometheus metrics at /metrics on the Admin API port.
| Metric | Type | Labels | Description |
|---|---|---|---|
mockd_requests_total |
Counter | method, path, status | Total mock requests |
mockd_request_duration_seconds |
Histogram | method, path | Request latency distribution |
mockd_match_hits_total |
Counter | mock_id | Mock match hits |
mockd_match_misses_total |
Counter | - | Requests that didn't match any mock |
| Metric | Type | Labels | Description |
|---|---|---|---|
mockd_mocks_total |
Gauge | type | Total configured mocks |
mockd_mocks_enabled |
Gauge | type | Enabled mocks |
mockd_active_connections |
Gauge | protocol | Active WebSocket/SSE connections |
| Metric | Type | Labels | Description |
|---|---|---|---|
mockd_admin_requests_total |
Counter | method, path, status | Admin API requests |
mockd_admin_request_duration_seconds |
Histogram | method, path | Admin API latency |
| Metric | Type | Description |
|---|---|---|
go_goroutines |
Gauge | Number of goroutines |
go_memstats_heap_alloc_bytes |
Gauge | Heap bytes allocated and in use |
go_memstats_heap_sys_bytes |
Gauge | Heap bytes obtained from system |
go_gc_duration_seconds |
Gauge | Total GC pause duration |
go_gc_cycles_total |
Gauge | Total number of completed GC cycles |
go_info |
Gauge | Go version information |
# Request rate over 5 minutes
sum(rate(mockd_requests_total[5m]))
# p95 latency
histogram_quantile(0.95, sum(rate(mockd_request_duration_seconds_bucket[5m])) by (le))
# Error rate (4xx and 5xx responses)
sum(rate(mockd_requests_total{status=~"[45].."}[5m])) / sum(rate(mockd_requests_total[5m])) * 100
# Top 10 most-hit mocks
topk(10, sum(mockd_match_hits_total) by (mock_id))
# Mock miss ratio
sum(rate(mockd_match_misses_total[5m])) / (sum(rate(mockd_match_hits_total[5m])) + sum(rate(mockd_match_misses_total[5m])))
# Memory usage
go_memstats_heap_alloc_bytes / 1024 / 1024 # MB
Pre-configured alerting rules are included in prometheus/rules/mockd-alerts.yml:
| Alert | Severity | Description |
|---|---|---|
MockdDown |
critical | mockd instance is down |
MockdNoTraffic |
warning | No requests received recently |
MockdHighServerErrorRate |
warning | >5% 5xx error rate |
MockdCriticalErrorRate |
critical | >20% 5xx error rate |
MockdHighLatencyP95 |
warning | p95 latency >500ms |
MockdHighLatencyP99 |
critical | p99 latency >2s |
MockdHighMissRate |
info | Many requests hitting unconfigured endpoints |
MockdHighMissRatio |
warning | >20% of requests not matching mocks |
To view active alerts:
- Open Prometheus at http://localhost:9090/alerts
- Or check in Grafana under Alerting
mockd supports OpenTelemetry tracing via OTLP HTTP. Configure with CLI flags:
# Send traces to Jaeger via OTLP
mockd serve --otlp-endpoint http://localhost:4318/v1/traces
# With sampling (e.g., sample 10% of traces)
mockd serve --otlp-endpoint http://localhost:4318/v1/traces --trace-sampler 0.1
# Combined with structured logging
mockd serve --otlp-endpoint http://localhost:4318/v1/traces --log-level debug --log-format jsonEach request span includes:
http.method- HTTP method (GET, POST, etc.)http.url- Request URLhttp.target- Request pathhttp.host- Host headerhttp.scheme- http or httpshttp.status_code- Response status codehttp.user_agent- Client user agentotel.status_code- OK or ERROR based on status code
The following paths are excluded from tracing to reduce noise:
/metrics- Prometheus scrape endpoint/health,/healthz,/ready,/readyz,/livez- Health checks/_/health,/__health- Alternative health checks
- Open Jaeger UI at http://localhost:16686
- Select "mockd" from the Service dropdown
- Click "Find Traces" to see recent requests
- Click on a trace to see the span details and timing
mockd supports structured logging with trace ID injection for log-to-trace correlation.
# JSON structured logging
mockd serve --log-format json --log-level debug
# With tracing enabled (trace_id will be injected into logs)
mockd serve --log-format json --otlp-endpoint http://localhost:4318/v1/traces| Level | Description |
|---|---|
error |
Errors only |
warn |
Warnings and errors |
info |
Normal operational messages (default) |
debug |
Detailed debugging information |
mockd can send logs directly to Loki for centralized log aggregation:
# Send logs to Loki
./mockd serve --loki-endpoint http://localhost:3100/loki/api/v1/push
# Combined with tracing and JSON format
./mockd serve \
--loki-endpoint http://localhost:3100/loki/api/v1/push \
--otlp-endpoint http://localhost:4318/v1/traces \
--log-format json \
--log-level debugLoki is configured as a datasource in Grafana with trace-to-log correlation. When viewing logs in Grafana, you can click on trace IDs to jump directly to the trace in Jaeger.
When using --loki-endpoint, logs are sent with these labels:
job=mockd- Fixed job nameservice=mockd- Service identifierport=<port>- The mock server port
- Open Grafana at http://localhost:3000
- Go to Explore
- Select "Loki" datasource
- Use LogQL queries:
{job="mockd"} # All mockd logs {job="mockd"} |= "error" # Logs containing "error" {job="mockd"} | json | level="ERROR" # JSON logs with ERROR level
The pre-configured mockd Overview dashboard includes:
- Requests (5m) - Request count in the last 5 minutes
- Mock Hits - Total mock match hits
- Mock Misses - Requests that didn't match any mock
- Error Rate - Percentage of 4xx/5xx responses
- p95 Latency - 95th percentile latency
- Uptime - Server uptime
- Request Rate by Method - Requests/second by HTTP method
- Request Rate by Status - Requests/second by status code
- Request Latency Percentiles - p50, p90, p95, p99 over time
- Latency Distribution - Histogram of request durations
- Requests by Status Code - Pie chart of response codes
- Active Connections - WebSocket/SSE connections over time
- Top 10 Mocks by Hits - Most frequently matched mocks
- Admin API Request Rate - Admin endpoint usage
- Admin API Latency - Admin endpoint performance
The dashboard supports filtering by:
- Method - Filter by HTTP method (GET, POST, etc.)
- Status - Filter by status code (200, 404, 500, etc.)
To prevent metric cardinality explosion, dynamic path segments are automatically normalized:
| Pattern | Replacement | Example |
|---|---|---|
| UUID | {uuid} |
/items/a1b2c3d4-... → /items/{uuid} |
| MongoDB ObjectID (24 hex) | {id} |
/docs/507f1f77bcf86cd799439011 → /docs/{id} |
| Numeric ID | {id} |
/users/123 → /users/{id} |
- API Key Authentication: Configure Prometheus with the mockd API key:
scrape_configs: - job_name: 'mockd' bearer_token_file: /path/to/api-key static_configs: - targets: ['mockd:4290']
-
Trace Sampling: In high-traffic environments, use sampling:
mockd serve --trace-sampler 0.1 # Sample 10% of traces -
Resource Limits: Adjust container resource limits based on your traffic
-
Prometheus Retention: Configure retention for your needs:
command: - '--storage.tsdb.retention.time=15d'
-
Loki Retention: Configure log retention in
loki-config.yml:limits_config: reject_old_samples_max_age: 168h # 7 days
- AlertManager: For production, configure AlertManager for alert routing:
# prometheus.yml alerting: alertmanagers: - static_configs: - targets: ['alertmanager:9093']
observability/
├── README.md # This file
├── prometheus-local.yml # Prometheus config for local dev
├── prometheus/
│ └── rules/
│ └── mockd-alerts.yml # Alerting rules
├── loki/
│ └── loki-config.yml # Loki configuration
├── promtail/
│ └── promtail-config.yml # Promtail config (optional, for file-based log shipping)
└── grafana/
└── provisioning/
├── datasources/
│ └── datasources.yml # Prometheus, Jaeger, Loki datasources
└── dashboards/
├── dashboards.yml # Dashboard provisioning config
└── json/
└── mockd-overview.json # Main dashboard
- Check mockd is running:
curl http://localhost:4290/health - Check metrics endpoint:
curl http://localhost:4290/metrics - Check Prometheus targets: http://localhost:9090/targets
- Verify OTLP endpoint is correct:
--otlp-endpoint http://localhost:4318/v1/traces - Check Jaeger is running:
docker ps | grep jaeger - Generate some traffic and wait a few seconds (traces are batched)
- Ensure Prometheus datasource is configured with UID
prometheus - Check the time range in Grafana (default: Last 15 minutes)
- Verify mockd has received requests
If you see cardinality warnings, you may have dynamic paths not being normalized. Check the path normalization in pkg/engine/metrics_middleware.go and add patterns as needed.