Chuyển đến nội dung chính

Lesson 3: Nginx Logging and Monitoring

A lesson on Nginx logging and monitoring with access log, error log, custom log formats, and log rotation. Guide to analyzing logs, troubleshooting, using logrotate, and basic metrics to track server performance. Includes practical examples and best practices.

🔒 DevSecOps — Lesson 3 Lesson 3: Nginx Logging and Monitoring

Nginx from Basics to Advanced

Part 1: Basics

xdev.asia

1. Access Log and Error Log

Nginx has two main log types to monitor server activity: Access log (records all requests) and Error log (records errors and warnings).

1.1. Access Log

Access log records every request to the server, including information about the client, request, response status, and processing time.

Default location:

# Ubuntu/Debian
/var/log/nginx/access.log

CentOS/RHEL

/var/log/nginx/access.log

macOS (Homebrew)

/usr/local/var/log/nginx/access.log

Basic configuration:

http {
# Access log for the entire HTTP context
access_log /var/log/nginx/access.log;

server {
    listen 80;
    server_name example.com;
    
    # Separate access log for virtual host
    access_log /var/log/nginx/example.com.access.log;
    
    location / {
        root /var/www/html;
    }
    
    # Disable access log for a specific location
    location /health-check {
        access_log off;
        return 200 "OK\n";
    }
}

}

Default format (combined):

192.168.1.100 - - [03/Dec/2024:10:30:45 +0700] "GET /index.html HTTP/1.1" 200 1234 "https://google.com" "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36"

Field explanations:

  • 192.168.1.100 - Client IP address
  • - - Remote user (usually - if no authentication)
  • - - Authenticated user
  • [03/Dec/2024:10:30:45 +0700] - Timestamp
  • "GET /index.html HTTP/1.1" - Request method, URI, and HTTP version
  • 200 - HTTP status code
  • 1234 - Response body size (bytes)
  • "https://google.com" - Referer
  • "Mozilla/5.0..." - User Agent

1.2. Error Log

Error log records errors, warnings, and debug information from Nginx.

Default location:

/var/log/nginx/error.log

Log levels (from least to most verbose):

  1. emerg - Emergency: system unusable
  2. alert - Alert: action must be taken immediately
  3. crit - Critical conditions
  4. error - Error conditions
  5. warn - Warning conditions
  6. notice - Normal but significant
  7. info - Informational
  8. debug - Debug messages

Configuration:

# Global error log
error_log /var/log/nginx/error.log warn;

http { # HTTP-level error log error_log /var/log/nginx/http-error.log error;

server {
    listen 80;
    server_name example.com;
    
    # Server-level error log
    error_log /var/log/nginx/example.com.error.log error;
    
    # Debug log for troubleshooting
    error_log /var/log/nginx/debug.log debug;
}

}

Example error log entries:

2024/12/03 10:30:45 [error] 1234#1234: *1 open() "/var/www/html/notfound.html" failed (2: No such file or directory), client: 192.168.1.100, server: example.com, request: "GET /notfound.html HTTP/1.1", host: "example.com"

2024/12/03 10:31:20 [warn] 1234#1234: *2 upstream server temporarily disabled while connecting to upstream, client: 192.168.1.101, server: api.example.com, request: "GET /api/users HTTP/1.1", upstream: "http://192.168.1.200:3000/api/users"

2024/12/03 10:32:05 [crit] 1234#1234: malloc() 8192 bytes failed (12: Cannot allocate memory)

1.3. Viewing and Monitoring Logs in Real-time

# View access log
sudo tail -f /var/log/nginx/access.log

View error log

sudo tail -f /var/log/nginx/error.log

View last 100 lines

sudo tail -n 100 /var/log/nginx/access.log

View both logs simultaneously

sudo tail -f /var/log/nginx/access.log /var/log/nginx/error.log

Filter logs

sudo tail -f /var/log/nginx/access.log | grep "404" sudo tail -f /var/log/nginx/access.log | grep "192.168.1.100"

View logs with less (scrollable)

sudo less +F /var/log/nginx/access.log

1.4. Basic Log Analysis

Count total requests:

# Total requests
wc -l /var/log/nginx/access.log

Requests in the past hour

sudo awk -v date="$(date -d '1 hour ago' '+%d/%b/%Y:%H')" '$4 > "["date' /var/log/nginx/access.log | wc -l

Top 10 IPs:

sudo awk '{print $1}' /var/log/nginx/access.log | sort | uniq -c | sort -rn | head -10

Top 10 most accessed URLs:

sudo awk '{print $7}' /var/log/nginx/access.log | sort | uniq -c | sort -rn | head -10

HTTP status code counts:

sudo awk '{print $9}' /var/log/nginx/access.log | sort | uniq -c | sort -rn

Top User Agents:

sudo awk -F'"' '{print $6}' /var/log/nginx/access.log | sort | uniq -c | sort -rn | head -10

Requests per hour:

sudo awk '{print $4}' /var/log/nginx/access.log | cut -d: -f1-2 | sort | uniq -c

2. Custom Log Formats

Nginx allows you to create custom log formats to collect exactly the information you need.

2.1. Basic Log Format

Define format:

http {
# Default format (combined)
log_format combined '$remote_addr - $remote_user [$time_local] '
'"$request" $status $body_bytes_sent '
'"$http_referer" "$http_user_agent"';

# Simpler format
log_format simple '$remote_addr - $request - $status';

# Detailed format
log_format detailed '$remote_addr - $remote_user [$time_local] '
                    '"$request" $status $body_bytes_sent '
                    '"$http_referer" "$http_user_agent" '
                    'rt=$request_time uct="$upstream_connect_time" '
                    'uht="$upstream_header_time" urt="$upstream_response_time"';

server {
    listen 80;
    
    # Use custom format
    access_log /var/log/nginx/access.log detailed;
}

}

2.2. Common Variables in Log Formats

Client information:

$remote_addr          # Client IP
$remote_user          # HTTP authenticated user
$http_x_forwarded_for # Real IP if behind proxy/CDN

Request information:

$time_local           # Local time
$time_iso8601         # ISO 8601 time
$request              # Full request line
$request_method       # GET, POST, etc.
$request_uri          # Request URI with arguments
$uri                  # Current URI
$args                 # Query string arguments
$query_string         # Same as $args
$scheme               # http or https
$server_protocol      # HTTP/1.1, HTTP/2.0
$host                 # Host header
$server_name          # Server name

Response information:

$status               # HTTP status code
$body_bytes_sent      # Response body size
$bytes_sent           # Total bytes sent (headers + body)
$request_length       # Request length (including headers)

Timing information:

$request_time         # Request processing time (seconds)
$upstream_response_time    # Backend response time
$upstream_connect_time     # Time to connect to upstream
$upstream_header_time      # Time to receive upstream headers

Upstream information:

$upstream_addr             # Upstream server address
$upstream_status           # Upstream response status
$upstream_cache_status     # Cache status (HIT, MISS, etc.)

Headers:

$http_user_agent      # User-Agent header
$http_referer         # Referer header
$http_cookie          # Cookie header
$http_<header_name>   # Any HTTP header (lowercase with underscores)

2.3. Real-world Log Format Examples

Performance monitoring format:

log_format performance '$remote_addr - [$time_local] "$request" '
'$status $body_bytes_sent '
'rt=$request_time '
'uct=$upstream_connect_time '
'uht=$upstream_header_time '
'urt=$upstream_response_time';

server { listen 80; access_log /var/log/nginx/performance.log performance; }

Output:

192.168.1.100 - [03/Dec/2024:10:30:45 +0700] "GET /api/users HTTP/1.1" 200 1234 rt=0.125 uct=0.005 uht=0.050 urt=0.120

JSON format (easy to parse):

log_format json_combined escape=json
'{'
'"time_local":"$time_local",'
'"remote_addr":"$remote_addr",'
'"request":"$request",'
'"status":$status,'
'"body_bytes_sent":$body_bytes_sent,'
'"request_time":$request_time,'
'"http_referer":"$http_referer",'
'"http_user_agent":"$http_user_agent"'
'}';

server { listen 80; access_log /var/log/nginx/access.json json_combined; }

Output:

{"time_local":"03/Dec/2024:10:30:45 +0700","remote_addr":"192.168.1.100","request":"GET /index.html HTTP/1.1","status":200,"body_bytes_sent":1234,"request_time":0.005,"http_referer":"https://google.com","http_user_agent":"Mozilla/5.0"}

Security monitoring format:

log_format security '$remote_addr - [$time_local] '
'"$request" $status '
'"$http_user_agent" '
'"$http_x_forwarded_for" '
'host=$host '
'args=$args';

server { listen 80; access_log /var/log/nginx/security.log security; }

CDN/Proxy format:

log_format cdn '$http_x_forwarded_for - $remote_addr - $remote_user [$time_local] '
'"$request" $status $body_bytes_sent '
'"$http_referer" "$http_user_agent" '
'cache=$upstream_cache_status';

server { listen 80; access_log /var/log/nginx/cdn.log cdn; }

2.4. Conditional Logging

Log only when condition is met:

http {
# Define map to check conditions
map $status $loggable {
~^[23]  0;  # Don't log 2xx and 3xx
default 1;  # Log everything else
}

server {
    listen 80;
    
    # Only log if $loggable = 1
    access_log /var/log/nginx/errors-only.log combined if=$loggable;
}

}

Don't log static files:

map $request_uri $log_static {
~*.(jpg|jpeg|png|gif|ico|css|js)$ 0;
default 1;
}

server { listen 80; access_log /var/log/nginx/access.log combined if=$log_static; }

Don't log health checks:

map $request_uri $log_health {
~^/health$ 0;
~^/ping$ 0;
default 1;
}

server { listen 80; access_log /var/log/nginx/access.log combined if=$log_health; }

Don't log bots:

map $http_user_agent $log_bots {
~bot 0;
~crawler 0;
~*spider 0;
default 1;
}

server { listen 80; access_log /var/log/nginx/access.log combined if=$log_bots; }

2.5. Multiple Access Logs

server {
listen 80;
server_name example.com;

# Log all requests
access_log /var/log/nginx/all.log combined;

# Log errors only
access_log /var/log/nginx/errors.log combined if=$loggable;

# Performance log
access_log /var/log/nginx/performance.log performance;

# JSON log for processing
access_log /var/log/nginx/json.log json_combined;

}


3. Log Rotation with Logrotate

Log files can grow very quickly. Log rotation helps manage disk space by automatically compressing and deleting old logs.

3.1. Basic Logrotate

Default configuration file:

# Ubuntu/Debian
/etc/logrotate.d/nginx

CentOS/RHEL

/etc/logrotate.d/nginx

Default contents:

/var/log/nginx/*.log {
daily
missingok
rotate 14
compress
delaycompress
notifempty
create 0640 www-data adm
sharedscripts
postrotate
if [ -f /var/run/nginx.pid ]; then
kill -USR1 cat /var/run/nginx.pid
fi
endscript
}

Directive explanations:

  • daily - Rotate every day
  • missingok - Don't error if log file doesn't exist
  • rotate 14 - Keep 14 backup copies
  • compress - Compress old logs with gzip
  • delaycompress - Wait until next rotation to compress
  • notifempty - Don't rotate if file is empty
  • create 0640 www-data adm - Create new file with permissions
  • sharedscripts - Run postrotate script once for all logs
  • postrotate/endscript - Script run after rotation

3.2. Custom Logrotate Configuration

Rotate hourly (for high-traffic sites):

sudo nano /etc/logrotate.d/nginx-hourly

Contents:

/var/log/nginx/high-traffic.log { hourly rotate 168 # 7 days * 24 hours compress delaycompress notifempty create 0640 www-data adm dateext dateformat -%Y%m%d-%H sharedscripts postrotate if [ -f /var/run/nginx.pid ]; then kill -USR1 cat /var/run/nginx.pid fi endscript }

Rotate by size:

/var/log/nginx/.log {
size 100M           # Rotate when reaching 100MB
rotate 10
compress
delaycompress
notifempty
create 0640 www-data adm
sharedscripts
postrotate
if [ -f /var/run/nginx.pid ]; then
kill -USR1 cat /var/run/nginx.pid
fi
endscript
}

Rotate with custom naming:

/var/log/nginx/.log {
daily
rotate 30
compress
delaycompress
notifempty
create 0640 www-data adm
dateext
dateformat -.%Y-%m-%d
extension .log
sharedscripts
postrotate
if [ -f /var/run/nginx.pid ]; then
kill -USR1 cat /var/run/nginx.pid
fi
endscript
}

Output: access.log-2024-12-03.log.gz

Separate rotation for each log:

# Performance logs - keep longer
/var/log/nginx/performance.log {
daily
rotate 90           # 3 months
compress
delaycompress
notifempty
create 0640 www-data adm
}

Error logs - keep very long

/var/log/nginx/error.log { weekly rotate 52 # 1 year compress delaycompress notifempty create 0640 www-data adm }

Access logs - rotate quickly

/var/log/nginx/access.log { daily rotate 7 # 1 week compress delaycompress notifempty create 0640 www-data adm }

3.3. Test and Force Rotation

# Test configuration (dry run)
sudo logrotate -d /etc/logrotate.d/nginx

Force rotation (execute immediately)

sudo logrotate -f /etc/logrotate.d/nginx

Check status

sudo cat /var/lib/logrotate/status

Manual rotation (without logrotate)

sudo mv /var/log/nginx/access.log /var/log/nginx/access.log.1 sudo nginx -s reopen sudo gzip /var/log/nginx/access.log.1

3.4. Troubleshooting Logrotate

Check if logrotate is running:

# Check cron job
ls -la /etc/cron.daily/logrotate

Check logrotate status

sudo cat /var/lib/logrotate/status | grep nginx

Run logrotate manually with verbose

sudo logrotate -v /etc/logrotate.d/nginx

Common errors:

# Error: Permission denied

Fix: Check ownership

ls -la /var/log/nginx/ sudo chown www-data:adm /var/log/nginx/*.log

Error: Nginx not reopening logs

Fix: Check PID file

ls -la /var/run/nginx.pid sudo systemctl restart nginx

Error: Logs not being compressed

Fix: Check gzip installed

which gzip sudo apt install gzip


4. Basic Metrics for Monitoring

4.1. Requests per Second (RPS)

Script to calculate RPS:

#!/bin/bash

rps.sh - Calculate requests per second

LOG_FILE="/var/log/nginx/access.log" INTERVAL=60 # seconds

while true; do START_COUNT=$(wc -l < "$LOG_FILE") sleep $INTERVAL END_COUNT=$(wc -l < "$LOG_FILE")

REQUESTS=$((END_COUNT - START_COUNT))
RPS=$(echo "scale=2; $REQUESTS / $INTERVAL" | bc)

echo "$(date '+%Y-%m-%d %H:%M:%S') - RPS: $RPS"

done

Run script:

chmod +x rps.sh
./rps.sh

4.2. Response Time Analysis

Script to analyze response times:

#!/bin/bash

response_time.sh - Analyze response times

LOG_FILE="/var/log/nginx/access.log"

echo "Response Time Statistics:" echo "========================"

Extract request_time (assuming it's logged)

awk '{print $NF}' "$LOG_FILE" |
awk '{ sum += $1; count++; if ($1 > max) max = $1; if (min == 0 || $1 < min) min = $1; } END { print "Average: " sum/count " seconds"; print "Min: " min " seconds"; print "Max: " max " seconds"; }'

4.3. Status Code Distribution

#!/bin/bash

status_codes.sh - Count HTTP status codes

LOG_FILE="/var/log/nginx/access.log"

echo "HTTP Status Code Distribution:" echo "=============================="

awk '{print $9}' "$LOG_FILE" | sort | uniq -c | sort -rn |
while read count code; do percentage=$(echo "scale=2; ($count * 100) / $(wc -l < $LOG_FILE)" | bc) printf "%3s: %6d requests (%5.2f%%)\n" "$code" "$count" "$percentage" done

4.4. Traffic by Hour

#!/bin/bash

traffic_by_hour.sh - Analyze traffic by hour

LOG_FILE="/var/log/nginx/access.log"

echo "Traffic by Hour:" echo "================"

awk '{print $4}' "$LOG_FILE" | cut -d: -f2 | sort | uniq -c |
while read count hour; do printf "Hour %02d: %6d requests\n" "$hour" "$count" done

4.5. Top Clients (IP Addresses)

#!/bin/bash

top_clients.sh - Find top clients by requests

LOG_FILE="/var/log/nginx/access.log" TOP_N=10

echo "Top $TOP_N Clients:" echo "=================="

awk '{print $1}' "$LOG_FILE" | sort | uniq -c | sort -rn | head -n $TOP_N |
while read count ip; do printf "%15s: %6d requests\n" "$ip" "$count" done

4.6. Bandwidth Usage

#!/bin/bash

bandwidth.sh - Calculate bandwidth usage

LOG_FILE="/var/log/nginx/access.log"

echo "Bandwidth Statistics:" echo "===================="

Assuming $body_bytes_sent is in position 10

awk '{sum += $10} END { gb = sum / 1024 / 1024 / 1024; mb = sum / 1024 / 1024; kb = sum / 1024; printf "Total: %.2f GB (%.2f MB, %.2f KB)\n", gb, mb, kb; }' "$LOG_FILE"

4.7. Real-time Dashboard Script

#!/bin/bash

dashboard.sh - Real-time Nginx monitoring dashboard

LOG_FILE="/var/log/nginx/access.log"

while true; do clear echo "===================================" echo " NGINX MONITORING DASHBOARD" echo "===================================" echo "Time: $(date '+%Y-%m-%d %H:%M:%S')" echo

# Total requests
TOTAL=$(wc -l &lt; "$LOG_FILE")
echo "Total Requests: $TOTAL"
echo

# Last minute requests
LAST_MINUTE=$(tail -n 1000 "$LOG_FILE" | wc -l)
echo "Last ~1000 requests"
echo

# Status codes (last 1000)
echo "Status Codes (recent):"
tail -n 1000 "$LOG_FILE" | awk '{print $9}' | sort | uniq -c | sort -rn
echo

# Top 5 IPs (recent)
echo "Top 5 IPs (recent):"
tail -n 1000 "$LOG_FILE" | awk '{print $1}' | sort | uniq -c | sort -rn | head -5
echo

# Top 5 URLs (recent)
echo "Top 5 URLs (recent):"
tail -n 1000 "$LOG_FILE" | awk '{print $7}' | sort | uniq -c | sort -rn | head -5

sleep 5

done

Run dashboard:

chmod +x dashboard.sh
./dashboard.sh

4.8. Sending Alerts When Issues Occur

#!/bin/bash

alert.sh - Send alert when error rate is high

LOG_FILE="/var/log/nginx/access.log" ERROR_THRESHOLD=10 # % of 5xx errors EMAIL="[email protected]"

Count last 100 requests

TOTAL=$(tail -n 100 "$LOG_FILE" | wc -l) ERRORS=$(tail -n 100 "$LOG_FILE" | awk '{print $9}' | grep "^5" | wc -l)

ERROR_RATE=$(echo "scale=2; ($ERRORS * 100) / $TOTAL" | bc)

if (( $(echo "$ERROR_RATE > $ERROR_THRESHOLD" | bc -l) )); then MESSAGE="ALERT: High error rate detected! $ERROR_RATE% of requests are 5xx errors" echo "$MESSAGE" | mail -s "Nginx Alert" "$EMAIL" echo "$MESSAGE" fi

4.9. Integration with Monitoring Tools

Export metrics for Prometheus:

# Install nginx-prometheus-exporter
wget https://github.com/nginxinc/nginx-prometheus-exporter/releases/download/v0.11.0/nginx-prometheus-exporter_0.11.0_linux_amd64.tar.gz
tar xzf nginx-prometheus-exporter_0.11.0_linux_amd64.tar.gz
sudo mv nginx-prometheus-exporter /usr/local/bin/

Run exporter

nginx-prometheus-exporter -nginx.scrape-uri=http://localhost:8080/stub_status

Configure Nginx stub_status:

server {
listen 8080;
server_name localhost;

location /stub_status {
    stub_status;
    access_log off;
    allow 127.0.0.1;
    deny all;
}

}


5. Practice Exercises

Exercise 1: Custom Log Format

  1. Create a custom log format named timing that includes:
    • Remote address
    • Request
    • Status
    • Request time
    • Upstream response time
  2. Apply this format to a virtual host
  3. Generate traffic and view logs

Exercise 2: JSON Logging

  1. Create a JSON log format
  2. Configure Nginx to log in JSON
  3. Parse JSON logs with jq:
cat /var/log/nginx/access.json | jq '.status'
cat /var/log/nginx/access.json | jq 'select(.status >= 400)'

Exercise 3: Log Rotation

  1. Create a custom logrotate config that rotates at 10MB
  2. Test with logrotate -d
  3. Force rotation and verify

Exercise 4: Traffic Analysis

  1. Generate 1000 requests with ab:
ab -n 1000 -c 10 http://localhost/
  1. Analyze logs to find:
    • Total requests
    • Average response time
    • Status code distribution
    • Top URLs

Exercise 5: Real-time Monitoring

  1. Set up the dashboard script
  2. Modify it to add:
    • Error rate (%)
    • Bandwidth usage
    • Slowest requests

Exercise 6: Conditional Logging

  1. Configure to not log:
    • Static files (.css, .js, .jpg, .png)
    • Health check endpoint (/health)
    • Bot traffic
  2. Verify that these requests don't appear in logs

6. Troubleshooting with Logs

6.1. Debug 404 Errors

# Find all 404s
grep " 404 " /var/log/nginx/access.log

Top URLs causing 404

grep " 404 " /var/log/nginx/access.log | awk '{print $7}' | sort | uniq -c | sort -rn | head -10

404s from a specific IP

grep "192.168.1.100" /var/log/nginx/access.log | grep " 404 "

6.2. Debug 500 Errors

# Find 5xx errors
grep " 50[0-9] " /var/log/nginx/access.log

Check error log for details

sudo tail -100 /var/log/nginx/error.log | grep "error"

5xx errors over time

grep " 50[0-9] " /var/log/nginx/access.log | awk '{print $4}' | cut -d: -f1-2 | uniq -c

6.3. Debug Slow Requests

# Find requests > 1 second (assuming request_time is logged)
awk '$NF > 1.0' /var/log/nginx/access.log

Top 10 slowest requests

awk '{print $NF, $7}' /var/log/nginx/access.log | sort -rn | head -10

6.4. Debug High Traffic

# Requests per minute
awk '{print $4}' /var/log/nginx/access.log | cut -d: -f1-3 | uniq -c

Identify traffic spikes

awk '{print $4}' /var/log/nginx/access.log | cut -d: -f1-3 | uniq -c | awk '$1 > 1000'

6.5. Debug Security Issues

# Find SQL injection attempts
grep -i "select.*from|union.*select" /var/log/nginx/access.log

Find path traversal attempts

grep ".." /var/log/nginx/access.log

Suspicious user agents

grep -i "sqlmap|nikto|nmap" /var/log/nginx/access.log


7. Best Practices

7.1. Log Management

  1. Separate logs per virtual host:
server {
server_name site1.com;
access_log /var/log/nginx/site1.access.log;
error_log /var/log/nginx/site1.error.log;
}
  1. Use appropriate log levels:
# Production: error or warn
error_log /var/log/nginx/error.log warn;

Development: info or debug

error_log /var/log/nginx/error.log debug;

  1. Don't log excessively:
# Disable for static files
location ~* .(jpg|jpeg|png|gif|ico|css|js)$ {
access_log off;
}

Disable for health checks

location /health { access_log off; return 200; }

7.2. Performance

  1. Buffer logs:
access_log /var/log/nginx/access.log combined buffer=32k;
  1. Async logging (Nginx 1.7.11+):
access_log /var/log/nginx/access.log combined buffer=32k flush=5s;

7.3. Security

  1. Protect log files:
sudo chmod 640 /var/log/nginx/.log
sudo chown www-data:adm /var/log/nginx/.log
  1. Rotate regularly:
# Daily rotation for high-traffic sites

Weekly for low-traffic sites

  1. Monitor and alert:
# Set up monitoring for error rates

Alert on abnormal spikes


Summary

In this lesson, you learned:

  • ✅ Access log and error log
  • ✅ Custom log formats and variables
  • ✅ Log rotation with logrotate
  • ✅ Log analysis and metrics
  • ✅ Troubleshooting with logs
  • ✅ Best practices for logging

Next lesson: We will explore Reverse Proxy — how to use Nginx as a reverse proxy for backend applications.