warning
8.4.1.1.GitLab Puma workers not running
GitLab Puma on {{ $labels.instance }} has {{ $value }} running workers out of expected total.
- alert: GitlabPumaWorkersNotRunning
expr: 'puma_running_workers < puma_workers'
for: 5m
labels:
severity: warning
annotations:
summary: GitLab Puma workers not running (instance {{ $labels.instance }})
description: "GitLab Puma on {{ $labels.instance }} has {{ $value }} running workers out of expected total.\n VALUE = {{ $value }}\n LABELS = {{ $labels }}"warning
8.4.1.2.GitLab high memory usage
GitLab process on {{ $labels.instance }} is using {{ $value | humanize1024 }}B of RSS memory.
# Threshold of 2GB may need adjustment based on your instance size.
# High memory usage can lead to OOM kills and service disruptions.
- alert: GitlabHighMemoryUsage
expr: 'process_resident_memory_bytes{job=~".*gitlab.*"} > 2e+9'
for: 10m
labels:
severity: warning
annotations:
summary: GitLab high memory usage (instance {{ $labels.instance }})
description: "GitLab process on {{ $labels.instance }} is using {{ $value | humanize1024 }}B of RSS memory.\n VALUE = {{ $value }}\n LABELS = {{ $labels }}"warning
8.4.1.3.GitLab Ruby heap fragmentation
GitLab Ruby heap fragmentation on {{ $labels.instance }} is {{ $value }}. High fragmentation wastes memory.
# Heap fragmentation above 50% means a significant amount of memory is wasted.
# A Puma worker restart may help reclaim memory.
- alert: GitlabRubyHeapFragmentation
expr: 'ruby_gc_stat_ext_heap_fragmentation{job=~".*gitlab.*"} > 0.5'
for: 15m
labels:
severity: warning
annotations:
summary: GitLab Ruby heap fragmentation (instance {{ $labels.instance }})
description: "GitLab Ruby heap fragmentation on {{ $labels.instance }} is {{ $value }}. High fragmentation wastes memory.\n VALUE = {{ $value }}\n LABELS = {{ $labels }}"warning
8.4.1.4.GitLab Puma high queued connections
GitLab Puma has {{ $value }} queued connections on {{ $labels.instance }}. Requests are waiting for an available worker thread.
# Queued connections indicate Puma workers are saturated.
# Consider increasing puma['worker_processes'] or puma['max_threads'] in gitlab.rb.
- alert: GitlabPumaHighQueuedConnections
expr: 'puma_queued_connections > 5'
for: 5m
labels:
severity: warning
annotations:
summary: GitLab Puma high queued connections (instance {{ $labels.instance }})
description: "GitLab Puma has {{ $value }} queued connections on {{ $labels.instance }}. Requests are waiting for an available worker thread.\n VALUE = {{ $value }}\n LABELS = {{ $labels }}"warning
8.4.1.5.GitLab database connection pool saturation
GitLab database connection pool on {{ $labels.instance }} ({{ $labels.class }}) is {{ $value }}% busy.
# When the pool is near saturation, requests may block waiting for a connection.
# Increase db_pool_size in gitlab.rb or investigate slow queries.
- alert: GitlabDatabaseConnectionPoolSaturation
expr: 'gitlab_database_connection_pool_busy / gitlab_database_connection_pool_size * 100 > 90 and gitlab_database_connection_pool_size > 0'
for: 5m
labels:
severity: warning
annotations:
summary: GitLab database connection pool saturation (instance {{ $labels.instance }})
description: "GitLab database connection pool on {{ $labels.instance }} ({{ $labels.class }}) is {{ $value }}% busy.\n VALUE = {{ $value }}\n LABELS = {{ $labels }}"warning
8.4.1.6.GitLab database connection pool dead connections
GitLab database connection pool on {{ $labels.instance }} ({{ $labels.class }}) has {{ $value }} dead connections.
- alert: GitlabDatabaseConnectionPoolDeadConnections
expr: 'gitlab_database_connection_pool_dead > 0'
for: 5m
labels:
severity: warning
annotations:
summary: GitLab database connection pool dead connections (instance {{ $labels.instance }})
description: "GitLab database connection pool on {{ $labels.instance }} ({{ $labels.class }}) has {{ $value }} dead connections.\n VALUE = {{ $value }}\n LABELS = {{ $labels }}"warning
8.4.1.7.GitLab database connection pool waiting
GitLab on {{ $labels.instance }} has {{ $value }} threads waiting for a database connection.
- alert: GitlabDatabaseConnectionPoolWaiting
expr: 'gitlab_database_connection_pool_waiting > 0'
for: 5m
labels:
severity: warning
annotations:
summary: GitLab database connection pool waiting (instance {{ $labels.instance }})
description: "GitLab on {{ $labels.instance }} has {{ $value }} threads waiting for a database connection.\n VALUE = {{ $value }}\n LABELS = {{ $labels }}"warning
8.4.1.8.GitLab high file descriptor usage
GitLab on {{ $labels.instance }} is using {{ $value }}% of available file descriptors.
- alert: GitlabHighFileDescriptorUsage
expr: 'process_open_fds{job=~".*gitlab.*"} / process_max_fds * 100 > 80 and process_max_fds > 0'
for: 5m
labels:
severity: warning
annotations:
summary: GitLab high file descriptor usage (instance {{ $labels.instance }})
description: "GitLab on {{ $labels.instance }} is using {{ $value }}% of available file descriptors.\n VALUE = {{ $value }}\n LABELS = {{ $labels }}"warning
8.4.1.9.GitLab Puma high thread pool utilization
GitLab Puma thread pool on {{ $labels.instance }} is {{ $value }}% utilized, indicating sustained saturation before it becomes fully blocked.
- alert: GitlabPumaHighThreadPoolUtilization
expr: 'puma_active_connections / puma_max_threads * 100 > 90 and puma_max_threads > 0'
for: 10m
labels:
severity: warning
annotations:
summary: GitLab Puma high thread pool utilization (instance {{ $labels.instance }})
description: "GitLab Puma thread pool on {{ $labels.instance }} is {{ $value }}% utilized, indicating sustained saturation before it becomes fully blocked.\n VALUE = {{ $value }}\n LABELS = {{ $labels }}"critical
8.4.1.10.GitLab Puma no available pool capacity
GitLab Puma pool capacity on {{ $labels.instance }} has been at 0 for 5 minutes. All threads are busy.
- alert: GitlabPumaNoAvailablePoolCapacity
expr: 'puma_pool_capacity == 0'
for: 5m
labels:
severity: critical
annotations:
summary: GitLab Puma no available pool capacity (instance {{ $labels.instance }})
description: "GitLab Puma pool capacity on {{ $labels.instance }} has been at 0 for 5 minutes. All threads are busy.\n VALUE = {{ $value }}\n LABELS = {{ $labels }}"critical
8.4.1.11.GitLab high HTTP error rate
GitLab is returning more than 5% HTTP 5xx errors on {{ $labels.instance }}.
# Threshold is 5% of all requests returning server errors.
# Check GitLab logs at /var/log/gitlab/ for root cause.
- alert: GitlabHighHttpErrorRate
expr: 'sum(rate(http_requests_total{status=~"5.."}[5m])) / sum(rate(http_requests_total[5m])) * 100 > 5 and sum(rate(http_requests_total[5m])) > 0'
for: 5m
labels:
severity: critical
annotations:
summary: GitLab high HTTP error rate (instance {{ $labels.instance }})
description: "GitLab is returning more than 5% HTTP 5xx errors on {{ $labels.instance }}.\n VALUE = {{ $value }}\n LABELS = {{ $labels }}"warning
8.4.1.12.GitLab high HTTP request latency
GitLab p95 HTTP request latency on {{ $labels.instance }} is above 10 seconds.
# Threshold of 10s may need adjustment based on your instance size and workload.
- alert: GitlabHighHttpRequestLatency
expr: 'histogram_quantile(0.95, sum(rate(http_request_duration_seconds_bucket[5m])) by (le)) > 10'
for: 5m
labels:
severity: warning
annotations:
summary: GitLab high HTTP request latency (instance {{ $labels.instance }})
description: "GitLab p95 HTTP request latency on {{ $labels.instance }} is above 10 seconds.\n VALUE = {{ $value }}\n LABELS = {{ $labels }}"warning
8.4.1.13.GitLab Sidekiq jobs failing
GitLab Sidekiq jobs are failing at a rate of {{ $value }} per second on {{ $labels.instance }}.
# This metric requires the emit_sidekiq_histogram_metrics feature flag to be enabled.
# A sustained failure rate indicates background processing issues.
- alert: GitlabSidekiqJobsFailing
expr: 'rate(sidekiq_jobs_failed_total[5m]) > 0.1'
for: 10m
labels:
severity: warning
annotations:
summary: GitLab Sidekiq jobs failing (instance {{ $labels.instance }})
description: "GitLab Sidekiq jobs are failing at a rate of {{ $value }} per second on {{ $labels.instance }}.\n VALUE = {{ $value }}\n LABELS = {{ $labels }}"warning
8.4.1.14.GitLab Sidekiq queue too large
GitLab Sidekiq has {{ $value }} running jobs, approaching concurrency limit on {{ $labels.instance }}.
# When running jobs approach the concurrency limit, new jobs will queue up.
# Consider scaling Sidekiq workers or increasing concurrency.
- alert: GitlabSidekiqQueueTooLarge
expr: 'sum(sidekiq_running_jobs) >= sum(sidekiq_concurrency) * 0.9'
for: 10m
labels:
severity: warning
annotations:
summary: GitLab Sidekiq queue too large (instance {{ $labels.instance }})
description: "GitLab Sidekiq has {{ $value }} running jobs, approaching concurrency limit on {{ $labels.instance }}.\n VALUE = {{ $value }}\n LABELS = {{ $labels }}"warning
8.4.1.15.GitLab Sidekiq high job completion time
More than 5% of observed GitLab Sidekiq jobs take longer than 300 seconds ({{ $value }}%).
# This metric requires the emit_sidekiq_histogram_metrics feature flag to be enabled.
# GitLab's largest finite completion bucket is 300 seconds, so p95 cannot exceed it.
# This rule measures the percentage of completed jobs above that boundary, not an interpolated p95.
- alert: GitlabSidekiqHighJobCompletionTime
expr: '100 * (sum by (worker) (rate(sidekiq_jobs_completion_seconds_bucket{le="+Inf"}[5m])) - sum by (worker) (rate(sidekiq_jobs_completion_seconds_bucket{le="300"}[5m]))) / sum by (worker) (rate(sidekiq_jobs_completion_seconds_bucket{le="+Inf"}[5m])) > 5'
for: 10m
labels:
severity: warning
annotations:
summary: GitLab Sidekiq high job completion time (instance {{ $labels.instance }})
description: "More than 5% of observed GitLab Sidekiq jobs take longer than 300 seconds ({{ $value }}%).\n VALUE = {{ $value }}\n LABELS = {{ $labels }}"warning
8.4.1.16.GitLab Sidekiq high queue latency
More than 5% of observed GitLab Sidekiq queue waits take longer than 60 seconds ({{ $value }}%).
# This metric requires the emit_sidekiq_histogram_metrics feature flag to be enabled.
# High queue latency means jobs are stuck waiting. Check Sidekiq concurrency and queue sizes.
# GitLab's largest finite queue-duration bucket is 60 seconds, so p95 cannot exceed it.
# This rule measures the percentage of observed waits above that boundary, not an interpolated p95.
- alert: GitlabSidekiqHighQueueLatency
expr: '100 * (sum(rate(sidekiq_jobs_queue_duration_seconds_bucket{le="+Inf"}[5m])) - sum(rate(sidekiq_jobs_queue_duration_seconds_bucket{le="60"}[5m]))) / sum(rate(sidekiq_jobs_queue_duration_seconds_bucket{le="+Inf"}[5m])) > 5'
for: 5m
labels:
severity: warning
annotations:
summary: GitLab Sidekiq high queue latency (instance {{ $labels.instance }})
description: "More than 5% of observed GitLab Sidekiq queue waits take longer than 60 seconds ({{ $value }}%).\n VALUE = {{ $value }}\n LABELS = {{ $labels }}"warning
8.4.1.17.GitLab CI pipeline creation slow
GitLab CI pipeline creation p95 latency on {{ $labels.instance }} is above 30 seconds.
# 30s threshold is a rough default; depends on your pipeline complexity and CI instance load — adjust based on your baseline.
- alert: GitlabCiPipelineCreationSlow
expr: 'histogram_quantile(0.95, sum(rate(gitlab_ci_pipeline_creation_duration_seconds_bucket[5m])) by (le)) > 30'
for: 5m
labels:
severity: warning
annotations:
summary: GitLab CI pipeline creation slow (instance {{ $labels.instance }})
description: "GitLab CI pipeline creation p95 latency on {{ $labels.instance }} is above 30 seconds.\n VALUE = {{ $value }}\n LABELS = {{ $labels }}"warning
8.4.1.18.GitLab CI pipeline failures increasing
GitLab CI pipeline failures are increasing on {{ $labels.instance }} ({{ $value }}/s).
# This metric may not exist in all GitLab versions. Verify against your GitLab installation.
- alert: GitlabCiPipelineFailuresIncreasing
expr: 'deriv(gitlab_ci_pipeline_failure_reasons[5m]) > 0.05'
for: 10m
labels:
severity: warning
annotations:
summary: GitLab CI pipeline failures increasing (instance {{ $labels.instance }})
description: "GitLab CI pipeline failures are increasing on {{ $labels.instance }} ({{ $value }}/s).\n VALUE = {{ $value }}\n LABELS = {{ $labels }}"warning
8.4.1.19.GitLab rack uncaught errors
GitLab is experiencing uncaught errors in the Rack layer on {{ $labels.instance }} ({{ $value }}/s).
- alert: GitlabRackUncaughtErrors
expr: 'rate(rack_uncaught_errors_total[5m]) > 0.05'
for: 5m
labels:
severity: warning
annotations:
summary: GitLab rack uncaught errors (instance {{ $labels.instance }})
description: "GitLab is experiencing uncaught errors in the Rack layer on {{ $labels.instance }} ({{ $value }}/s).\n VALUE = {{ $value }}\n LABELS = {{ $labels }}"warning
8.4.1.20.GitLab version mismatch
Multiple GitLab versions are running across the fleet.
# This may happen during a rolling deployment. If it persists, investigate incomplete upgrades.
- alert: GitlabVersionMismatch
expr: 'count(count by (version) (gitlab_build_info)) > 1'
for: 0m
labels:
severity: warning
annotations:
summary: GitLab version mismatch (instance {{ $labels.instance }})
description: "Multiple GitLab versions are running across the fleet.\n VALUE = {{ $value }}\n LABELS = {{ $labels }}"warning
8.4.1.21.GitLab Ruby threads saturated
GitLab running threads on {{ $labels.instance }} have exceeded the expected maximum ({{ $value }}).
- alert: GitlabRubyThreadsSaturated
expr: 'sum by (instance) (gitlab_ruby_threads_running_threads) > on(instance) gitlab_ruby_threads_max_expected_threads * 1.5'
for: 10m
labels:
severity: warning
annotations:
summary: GitLab Ruby threads saturated (instance {{ $labels.instance }})
description: "GitLab running threads on {{ $labels.instance }} have exceeded the expected maximum ({{ $value }}).\n VALUE = {{ $value }}\n LABELS = {{ $labels }}"warning
8.4.1.22.GitLab Sidekiq queue not draining
GitLab Sidekiq queue {{ $labels.name }} has had a persistent backlog of {{ $value }} jobs for 30 minutes, suggesting a stuck or crashed worker class rather than transient load.
- alert: GitlabSidekiqQueueNotDraining
expr: 'sum by (name) (sidekiq_queue_size) > 0'
for: 30m
labels:
severity: warning
annotations:
summary: GitLab Sidekiq queue not draining (instance {{ $labels.instance }})
description: "GitLab Sidekiq queue {{ $labels.name }} has had a persistent backlog of {{ $value }} jobs for 30 minutes, suggesting a stuck or crashed worker class rather than transient load.\n VALUE = {{ $value }}\n LABELS = {{ $labels }}"warning
8.4.1.23.GitLab CI runner authentication failures
GitLab CI runners are experiencing authentication failures on {{ $labels.instance }} ({{ $value }} failures).
# Frequent runner auth failures may indicate expired tokens or misconfigured runners.
- alert: GitlabCiRunnerAuthenticationFailures
expr: 'increase(gitlab_ci_runner_authentication_failure_total[5m]) > 5'
for: 5m
labels:
severity: warning
annotations:
summary: GitLab CI runner authentication failures (instance {{ $labels.instance }})
description: "GitLab CI runners are experiencing authentication failures on {{ $labels.instance }} ({{ $value }} failures).\n VALUE = {{ $value }}\n LABELS = {{ $labels }}"