Alertmanager

Prometheus가 경보를 "감지"한다면, Alertmanager는 그 경보를 받아 중복 제거·그룹화·라우팅한 뒤 receiver(webhook/Discord/Slack/email)로 발송하는 시스템이다.[1] 핵심 역할은 경보 발송 자체가 아니라 "어떤 경보를 어디로, 어떻게 묶어서 보낼 것인가"를 결정하는 것이다.

flowchart LR
  Prom["Prometheus
alert rule 평가"] Am["Alertmanager"] Route["route
severity 분기"] Recv["receivers
Discord/Webhook/Slack"] Prom -->|fire alerts| Am Am --> Route Route --> Recv

핵심 요소

요소 설명
route 알림이 들어왔을 때 가장 먼저 적용되는 라우팅 규칙[1:1]
receivers 실제 발송 대상 (webhook, Discord, Slack, email 등)[1:2]
grouping 동일한 경보 여러 개를 하나로 묶어 전송[1:3]
inhibit_rules 특정 조건일 때 다른 경보 무시 (예: API DOWN → 5xx 알람 억제)[1:4]
silence 일정 기간·조건에 맞는 경보의 notification을 사람이 명시적으로 억제
template 메시지 내용 커스터마이징 (Go 템플릿 문법)[1:5]

inhibition은 다른 firing alert의 존재로 notification을 억제하는 규칙 기반 관계이고, silence는 운영자가 matcher와 기간을 지정해 잠시 notification을 억제하는 명시적 운영 조치라고 볼 수 있다. 둘을 같은 기능으로 취급하면 장애 대응 중 억제 원인을 추적하기 어렵다.

기본 설정 구조

severity 기반 라우팅 + Discord 패턴

severity 라벨로 라우팅 분기를 하면 Critical/Warning을 서로 다른 receiver로 보낼 수 있다. 다만 severity는 Alert Rule에서 추가한 값 기준으로 분기된다는 점이 핵심이다 — Alert rule에서 severity: critical을 달지 않으면 라우팅이 동작하지 않는다.[3]

route:
  receiver: "webhook-handler"
  group_by: ["alertname", "instance"]
  group_wait: 30s
  group_interval: 3m
  repeat_interval: 2h
  routes:
    - match:
        severity: critical
      receiver: "discord-critical"
    - match:
        severity: warning
      receiver: "discord-warning"

receivers:
  - name: "discord-critical"
    discord_configs:
      - webhook_url: "https://discord.com/api/webhooks/XXXXX"
  - name: "discord-warning"
    discord_configs:
      - webhook_url: "https://discord.com/api/webhooks/YYYYY"

Webhook receiver는 send_resolved: true로 해결됨 알림도 전송할지 제어할 수 있다. Alertmanager가 보내는 JSON은 version, status(firing/resolved), alerts[](labels + annotations) 구조다.[4]

Inhibit Rule 예시

inhibit_rules:
  - source_match:
      alertname: "ApiDown"
    target_match:
      alertname: "High5xxRate"
    equal: ["instance"]

ApiDown이 firing이면 같은 instance의 High5xxRate 경보가 자동으로 억제된다.[5] 이는 상위 에러 발생 시 하위 경보를 억제하는 예시이며,[5:1] 하위 증상 알람이 사용자를 혼란시키는 것을 막는다는 목적은 이 page의 해석이다.

실행 전 검증

alertmanager --config.file=alertmanager.yml --log.level=debug
# 또는
amtool check-config alertmanager.yml

amtool check-config로 YAML 문법을 사전 검증하면 설정 오류를 배포 전 잡을 수 있다.[6]

설정 예시의 경계

Prometheus가 Alert rule 평가만 담당하고, 발송 라우팅·중복 억제·해결 알림은 Alertmanager가 별도 스펙(alertmanager.yml)으로 다룬다. 이 분리 덕분에 알림 채널 변경이 Prometheus 재설정 없이 가능하다고 볼 수 있다.

흔한 실수

관련

출처

테스트 질문


  1. Altermanager.md — "역할: Prometheus가 감지한 경보를 '어디로, 어떻게 보낼지' 결정하는 라우팅 시스템", 핵심 요소 표(route/receivers/grouping/inhibit_rules/template) ↩︎ ↩︎ ↩︎ ↩︎ ↩︎ ↩︎

  2. Altermanager.md — 기본 구조 템플릿의 global.resolve_timeout, route.group_wait/group_interval/repeat_interval 주석 ↩︎ ↩︎

  3. Altermanager.md — Discord+Webhook+Severity 라우팅 예시, "severity 라벨은 Alert Rule에서 추가한 값 기준으로 분기됨" ↩︎

  4. Altermanager.md — Webhook 기반 알림 설명과 JSON 예시(version/status/alerts[]) ↩︎

  5. Altermanager.md — "## 5) Inhibit Rule 예시 (상위 에러 발생 시 하위 경보 억제)"(L132)의 ApiDown → High5xxRate 억제 설정 ↩︎ ↩︎

  6. Altermanager.md — 실행 전 검증(amtool check-config) ↩︎