Prometheus

Prometheus는 "데이터베이스"가 아니라 "수집 + 저장 + 평가" 계층으로 이해하는 것이 정확하다. 시각화는 Grafana에, 알림 발송은 Alertmanager에 위임하고, Prometheus는 수집 주기와 알림 규칙 평가만 담당한다. 따라서 Prometheus 단일 설정(prometheus.yml) 안에서 수집 대상과 알림 규칙, Alertmanager 위치가 모두 선언되어야 한다.

pull 기반 수집 모델

Prometheus는 scrape_configs에 정의한 수집 대상(Exporter, App 등)을 global의 기본 수집 주기(scrape_interval)마다 수집한다.[1][2] 공식 설정 문서에서 scrape_interval은 대상을 scrape하는 주기이고, metrics_path는 대상에서 metric을 가져오는 HTTP 경로이며 기본값은 /metrics다.[3] 이를 Prometheus가 대상을 직접 방문해 가져오는 pull 기반 시계열 수집 시스템이고 수집한 값을 TSDB에 저장한다고 정리하는 것은 이 page의 해석이다.

flowchart LR
  Target["Target (Exporter/App)
/metrics"] Prom["Prometheus
scrape + TSDB"] Rule["rule_files
alert/recording rules"] Am["Alertmanager"] Target -->|pull| Prom Prom --> Rule Rule --> Am

prometheus.yml의 네 핵심 섹션

prometheus.yml은 Prometheus가 어떤 데이터를 수집할지, 얼마나 자주 수집할지, 어떤 규칙 파일을 불러올지 정의하는 핵심 설정 파일이다.[1:1]

섹션 설명 필수
global 기본 수집 주기(scrape_interval), rule 평가 주기(evaluation_interval), 선택적 scrape_timeout 필수
scrape_configs 수집 대상 목록 정의 (job_name + targets) 핵심
rule_files Alert Rule / Recording Rule 파일 경로 등록. 안 하면 rules.yml 작성해도 경보 작동 안 함 경보 사용 시 필수
alerting Alertmanager 연결 (alertmanagers → static_configs.targets) 경보 사용 시 필수

설정 파일과 command-line flag는 다른 관심사를 담당한다

공식 문서는 configuration file과 command-line flag를 명확히 구분한다 — command-line flag는 storage location, 디스크/메모리 보관량처럼 immutable system parameter를 설정하고, configuration file은 scrape job/instance, 불러올 rule file처럼 재구성 가능한 동작을 정의한다.[7] scrape_configs, rule_files, alerting이 설정 파일의 관심사이고, storage location·retention처럼 프로세스 재시작이 필요한 값은 flag의 관심사인 이유가 여기에 있다.

Prometheus는 실행 중 configuration reload를 지원하지만, 새 설정이 유효하지 않으면 기존 설정을 계속 사용한다.[8] 이 page의 운영 해석으로는, 변경 전에 syntax/semantic validation을 거치고 reload 성공 여부를 확인해야 하며, reload가 가능한 설정과 재시작이 필요한 flag를 혼동하면 retention/storage 변경이 반영됐다고 잘못 가정하게 된다.

최소 실행 템플릿 구조

global:
  scrape_interval: 15s
  evaluation_interval: 15s

rule_files:
  - /etc/prometheus/rules.yml

alerting:
  alertmanagers:
    - static_configs:
        - targets: ['alertmanager:9093']

scrape_configs:
  - job_name: 'host'
    static_configs:
      - targets: ['node-exporter:9100']
  - job_name: 'containers'
    static_configs:
      - targets: ['cadvisor:8080']
  - job_name: 'api-probe'
    metrics_path: /probe
    params:
      module: [http_2xx]
    static_configs:
      - targets: ['http://app:3000/health']
    relabel_configs:
      - source_labels: [__address__]
        target_label: __param_target
      - target_label: instance
        replacement: api
      - target_label: __address__
        replacement: blackbox:9115

Recording Rule: 자주 쓰는 계산을 미리 저장한다

groups:
  - name: recording_rules
    interval: 1m          # 평가 주기 (prometheus.yml의 evaluation_interval 기본값 사용)
    rules:
      - record: job:http_requests:rate5m
        expr: sum by (job)(rate(http_requests_total[5m]))
구분 Recording Rule Alert Rule
목적 새 metric 생성 조건 만족 시 경보 발생
결과 TSDB에 저장 Alertmanager로 전달
핵심 키 record alert

설정 예시의 경계

실패 모드

관련

출처

테스트 질문


  • Prometheus.md — 0. 목적, 1. 구성 요소 구조 ↩︎ ↩︎

  • Prometheus.md — 2. 핵심 문법 정리 - scrape_configs ↩︎ ↩︎

  • prometheus-configuration.md — scrape_config(L299-323): "# How frequently to scrape targets from this job.", "# The HTTP resource path on which to fetch metrics from targets.", "[ metrics_path: | default = /metrics ]" ↩︎
  • Prometheus.md — 2. 핵심 문법 정리 - global ↩︎

  • Prometheus.md — 2. 핵심 문법 정리 - rule_files, "등록 안 하면 rules.yml 작성해도 경보 작동 안 함" ↩︎

  • Prometheus.md — 2. 핵심 문법 정리 - alerting ↩︎

  • prometheus-configuration.md — "While the command-line flags configure immutable system parameters (such as storage locations, amount of data to keep on disk and in memory, etc.), the configuration file defines everything related to scraping jobs and their instances, as well as which rule files to load." ↩︎

  • prometheus-configuration.md — "Prometheus can reload its configuration at runtime. If the new configuration is not well-formed, the changes will not be applied." ↩︎

  • Recording Rules.md — "PromQL 표현식을 미리 계산해 새 메트릭으로 저장 — 스파이크 제거 및 복잡한 쿼리 최적화", 목적 표(노이즈 제거, 쿼리 최적화, 대시보드 성능), 기본 구조 YAML, "record 이름 규칙: level:metric:operations 형식 권장", "60fps 환경에서 프레임 1개 처리 예산은 16.666ms", avg_over_time(apollo_frame_processing_latency_avg[1m]) / 16.666, "max 값을 기준으로 잡으면 200ms 같은 이상값이 나올 수 있음 — avg 사용 권장", "Recording Rules + Alert Rules 동일 파일 사용 가능", Alert Rule과의 차이 표, Alert Rule 예시 expr: apollo:frame_latency:budget_ratio1m > 1.0·for: 2m ↩︎ ↩︎ ↩︎ ↩︎ ↩︎