Prometheus는 "데이터베이스"가 아니라 "수집 + 저장 + 평가" 계층으로 이해하는 것이 정확하다. 시각화는 Grafana에, 알림 발송은 Alertmanager에 위임하고, Prometheus는 수집 주기와 알림 규칙 평가만 담당한다. 따라서 Prometheus 단일 설정(prometheus.yml) 안에서 수집 대상과 알림 규칙, Alertmanager 위치가 모두 선언되어야 한다.
Prometheus는 scrape_configs에 정의한 수집 대상(Exporter, App 등)을 global의 기본 수집 주기(scrape_interval)마다 수집한다.[1][2] 공식 설정 문서에서 scrape_interval은 대상을 scrape하는 주기이고, metrics_path는 대상에서 metric을 가져오는 HTTP 경로이며 기본값은 /metrics다.[3] 이를 Prometheus가 대상을 직접 방문해 가져오는 pull 기반 시계열 수집 시스템이고 수집한 값을 TSDB에 저장한다고 정리하는 것은 이 page의 해석이다.
flowchart LR Target["Target (Exporter/App)
/metrics"] Prom["Prometheus
scrape + TSDB"] Rule["rule_files
alert/recording rules"] Am["Alertmanager"] Target -->|pull| Prom Prom --> Rule Rule --> Am
prometheus.yml의 네 핵심 섹션prometheus.yml은 Prometheus가 어떤 데이터를 수집할지, 얼마나 자주 수집할지, 어떤 규칙 파일을 불러올지 정의하는 핵심 설정 파일이다.[1:1]
| 섹션 | 설명 | 필수 |
|---|---|---|
global |
기본 수집 주기(scrape_interval), rule 평가 주기(evaluation_interval), 선택적 scrape_timeout |
필수 |
scrape_configs |
수집 대상 목록 정의 (job_name + targets) | 핵심 |
rule_files |
Alert Rule / Recording Rule 파일 경로 등록. 안 하면 rules.yml 작성해도 경보 작동 안 함 | 경보 사용 시 필수 |
alerting |
Alertmanager 연결 (alertmanagers → static_configs.targets) |
경보 사용 시 필수 |
global 섹션은 scrape_interval, evaluation_interval, scrape_timeout을 정의한다.[4]scrape_configs의 job_name + static_configs.targets로 수집 대상을 그룹 단위로 정의하며, 특정 job에만 별도 scrape_interval을 적용할 수 있다.[2:1]rule_files에 Alert Rule 파일(예: /etc/prometheus/rules.yml)을 등록하지 않으면 rules.yml을 작성해도 경보가 작동하지 않는다.[5]alerting.alertmanagers.static_configs.targets로 Alertmanager 주소를 연결한다.[6]공식 문서는 configuration file과 command-line flag를 명확히 구분한다 — command-line flag는 storage location, 디스크/메모리 보관량처럼 immutable system parameter를 설정하고, configuration file은 scrape job/instance, 불러올 rule file처럼 재구성 가능한 동작을 정의한다.[7] scrape_configs, rule_files, alerting이 설정 파일의 관심사이고, storage location·retention처럼 프로세스 재시작이 필요한 값은 flag의 관심사인 이유가 여기에 있다.
Prometheus는 실행 중 configuration reload를 지원하지만, 새 설정이 유효하지 않으면 기존 설정을 계속 사용한다.[8] 이 page의 운영 해석으로는, 변경 전에 syntax/semantic validation을 거치고 reload 성공 여부를 확인해야 하며, reload가 가능한 설정과 재시작이 필요한 flag를 혼동하면 retention/storage 변경이 반영됐다고 잘못 가정하게 된다.
global:
scrape_interval: 15s
evaluation_interval: 15s
rule_files:
- /etc/prometheus/rules.yml
alerting:
alertmanagers:
- static_configs:
- targets: ['alertmanager:9093']
scrape_configs:
- job_name: 'host'
static_configs:
- targets: ['node-exporter:9100']
- job_name: 'containers'
static_configs:
- targets: ['cadvisor:8080']
- job_name: 'api-probe'
metrics_path: /probe
params:
module: [http_2xx]
static_configs:
- targets: ['http://app:3000/health']
relabel_configs:
- source_labels: [__address__]
target_label: __param_target
- target_label: instance
replacement: api
- target_label: __address__
replacement: blackbox:9115
level:metric:operations 형식을 권장한다.[9:1]groups:
- name: recording_rules
interval: 1m # 평가 주기 (prometheus.yml의 evaluation_interval 기본값 사용)
rules:
- record: job:http_requests:rate5m
expr: sum by (job)(rate(http_requests_total[5m]))
avg_over_time(...[1m]) / 16.666을 예산 대비 비율 metric으로 저장한다. 원문은 max를 기준으로 잡으면 200ms 같은 이상값이 나올 수 있어 avg를 권장한다.[9:2]rule_files에 함께 둘 수 있지만 목적과 결과가 다르다.[9:3]| 구분 | Recording Rule | Alert Rule |
|---|---|---|
| 목적 | 새 metric 생성 | 조건 만족 시 경보 발생 |
| 결과 | TSDB에 저장 | Alertmanager로 전달 |
| 핵심 키 | record |
alert |
apollo:frame_latency:budget_ratio1m > 1.0이 for: 2m 동안 지속되면 경보를 낸다.[9:4] 스파이크를 먼저 평균으로 누른 값에 지속 시간 조건까지 걸어 오탐을 이중으로 줄이는 구성으로 읽을 수 있다.prometheus.yml 예시는 scrape 대상, 파일 경로, rule file 이름이 현재 운영 환경에 맞는 경우에만 적용된다.rule_files 등록 없이 rules.yml만 작성해 놓고 경보가 안 울린다고 착각하는 경우.alerting 섹션 없이 Alertmanager만 실행해 두고 알림이 발송되지 않는다고 가정하는 경우.static_configs만 쓰면서 대상이 늘어날 때마다 수동으로 yml을 수정하는 경우 (→ docker_sd_configs/kubernetes_sd_configs 도입 검토).basic_auth/bearer_token_file 없이 scrape 시도 → 401 반환./metrics가 unreachable인 경우./metrics 노출 주체들 (node-exporter, cAdvisor, blackbox 등).prometheus.yml의 필수 섹션 4가지와 각 역할은 무엇인가?rule_files를 등록하지 않으면 어떤 문제가 발생하는가?Prometheus.md — 0. 목적, 1. 구성 요소 구조 ↩︎ ↩︎
Prometheus.md — 2. 핵심 문법 정리 - scrape_configs ↩︎ ↩︎
Prometheus.md — 2. 핵심 문법 정리 - global ↩︎
Prometheus.md — 2. 핵심 문법 정리 - rule_files, "등록 안 하면 rules.yml 작성해도 경보 작동 안 함" ↩︎
Prometheus.md — 2. 핵심 문법 정리 - alerting ↩︎
prometheus-configuration.md — "While the command-line flags configure immutable system parameters (such as storage locations, amount of data to keep on disk and in memory, etc.), the configuration file defines everything related to scraping jobs and their instances, as well as which rule files to load." ↩︎
prometheus-configuration.md — "Prometheus can reload its configuration at runtime. If the new configuration is not well-formed, the changes will not be applied." ↩︎
Recording Rules.md — "PromQL 표현식을 미리 계산해 새 메트릭으로 저장 — 스파이크 제거 및 복잡한 쿼리 최적화", 목적 표(노이즈 제거, 쿼리 최적화, 대시보드 성능), 기본 구조 YAML, "record 이름 규칙: level:metric:operations 형식 권장", "60fps 환경에서 프레임 1개 처리 예산은 16.666ms", avg_over_time(apollo_frame_processing_latency_avg[1m]) / 16.666, "max 값을 기준으로 잡으면 200ms 같은 이상값이 나올 수 있음 — avg 사용 권장", "Recording Rules + Alert Rules 동일 파일 사용 가능", Alert Rule과의 차이 표, Alert Rule 예시 expr: apollo:frame_latency:budget_ratio1m > 1.0·for: 2m ↩︎ ↩︎ ↩︎ ↩︎ ↩︎