モニタリング¶
重要
PrometheusとGrafanaのインストールは、このプロジェクトの範囲を超えています。それらがシステムに正しくインストールされていると仮定します。ただし、実験のために、 Part 4 of the Quickstart で手順を提供します。
モニタリングインスタンス¶
PostgreSQLインスタンスごとに、オペレーターは、metrics
という名前のポート9187で、HTTPを介して Prometheusオペレーターの例 のメトリックのエクスポーターを提供します。オペレーターには
事前定義されたメトリックのセット と、1つ以上の
ConfigMapまたはSecret
リソースを介して追加のクエリを定義するための高度に構成およびカスタマイズ可能なシステムが付属しています(詳細については、以下の ユーザー定義のメトリック を参照)。
重要
バージョン1.11以降、CloudNativePGは`default-monitoring` と呼ばれる`ConfigMap` に by default a set of predefined metrics を既にインストールしています。
注釈
エクスポートされたメトリックを検査する方法 の指示に従って、エクスポートされたメトリックを検査できます
以下のセクション。
PostgreSQLで実行されるすべてのモニタリングクエリは次のとおりです。
アトミック(クエリごとに1つのトランザクション)
pg_monitorロールで実行application_nameをcnpg_metrics_exporterに設定して実行ユーザー
postgresとして実行
- PostgreSQL
documentation の「デフォルトのロール」セクションを参照してください
pg_monitor ロールの詳細については。
クエリは、デフォルトで、次のロジックに従って、 Cluster
リソースの指定されたbootstrap メソッドで定義された
メインデータベース に対して実行されます。
initdbを使用:クエリはデフォルトでinitdb.database、または指定されていない場合appの指定されたデータベースに対して実行されますrecoveryを使用:クエリはデフォルトでrecovery.database、または指定されていない場合postgresの指定されたデータベースに対して実行されますpg_basebackupを使用:クエリはデフォルトでpg_basebackup.database、または指定されていない場合postgresの指定されたデータベースに対して実行されます
target_databases
オプションで1つ以上のデータベースのリストを指定することにより、デフォルトのデータベースを特定のユーザー定義メトリックに対していつでもオーバーライドできます。
Prometheusオペレーターの例¶
を使用して、特定のPostgreSQLクラスターを監視できます。クラスターを正しく指すPodMonitorは、クラスターリソース自分自身で.spec.monitoring.enablePodMonitor
をtrue
に設定することにより、オペレーターが自動的に作成できます(デフォルト:false)。
重要
自動的に作成された`PodMonitor` への変更は、次の調整サイクルでオペレーターによってオーバーライドされます。カスタマイズが必要な場合は、以下に説明するように行うことができます。
特定のクラスターにPodMonitor
を手動でデプロイするには、次のように定義し、必要に応じて変更します。
apiVersion: monitoring.coreos.com/v1
kind: PodMonitor
metadata:
name: cluster-example
spec:
selector:
matchLabels:
"cnpg.io/cluster": cluster-example
podMetricsEndpoints:
- port: metrics
重要
一意の名前、正しいクラスターの名前空間とラベルを使用して、上記の例を変更してください(cluster-example を使用しています)。
重要
このドキュメントの以前のバージョンで使用されたラベル`postgresql` は非推奨であり、将来的に削除されます。インスタンスを選択するには、代わりにラベル`cnpg.io/cluster` を使用してください。
事前定義されたメトリックのセット¶
すべてのPostgreSQLインスタンスエクスポーターは、事前定義された一連のメトリックを自動的に公開します。これは、2つの主要なカテゴリに分類できます。
cnpg_collector_*で始まるPostgreSQL関連のメトリック。ディスク上のWALファイルの数と合計サイズ - アーカイブステータスフォルダー内の
.readyおよび.doneファイルの数 - 要求された同期レプリカの最小数と最大数、ならびに予想および実際に観測された値 - インスタンスを収容する個別のノードの数 -最後に失敗したバックアップと最後に利用可能なバックアップ、およびクラスターの最初の回復可能性を示すタイムスタンプ - レプリカクラスターモードが有効または無効かどうかを示すフラグ - 手動スイッチオーバーが必要かどうかを示すフラグ - フェンシングが有効か無効かを示すフラグgo_*で始まるGoランタイム関連のメトリック
以下は、インスタンスのlocalhost:9187/metrics
エンドポイントによって返されるメトリックのサンプルです。ご覧のとおり、Prometheusフォーマットは自己文書化です。
# HELP cnpg_collector_collection_duration_seconds Collection time duration in seconds
# TYPE cnpg_collector_collection_duration_seconds gauge
cnpg_collector_collection_duration_seconds{collector="Collect.up"} 0.0031393
# HELP cnpg_collector_collections_total Total number of times PostgreSQL was accessed for metrics.
# TYPE cnpg_collector_collections_total counter
cnpg_collector_collections_total 2
# HELP cnpg_collector_fencing_on 1 if the instance is fenced, 0 otherwise
# TYPE cnpg_collector_fencing_on gauge
cnpg_collector_fencing_on 0
# HELP cnpg_collector_nodes_used NodesUsed represents the count of distinct nodes accommodating the instances. A value of -1 suggests that the metric is not available. A value of 1 suggests that all instances are hosted on a single node, implying the absence of High Availability (HA). Ideally this value should match the number of instances in the cluster.
# TYPE cnpg_collector_nodes_used gauge
cnpg_collector_nodes_used 3
# HELP cnpg_collector_last_collection_error 1 if the last collection ended with error, 0 otherwise.
# TYPE cnpg_collector_last_collection_error gauge
cnpg_collector_last_collection_error 0
# HELP cnpg_collector_manual_switchover_required 1 if a manual switchover is required, 0 otherwise
# TYPE cnpg_collector_manual_switchover_required gauge
cnpg_collector_manual_switchover_required 0
# HELP cnpg_collector_pg_wal Total size in bytes of WAL segments in the /var/lib/postgresql/data/pgdata/pg_wal directory computed as (wal_segment_size * count)
# TYPE cnpg_collector_pg_wal gauge
cnpg_collector_pg_wal{value="count"} 9
cnpg_collector_pg_wal{value="slots_max"} NaN
cnpg_collector_pg_wal{value="keep"} 32
cnpg_collector_pg_wal{value="max"} 64
cnpg_collector_pg_wal{value="min"} 5
cnpg_collector_pg_wal{value="size"} 1.50994944e+08
cnpg_collector_pg_wal{value="volume_max"} 128
cnpg_collector_pg_wal{value="volume_size"} 2.147483648e+09
# HELP cnpg_collector_pg_wal_archive_status Number of WAL segments in the /var/lib/postgresql/data/pgdata/pg_wal/archive_status directory (ready, done)
# TYPE cnpg_collector_pg_wal_archive_status gauge
cnpg_collector_pg_wal_archive_status{value="done"} 6
cnpg_collector_pg_wal_archive_status{value="ready"} 0
# HELP cnpg_collector_replica_mode 1 if the cluster is in replica mode, 0 otherwise
# TYPE cnpg_collector_replica_mode gauge
cnpg_collector_replica_mode 0
# HELP cnpg_collector_sync_replicas Number of requested synchronous replicas (synchronous_standby_names)
# TYPE cnpg_collector_sync_replicas gauge
cnpg_collector_sync_replicas{value="expected"} 0
cnpg_collector_sync_replicas{value="max"} 0
cnpg_collector_sync_replicas{value="min"} 0
cnpg_collector_sync_replicas{value="observed"} 0
# HELP cnpg_collector_up 1 if PostgreSQL is up, 0 otherwise.
# TYPE cnpg_collector_up gauge
cnpg_collector_up{cluster="cluster-example"} 1
# HELP cnpg_collector_postgres_version Postgres version
# TYPE cnpg_collector_postgres_version gauge
cnpg_collector_postgres_version{cluster="cluster-example",full="15.3"} 15.3
# HELP cnpg_collector_last_failed_backup_timestamp The last failed backup as a unix timestamp
# TYPE cnpg_collector_last_failed_backup_timestamp gauge
cnpg_collector_last_failed_backup_timestamp 0
# HELP cnpg_collector_last_available_backup_timestamp The last available backup as a unix timestamp
# TYPE cnpg_collector_last_available_backup_timestamp gauge
cnpg_collector_last_available_backup_timestamp 1.63238406e+09
# HELP cnpg_collector_first_recoverability_point The first point of recoverability for the cluster as a unix timestamp
# TYPE cnpg_collector_first_recoverability_point gauge
cnpg_collector_first_recoverability_point 1.63238406e+09
# HELP cnpg_collector_lo_pages Estimated number of pages in the pg_largeobject table
# TYPE cnpg_collector_lo_pages gauge
cnpg_collector_lo_pages{datname="app"} 0
cnpg_collector_lo_pages{datname="postgres"} 78
# HELP cnpg_collector_wal_buffers_full Number of times WAL data was written to disk because WAL buffers became full. Only available on PG 14+
# TYPE cnpg_collector_wal_buffers_full gauge
cnpg_collector_wal_buffers_full{stats_reset="2023-06-19T10:51:27.473259Z"} 6472
# HELP cnpg_collector_wal_bytes Total amount of WAL generated in bytes. Only available on PG 14+
# TYPE cnpg_collector_wal_bytes gauge
cnpg_collector_wal_bytes{stats_reset="2023-06-19T10:51:27.473259Z"} 1.0035147e+07
# HELP cnpg_collector_wal_fpi Total number of WAL full page images generated. Only available on PG 14+
# TYPE cnpg_collector_wal_fpi gauge
cnpg_collector_wal_fpi{stats_reset="2023-06-19T10:51:27.473259Z"} 1474
# HELP cnpg_collector_wal_records Total number of WAL records generated. Only available on PG 14+
# TYPE cnpg_collector_wal_records gauge
cnpg_collector_wal_records{stats_reset="2023-06-19T10:51:27.473259Z"} 26178
# HELP cnpg_collector_wal_sync Number of times WAL files were synced to disk via issue_xlog_fsync request (if fsync is on and wal_sync_method is either fdatasync, fsync or fsync_writethrough, otherwise zero). Only available on PG 14+
# TYPE cnpg_collector_wal_sync gauge
cnpg_collector_wal_sync{stats_reset="2023-06-19T10:51:27.473259Z"} 37
# HELP cnpg_collector_wal_sync_time Total amount of time spent syncing WAL files to disk via issue_xlog_fsync request, in milliseconds (if track_wal_io_timing is enabled, fsync is on, and wal_sync_method is either fdatasync, fsync or fsync_writethrough, otherwise zero). Only available on PG 14+
# TYPE cnpg_collector_wal_sync_time gauge
cnpg_collector_wal_sync_time{stats_reset="2023-06-19T10:51:27.473259Z"} 0
# HELP cnpg_collector_wal_write Number of times WAL buffers were written out to disk via XLogWrite request. Only available on PG 14+
# TYPE cnpg_collector_wal_write gauge
cnpg_collector_wal_write{stats_reset="2023-06-19T10:51:27.473259Z"} 7243
# HELP cnpg_collector_wal_write_time Total amount of time spent writing WAL buffers to disk via XLogWrite request, in milliseconds (if track_wal_io_timing is enabled, otherwise zero). This includes the sync time when wal_sync_method is either open_datasync or open_sync. Only available on PG 14+
# TYPE cnpg_collector_wal_write_time gauge
cnpg_collector_wal_write_time{stats_reset="2023-06-19T10:51:27.473259Z"} 0
# HELP cnpg_last_error 1 if the last collection ended with error, 0 otherwise.
# TYPE cnpg_last_error gauge
cnpg_last_error 0
# HELP go_gc_duration_seconds A summary of the pause duration of garbage collection cycles.
# TYPE go_gc_duration_seconds summary
go_gc_duration_seconds{quantile="0"} 5.01e-05
go_gc_duration_seconds{quantile="0.25"} 7.27e-05
go_gc_duration_seconds{quantile="0.5"} 0.0001748
go_gc_duration_seconds{quantile="0.75"} 0.0002959
go_gc_duration_seconds{quantile="1"} 0.0012776
go_gc_duration_seconds_sum 0.0035741
go_gc_duration_seconds_count 13
# HELP go_goroutines Number of goroutines that currently exist.
# TYPE go_goroutines gauge
go_goroutines 25
# HELP go_info Information about the Go environment.
# TYPE go_info gauge
go_info{version="go1.20.5"} 1
# HELP go_memstats_alloc_bytes Number of bytes allocated and still in use.
# TYPE go_memstats_alloc_bytes gauge
go_memstats_alloc_bytes 4.493744e+06
# HELP go_memstats_alloc_bytes_total Total number of bytes allocated, even if freed.
# TYPE go_memstats_alloc_bytes_total counter
go_memstats_alloc_bytes_total 2.1698216e+07
# HELP go_memstats_buck_hash_sys_bytes Number of bytes used by the profiling bucket hash table.
# TYPE go_memstats_buck_hash_sys_bytes gauge
go_memstats_buck_hash_sys_bytes 1.456234e+06
# HELP go_memstats_frees_total Total number of frees.
# TYPE go_memstats_frees_total counter
go_memstats_frees_total 172118
# HELP go_memstats_gc_cpu_fraction The fraction of this programs available CPU time used by the GC since the program started.
# TYPE go_memstats_gc_cpu_fraction gauge
go_memstats_gc_cpu_fraction 1.0749468700447189e-05
# HELP go_memstats_gc_sys_bytes Number of bytes used for garbage collection system metadata.
# TYPE go_memstats_gc_sys_bytes gauge
go_memstats_gc_sys_bytes 5.530048e+06
# HELP go_memstats_heap_alloc_bytes Number of heap bytes allocated and still in use.
# TYPE go_memstats_heap_alloc_bytes gauge
go_memstats_heap_alloc_bytes 4.493744e+06
# HELP go_memstats_heap_idle_bytes Number of heap bytes waiting to be used.
# TYPE go_memstats_heap_idle_bytes gauge
go_memstats_heap_idle_bytes 5.8236928e+07
# HELP go_memstats_heap_inuse_bytes Number of heap bytes that are in use.
# TYPE go_memstats_heap_inuse_bytes gauge
go_memstats_heap_inuse_bytes 7.528448e+06
# HELP go_memstats_heap_objects Number of allocated objects.
# TYPE go_memstats_heap_objects gauge
go_memstats_heap_objects 26306
# HELP go_memstats_heap_released_bytes Number of heap bytes released to OS.
# TYPE go_memstats_heap_released_bytes gauge
go_memstats_heap_released_bytes 5.7401344e+07
# HELP go_memstats_heap_sys_bytes Number of heap bytes obtained from system.
# TYPE go_memstats_heap_sys_bytes gauge
go_memstats_heap_sys_bytes 6.5765376e+07
# HELP go_memstats_last_gc_time_seconds Number of seconds since 1970 of last garbage collection.
# TYPE go_memstats_last_gc_time_seconds gauge
go_memstats_last_gc_time_seconds 1.6311727586032727e+09
# HELP go_memstats_lookups_total Total number of pointer lookups.
# TYPE go_memstats_lookups_total counter
go_memstats_lookups_total 0
# HELP go_memstats_mallocs_total Total number of mallocs.
# TYPE go_memstats_mallocs_total counter
go_memstats_mallocs_total 198424
# HELP go_memstats_mcache_inuse_bytes Number of bytes in use by mcache structures.
# TYPE go_memstats_mcache_inuse_bytes gauge
go_memstats_mcache_inuse_bytes 14400
# HELP go_memstats_mcache_sys_bytes Number of bytes used for mcache structures obtained from system.
# TYPE go_memstats_mcache_sys_bytes gauge
go_memstats_mcache_sys_bytes 16384
# HELP go_memstats_mspan_inuse_bytes Number of bytes in use by mspan structures.
# TYPE go_memstats_mspan_inuse_bytes gauge
go_memstats_mspan_inuse_bytes 191896
# HELP go_memstats_mspan_sys_bytes Number of bytes used for mspan structures obtained from system.
# TYPE go_memstats_mspan_sys_bytes gauge
go_memstats_mspan_sys_bytes 212992
# HELP go_memstats_next_gc_bytes Number of heap bytes when next garbage collection will take place.
# TYPE go_memstats_next_gc_bytes gauge
go_memstats_next_gc_bytes 8.689632e+06
# HELP go_memstats_other_sys_bytes Number of bytes used for other system allocations.
# TYPE go_memstats_other_sys_bytes gauge
go_memstats_other_sys_bytes 2.566622e+06
# HELP go_memstats_stack_inuse_bytes Number of bytes in use by the stack allocator.
# TYPE go_memstats_stack_inuse_bytes gauge
go_memstats_stack_inuse_bytes 1.343488e+06
# HELP go_memstats_stack_sys_bytes Number of bytes obtained from system for stack allocator.
# TYPE go_memstats_stack_sys_bytes gauge
go_memstats_stack_sys_bytes 1.343488e+06
# HELP go_memstats_sys_bytes Number of bytes obtained from system.
# TYPE go_memstats_sys_bytes gauge
go_memstats_sys_bytes 7.6891144e+07
# HELP go_threads Number of OS threads created.
# TYPE go_threads gauge
go_threads 18
注釈
cnpg_collector_postgres_version は、PostgreSQLの`Major.Minor` バージョンを含むGaugeVecメトリックです。完全なセマンティックバージョン`Major.Minor.Patch` は、full という名前のラベルフィールドの1つの中にあります。
注釈
cnpg_collector_first_recoverability_point および`cnpg_collector_last_available_backup_timestamp` は、オブジェクトストアへの最初のバックアップまでゼロです。これはWALアーカイブとは別のものです。
ユーザー定義のメトリック¶
この機能は現在 ベータ 状態であり、形式はPostgreSQL Prometheus Exporterの queries.yaml file に触発されています。
ユーザーは、次の例のように、 .spec.monitoring.customQueriesConfigMap
またはcustomQueriesSecret セクションの下のCluster
定義で作成されたConfigmap /Secret
を参照することにより、カスタムメトリックを定義できます。
apiVersion: postgresql.cnpg.io/v1
kind: Cluster
metadata:
name: cluster-example
namespace: test
spec:
instances: 3
storage:
size: 1Gi
monitoring:
customQueriesConfigMap:
- name: example-monitoring
key: custom-queries
customQueriesConfigMap /customQueriesSecret
セクションには、カスタムクエリが定義されるキーを指定するConfigMap
/Secret 参照のリストが含まれています。参照されるリソースは、
クラスター
リソースと同じ名前空間に作成する必要があることに注意してください。
注釈
ConfigMapとシークレットをインスタンスによって**自動的に**リロードしたい場合は、キー cnpg.io/reload を持つラベルをそれに追加できます。そうでない場合、 kubectl cnpg reload サブコマンドを使用してインスタンスをリロードする必要があります。
重要
ユーザー定義のメトリックが既存のメトリックを上書きすると、インスタンスマネージャーはメッセージを含むjson警告ログを出力します。Query with the same name already found. Overwriting the existing one. と、上書きされたクエリ名を含むキー queryName 。
ユーザー定義メトリックの例¶
ここでは、上記のCluster
の例で参照される、単一のカスタムクエリを含むConfigMap
の例を見ることができます。
apiVersion: v1
kind: ConfigMap
metadata:
name: example-monitoring
namespace: test
labels:
cnpg.io/reload: ""
data:
custom-queries: |
pg_replication:
query: "SELECT CASE WHEN NOT pg_is_in_recovery()
THEN 0
ELSE GREATEST (0,
EXTRACT(EPOCH FROM (now() - pg_last_xact_replay_timestamp())))
END AS lag,
pg_is_in_recovery() AS in_recovery,
EXISTS (TABLE pg_stat_wal_receiver) AS is_wal_receiver_up,
(SELECT count(*) FROM pg_stat_replication) AS streaming_replicas"
metrics:
- lag:
usage: "GAUGE"
description: "Replication lag behind primary in seconds"
- in_recovery:
usage: "GAUGE"
description: "Whether the instance is in recovery"
- is_wal_receiver_up:
usage: "GAUGE"
description: "Whether the instance wal_receiver is up"
- streaming_replicas:
usage: "GAUGE"
description: "Number of streaming replicas connected to the instance"
基本的なモニタリングクエリのリストは、 default-monitoring.yaml
- CloudNativePGデプロイメントに既にインストールされています(
メトリックのデフォルトセット を参照)。
複数のデータベースで実行されるユーザー定義メトリックの例¶
target_databases
オプションに複数のデータベースがリストされている場合、メトリックはそれぞれから収集されます。
target_databases のリストに シェルのようなパターン (つまり、
* 、? または[]
を含む)を指定することにより、特定のクエリに対してデータベースの自動検出を有効にできます。提供されている場合、オペレーターは
SELECT datname FROM pg_database WHERE datallowconn AND NOT datistemplate
の実行によって返されたすべてのデータベースを追加し、
path.Match() ルールに従ってパターンを照合することにより、ターゲットデータベースのリストを展開します。
注釈
* 文字にはyamlに special meaning があるため、そのようなパターンが含まれる場合は`target_databases` 値をクォート("*" )する必要があります。
たとえば、次の例のようにcurrent_database()
ファンクションを使用して、返されるラベルにデータベースの名前を常に含めることをお勧めします。
some_query: |
query: |
SELECT
current_database() as datname,
count(*) as rows
FROM some_table
metrics:
- datname:
usage: "LABEL"
description: "Name of current database"
- rows:
usage: "GAUGE"
description: "number of rows"
target_databases:
- albert
- bb
- freddie
これにより、次のメトリックが公開されます。
cnpg_some_query_rows{datname="albert"} 2
cnpg_some_query_rows{datname="bb"} 5
cnpg_some_query_rows{datname="freddie"} 10
これは、 template1
データベースでも実行される自動検出を有効にしたクエリの例です(それ以外の場合、前述のクエリで返されません)。
some_query: |
query: |
SELECT
current_database() as datname,
count(*) as rows
FROM some_table
metrics:
- datname:
usage: "LABEL"
description: "Name of current database"
- rows:
usage: "GAUGE"
description: "number of rows"
target_databases:
- "*"
- "template1"
上記の例では、次のメトリックが生成されます(データベースが存在する場合)。
cnpg_some_query_rows{datname="albert"} 2
cnpg_some_query_rows{datname="bb"} 5
cnpg_some_query_rows{datname="freddie"} 10
cnpg_some_query_rows{datname="template1"} 7
cnpg_some_query_rows{datname="postgres"} 42
ユーザー定義メトリックの構造¶
すべてのカスタムクエリには、次の基本構造があります。
<MetricName>:
query: "<SQLQuery>"
metrics:
- <ColumnName>:
usage: "<MetricType>"
description: "<MetricDescription>"
利用可能なすべてのフィールドの簡単な説明は次のとおりです。
<MetricName>:プロメテウスメトリックの名前 -query:メトリックを生成するためにターゲットデータベースで実行するSQLクエリ -primary:プライマリインスタンスでのみクエリを実行するかどうか -master:primaryと同じ(Prometheus PostgreSQLエクスポーターの構文との互換性のため - 非推奨) -runonserver:クエリを実行するPostgreSQLのバージョンを制限するセマンティックバージョン範囲(例">=11.0.0"または">=12.0.0 <=15.0.0") -target_databases:queryを実行するデータベースのリスト、または shell-like pattern
自動検出を有効にします。提供されている場合、デフォルトのデータベースを上書きします。
- metrics
:エクスポートされたすべての列のリストを含むセクション。次のように定義されます。
- <ColumnName> :クエリによって返される列の名前usage
がMAPPEDMETRIC に設定されている場合
usage の可能な値は次のとおりです。
DISCARD |
この列は無視する必要があります |
LABEL |
この列をラベルとして使用する |
COUNTER |
この列をカウンターとして使用する |
GAUGE |
この列をゲージとして使用する |
MAPPEDMETRIC |
この列を指定されたテキスト値のマッピングで使用します |
DURATION |
この列をテキスト期間として使用します(ミリ秒単位) |
HISTOGRAM |
この列をヒストグラムとして使用する |
Metric Types をご覧ください |
詳細については、Prometheusのドキュメントを参照してください。
ユーザー定義メトリックの出力¶
カスタム定義のメトリックは、次の形式でPrometheusエクスポーターエンドポイント(:9187/metrics
)によって返されます。
cnpg_<MetricName>_<ColumnName>{<LabelColumnName>=<LabelColumnValue> ... } <ColumnValue>
注釈
LabelColumnName は、 usage が`LABEL` に設定されたメトリックとその`Value` です
上記のpg_replication
の例を考慮すると、エクスポーターのエンドポイントは、呼び出されると次の出力を返します。
# HELP cnpg_pg_replication_in_recovery Whether the instance is in recovery
# TYPE cnpg_pg_replication_in_recovery gauge
cnpg_pg_replication_in_recovery 0
# HELP cnpg_pg_replication_lag Replication lag behind primary in seconds
# TYPE cnpg_pg_replication_lag gauge
cnpg_pg_replication_lag 0
# HELP cnpg_pg_replication_streaming_replicas Number of streaming replicas connected to the instance
# TYPE cnpg_pg_replication_streaming_replicas gauge
cnpg_pg_replication_streaming_replicas 2
# HELP cnpg_pg_replication_is_wal_receiver_up Whether the instance wal_receiver is up
# TYPE cnpg_pg_replication_is_wal_receiver_up gauge
cnpg_pg_replication_is_wal_receiver_up 0
メトリックのデフォルトセット¶
- オペレーターは、オペレーターの名前空間内のConfigMapまたはSecretで定義された一連のモニタリングクエリをクラスターに自動的に注入するように構成できます。
オペレーター設定 の
MONITORING_QUERIES_CONFIGMAP
またはMONITORING_QUERIES_SECRET
キーを、それぞれConfigMapまたはSecretの名前に設定する必要があります。オペレーターは
queries キーの内容を使用します。
queries
コンテンツへの変更は、それを使用するデプロイされたすべてのクラスターにすぐに反映されます。
オペレーターインストールマニフェストには、すべてのクラスターで使用されるcnpg-default-monitoring
と呼ばれる事前定義されたConfigMapが付属しています。
MONITORING_QUERIES_CONFIGMAP
は、オペレーター構成でデフォルトでcnpg-default-monitoring
に設定されます。
デフォルトのメトリックセットを無効にする場合は、次のことができます。
オペレーターレベルで無効にします。オペレーターConfigMapで、
MONITORING_QUERIES_CONFIGMAP/MONITORING_QUERIES_SECRETキーを""(空の文字列)に設定します。オペレーターConfigMapを変更するには、オペレーターの再起動が必要です。特定のクラスターで無効にします。クラスターで
.spec.monitoring.disableDefaultQueriesをtrueに設定します。
重要
MONITORING_QUERIES_CONFIGMAP /MONITORING_QUERIES_SECRET を介して指定されたConfigMapまたはSecretは、常に固定名前`cnpg-default-monitoring` でクラスターの名前空間にコピーされます。そのため、デフォルトのメトリックを使用する場合は、クラスターの名前空間にこの名前のConfigMapを作成しないでください。
Prometheus Postgresエクスポーターとの違い¶
CloudNativePGはPostgreSQL Prometheus
Exporterに触発されていますが、いくつかの違いがあります。特に、
cache_seconds
フィールドはCloudNativePGのエクスポーターに実装されていません。
オペレーターの監視¶
オペレーターは、metrics
という名前のポート8080でHTTPを介して Prometheusオペレーターの例 メトリックを内部的に公開します。
注釈
エクスポートされたメトリックを検査する方法 の指示に従って、エクスポートされたメトリックを検査できます
以下のセクション。
現在、オペレーターはデフォルトのkubebuilder
メトリックを公開しています。詳細については、
kubebuilder documentation を参照してください。
Prometheusオペレーターの例¶
- オペレーターの展開は、次の PodMonitor を定義することにより、
Prometheusオペレーターの例 を使用して監視できます
リソース:
apiVersion: monitoring.coreos.com/v1
kind: PodMonitor
metadata:
name: cnpg-controller-manager
spec:
selector:
matchLabels:
app.kubernetes.io/name: cloudnative-pg
podMetricsEndpoints:
- port: metrics
エクスポートされたメトリックを検査する方法¶
このセクションでは、同じ名前空間で curl
を実行している一時ポッドを使用して、特定のPostgreSQLインスタンスマネージャー(プライマリまたはレプリカ)またはオペレーターによってエクスポートされたメトリックを検査する方法に関する基本的な手順を提供します。
注釈
以下の例では、PostgreSQLクラスターと一緒にデフォルトの名前空間で作業していると仮定します。 Kubernetesの基本的な知識を適用して、この例を自分のユースケースに自由に適応させてください。
次の内容でcurl.yaml ファイルを作成します。
apiVersion: v1
kind: Pod
metadata:
name: curl
spec:
containers:
- name: curl
image: curlimages/curl:7.84.0
command: [sleep, 3600]
次に、ポッドを作成します。
kubectl apply -f curl.yaml
インスタンスによってエクスポートされたメトリックを検査する場合は、ターゲットポッドのポート9187に接続する必要があります。これは、実行する汎用コマンドです(ポッドに正しいIPを使用していることを確認してください)。
kubectl exec -ti curl -- curl -s <pod_ip>:9187/metrics
たとえば、PostgreSQLクラスターの名前がcluster-example
で、クラスター内の最初のポッドのエクスポートされたメトリックを取得する場合、次のコマンドを実行して、そのポッドのIPをプログラムで取得できます。
POD_IP=$(kubectl get pod cluster-example-1 --template {{.status.podIP}})
そして、実行します。
kubectl exec -ti curl -- curl -s ${POD_IP}:9187/metrics
オペレーターのメトリックにアクセスする場合は、オペレーターが実行されているポッドをポイントし、TCPポート8080をターゲットとして使用する必要があります。
検査の最後に、 curl podを削除してください。
kubectl delete -f curl.yaml
補助リソース¶
- ディレクトリには、可観測性のための一連のサンプルファイルがあります。
Part 4 of the quickstart を参照してください
コンテキストのセクション:
kube-stack-config.yaml:kube-stackヘルムチャートインストール用の構成ファイル。これにより、PrometheusがすべてのPodMonitorリソースをリッスンします。cnpg-prometheusrule.yaml:CloudNativePGのアラートを含むPrometheusRule。注:これには、通知サービスとの相互運用は含まれません。Prometheus documentation を参照してください。
grafana-configmap.yaml:サンプルCloudNativePGダッシュボードの定義を含むConfigMap。定義のラベルに注意してください。これにより、GrafanaデプロイメントがConfigMapを見つけることが保証されます。
さらに、ご参考までに、GrafanaダッシュボードとPrometheusアラートルールの「生の」ソースを提供します。
alerts.yaml:アラート付きのPrometheusルールgrafana-dashboard.json:ネイティブGrafana JSONとしてのCloudNativePGダッシュボード。
kube-prometheus-stack の構成では、kube-stack-config.yaml
で提供するものよりも他のフィールドと設定を使用できることに注意してください。
helm show values prometheus-community/kube-prometheus-stack
を実行して表示できます。詳細については、
kube-prometheus-stack を参照してください
ページ。