モニタリング

インスタンスのモニタリング

オペレーターは、PostgreSQLインスタンスごとに、ポート9187でmetrics というHTTPを介して Prometheusオペレーターの例 のメトリックのエクスポーターを提供します。オペレーターには

事前定義されたメトリックセット と、1つ以上の ConfigMap または Secret

リソースを介して追加のクエリを定義するための高度に構成およびカスタマイズ可能なシステムが付属しています(詳細については、以下の ユーザー定義のメトリック を参照)。

重要

バージョン1.11以降、CloudNativePGは既に`default-monitoring` と呼ばれる`ConfigMap` に by default a set of predefined metrics をインストールしています。

メトリックには次のようにアクセスできます。

curl http://<pod_ip>:9187/metrics

PostgreSQLで実行されるすべての監視クエリは次のとおりです。

  • トランザクション的にアトミック(クエリごとに1つのトランザクション)

  • pg_monitor ロールで実行

  • application_name をcnpg_metrics_exporter に設定して実行

  • ユーザーpostgres として実行

pg_monitor ロールの詳細については、PostgreSQL

documentation の「デフォルトのロール」セクションを参照してください。

デフォルトでは、クエリは次のロジックに従って、 Cluster リソースの指定されたbootstrap メソッドで定義されたメインデータベースに対して実行されます。

  • initdb を使用:クエリはinitdb.database 、または指定されていない場合はapp に対してデフォルトで実行されます

  • recovery を使用:クエリはrecovery.database 、または指定されていない場合はpostgres に対してデフォルトで実行されます

  • pg_basebackup を使用:クエリはpg_basebackup.database 、または指定されていない場合はpostgres に対してデフォルトで実行されます

target_databases オプションで1つ以上のデータベースのリストを指定することにより、特定のユーザー定義メトリックのデフォルトのデータベースを常にオーバーライドできます。

Prometheusオペレーターの例

を使用して、特定のPostgreSQLクラスターを監視できます。クラスターを正しく指すPodMonitorは、クラスターリソース自分自身で.spec.monitoring.enablePodMonitor をtrue に設定することにより、オペレーターが自動的に作成できます(デフォルト:false)。

重要

自動的に作成された`PodMonitor` への変更は、次の調整サイクルでオペレーターによってオーバーライドされます。カスタマイズが必要な場合は、以下で説明します。

特定のクラスターのPodMonitor を手動でデプロイするには、次のように定義し、必要に応じて変更します。

重要

一意の名前と、正しいクラスターの名前空間とラベルを使用して、上の例を変更してください(cluster-example を使用しています)。

事前定義されたメトリックセット

すべてのPostgreSQLインスタンスエクスポーターは、事前定義された一連のメトリックを自動的に公開します。これは、2つの主要なカテゴリに分類できます。

  • cnpg_collector_* で始まるPostgreSQL関連のメトリック。

  • ディスク上のWALファイルの数と合計サイズ-アーカイブステータスフォルダー内の.ready および.done ファイルの数-要求された同期レプリカの数、予想および実際に観測された値-レプリカクラスターモードが有効かどうかを示すフラグまたはdisabled - 手動スイッチオーバーが必要かどうかを示すフラグ

  • go_* で始まる、ランタイム関連のメトリック

以下は、インスタンスの localhost:9187/metrics エンドポイントによって返されるメトリックのサンプルです。ご覧のとおり、Prometheusフォーマットは自己文書化です。

#  HELP cnpg_collector_collection_duration_seconds Collection time duration in seconds
#  TYPE cnpg_collector_collection_duration_seconds gauge
cnpg_collector_collection_duration_seconds{collector="Collect.up"} 0.0031393

#  HELP cnpg_collector_collections_total Total number of times PostgreSQL was accessed for metrics.
#  TYPE cnpg_collector_collections_total counter
cnpg_collector_collections_total 2

#  HELP cnpg_collector_last_collection_error 1 if the last collection ended with error, 0 otherwise.
#  TYPE cnpg_collector_last_collection_error gauge
cnpg_collector_last_collection_error 0

#  HELP cnpg_collector_manual_switchover_required 1 if a manual switchover is required, 0 otherwise
#  TYPE cnpg_collector_manual_switchover_required gauge
cnpg_collector_manual_switchover_required 0

#  HELP cnpg_collector_pg_wal Total size in bytes of WAL segments in the /var/lib/postgresql/data/pgdata/pg_wal directory  computed as (wal_segment_size * count)
#  TYPE cnpg_collector_pg_wal gauge
cnpg_collector_pg_wal{value="count"} 7
cnpg_collector_pg_wal{value="size"} 1.17440512e+08

#  HELP cnpg_collector_pg_wal_archive_status Number of WAL segments in the /var/lib/postgresql/data/pgdata/pg_wal/archive_status directory (ready, done)
#  TYPE cnpg_collector_pg_wal_archive_status gauge
cnpg_collector_pg_wal_archive_status{value="done"} 6
cnpg_collector_pg_wal_archive_status{value="ready"} 0

#  HELP cnpg_collector_replica_mode 1 if the cluster is in replica mode, 0 otherwise
#  TYPE cnpg_collector_replica_mode gauge
cnpg_collector_replica_mode 0

#  HELP cnpg_collector_sync_replicas Number of requested synchronous replicas (synchronous_standby_names)
#  TYPE cnpg_collector_sync_replicas gauge
cnpg_collector_sync_replicas{value="expected"} 0
cnpg_collector_sync_replicas{value="max"} 0
cnpg_collector_sync_replicas{value="min"} 0
cnpg_collector_sync_replicas{value="observed"} 0

#  HELP cnpg_collector_up 1 if PostgreSQL is up, 0 otherwise.
#  TYPE cnpg_collector_up gauge
cnpg_collector_up{cluster="cluster-example"} 1

#  HELP cnpg_collector_postgres_version Postgres version
#  TYPE cnpg_collector_postgres_version gauge
cnpg_collector_postgres_version{cluster="cluster-example",full="13.4.0"} 13.4

#  HELP cnpg_collector_first_recoverability_point The first point of recoverability for the cluster as a unix timestamp
#  TYPE cnpg_collector_first_recoverability_point gauge
cnpg_collector_first_recoverability_point 1.63238406e+09

#  HELP cnpg_collector_lo_pages Estimated number of pages in the pg_largeobject table
#  TYPE cnpg_collector_lo_pages gauge
cnpg_collector_lo_pages{datname="app"} 0
cnpg_collector_lo_pages{datname="postgres"} 78

#  HELP go_gc_duration_seconds A summary of the pause duration of garbage collection cycles.
#  TYPE go_gc_duration_seconds summary
go_gc_duration_seconds{quantile="0"} 5.01e-05
go_gc_duration_seconds{quantile="0.25"} 7.27e-05
go_gc_duration_seconds{quantile="0.5"} 0.0001748
go_gc_duration_seconds{quantile="0.75"} 0.0002959
go_gc_duration_seconds{quantile="1"} 0.0012776
go_gc_duration_seconds_sum 0.0035741
go_gc_duration_seconds_count 13

#  HELP go_goroutines Number of goroutines that currently exist.
#  TYPE go_goroutines gauge
go_goroutines 25

#  HELP go_info Information about the Go environment.
#  TYPE go_info gauge
go_info{version="go1.17.1"} 1

#  HELP go_memstats_alloc_bytes Number of bytes allocated and still in use.
#  TYPE go_memstats_alloc_bytes gauge
go_memstats_alloc_bytes 4.493744e+06

#  HELP go_memstats_alloc_bytes_total Total number of bytes allocated, even if freed.
#  TYPE go_memstats_alloc_bytes_total counter
go_memstats_alloc_bytes_total 2.1698216e+07

#  HELP go_memstats_buck_hash_sys_bytes Number of bytes used by the profiling bucket hash table.
#  TYPE go_memstats_buck_hash_sys_bytes gauge
go_memstats_buck_hash_sys_bytes 1.456234e+06

#  HELP go_memstats_frees_total Total number of frees.
#  TYPE go_memstats_frees_total counter
go_memstats_frees_total 172118

#  HELP go_memstats_gc_cpu_fraction The fraction of this programs available CPU time used by the GC since the program started.
#  TYPE go_memstats_gc_cpu_fraction gauge
go_memstats_gc_cpu_fraction 1.0749468700447189e-05

#  HELP go_memstats_gc_sys_bytes Number of bytes used for garbage collection system metadata.
#  TYPE go_memstats_gc_sys_bytes gauge
go_memstats_gc_sys_bytes 5.530048e+06

#  HELP go_memstats_heap_alloc_bytes Number of heap bytes allocated and still in use.
#  TYPE go_memstats_heap_alloc_bytes gauge
go_memstats_heap_alloc_bytes 4.493744e+06

#  HELP go_memstats_heap_idle_bytes Number of heap bytes waiting to be used.
#  TYPE go_memstats_heap_idle_bytes gauge
go_memstats_heap_idle_bytes 5.8236928e+07

#  HELP go_memstats_heap_inuse_bytes Number of heap bytes that are in use.
#  TYPE go_memstats_heap_inuse_bytes gauge
go_memstats_heap_inuse_bytes 7.528448e+06

#  HELP go_memstats_heap_objects Number of allocated objects.
#  TYPE go_memstats_heap_objects gauge
go_memstats_heap_objects 26306

#  HELP go_memstats_heap_released_bytes Number of heap bytes released to OS.
#  TYPE go_memstats_heap_released_bytes gauge
go_memstats_heap_released_bytes 5.7401344e+07

#  HELP go_memstats_heap_sys_bytes Number of heap bytes obtained from system.
#  TYPE go_memstats_heap_sys_bytes gauge
go_memstats_heap_sys_bytes 6.5765376e+07

#  HELP go_memstats_last_gc_time_seconds Number of seconds since 1970 of last garbage collection.
#  TYPE go_memstats_last_gc_time_seconds gauge
go_memstats_last_gc_time_seconds 1.6311727586032727e+09

#  HELP go_memstats_lookups_total Total number of pointer lookups.
#  TYPE go_memstats_lookups_total counter
go_memstats_lookups_total 0

#  HELP go_memstats_mallocs_total Total number of mallocs.
#  TYPE go_memstats_mallocs_total counter
go_memstats_mallocs_total 198424

#  HELP go_memstats_mcache_inuse_bytes Number of bytes in use by mcache structures.
#  TYPE go_memstats_mcache_inuse_bytes gauge
go_memstats_mcache_inuse_bytes 14400

#  HELP go_memstats_mcache_sys_bytes Number of bytes used for mcache structures obtained from system.
#  TYPE go_memstats_mcache_sys_bytes gauge
go_memstats_mcache_sys_bytes 16384

#  HELP go_memstats_mspan_inuse_bytes Number of bytes in use by mspan structures.
#  TYPE go_memstats_mspan_inuse_bytes gauge
go_memstats_mspan_inuse_bytes 191896

#  HELP go_memstats_mspan_sys_bytes Number of bytes used for mspan structures obtained from system.
#  TYPE go_memstats_mspan_sys_bytes gauge
go_memstats_mspan_sys_bytes 212992

#  HELP go_memstats_next_gc_bytes Number of heap bytes when next garbage collection will take place.
#  TYPE go_memstats_next_gc_bytes gauge
go_memstats_next_gc_bytes 8.689632e+06

#  HELP go_memstats_other_sys_bytes Number of bytes used for other system allocations.
#  TYPE go_memstats_other_sys_bytes gauge
go_memstats_other_sys_bytes 2.566622e+06

#  HELP go_memstats_stack_inuse_bytes Number of bytes in use by the stack allocator.
#  TYPE go_memstats_stack_inuse_bytes gauge
go_memstats_stack_inuse_bytes 1.343488e+06

#  HELP go_memstats_stack_sys_bytes Number of bytes obtained from system for stack allocator.
#  TYPE go_memstats_stack_sys_bytes gauge
go_memstats_stack_sys_bytes 1.343488e+06

#  HELP go_memstats_sys_bytes Number of bytes obtained from system.
#  TYPE go_memstats_sys_bytes gauge
go_memstats_sys_bytes 7.6891144e+07

#  HELP go_threads Number of OS threads created.
#  TYPE go_threads gauge
go_threads 18

注釈

cnpg_collector_postgres_version は、PostgreSQLの`Major.Minor` バージョンを含むGaugeVecメトリックです。完全なセマンティックバージョン`Major.Minor.Patch` は、 full という名前のラベルフィールドの中にあります。

ユーザー定義のメトリック

この機能は現在ベータ状態であり、形式はPostgreSQL Prometheusエクスポーターの queries.yaml file に触発されています。

カスタムメトリックは、次の例のように、 .spec.monitoring.customQueriesConfigMap またはcustomQueriesSecret セクションの下のCluster 定義で作成されたConfigmap /Secret を参照することにより、定義できます。

apiVersion: postgresql.cnpg.io/v1
kind: Cluster
metadata:
  name: cluster-example
  namespace: test
spec:
  instances: 3

  storage:
    size: 1Gi

  monitoring:
    customQueriesConfigMap:
      - name: example-monitoring
        key: custom-queries

customQueriesConfigMap /customQueriesSecret セクションには、カスタムクエリが定義されているキーを指定するConfigMap /Secret 参照のリストが含まれています。参照されるリソースは、クラスターリソースと同じ名前空間に作成する必要があることに注意してください。

注釈

ConfigMapとシークレットをインスタンスによって**自動的に**リロードしたい場合は、キー cnpg.io/reload でラベルを追加できます。それ以外の場合は、 kubectl cnpg reload サブコマンドを使用してインスタンスをリロードする必要があります。

重要

ユーザー定義のメトリックが既存のメトリックを上書きすると、インスタンスマネージャーはメッセージを含むjson警告ログを出力します:Query with the same name already found. Overwriting the existing one. と、上書きされたクエリ名を含むキー。

ユーザー定義のメトリックの例

ここでは、上記のCluster の例で参照される、単一のカスタムクエリを含むConfigMap の例を見ることができます。

apiVersion: v1
kind: ConfigMap
metadata:
  name: example-monitoring
  namespace: test
  labels:
    cnpg.io/reload: ""
data:
  custom-queries: |
    pg_replication:
      query: "SELECT CASE WHEN NOT pg_is_in_recovery()
              THEN 0
              ELSE GREATEST (0,
                EXTRACT(EPOCH FROM (now() - pg_last_xact_replay_timestamp())))
              END AS lag,
              pg_is_in_recovery() AS in_recovery,
              EXISTS (TABLE pg_stat_wal_receiver) AS is_wal_receiver_up,
              (SELECT count(*) FROM pg_stat_replication) AS streaming_replicas"

      metrics:
        - lag:
            usage: "GAUGE"
            description: "Replication lag behind primary in seconds"
        - in_recovery:
            usage: "GAUGE"
            description: "Whether the instance is in recovery"
        - is_wal_receiver_up:
            usage: "GAUGE"
            description: "Whether the instance wal_receiver is up"
        - streaming_replicas:
            usage: "GAUGE"
            description: "Number of streaming replicas connected to the instance"

基本的な監視クエリのリストは、 `cnpg-basic-monitoring.yaml file <./samples/cnpg-basic-monitoring.yaml>`__.

複数のデータベースで実行されるユーザー定義のメトリックの例

target_databases オプションに複数のデータベースがリストされている場合、メトリックはそれぞれから収集されます。

target_databases のリストにシェルのようなパターン(つまり、 * 、? または[] を含む)を指定することにより、特定のクエリに対してデータベースの自動検出を有効にできます。提供されている場合、オペレーターは SELECT datname FROM pg_database WHERE datallowconn AND NOT datistemplate の実行によって返されたすべてのデータベースを追加し、

path.Match() ルールに従ってパターンをマッチングすることにより、ターゲットデータベースのリストを展開します。

注釈

* のキャラクターはyamlに special meaning があるため、そのようなパターンが含まれる場合は`target_databases` の値をクォート("*" )する必要があります。

たとえば、次の例のようにcurrent_database() ファンクションを使用して、返されるラベルにデータベースの名前を常に含めることをお勧めします。

some_query:
  query: |
    SELECT
     current_database() as datname,
     count(*) as rows
    FROM some_table
  metrics:
    - datname:
        usage: "LABEL"
        description: "Name of current database"
    - rows:
        usage: "GAUGE"
        description: "number of rows"
  target_databases:
    - albert
    - bb
    - freddie

これにより、次のメトリックが公開されます。

cnpg_some_query_rows{datname="albert"} 2
cnpg_some_query_rows{datname="bb"} 5
cnpg_some_query_rows{datname="freddie"} 10

これは、 template1 データベースでも実行される自動検出を有効にしたクエリの例です(それ以外の場合、前述のクエリでは返されません)。

some_query:
  query: |
    SELECT
     current_database() as datname,
     count(*) as rows
    FROM some_table
  metrics:
    - datname:
        usage: "LABEL"
        description: "Name of current database"
    - rows:
        usage: "GAUGE"
        description: "number of rows"
  target_databases:
    - "*"
    - "template1"

上記の例では、次のメトリックが生成されます(データベースが存在する場合)。

cnpg_some_query_rows{datname="albert"} 2
cnpg_some_query_rows{datname="bb"} 5
cnpg_some_query_rows{datname="freddie"} 10
cnpg_some_query_rows{datname="template1"} 7
cnpg_some_query_rows{datname="postgres"} 42

ユーザー定義メトリックの構造

すべてのカスタムクエリには、次の基本構造があります。

<MetricName>:
      query: "<SQLQuery>"
      metrics:
        - <ColumnName>:
            usage: "<MetricType>"
            description: "<MetricDescription>"

使用可能なすべてのフィールドの簡単な説明は次のとおりです。

  • <MetricName> :Prometheusメトリックの名前 - query :メトリックを生成するためにターゲットデータベースで実行するSQLクエリー - primary :プライマリインスタンスでのみクエリーを実行するかどうか - master :primary と同じ(Prometheus PostgreSQLエクスポーターの構文との互換性のため - 非推奨) - runonserver :クエリを実行するPostgreSQLのバージョンを制限するセマンティックバージョン範囲(例">=10.0.0" または">=12.0.0 <=14.4.0" )- target_databases :query を実行するデータベースのリスト、または自動検出を有効にする shell-like pattern 。提供されている場合、デフォルトのデータベースを上書きします。 - metrics :エクスポートされたすべての列のリストを含むセクション。次のように定義されます。 - <ColumnName> :クエリによって返される列の名前usage がMAPPEDMETRIC に設定されている場合

usage の可能な値は次のとおりです。

Column Usage Label

Description

DISCARD

この列は無視する必要があります

LABEL

この列をラベルとして使用する

COUNTER

この列をカウンターとして使用する

GAUGE

この列をゲージとして使用する

MAPPEDMETRIC

この列を指定されたテキスト値のマッピングで使用します

DURATION

この列をテキスト期間として使用します(ミリ秒単位)

HISTOGRAM

この列をヒストグラムとして使用する

詳細については、Prometheusのドキュメントの Metric Types にアクセスしてください。

ユーザー定義のメトリックの出力

カスタム定義のメトリックは、Prometheusエクスポーターエンドポイント(:9187/metrics )によって次の形式で返されます。

cnpg_<MetricName>_<ColumnName>{<LabelColumnName>=<LabelColumnValue> ... } <ColumnValue>

注釈

LabelColumnName は、 usage が`LABEL` に設定されたメトリックとその`Value`

上記のpg_replication の例を考慮すると、エクスポーターのエンドポイントは呼び出されると次の出力を返します。

#  HELP cnpg_pg_replication_in_recovery Whether the instance is in recovery
#  TYPE cnpg_pg_replication_in_recovery gauge
cnpg_pg_replication_in_recovery 0
#  HELP cnpg_pg_replication_lag Replication lag behind primary in seconds
#  TYPE cnpg_pg_replication_lag gauge
cnpg_pg_replication_lag 0
#  HELP cnpg_pg_replication_streaming_replicas Number of streaming replicas connected to the instance
#  TYPE cnpg_pg_replication_streaming_replicas gauge
cnpg_pg_replication_streaming_replicas 2
#  HELP cnpg_pg_replication_is_wal_receiver_up Whether the instance wal_receiver is up
#  TYPE cnpg_pg_replication_is_wal_receiver_up gauge
cnpg_pg_replication_is_wal_receiver_up 0

メトリックのデフォルトセット

オペレーターは、オペレーターの名前空間内のConfigMapまたはSecretで定義された一連のモニタリングクエリをクラスターに自動的に注入するように構成できます。

オペレーター設定 の MONITORING_QUERIES_CONFIGMAP または

MONITORING_QUERIES_SECRET キーを、それぞれConfigMapまたはSecretの名前に設定する必要があります。オペレーターは queries キーの内容を使用します。

queries コンテンツへの変更は、それを使用するデプロイされたすべてのクラスターにすぐに反映されます。

オペレーターインストールマニフェストには、すべてのクラスターで使用される cnpg-default-monitoring と呼ばれる事前定義された ConfigMap が付属しています。 MONITORING_QUERIES_CONFIGMAP は、デフォルトでオペレーター構成でcnpg-default-monitoring に設定されます。

デフォルトのメトリックセットを無効にする場合は、次のようにします。

  • オペレーターレベルで無効にします。オペレーターConfigMapでMONITORING_QUERIES_CONFIGMAP /MONITORING_QUERIES_SECRET キーを"" (空の文字列)に設定します。オペレーターの ConfigMap を変更するには、オペレーターの再起動が必要です。

  • 特定のクラスターで無効にします。クラスターで.spec.monitoring.disableDefaultQueries をtrue に設定します。

重要

MONITORING_QUERIES_CONFIGMAP /MONITORING_QUERIES_SECRET を介して指定されたConfigMapまたはSecretは、常に固定名`cnpg-default-monitoring` でクラスターの名前空間にコピーされます。そのため、デフォルトのメトリックを使用する場合は、クラスターの名前空間にこの名前で ConfigMap を作成しないでください。

Prometheus Postgresエクスポーターとの違い

CloudNativePGはPostgreSQL Prometheus Exporterに触発されていますが、いくつかの違いがあります。特に、 cache_seconds フィールドはCloudNativePGのエクスポーターに実装されていません。

オペレーターの監視

オペレーターは、 metrics という名前のポート8080でHTTPを介して Prometheusオペレーターの例 メトリックを内部的に公開します。

メトリックには次のようにアクセスできます。

curl http://<pod_ip>:8080/metrics

現在、オペレーターはデフォルトのkubebuilder メトリックを公開しています。詳細については、

kubebuilder documentation を参照してください。

Prometheusオペレーターの例

オペレーターのデプロイメントは、次の PodMonitor リソースを定義することにより、

Prometheusオペレーターの例 を使用して監視できます。

apiVersion: monitoring.coreos.com/v1
kind: PodMonitor
metadata:
  name: cnpg-controller-manager
spec:
  selector:
    matchLabels:
      app.kubernetes.io/name: cloudnative-pg
  podMetricsEndpoints:
    - port: metrics