Operator Capability Levels

This section provides a summary of the capabilities implemented by CloudNativePG,classified using the Operator SDK definition of Capability Levels framework.

Operator Capability Levels

Operator Capability Levels

Important

Based on the Operator Capability Levels model , You can expect a “Level V - Auto Pilot” set of capabilities from the CloudNativePG Operator.

Each capability level is associated with a certain set of management features the operator offers:

  1. Basic Install2. Seamless Upgrades3. Full Lifecycle4. Deep Insights5. Auto Pilot

Note

We consider this framework as a guide for future work and implementations in the operator.

Level 1 - Basic Install

Capability level 1 involves installation and configuration of theoperator. This category includes usability and user experienceenhancements, such as improvements in how you interact with theoperator and a PostgreSQL cluster configuration.

Important

We consider Information Security part of this level.

Operator deployment via declarative configuration

The operator is installed in a declarative way using a Kubernetes manifestwhich defines 4 major CustomResourceDefinition objects: Cluster , Pooler , Backup , and ScheduledBackup .

PostgreSQL cluster deployment via declarative configuration

A PostgreSQL cluster (operand) is defined using the Cluster custom resourcein a fully declarative way. The PostgreSQL version is determined by theoperand container image defined in the CR, which is automatically fetchedfrom the requested registry. When deploying an operand, the operator alsoautomatically creates the following resources: Pod , Service , Secret , ConfigMap , PersistentVolumeClaim , PodDisruptionBudget , ServiceAccount , RoleBinding , Role .

Override of operand images through the CRD

The operator is designed to support any operand container image withPostgreSQL inside.By default, the operator uses the latest available minorversion of the latest stable major version supported by the PostgreSQLCommunity and published on ghcr.io.You can use any compatible image of PostgreSQL supporting theprimary/standby architecture directly by setting the imageName attribute in the CR. The operator also supports imagePullSecrets to access private container registries, as well as digests in addition totags for finer control of container image immutability.

Labels and annotations

The operator can be configured to support inheritance of labels and annotationsthat are defined in a cluster’s metadata, with the goal to improve organizationsof CloudNativePG deployment in your Kubernetes infrastructure.

Self-contained instance manager

Instead of relying on an external tool such as Patroni or Stolon tocoordinate PostgreSQL instances in the Kubernetes cluster pods, the operatorinjects the operator executable inside each pod, in a file named /controller/manager . The application is used to control the underlyingPostgreSQL instance and to reconcile the pod status with the instance itselfbased on the PostgreSQL cluster topology. The instance manager also starts aweb server that is invoked by the kubelet for probes. Unix signals invokedby the kubelet are filtered by the instance manager and, where appropriate,forwarded to the postgres process for fast and controlled reactions toexternal events. The instance manager is written in Go and has no externaldependencies.

Storage configuration

Storage is a critical component in a database workload. Taking advantage ofKubernetes native capabilities and resources in terms of storage, theoperator gives you enough flexibility to choose the right storage for yourworkload requirements, based on what the underlying Kubernetes environmentcan offer. This implies choosing a particular storage class ina public cloud environment or fine-tuning the generated PVC through aPVC template in the CR’s storage parameter.The cnp-bench open sourceproject can be used to benchmark both the storage and the database prior toproduction.

Replica configuration

The operator automatically detects replicas in a clusterthrough a single parameter called instances . If set to 1 , the clustercomprises a single primary PostgreSQL instance with no replica. If higherthan 1 , the operator manages instances -1 replicas, including highavailability through automated failover and rolling updates throughswitchover operations.

Database configuration

The operator is designed to manage a PostgreSQL cluster with a singledatabase. The operator transparently manages access to the database throughthree Kubernetes services automatically provisioned and managed for read-write,read, and read-only workloads.Using the convention over configuration approach, the operator creates adatabase called app , by default owned by a regular Postgres user with thesame name. Both the database name and the user name can be specified ifrequired.Although no configuration is required to run the cluster, you can customizeboth PostgreSQL run-time configuration and PostgreSQL Host-BasedAuthentication rules in the postgresql section of the CR.

Pod Security Policies

For InfoSec requirements, the operator does not require privileged mode forany container and enforces read only root filesystem to guarantee containersimmutability for both the operator and the operand pods. It also explicitlysets the required security contexts.

Affinity

The cluster’s affinity section enables fine-tuning of how pods and relatedresources such as persistent volumes are scheduled across the nodes of aKubernetes cluster. In particular, the operator supports:

  • pod affinity and anti-affinity

  • node selector

  • taints and tolerations

Command line interface

CloudNativePG does not have its own command line interface.It simply relies on the best command line interface for Kubernetes, kubectl ,by providing a plugin called cnpg to enhance and simplify your PostgreSQLcluster management experience.

Current status of the cluster

The operator continuously updates the status section of the CR with theobserved status of the cluster. The entire PostgreSQL cluster status iscontinuously monitored by the instance manager running in each pod: theinstance manager is responsible for applying the required changes to thecontrolled PostgreSQL instance to converge to the required status ofthe cluster (for example: if the cluster status reports that pod -1 is theprimary, pod -1 needs to promote itself while the other pods need to followpod -1 ). The same status is used by the cnpg plugin for kubectl to providedetails.

Operator’s certification authority

The operator automatically creates a certification authority for itself.It creates and signs with the operator certification authority a leaf certificateto be used by the webhook server, to ensure safe communication between theKubernetes API Server and the operator itself.

Cluster’s certification authority

The operator automatically creates a certification authority for every PostgreSQLcluster, which is used to issue and renew TLS certificates for clients’ authentication,including streaming replication standby servers (instead of passwords).Support for a custom certification authority for client certificates isavailable through secrets: this also includes integration with cert-manager.Certificates can be issued with the cnpg plugin for kubectl .

TLS connections

The operator transparently and natively supports TLS/SSL connectionsto encrypt client/server communications for increased security using thecluster’s certification authority.Support for custom server certificates is available through secrets: this alsoincludes integration with cert-manager.

Certificate authentication for streaming replication

The operator relies on TLS client certificate authentication to authorize streamingreplication connections from the standby servers, instead of relying on a password(and therefore a secret).

Continuous configuration management

The operator enables you to apply changes to the Cluster resource YAMLsection of the PostgreSQL configuration and makes sure that all instancesare properly reloaded or restarted, depending on the configuration option.Current limitation: changes with ALTER SYSTEM are not detected, meaningthat the cluster state is not enforced.

Basic LDAP authentication for PostgreSQL

The operator allows you to configure LDAP authentication for your PostgreSQLclients, using either the simple bind or search+bind mode, as described inthe PostgreSQL documentation: LDAP authentication .

Multiple installation methods

The operator can be installed through a Kubernetes manifest via kubectlapply , to be used in a traditional Kubernetes installation in publicand private cloud environments. Additionally, a Helm Chart for the operator isalso available.

Convention over configuration

The operator supports the convention over configuration paradigm, decidingstandard default values while allowing you to override them and customizethem. You can specify a deployment of a PostgreSQL cluster usingthe Cluster CRD in a couple of YAML code lines.

Level 2 - Seamless Upgrades

Capability level 2 is about enabling updatesoftheoperatorandtheactualworkload , in our case PostgreSQL servers. This includes PostgreSQLminorreleaseupdates (security and bug fixes normally) as well as majoronlineupgrades .

Upgrade of the operator

You can upgrade the operator seamlessly as a new deployment. A change in theoperator does not require a change in the operand - thanks to the instancemanager’s injection. The operator can manage older versions of the operand.

CloudNativePG also supports In-place updates of the instance manager following an upgrade of the operator: in-place updates do not require a rollingupdate - and subsequent switchover - of the cluster.

Upgrade of the managed workload

The operand can be upgraded using a declarative configuration approach aspart of changing the CR and, in particular, the imageName parameter. Theoperator prevents major upgrades of PostgreSQL while making it possible to goin both directions in terms of minor PostgreSQL releases within a majorversion (enabling updates and rollbacks).

In the presence of standby servers, the operator performs rolling updatesstarting from the replicas by dropping the existing pod and creating a newone with the new requested operand image that reuses the underlying storage.Depending on the value of the primaryUpdateStrategy , the operator proceedswith a switchover before updating the former primary ( unsupervised ) or waitsfor the user to manually issue the switchover procedure ( supervised ) via the cnpg plugin for kubectl .Which setting to use depends on the business requirements as the operationmight generate some downtime for the applications, from a few seconds tominutes based on the actual database workload.

Display cluster availability status during upgrade

At any time, convey the cluster’s high availability status, for example, Setting up primary , Creating a new replica , Cluster in healthy state , Switchover in progress , Failing over , Upgrading cluster , etc.

Level 3 - Full Lifecycle

Capability level 3 requires the operator to manage aspects of businesscontinuity and scalability . Disasterrecovery is a business continuity component that requiresthat both backup and recovery of a database work correctly. While as astarting point, the goal is to achieve RPO < 5 minutes, the long term goal isto implement RPO=0 backup solutions. HighAvailability is the otherimportant component of business continuity that, through PostgreSQL nativephysical replication and hot standby replicas, allows the operator to performfailover and switchover operations. This area includes enhancements in:

PostgreSQL Backups

The operator has been designed to provide application-level backups usingPostgreSQL’s native continuous backup technology based onphysical base backups and continuous WAL archiving. Specifically,the operator currently supports only backups on object stores (AWS S3 andS3-compatible, Azure Blob Storage, Google Cloud Storage, and gateways likeMinIO).

WAL archiving and base backups are defined at the cluster level, declaratively,through the backup parameter in the cluster definition, by specifyingan S3 protocol destination URL (for example, to point to a specific folder inan AWS S3 bucket) and, optionally, a generic endpoint URL. WAL archiving,a prerequisite for continuous backup, does not require any furtheraction from the user: the operator will automatically and transparently setthe archive_command to rely on barman-cloud-wal-archive to ship WALfiles to the defined endpoint. Users can decide the compression algorithm,as well as the number of parallel jobs to concurrently upload WAL filesin the archive. In addition to that Instance Manager automatically checks the correctness of the archive destination, by performing barman-cloud-check-wal-archive command before beginning to ship the very first set of WAL files.

You can define base backups in two ways: on-demand (through the Backup custom resource definition) or scheduled (through the ScheduledBackup customer resource definition, using a cron-like syntax). They both rely on barman-cloud-backup for the job (distributed as part of the applicationcontainer image) to relay backups in the same endpoint, alongside WAL files.

Both barman-cloud-wal-restore and barman-cloud-backup are distributed inthe application container image under GNU GPL 3 terms.

Full restore from a backup

The operator enables you to bootstrap a new cluster (with its settings)starting from an existing and accessible backup taken using barman-cloud-backup . Once the bootstrap process is completed, the operatorinitiates the instance in recovery mode and replays all available WAL filesfrom the specified archive, exiting recovery and starting as a primary.Subsequently, the operator will clone the requested number of standby instancesfrom the primary.CloudNativePG supports parallel WAL fetching from the archive.

Point-In-Time Recovery (PITR) from a backup

The operator enables you to create a new PostgreSQL cluster by recoveringan existing backup to a specific point-in-time, defined with a timestamp, alabel or a transaction ID. This capability is built on top of the full restoreone and supports all the options available in

Zero Data Loss clusters through synchronous replication

Achieve Zero Data Loss (RPO=0) in your local High Availability CloudNativePGcluster through quorum based synchronous replication support. The operator providestwo configuration options that control the minimum and maximum number ofexpected synchronous standby replicas available at any time. The operator willreact accordingly, based on the number of available and ready PostgreSQLinstances in the cluster, through the following formula for the quorum ( q ):

1 <= minSyncReplicas <= q <= maxSyncReplicas <= readyReplicas

Replica clusters

Define a cross Kubernetes cluster topology of PostgreSQL clusters, by takingadvantage of PostgreSQL native streaming and cascading replication.Through the replica option, you can setup an independent cluster to becontinuously replicating data from another PostgreSQL source of the same majorversion: such a source can be anywhere, as long as a direct streamingconnection via TLS is allowed from the two endpoints.Moreover, the source can be even outside Kubernetes, running in a physical orvirtual environment.Replica clusters can be created from a recovery object store (backup in BarmanCloud format) or via streaming through pg_basebackup . Both WAL file shippingand WAL streaming are allowed.Replica clusters dramatically improve the business continuity posture of yourPostgreSQL databases in Kubernetes, spanning over multiple datacenters andopening up for hybrid and multi-cloud setups (currently, manual switchoveracross data centers is required, while waiting for Kubernetes federationnative capabilities).

Liveness and readiness probes

The operator defines liveness and readiness probes for the PostgresContainers that are then invoked by the kubelet. They are mapped respectivelyto the /healthz and /readyz endpoints of the web server manageddirectly by the instance manager.The liveness probe is based on the pg_isready executable, and the pod isconsidered healthy with exit codes 0 (server accepting connections normally)and 1 (server is rejecting connections, for example during startup). Thereadiness probe issues a simple query ( ; ) to verify that the server isready to accept connections.

Rolling deployments

The operator supports rolling deployments to minimize the downtime and, if aPostgreSQL cluster is exposed publicly, the Service will load-balance theread-only traffic only to available pods during the initialization or theupdate.

Scale up and down of replicas

The operator allows you to scale up and down the number of instances in aPostgreSQL cluster. New replicas are automatically started up from theprimary server and will participate in the cluster’s HA infrastructure.The CRD declares a “scale” subresource that allows the user to use the kubectl scale command.

Maintenance window and PodDisruptionBudget for Kubernetes nodes

The operator creates a PodDisruptionBudget resource to limit the number ofconcurrent disruptions to one primary instance. This configuration prevents themaintenance operation from deleting all the pods in a cluster, allowing thespecified number of instances to be created. The PodDisruptionBudget will beapplied during the node draining operation, preventing any disruption of thecluster service.

While this strategy is correct for Kubernetes Clusters wherestorage is shared among all the worker nodes, it may not be the best solutionfor clusters using Local Storage or for clusters installed in a privatecloud. The operator allows you to specify a Maintenance Window andconfigure the reaction to any underlying node eviction. The ReusePVC optionin the maintenance window section enables to specify the strategy to be used:allocate new storage in a different PVC for the evicted instance or waitfor the underlying node to be available again.

Fencing

Fencing is the process of protecting the data in one, more, or even allinstances of a PostgreSQL cluster when they appear to be malfunctioning.When an instance is fenced, the PostgreSQL server process isguaranteed to be shut down, while the pod is kept running. This makes surethat, until the fence is lifted, data on the pod is not modified by PostgreSQLand that the file system can be investigated for debugging and troubleshootingpurposes.

Reuse of Persistent Volumes storage in Pods

When the operator needs to create a pod that has been deleted by the user orhas been evicted by a Kubernetes maintenance operation, it reuses the PersistentVolumeClaim if available, avoiding the needto re-clone the data from the primary.

CPU and memory requests and limits

The operator allows administrators to control and manage resource usage bythe cluster’s pods, through the resources section of the manifest. Inparticular requests and limits values can be set for both CPU and RAM.

Connection pooling with PgBouncer

CloudNativePG provides native support for connection pooling with

PgBouncerIntegrationStatus , one of the most popular open sourceconnection poolers

for PostgreSQL. From an architectural point of view, thenative implementation of a PgBouncer connection pooler introduces a new layerto access the database which optimizes the query flow towards the instancesand makes the usage of the underlying PostgreSQL resources more efficient.Instead of connecting directly to a PostgreSQL service, applications can nowconnect to the PgBouncer service and start reusing any existing connection.

Level 4 - Deep Insights

Capability level 4 is about observability : in particular, monitoring,alerting, trending, log processing. This might involve the use of external toolssuch as Prometheus, Grafana, Fluent Bit, as well as extensions in thePostgreSQL engine for the output of error logs directly in JSON format.

CloudNativePG has been designed to provide everything that is neededto easily integrate with industry-standard and community accepted tools forflexible monitoring and logging.

Prometheus exporter with configurable queries

The instance manager provides a pluggable framework and, via its own web serverlistening on the metrics port (9187), exposes an endpoint to export metricsfor the Prometheus Operator example monitoring and alerting tool.The operator supports custom monitoring queries defined as ConfigMap and/or Secret objects using a syntax that is compatible with the

postgres_exporter .CloudNativePG provides a set of basic monitoring queries

forPostgreSQL that can be integrated and adapted to your context.The [cnp-sandbox project] is an open source Helm chart that demonstrateshow to integrate CloudNativePG with Prometheus and Grafana, by providingsome basic metrics and an example of dashboard.

Standard output logging of PostgreSQL error messages in JSON format

Every log message is delivered to standard output in JSON format, with the first leveldefinition of the timestamp, the log level and the type of log entry, such as postgres for the canonical PostgreSQL error message channel.As a result, every Pod managed by CloudNativePG can be easily and directlyintegrated with any downstream log processing stack that supports JSON as sourcedata type.

Real-time query monitoring

CloudNativePG transparently and natively supports:

Audit

CloudNativePG allows database and security administrators, auditors,and operators to track and analyze database activities using PGAudit (forPostgreSQL).Such activities flow directly in the JSON log and can be properly routed to thecorrect downstream target using common log brokers like Fluentd.

Kubernetes events

Record major events as expected by the Kubernetes API, such as creating resources,removing nodes, upgrading, and so on. Events can be displayed throughthe kubectl describe and kubectl get events command.

Level 5 - Auto Pilot

Capability level 5 is focused on automatedscaling , healing and tuning - through the discovery of anomalies and insights that emergedfrom the observability layer.

Automated Failover for self-healing

In case of detected failure on the primary, the operator will change thestatus of the cluster by setting the most aligned replica as the new targetprimary. As a consequence, the instance manager in each alive pod willinitiate the required procedures to align itself with the requested status ofthe cluster, by either becoming the new primary or by following it.In case the former primary comes back up, the same mechanism will avoid asplit-brain by preventing applications from reaching it, running pg_rewind onthe server and restarting it as a standby.

Automated recreation of a standby

In case the pod hosting a standby has been removed, the operator initiatesthe procedure to recreate a standby server.