Operator Capability Levels¶
This section provides a summary of the capabilities implemented by CloudNativePG,classified using the Operator SDK definition of Capability Levels framework.
Operator Capability Levels¶
Important
Based on the Operator Capability Levels model , You can expect a “Level V - Auto Pilot” set of capabilities from the CloudNativePG Operator.
Each capability level is associated with a certain set of management features the operator offers:
Basic Install2. Seamless Upgrades3. Full Lifecycle4. Deep Insights5. Auto Pilot
Note
We consider this framework as a guide for future work and implementations in the operator.
Level 1 - Basic Install¶
Capability level 1 involves installation and configuration of theoperator. This category includes usability and user experienceenhancements, such as improvements in how you interact with theoperator and a PostgreSQL cluster configuration.
Important
We consider Information Security part of this level.
Operator deployment via declarative configuration¶
The operator is installed in a declarative way using a Kubernetes
manifestwhich defines 4 major CustomResourceDefinition objects:
Cluster , Pooler , Backup , and ScheduledBackup .
PostgreSQL cluster deployment via declarative configuration¶
A PostgreSQL cluster (operand) is defined using the Cluster custom
resourcein a fully declarative way. The PostgreSQL version is determined
by theoperand container image defined in the CR, which is automatically
fetchedfrom the requested registry. When deploying an operand, the
operator alsoautomatically creates the following resources: Pod ,
Service , Secret , ConfigMap , PersistentVolumeClaim ,
PodDisruptionBudget , ServiceAccount , RoleBinding ,
Role .
Override of operand images through the CRD¶
The operator is designed to support any operand container image
withPostgreSQL inside.By default, the operator uses the latest available
minorversion of the latest stable major version supported by the
PostgreSQLCommunity and published on ghcr.io.You can use any compatible
image of PostgreSQL supporting theprimary/standby architecture directly
by setting the imageName attribute in the CR. The operator also
supports imagePullSecrets to access private container registries, as
well as digests in addition totags for finer control of container image
immutability.
Labels and annotations¶
The operator can be configured to support inheritance of labels and annotationsthat are defined in a cluster’s metadata, with the goal to improve organizationsof CloudNativePG deployment in your Kubernetes infrastructure.
Self-contained instance manager¶
Instead of relying on an external tool such as Patroni or Stolon
tocoordinate PostgreSQL instances in the Kubernetes cluster pods, the
operatorinjects the operator executable inside each pod, in a file named
/controller/manager . The application is used to control the
underlyingPostgreSQL instance and to reconcile the pod status with the
instance itselfbased on the PostgreSQL cluster topology. The instance
manager also starts aweb server that is invoked by the kubelet for
probes. Unix signals invokedby the kubelet are filtered by the
instance manager and, where appropriate,forwarded to the postgres
process for fast and controlled reactions toexternal events. The
instance manager is written in Go and has no externaldependencies.
Storage configuration¶
Storage is a critical component in a database workload. Taking advantage
ofKubernetes native capabilities and resources in terms of storage,
theoperator gives you enough flexibility to choose the right storage for
yourworkload requirements, based on what the underlying Kubernetes
environmentcan offer. This implies choosing a particular storage class
ina public cloud environment or fine-tuning the generated PVC through
aPVC template in the CR’s storage parameter.The cnp-bench open
sourceproject can be used to benchmark both the storage and the database
prior toproduction.
Replica configuration¶
The operator automatically detects replicas in a clusterthrough a single
parameter called instances . If set to 1 , the clustercomprises
a single primary PostgreSQL instance with no replica. If higherthan
1 , the operator manages instances -1 replicas, including
highavailability through automated failover and rolling updates
throughswitchover operations.
Database configuration¶
The operator is designed to manage a PostgreSQL cluster with a
singledatabase. The operator transparently manages access to the
database throughthree Kubernetes services automatically provisioned and
managed for read-write,read, and read-only workloads.Using the
convention over configuration approach, the operator creates adatabase
called app , by default owned by a regular Postgres user with
thesame name. Both the database name and the user name can be specified
ifrequired.Although no configuration is required to run the cluster, you
can customizeboth PostgreSQL run-time configuration and PostgreSQL
Host-BasedAuthentication rules in the postgresql section of the CR.
Pod Security Policies¶
For InfoSec requirements, the operator does not require privileged mode forany container and enforces read only root filesystem to guarantee containersimmutability for both the operator and the operand pods. It also explicitlysets the required security contexts.
Affinity¶
The cluster’s affinity section enables fine-tuning of how pods and
relatedresources such as persistent volumes are scheduled across the
nodes of aKubernetes cluster. In particular, the operator supports:
pod affinity and anti-affinity
node selector
taints and tolerations
Command line interface¶
CloudNativePG does not have its own command line interface.It simply
relies on the best command line interface for Kubernetes, kubectl
,by providing a plugin called cnpg to enhance and simplify your
PostgreSQLcluster management experience.
Current status of the cluster¶
The operator continuously updates the status section of the CR with
theobserved status of the cluster. The entire PostgreSQL cluster status
iscontinuously monitored by the instance manager running in each pod:
theinstance manager is responsible for applying the required changes to
thecontrolled PostgreSQL instance to converge to the required status
ofthe cluster (for example: if the cluster status reports that pod
-1 is theprimary, pod -1 needs to promote itself while the other
pods need to followpod -1 ). The same status is used by the cnpg
plugin for kubectl to providedetails.
TLS connections¶
The operator transparently and natively supports TLS/SSL connectionsto encrypt client/server communications for increased security using thecluster’s certification authority.Support for custom server certificates is available through secrets: this alsoincludes integration with cert-manager.
Certificate authentication for streaming replication¶
The operator relies on TLS client certificate authentication to authorize streamingreplication connections from the standby servers, instead of relying on a password(and therefore a secret).
Continuous configuration management¶
The operator enables you to apply changes to the Cluster resource
YAMLsection of the PostgreSQL configuration and makes sure that all
instancesare properly reloaded or restarted, depending on the
configuration option.Current limitation: changes with
ALTER SYSTEM are not detected, meaningthat the cluster state is not
enforced.
Basic LDAP authentication for PostgreSQL¶
The operator allows you to configure LDAP authentication for your PostgreSQLclients, using either the simple bind or search+bind mode, as described inthe PostgreSQL documentation: LDAP authentication .
Multiple installation methods¶
The operator can be installed through a Kubernetes manifest via
kubectlapply , to be used in a traditional Kubernetes installation
in publicand private cloud environments. Additionally, a Helm Chart for
the operator isalso available.
Convention over configuration¶
The operator supports the convention over configuration paradigm,
decidingstandard default values while allowing you to override them and
customizethem. You can specify a deployment of a PostgreSQL cluster
usingthe Cluster CRD in a couple of YAML code lines.
Level 2 - Seamless Upgrades¶
Capability level 2 is about enabling updatesoftheoperatorandtheactualworkload , in our case PostgreSQL servers. This includes PostgreSQLminorreleaseupdates (security and bug fixes normally) as well as majoronlineupgrades .
Upgrade of the operator¶
You can upgrade the operator seamlessly as a new deployment. A change in theoperator does not require a change in the operand - thanks to the instancemanager’s injection. The operator can manage older versions of the operand.
CloudNativePG also supports In-place updates of the instance manager following an upgrade of the operator: in-place updates do not require a rollingupdate - and subsequent switchover - of the cluster.
Upgrade of the managed workload¶
The operand can be upgraded using a declarative configuration approach
aspart of changing the CR and, in particular, the imageName
parameter. Theoperator prevents major upgrades of PostgreSQL while
making it possible to goin both directions in terms of minor PostgreSQL
releases within a majorversion (enabling updates and rollbacks).
In the presence of standby servers, the operator performs rolling
updatesstarting from the replicas by dropping the existing pod and
creating a newone with the new requested operand image that reuses the
underlying storage.Depending on the value of the
primaryUpdateStrategy , the operator proceedswith a switchover
before updating the former primary ( unsupervised ) or waitsfor the
user to manually issue the switchover procedure ( supervised ) via
the cnpg plugin for kubectl .Which setting to use depends on the
business requirements as the operationmight generate some downtime for
the applications, from a few seconds tominutes based on the actual
database workload.
Display cluster availability status during upgrade¶
At any time, convey the cluster’s high availability status, for example,
Setting up primary , Creating a new replica ,
Cluster in healthy state , Switchover in progress ,
Failing over , Upgrading cluster , etc.
Level 3 - Full Lifecycle¶
Capability level 3 requires the operator to manage aspects of businesscontinuity and scalability . Disasterrecovery is a business continuity component that requiresthat both backup and recovery of a database work correctly. While as astarting point, the goal is to achieve RPO < 5 minutes, the long term goal isto implement RPO=0 backup solutions. HighAvailability is the otherimportant component of business continuity that, through PostgreSQL nativephysical replication and hot standby replicas, allows the operator to performfailover and switchover operations. This area includes enhancements in:
PostgreSQL Backups¶
The operator has been designed to provide application-level backups usingPostgreSQL’s native continuous backup technology based onphysical base backups and continuous WAL archiving. Specifically,the operator currently supports only backups on object stores (AWS S3 andS3-compatible, Azure Blob Storage, Google Cloud Storage, and gateways likeMinIO).
WAL archiving and base backups are defined at the cluster level,
declaratively,through the backup parameter in the cluster
definition, by specifyingan S3 protocol destination URL (for example, to
point to a specific folder inan AWS S3 bucket) and, optionally, a
generic endpoint URL. WAL archiving,a prerequisite for continuous
backup, does not require any furtheraction from the user: the operator
will automatically and transparently setthe archive_command to rely
on barman-cloud-wal-archive to ship WALfiles to the defined
endpoint. Users can decide the compression algorithm,as well as the
number of parallel jobs to concurrently upload WAL filesin the archive.
In addition to that Instance Manager automatically checks the
correctness of the archive destination, by performing
barman-cloud-check-wal-archive command before beginning to ship the
very first set of WAL files.
You can define base backups in two ways: on-demand (through the
Backup custom resource definition) or scheduled (through the
ScheduledBackup customer resource definition, using a cron-like
syntax). They both rely on barman-cloud-backup for the job
(distributed as part of the applicationcontainer image) to relay backups
in the same endpoint, alongside WAL files.
Both barman-cloud-wal-restore and barman-cloud-backup are
distributed inthe application container image under GNU GPL 3 terms.
Full restore from a backup¶
The operator enables you to bootstrap a new cluster (with its
settings)starting from an existing and accessible backup taken using
barman-cloud-backup . Once the bootstrap process is completed, the
operatorinitiates the instance in recovery mode and replays all
available WAL filesfrom the specified archive, exiting recovery and
starting as a primary.Subsequently, the operator will clone the
requested number of standby instancesfrom the primary.CloudNativePG
supports parallel WAL fetching from the archive.
Point-In-Time Recovery (PITR) from a backup¶
The operator enables you to create a new PostgreSQL cluster by recoveringan existing backup to a specific point-in-time, defined with a timestamp, alabel or a transaction ID. This capability is built on top of the full restoreone and supports all the options available in
Zero Data Loss clusters through synchronous replication¶
Achieve Zero Data Loss (RPO=0) in your local High Availability
CloudNativePGcluster through quorum based synchronous replication
support. The operator providestwo configuration options that control the
minimum and maximum number ofexpected synchronous standby replicas
available at any time. The operator willreact accordingly, based on the
number of available and ready PostgreSQLinstances in the cluster,
through the following formula for the quorum ( q ):
1 <= minSyncReplicas <= q <= maxSyncReplicas <= readyReplicas
Replica clusters¶
Define a cross Kubernetes cluster topology of PostgreSQL clusters, by
takingadvantage of PostgreSQL native streaming and cascading
replication.Through the replica option, you can setup an independent
cluster to becontinuously replicating data from another PostgreSQL
source of the same majorversion: such a source can be anywhere, as long
as a direct streamingconnection via TLS is allowed from the two
endpoints.Moreover, the source can be even outside Kubernetes, running
in a physical orvirtual environment.Replica clusters can be created from
a recovery object store (backup in BarmanCloud format) or via streaming
through pg_basebackup . Both WAL file shippingand WAL streaming are
allowed.Replica clusters dramatically improve the business continuity
posture of yourPostgreSQL databases in Kubernetes, spanning over
multiple datacenters andopening up for hybrid and multi-cloud setups
(currently, manual switchoveracross data centers is required, while
waiting for Kubernetes federationnative capabilities).
Liveness and readiness probes¶
The operator defines liveness and readiness probes for the
PostgresContainers that are then invoked by the kubelet. They are mapped
respectivelyto the /healthz and /readyz endpoints of the web
server manageddirectly by the instance manager.The liveness probe is
based on the pg_isready executable, and the pod isconsidered healthy
with exit codes 0 (server accepting connections normally)and 1 (server
is rejecting connections, for example during startup). Thereadiness
probe issues a simple query ( ; ) to verify that the server isready
to accept connections.
Rolling deployments¶
The operator supports rolling deployments to minimize the downtime and, if aPostgreSQL cluster is exposed publicly, the Service will load-balance theread-only traffic only to available pods during the initialization or theupdate.
Scale up and down of replicas¶
The operator allows you to scale up and down the number of instances in
aPostgreSQL cluster. New replicas are automatically started up from
theprimary server and will participate in the cluster’s HA
infrastructure.The CRD declares a “scale” subresource that allows the
user to use the kubectl scale command.
Maintenance window and PodDisruptionBudget for Kubernetes nodes¶
The operator creates a PodDisruptionBudget resource to limit the
number ofconcurrent disruptions to one primary instance. This
configuration prevents themaintenance operation from deleting all the
pods in a cluster, allowing thespecified number of instances to be
created. The PodDisruptionBudget will beapplied during the node draining
operation, preventing any disruption of thecluster service.
While this strategy is correct for Kubernetes Clusters wherestorage is
shared among all the worker nodes, it may not be the best solutionfor
clusters using Local Storage or for clusters installed in a
privatecloud. The operator allows you to specify a Maintenance Window
andconfigure the reaction to any underlying node eviction. The
ReusePVC optionin the maintenance window section enables to specify
the strategy to be used:allocate new storage in a different PVC for the
evicted instance or waitfor the underlying node to be available again.
Fencing¶
Fencing is the process of protecting the data in one, more, or even allinstances of a PostgreSQL cluster when they appear to be malfunctioning.When an instance is fenced, the PostgreSQL server process isguaranteed to be shut down, while the pod is kept running. This makes surethat, until the fence is lifted, data on the pod is not modified by PostgreSQLand that the file system can be investigated for debugging and troubleshootingpurposes.
Reuse of Persistent Volumes storage in Pods¶
When the operator needs to create a pod that has been deleted by the
user orhas been evicted by a Kubernetes maintenance operation, it reuses
the PersistentVolumeClaim if available, avoiding the needto re-clone
the data from the primary.
CPU and memory requests and limits¶
The operator allows administrators to control and manage resource usage
bythe cluster’s pods, through the resources section of the manifest.
Inparticular requests and limits values can be set for both CPU
and RAM.
Connection pooling with PgBouncer¶
- CloudNativePG provides native support for connection pooling with
PgBouncerIntegrationStatus , one of the most popular open sourceconnection poolers
for PostgreSQL. From an architectural point of view, thenative implementation of a PgBouncer connection pooler introduces a new layerto access the database which optimizes the query flow towards the instancesand makes the usage of the underlying PostgreSQL resources more efficient.Instead of connecting directly to a PostgreSQL service, applications can nowconnect to the PgBouncer service and start reusing any existing connection.
Level 4 - Deep Insights¶
Capability level 4 is about observability : in particular, monitoring,alerting, trending, log processing. This might involve the use of external toolssuch as Prometheus, Grafana, Fluent Bit, as well as extensions in thePostgreSQL engine for the output of error logs directly in JSON format.
CloudNativePG has been designed to provide everything that is neededto easily integrate with industry-standard and community accepted tools forflexible monitoring and logging.
Prometheus exporter with configurable queries¶
The instance manager provides a pluggable framework and, via its own web
serverlistening on the metrics port (9187), exposes an endpoint to
export metricsfor the Prometheus Operator example monitoring and alerting tool.The
operator supports custom monitoring queries defined as ConfigMap
and/or Secret objects using a syntax that is compatible with the
postgres_exporter .CloudNativePG provides a set of basic monitoring queries
forPostgreSQL that can be integrated and adapted to your context.The [cnp-sandbox project] is an open source Helm chart that demonstrateshow to integrate CloudNativePG with Prometheus and Grafana, by providingsome basic metrics and an example of dashboard.
Standard output logging of PostgreSQL error messages in JSON format¶
Every log message is delivered to standard output in JSON format, with
the first leveldefinition of the timestamp, the log level and the type
of log entry, such as postgres for the canonical PostgreSQL error
message channel.As a result, every Pod managed by CloudNativePG can be
easily and directlyintegrated with any downstream log processing stack
that supports JSON as sourcedata type.
Real-time query monitoring¶
CloudNativePG transparently and natively supports:
Audit¶
CloudNativePG allows database and security administrators, auditors,and operators to track and analyze database activities using PGAudit (forPostgreSQL).Such activities flow directly in the JSON log and can be properly routed to thecorrect downstream target using common log brokers like Fluentd.
Kubernetes events¶
Record major events as expected by the Kubernetes API, such as creating
resources,removing nodes, upgrading, and so on. Events can be displayed
throughthe kubectl describe and kubectl get events command.
Level 5 - Auto Pilot¶
Capability level 5 is focused on automatedscaling , healing and tuning - through the discovery of anomalies and insights that emergedfrom the observability layer.
Automated Failover for self-healing¶
In case of detected failure on the primary, the operator will change
thestatus of the cluster by setting the most aligned replica as the new
targetprimary. As a consequence, the instance manager in each alive pod
willinitiate the required procedures to align itself with the requested
status ofthe cluster, by either becoming the new primary or by following
it.In case the former primary comes back up, the same mechanism will
avoid asplit-brain by preventing applications from reaching it, running
pg_rewind onthe server and restarting it as a standby.
Automated recreation of a standby¶
In case the pod hosting a standby has been removed, the operator initiatesthe procedure to recreate a standby server.