Replication

Physical replication is one of the strengths of PostgreSQL and one of thereasons why some of the largest organizations in the world have chosenit for the management of their data in business continuity contexts.Primarily used to achieve high availability, physical replication also allowsscale-out of read-only workloads and offloading some work from the primary.

Application-level replication

Having contributed throughout the years to the replication feature in PostgreSQL,we have decided to build high availability in CloudNativePG on top ofthe native physical replication technology, and integrate itdirectly in the Kubernetes API.

In Kubernetes terms, this is referred to as application-levelreplication , incontrast with storage-level replication.

A very mature technology

PostgreSQL has a very robust and mature native framework for replicating datafrom the primary instance to one or more replicas, built around theconcept of transactional changes continuously stored in the WAL (Write Ahead Log).

Started as the evolution of crash recovery and point in time recoverytechnologies, physical replication was first introduced in PostgreSQL 8.2(2006) through WAL shipping from the primary to a warm standby incontinuous recovery.

PostgreSQL 9.0 (2010) enhanced it with WAL streaming and read-only replicas viahot standby, while 9.1 (2011) introduced synchronous replication at thetransaction level (for RPO=0 clusters). Cascading replication was released withPostgreSQL 9.2 (2012). The foundations of logical replication were laid inPostgreSQL 9.4, while version 10 (2017) introduced native support for thepublisher/subscriber pattern to replicate data from an origin to a destination.

Replication within a PostgreSQL cluster

Streaming replication support

At the moment, CloudNativePG natively and transparently managesphysical streaming replicas within a cluster in a declarative way, based onthe number of provided instances in the spec :

replicas = instances - 1 (where  instances > 0)

Immediately after the initialization of a cluster, the operator creates a usercalled streaming_replica as follows:

CREATE USER streaming_replica WITH REPLICATION;
   -- NOSUPERUSER INHERIT NOCREATEROLE NOCREATEDB NOBYPASSRLS

Note

Due to a pg_rewind requirement, in PostgreSQL 10 the streaming_replica user is created with SUPERUSER privileges.

Out of the box, the operator automatically sets up streaming replication withinthe cluster over an encrypted channel and enforces TLS client certificateauthentication for the streaming_replica user - as highlighted by the followingexcerpt taken from pg_hba.conf :

#  Require client certificate authentication for the streaming_replica user
hostssl postgres streaming_replica all cert
hostssl replication streaming_replica all cert

Continuous backup integration

In case continuous backup is configured in the cluster, CloudNativePGtransparently configures replicas to take advantage of restore_command whenin continuous recovery. As a result, PostgreSQL is able to use the WAL archiveas a fallback option whenever pulling WALs via streaming replication fails.

Synchronous replication

CloudNativePG supports configuration of quorum-basedsynchronousstreamingreplication via two configuration options called minSyncReplicas and maxSyncReplicas which are the minimum and maximum number of expectedsynchronous standby replicas available at any time.For self-healing purposes, the operator always compares these two values withthe number of available replicas in order to determine the quorum.

Synchronous replication is disabled by default ( minSyncReplicas and maxSyncReplicas are not defined).In case both minSyncReplicas and maxSyncReplicas are set, CloudNativePGautomatically updates the synchronous_standby_names option inPostgreSQL to the following value:

ANY q (pod1, pod2, ...)

Where:

Warning

To provide self-healing capabilities, the operator has the power to ignore minSyncReplicas in case such value is higher than the currently available number of replicas. Synchronous replication is automatically disabled when readyReplicas is 0 .

As stated in the PostgreSQL documentation ,the method ``ANY`` specifies a quorum-based synchronous replication and makestransaction commits wait until their WAL records are replicated to at least therequested number of synchronous standbys in the list.

Important

Even though the operator chooses self-healing over enforcement of synchronous replication settings, our recommendation is to plan for synchronous replication only in clusters with 3+ instances or, more generally, when maxSyncReplicas < (instances - 1) .

Replication from an external PostgreSQL cluster

CloudNativePG relies on the foundations of the PostgreSQL replicationframework even when a PostgreSQL cluster is created from an existing one (source)and kept synchronized through the Replica clusters feature. The sourcecan be a primary cluster or another replica cluster (cascading replica cluster).

The available options in terms of replication, both at bootstrap and continuousrecovery level, are:

All you have to do is actually define an external cluster.Please refer to the BootstrapConfiguration for information on how to clone a PostgreSQL server using either pg_basebackup (streaming) or recovery (object store).

If the external cluster contains a barmanObjectStore section:

If the external cluster contains a connectionParameters section:

The created replica cluster can perform backups in a reserved object store fromthe designated primary, enabling symmetric architectures in a distributedfashion.

You have full flexibility and freedom to decide your favouritedistributed architecture for a PostgreSQL database, by choosing:

Setting up a replica cluster

To setup a replica cluster from a source cluster, we need to create a cluster yamlfile and define the following parts accordingly:

This firstexample defines a replica cluster using streaming replication inboth bootstrap and continuous recovery. The replica cluster connects to thesource cluster using TLS authentication.

You can check the sample YAML in the samples/ subdirectory.

Note the bootstrap and replica sections pointing to the source cluster.

bootstrap:
  pg_basebackup:
    source: cluster-example

replica:
  enabled: true
  source: cluster-example

In the externalClusters section, remember to use the right namespace for thehost in the connectionParameters sub-section.The -replication and -ca secrets should have been copied over if necessary,in case the replica cluster is in a separate namespace.

externalClusters:
- name: <MAIN-CLUSTER>
  connectionParameters:
    host: <MAIN-CLUSTER>-rw.<NAMESPACE>.svc
    user: streaming_replica
    sslmode: verify-full
    dbname: postgres
  sslKey:
    name: <MAIN-CLUSTER>-replication
    key: tls.key
  sslCert:
    name: <MAIN-CLUSTER>-replication
    key: tls.crt
  sslRootCert:
    name: <MAIN-CLUSTER>-ca
    key: ca.crt

The secondexample defines a replica cluster which bootstraps from an objectstore using the recovery section, and continuous recovery using both streamingreplication and the given object store. For streaming replication, the replicacluster connects to the source cluster using basic authentication.

You can check the sample YAML for it in the samples/ subdirectory.

Note the bootstrap and replica sections pointing to the source cluster.

bootstrap:
  recovery:
    source: cluster-example

replica:
  enabled: true
  source: cluster-example

In the externalClusters section, take care to use the right namespace in the endpointURL and the connectionParameters.host .And do ensure that the necessary secrets have been copied if necessary, and thata backup of the source cluster has been created already.

externalClusters:
- name: <MAIN-CLUSTER>
  barmanObjectStore:
    destinationPath: s3://backups/
    endpointURL: http://minio:9000
    s3Credentials:
      …
  connectionParameters:
    host: <MAIN-CLUSTER>-rw.default.svc
    user: postgres
    dbname: postgres
  password:
    name: <MAIN-CLUSTER>-superuser
    key: password

Note

To use streaming replication between the source cluster and the replica cluster, we need to make sure there is network connectivity between the two clusters, and that all the necessary secrets which hold passwords or certificates are properly created in advance.

Promoting the designated primary in the replica cluster

To promote the designatedprimary to primary , all we need to do is todisable the replica mode in the replica cluster through the option spec.replica.enabled

replica:
  enabled: false
  source: cluster-example

Once the replica mode is disabled, the replica cluster and the source clusterwill become two separate clusters, and the designatedprimary in the replicacluster will be promoted to be that cluster’s primary . We can verify the rolechange using the cnpg plugin, checking the status of the cluster which waspreviously the replica:

kubectl cnpg -n <cluster-name-space> status cluster-replica-example

Note

Disabling replication is an irreversible operation: once replication is disabled and the designated primary is promoted to primary, the replica cluster and the source cluster will become two independent clusters definitively.