Replication¶
Physical replication is one of the strengths of PostgreSQL and one of thereasons why some of the largest organizations in the world have chosenit for the management of their data in business continuity contexts.Primarily used to achieve high availability, physical replication also allowsscale-out of read-only workloads and offloading some work from the primary.
Application-level replication¶
Having contributed throughout the years to the replication feature in PostgreSQL,we have decided to build high availability in CloudNativePG on top ofthe native physical replication technology, and integrate itdirectly in the Kubernetes API.
In Kubernetes terms, this is referred to as application-levelreplication , incontrast with storage-level replication.
A very mature technology¶
PostgreSQL has a very robust and mature native framework for replicating datafrom the primary instance to one or more replicas, built around theconcept of transactional changes continuously stored in the WAL (Write Ahead Log).
Started as the evolution of crash recovery and point in time recoverytechnologies, physical replication was first introduced in PostgreSQL 8.2(2006) through WAL shipping from the primary to a warm standby incontinuous recovery.
PostgreSQL 9.0 (2010) enhanced it with WAL streaming and read-only replicas viahot standby, while 9.1 (2011) introduced synchronous replication at thetransaction level (for RPO=0 clusters). Cascading replication was released withPostgreSQL 9.2 (2012). The foundations of logical replication were laid inPostgreSQL 9.4, while version 10 (2017) introduced native support for thepublisher/subscriber pattern to replicate data from an origin to a destination.
Replication within a PostgreSQL cluster¶
Streaming replication support¶
At the moment, CloudNativePG natively and transparently managesphysical
streaming replicas within a cluster in a declarative way, based onthe
number of provided instances in the spec :
replicas = instances - 1 (where instances > 0)
Immediately after the initialization of a cluster, the operator creates
a usercalled streaming_replica as follows:
CREATE USER streaming_replica WITH REPLICATION;
-- NOSUPERUSER INHERIT NOCREATEROLE NOCREATEDB NOBYPASSRLS
Note
Due to a pg_rewind requirement, in PostgreSQL 10 the streaming_replica user is created with SUPERUSER privileges.
Out of the box, the operator automatically sets up streaming replication
withinthe cluster over an encrypted channel and enforces TLS client
certificateauthentication for the streaming_replica user - as
highlighted by the followingexcerpt taken from pg_hba.conf :
# Require client certificate authentication for the streaming_replica user
hostssl postgres streaming_replica all cert
hostssl replication streaming_replica all cert
Continuous backup integration¶
In case continuous backup is configured in the cluster,
CloudNativePGtransparently configures replicas to take advantage of
restore_command whenin continuous recovery. As a result, PostgreSQL
is able to use the WAL archiveas a fallback option whenever pulling WALs
via streaming replication fails.
Synchronous replication¶
CloudNativePG supports configuration of
quorum-basedsynchronousstreamingreplication via two configuration
options called minSyncReplicas and maxSyncReplicas which are the
minimum and maximum number of expectedsynchronous standby replicas
available at any time.For self-healing purposes, the operator always
compares these two values withthe number of available replicas in order
to determine the quorum.
Synchronous replication is disabled by default ( minSyncReplicas and
maxSyncReplicas are not defined).In case both minSyncReplicas
and maxSyncReplicas are set, CloudNativePGautomatically updates the
synchronous_standby_names option inPostgreSQL to the following
value:
ANY q (pod1, pod2, ...)
Where:
Warning
To provide self-healing capabilities, the operator has the power to ignore minSyncReplicas in case such value is higher than the currently available number of replicas. Synchronous replication is automatically disabled when readyReplicas is 0 .
As stated in the PostgreSQL documentation ,the method ``ANY`` specifies a quorum-based synchronous replication and makestransaction commits wait until their WAL records are replicated to at least therequested number of synchronous standbys in the list.
Important
Even though the operator chooses self-healing over enforcement of synchronous replication settings, our recommendation is to plan for synchronous replication only in clusters with 3+ instances or, more generally, when maxSyncReplicas < (instances - 1) .
Replication from an external PostgreSQL cluster¶
CloudNativePG relies on the foundations of the PostgreSQL replicationframework even when a PostgreSQL cluster is created from an existing one (source)and kept synchronized through the Replica clusters feature. The sourcecan be a primary cluster or another replica cluster (cascading replica cluster).
The available options in terms of replication, both at bootstrap and continuousrecovery level, are:
All you have to do is actually define an external cluster.Please refer
to the BootstrapConfiguration for information on how to clone a PostgreSQL server
using either pg_basebackup (streaming) or recovery (object
store).
If the external cluster contains a barmanObjectStore section:
If the external cluster contains a connectionParameters section:
The created replica cluster can perform backups in a reserved object store fromthe designated primary, enabling symmetric architectures in a distributedfashion.
You have full flexibility and freedom to decide your favouritedistributed architecture for a PostgreSQL database, by choosing:
Setting up a replica cluster¶
To setup a replica cluster from a source cluster, we need to create a cluster yamlfile and define the following parts accordingly:
This firstexample defines a replica cluster using streaming replication inboth bootstrap and continuous recovery. The replica cluster connects to thesource cluster using TLS authentication.
You can check the sample YAML in the samples/ subdirectory.
Note the bootstrap and replica sections pointing to the source
cluster.
bootstrap:
pg_basebackup:
source: cluster-example
replica:
enabled: true
source: cluster-example
In the externalClusters section, remember to use the right namespace
for thehost in the connectionParameters sub-section.The
-replication and -ca secrets should have been copied over if
necessary,in case the replica cluster is in a separate namespace.
externalClusters:
- name: <MAIN-CLUSTER>
connectionParameters:
host: <MAIN-CLUSTER>-rw.<NAMESPACE>.svc
user: streaming_replica
sslmode: verify-full
dbname: postgres
sslKey:
name: <MAIN-CLUSTER>-replication
key: tls.key
sslCert:
name: <MAIN-CLUSTER>-replication
key: tls.crt
sslRootCert:
name: <MAIN-CLUSTER>-ca
key: ca.crt
The secondexample defines a replica cluster which bootstraps from an
objectstore using the recovery section, and continuous recovery
using both streamingreplication and the given object store. For
streaming replication, the replicacluster connects to the source cluster
using basic authentication.
You can check the sample YAML for it in the samples/ subdirectory.
Note the bootstrap and replica sections pointing to the source
cluster.
bootstrap:
recovery:
source: cluster-example
replica:
enabled: true
source: cluster-example
In the externalClusters section, take care to use the right
namespace in the endpointURL and the connectionParameters.host
.And do ensure that the necessary secrets have been copied if necessary,
and thata backup of the source cluster has been created already.
externalClusters:
- name: <MAIN-CLUSTER>
barmanObjectStore:
destinationPath: s3://backups/
endpointURL: http://minio:9000
s3Credentials:
…
connectionParameters:
host: <MAIN-CLUSTER>-rw.default.svc
user: postgres
dbname: postgres
password:
name: <MAIN-CLUSTER>-superuser
key: password
Note
To use streaming replication between the source cluster and the replica cluster, we need to make sure there is network connectivity between the two clusters, and that all the necessary secrets which hold passwords or certificates are properly created in advance.
Promoting the designated primary in the replica cluster¶
To promote the designatedprimary to primary , all we need to do
is todisable the replica mode in the replica cluster through the option
spec.replica.enabled
replica:
enabled: false
source: cluster-example
Once the replica mode is disabled, the replica cluster and the source clusterwill become two separate clusters, and the designatedprimary in the replicacluster will be promoted to be that cluster’s primary . We can verify the rolechange using the cnpg plugin, checking the status of the cluster which waspreviously the replica:
kubectl cnpg -n <cluster-name-space> status cluster-replica-example
Note
Disabling replication is an irreversible operation: once replication is disabled and the designated primary is promoted to primary, the replica cluster and the source cluster will become two independent clusters definitively.