Creating a Failover Manager cluster¶
Failover Manager is a high-availability module from EnterpriseDB that enables a Postgres Primary node to automatically failover to a standby node in the event of a software or hardware failure on the primary.
This quick-start guide describes configuring a Failover Manager cluster in a test environment. Read and understand the Failover Manager User Guide before configuring Failover Manager for a production deployment.
Perform these basic installation and configuration steps before starting the tutorial:
Install and initialize a database server on one primary and one or two standby nodes. For information about installing Advanced Server, see [EDB Postgres Advanced Server](/epas/latest/) documentation.
Postgres streaming replication must be configured and running between the primary and standby nodes. For detailed information about configuring streaming replication, see [PostgreSQL Streaming Replication documentation](https://www.postgresql.org/docs/current/warm-standby.html#STREAMING-REPLICATION).
Install Failover Manager on each primary and standby node. During Advanced Server installation, you configured an EnterpriseDB repository on each database host. You can use the EnterpriseDB repository and the `yum install` command to install Failover Manager on each node of the cluster:
* yum install edb-efm43
During the installation process, the installer creates a user named efm
that has privileges to invoke scripts that control the Failover Manager
service for clusters owned by enterprisedb or postgres. The example that
follows creates a cluster named efm.
Start the configuration process on a primary or standby node. Then, copy the configuration files to other nodes to save time.
Create working configuration files. Copy the provided sample files to create Failover Manager configuration files, and correct the ownership:
* cd /etc/edb/efm-4.4
* cp efm.properties.in efm.properties
* cp efm.nodes.in efm.nodes
* chown efm:efm efm.properties
* chown efm:efm efm.nodes
Create an encrypted password needed for the properties file:
* /usr/edb/efm-4.4/bin/efm encrypt efm
Follow the onscreen instructions to produce the encrypted version of your database password.
Update
efm.properties. The<cluster_name>.propertiesfile (efm.propertiesin this example) contains parameters that specify connection properties and behaviors for your Failover Manager cluster. Modifications to property settings are applied when Failover Manager starts.The properties mentioned in this tutorial are the minimal properties required to configure a Failover Manager cluster. If you’re configuring a production system, review the Failover Manager User Guide for detailed information about Failover Manager options.
Provide values for the following properties on all cluster nodes:
- Property | Description | | ———————– | ——————————————————————————————————————————– | |
db.user| The name of the database user. | |db.password.encrypted| The encrypted password of the database user. | |db.port| The port monitored by the database. | |db.database| The name of the database. | |db.service.owner| The owner of thedatadirectory (usuallypostgresorenterprisedb). Required only if the database is running as a service. | |db.service.name| The name of the database service (used to restart the server). Required only if the database is running as a service. | |db.bin| The path to thebindirectory (used for calls topg_ctl). | |db.data.dir| Thedatadirectory in which EFM will find or create therecovery.conffile or thestandby.signalfile. | |user.email| An email address at which to receive email notifications (notification text is also in the agent log file). | |bind.address| The local address of the node and the port to use for Failover Manager. The format is:bind.address=1.2.3.4:7800| |is.witness|trueon a witness node andfalseif it is a primary or standby. | |ping.server.ip| If you are running on a network without Internet access, setping.server.ipto an address that is available on your network. | |auto.allow.hosts| On a test cluster, set totrueto simplify startup; for production usage, consult the user’s guide. | |stable.nodes.file| On a test cluster, set totrueto simplify startup; for production usage, consult the user’s guide. |
Update
efm.nodes. The<cluster_name>.nodesfile (efm.nodesin this example) is read at startup to tell an agent how to find the rest of the cluster or, in the case of the first node started, can be used to simplify authorization of subsequent nodes. Add the addresses and ports of each node in the cluster to this file. One node acts as the membership coordinator. Include in the list at least the membership coordinator’s address. For example:
1.2.3.4:7800
1.2.3.5:7800
1.2.3.6:7800
The Failover Manager agent doesn’t validate the addresses in the
efm.nodesfile. The agent expects that some of the addresses in the file can’t be reached (for example, that another agent hasn’t been started yet).
Configure the other nodes. Copy the
efm.propertiesandefm.nodesfiles to/etc/edb/efm-4.4on the other nodes in your sample cluster. After copying the files, change the file ownership so the files are owned by efm:efm. Theefm.propertiesfile can be the same on every node, except for the following properties:Modify the
bind.addressproperty to use the node’s local address. - Setis.witnesstotrueif the node is a witness node. If the node is a witness node, the properties relating to a local database installation are ignored.
Start the Failover Manager cluster. On any node, start the Failover Manager agent. The agent is named
edb-efm-4.4; you can use your platform-specific service command to control the service. For example, on a CentOS/RHEL 7.x or CentOS/RHEL 8.x host, use the command:
* systemctl start edb-efm-4.4
On a CentOS or RHEL 6.x host, use the command:
* service edb-efm-4.4 start
After the agent starts, run the following command to see the status of the single-node cluster. The addresses of the other nodes appear in the
Allowed node hostlist.
* /usr/edb/efm-4.4/bin/efm cluster-status efm
Start the agent on the other nodes. Run the
efm cluster-status efmcommand on any node to see the cluster status.If any agent fails to start, see the startup log for information about what went wrong:
* cat /var/log/efm-4.4/startup-efm.log
Perform a switchover¶
If the cluster status output shows that the primary and standby nodes are in sync, you can perform a switchover:
* /usr/edb/efm-4.4/bin/efm promote efm -switchover
The command promotes a standby and reconfigures the primary database as a new standby in the cluster. To switch back, run the command again.
Access online help¶
For quick access to online help, use:
* /usr/edb/efm-4.4/bin/efm --help