HARP Functionality Overview¶
HARP is a new approach to High Availability for BDRclusters. Itleverages consensus-driven Quorum to determine the correct connection end-pointin a semi-exclusive manner to prevent unintended multi-node writes from anapplication.
The Importance of Quorum¶
The central purpose of HARP is to enforce full Quorum on any Postgres clusterit manages. Quorum is merely a term generally applied to a voting body thatmandates a certain minimum of attendees are available to make a decision. Orperhaps even more simply: Majority Rules.
In order for any vote to end in a result other than a tie, an odd number ofnodes must constitute the full cluster membership. Quorum however does notstrictly demand this restriction; a simple majority will suffice. This meansthat in a cluster of N nodes, Quorum requires a minimum of N/2+1 nodes to holda meaningful vote.
All of this ensures the cluster is always in agreement regarding which nodeshould be “in charge”. For a BDR cluster consisting of multiple nodes, thisdetermines which node is the primary write target. HARP designates this nodeas the Lead Master.
Reducing Write Targets¶
The consequence of ignoring the concept of Quorum, or applying itinsufficiently, may lead to a Split Brain scenario where the “correct” writetarget is ambiguous or unknowable. In a standard Postgres cluster, it isimportant that only a single node is ever writable and sending replicationtraffic to the remaining nodes.
Even in Multi-Master capable approaches such as BDR, it can be beneficial toreduce the amount of necessary conflict management to derive identical dataacross the cluster. In clusters that consist of multiple BDR nodes per physicallocation or region, this usually means a single BDR node acts as a “Leader” andremaining nodes are “Shadows”. These Shadow nodes are still writable, but doingso is discouraged unless absolutely necessary.
By leveraging Quorum, it’s possible for all nodes to agree exactly whichPostgres node should represent the entire cluster, or a local BDR region. Anynodes that lose contact with the remainder of the Quorum, or are overruled byit, by definition cannot become the cluster Leader.
This prevents Split Brain situations where writes unintentionally reach twoPostgres nodes. Unlike technologies such as VPNs, Proxies, load balancers, orDNS, a Quorum-derived consensus cannot be circumvented by mis-configuration ornetwork partitions. So long as it’s possible to contact the Consensus layer todetermine the state of the Quorum maintained by HARP, only one target is evervalid.
Basic Architecture¶
The design of HARP comes in essentially two parts consisting of a Manager anda Proxy. The following diagram describes how these interact with a singlePostgres instance:
HARP Unit¶
The Consensus Layer is an external entity where Harp Manager maintains information it learns about its assigned Postgres node, and HARP Proxy translates this information to a valid Postgres node target. Because Proxyobtains the node target from the Consensus Layer, several such instances mayexist independently.
While using BDR itself as the Consensus Layer, each server node resembles thisvariant instead.
HARP Unit w/BDR Consensus¶
In either case, each unit consists of the following elements:
A Postgres or EDB instance
A Consensus Layer resource, meant to track various attributes of the Postgres
instance
A HARP Manager process to convey the state of the Postgres node to the
Consensus Layer
A HARP Proxy service that directs traffic to the proper Lead Master node,
as derived from the Consensus Layer
Not every application stack has access to additional node resources specifically for the Proxy component, so it can be combined with the application server to simplify the stack itself.
This is a typical design using two BDR nodes in a single Data Center organized in a Lead Master / Shadow Master configuration:
HARP Cluster¶
Note that when using BDR itself as the HARP Consensus Layer, at least threefully qualified BDR nodes must be present to ensure a quorum majority.
HARP Cluster w/BDR Consensus¶
(Not shown in the above diagram are connections between BDR nodes.)
How it Works¶
When managing a BDR cluster, HARP maintains at most one “Leader” node perdefined Location. Canonically this is referred to as the Lead Master. Other BDRnodes which are eligible to take this position are Shadow Master state untilsuch a time they take the Leader role.
Applications may contact the current Leader only through the Proxy service. Since the Consensus Layer requires Quorum agreement before conveying Leader state, any and all Proxy services will direct traffic to that node.
At a high level, this is ultimately what prevents application interaction withmultiple nodes simultaneously.
Determining a Leader¶
As an example, consider the role of Lead Master within a locally subdividedBDR Always-On group as may exist within a single data center. When anyPostgres or Manager resource is started, and after a configurable refreshinterval, the following must occur:
The Manager checks the status of its assigned Postgres resource. - If Postgres is not running, try again after configurable timeout. - If Postgres is running, continue.2. The Manager checks the status of the Leader lease in the Consensus Layer. - If the lease is unclaimed, acquire it and assign the identity of the Postgres instance assigned to this Manager. This lease duration is configurable, but setting it too low may result in unexpected leadership transitions. - If the lease is already claimed by us, renew the lease TTL. - Otherwise do nothing.
Obviously a lot more happens here, but this simplified version should explainwhat’s happening. The Leader lease can only be held by one node, and if it’sheld elsewhere, HARP Manager gives up and tries again later.
!!! Note * Depending on the chosen Consensus Layer, rather than repeatedly looping to check the status of the Leader lease, HARP will subscribe to notifications instead. In this case, it can respond immediately any time the state of the lease changes, rather than polling. Currently this functionality is restricted to the etcd Consensus Layer.
This means HARP itself does not hold elections or manage Quorum; this isdelegated to the Consensus Layer. The act of obtaining the lease must beacknowledged by a Quorum of the Consensus Layer, so if the request succeeds,that node leads the cluster in that Location.
Connection Routing¶
Once the role of the Lead Master is established, connections are handledwith a similar deterministic result as reflected by HARP Proxy. Consider a casewhere HAProxy needs to determine the connection target for a particular backendresource:
HARP Proxy interrogates the Consensus layer for the current Lead Master in its configured location.2. If this is unset or in transition; - New client connections to Postgres are barred, but clients will accumulate and be in a paused state until a Lead Master appears. - Existing client connections are allowed to complete current transaction, and are then reverted to a similar pending state as new connections.3. Client connections are forwarded to the Lead Master.
Note that the interplay demonstrated in this case does not require anyinteraction with either HARP Manager or Postgres. The Consensus Layer itselfis the source of all truth from the Proxy’s perspective.
Colocation¶
The arrangement of the work units is such that their organization is requiredto follow these principles:
The Manager and Postgres units must exist concomitantly within the same node.2. The contents of the Consensus Layer dictate the prescriptive role of all operational work units.
This delegates cluster Quorum responsibilities to the Consensus Layer itself, while HARP leverages it for critical role assignments and key/value storage. Neither storage or retrieval will succeed if the Consensus Layer is inoperable or unreachable, thus preventing rogue Postgres nodes from accepting connections.
As a result, the Consensus Layer should generally exist outside of HARP or HARP managed nodes for maximum safety. Our reference diagrams reflect this in orderto encourage such separation, though it is not required.
!!! Note * In order to operate and manage cluster state, BDR contains its own implementation of the Raft Consensus model. HARP may be configured to leverage this same layer to reduce reliance on external dependencies and to preserve server resources. However, there are certain drawbacks to this approach that are discussed in further depth in the section on the Consensus Layer.
Recommended Architecture and Use¶
HARP was primarily designed to represent a BDR Always-On architecture whichresides within two (or more) Data Centers and consists of at least five BDRnodes. This does not count any Logical Standby nodes.
The current and standard representation of this can be seen in the followingdiagram:
BDR Always-On Reference Architecture¶
In this diagram, HARP Manager would exist on BDR Nodes 1-4. The initial state of the cluster would be that BDR Node 1 is the Lead master of DC A, and BDR Node 3 is the Lead Master of DC B.
This would result in any HARP Proxy resource in DC A connecting to BDR Node 1, and likewise the HARP Proxy resource in DC B connecting to BDR Node 3.
!!! Note * While this diagram only shows a single HARP Proxy per DC, this is merely illustrative and should not be considered a Single Point of Failure. Any number of HARP Proxy nodes may exist, and they will all direct application traffic to the same node.
Location Configuration¶
In order for multiple BDR nodes to be eligible to take the Lead Master
lock ina location, a Location must be defined within the config.yml
configurationfile.
To reproduce the diagram above, we would have these lines in the
config.ymlconfiguration for BDR Nodes 1 and 2:
location: dca
And for BDR Nodes 3 and 4:
location: dcb
This applies to any HARP Proxy nodes which are designated in those respectivedata centers as well.
BDR 3.7 Compatibility¶
BDR 3.7 and above offers more direct Location definition by assigning aLocation to the BDR node itself. This is done by calling the following SQLAPI function while connected to the BDR node. So for BDR Nodes 1 and 2, wemight do this:
SELECT bdr.set_node_location('dca');
And for BDR Nodes 3 and 4:
SELECT bdr.set_node_location('dcb');
Afterwards, future versions of HARP Manager would derive the
location fielddirectly from BDR itself. This HARP functionality is
not available yet, so werecommend using this and the setting in
config.yml until HARP reportscompatibility with this BDR API method.