HARP Functionality Overview

HARP is a new approach to High Availability for BDRclusters. Itleverages consensus-driven Quorum to determine the correct connection end-pointin a semi-exclusive manner to prevent unintended multi-node writes from anapplication.

The Importance of Quorum

The central purpose of HARP is to enforce full Quorum on any Postgres clusterit manages. Quorum is merely a term generally applied to a voting body thatmandates a certain minimum of attendees are available to make a decision. Orperhaps even more simply: Majority Rules.

In order for any vote to end in a result other than a tie, an odd number ofnodes must constitute the full cluster membership. Quorum however does notstrictly demand this restriction; a simple majority will suffice. This meansthat in a cluster of N nodes, Quorum requires a minimum of N/2+1 nodes to holda meaningful vote.

All of this ensures the cluster is always in agreement regarding which nodeshould be “in charge”. For a BDR cluster consisting of multiple nodes, thisdetermines which node is the primary write target. HARP designates this nodeas the Lead Master.

Reducing Write Targets

The consequence of ignoring the concept of Quorum, or applying itinsufficiently, may lead to a Split Brain scenario where the “correct” writetarget is ambiguous or unknowable. In a standard Postgres cluster, it isimportant that only a single node is ever writable and sending replicationtraffic to the remaining nodes.

Even in Multi-Master capable approaches such as BDR, it can be beneficial toreduce the amount of necessary conflict management to derive identical dataacross the cluster. In clusters that consist of multiple BDR nodes per physicallocation or region, this usually means a single BDR node acts as a “Leader” andremaining nodes are “Shadows”. These Shadow nodes are still writable, but doingso is discouraged unless absolutely necessary.

By leveraging Quorum, it’s possible for all nodes to agree exactly whichPostgres node should represent the entire cluster, or a local BDR region. Anynodes that lose contact with the remainder of the Quorum, or are overruled byit, by definition cannot become the cluster Leader.

This prevents Split Brain situations where writes unintentionally reach twoPostgres nodes. Unlike technologies such as VPNs, Proxies, load balancers, orDNS, a Quorum-derived consensus cannot be circumvented by mis-configuration ornetwork partitions. So long as it’s possible to contact the Consensus layer todetermine the state of the Quorum maintained by HARP, only one target is evervalid.

Basic Architecture

The design of HARP comes in essentially two parts consisting of a Manager anda Proxy. The following diagram describes how these interact with a singlePostgres instance:

HARP Unit

HARP Unit

The Consensus Layer is an external entity where Harp Manager maintains information it learns about its assigned Postgres node, and HARP Proxy translates this information to a valid Postgres node target. Because Proxyobtains the node target from the Consensus Layer, several such instances mayexist independently.

While using BDR itself as the Consensus Layer, each server node resembles thisvariant instead.

HARP Unit w/BDR Consensus

HARP Unit w/BDR Consensus

In either case, each unit consists of the following elements:

  • A Postgres or EDB instance

  • A Consensus Layer resource, meant to track various attributes of the Postgres

  • instance

  • A HARP Manager process to convey the state of the Postgres node to the

  • Consensus Layer

  • A HARP Proxy service that directs traffic to the proper Lead Master node,

  • as derived from the Consensus Layer

Not every application stack has access to additional node resources specifically for the Proxy component, so it can be combined with the application server to simplify the stack itself.

This is a typical design using two BDR nodes in a single Data Center organized in a Lead Master / Shadow Master configuration:

HARP Cluster

HARP Cluster

Note that when using BDR itself as the HARP Consensus Layer, at least threefully qualified BDR nodes must be present to ensure a quorum majority.

HARP Cluster w/BDR Consensus

HARP Cluster w/BDR Consensus

(Not shown in the above diagram are connections between BDR nodes.)

How it Works

When managing a BDR cluster, HARP maintains at most one “Leader” node perdefined Location. Canonically this is referred to as the Lead Master. Other BDRnodes which are eligible to take this position are Shadow Master state untilsuch a time they take the Leader role.

Applications may contact the current Leader only through the Proxy service. Since the Consensus Layer requires Quorum agreement before conveying Leader state, any and all Proxy services will direct traffic to that node.

At a high level, this is ultimately what prevents application interaction withmultiple nodes simultaneously.

Determining a Leader

As an example, consider the role of Lead Master within a locally subdividedBDR Always-On group as may exist within a single data center. When anyPostgres or Manager resource is started, and after a configurable refreshinterval, the following must occur:

  1. The Manager checks the status of its assigned Postgres resource. - If Postgres is not running, try again after configurable timeout. - If Postgres is running, continue.2. The Manager checks the status of the Leader lease in the Consensus Layer. - If the lease is unclaimed, acquire it and assign the identity of the Postgres instance assigned to this Manager. This lease duration is configurable, but setting it too low may result in unexpected leadership transitions. - If the lease is already claimed by us, renew the lease TTL. - Otherwise do nothing.

Obviously a lot more happens here, but this simplified version should explainwhat’s happening. The Leader lease can only be held by one node, and if it’sheld elsewhere, HARP Manager gives up and tries again later.

!!! Note * Depending on the chosen Consensus Layer, rather than repeatedly looping to check the status of the Leader lease, HARP will subscribe to notifications instead. In this case, it can respond immediately any time the state of the lease changes, rather than polling. Currently this functionality is restricted to the etcd Consensus Layer.

This means HARP itself does not hold elections or manage Quorum; this isdelegated to the Consensus Layer. The act of obtaining the lease must beacknowledged by a Quorum of the Consensus Layer, so if the request succeeds,that node leads the cluster in that Location.

Connection Routing

Once the role of the Lead Master is established, connections are handledwith a similar deterministic result as reflected by HARP Proxy. Consider a casewhere HAProxy needs to determine the connection target for a particular backendresource:

  1. HARP Proxy interrogates the Consensus layer for the current Lead Master in its configured location.2. If this is unset or in transition; - New client connections to Postgres are barred, but clients will accumulate and be in a paused state until a Lead Master appears. - Existing client connections are allowed to complete current transaction, and are then reverted to a similar pending state as new connections.3. Client connections are forwarded to the Lead Master.

Note that the interplay demonstrated in this case does not require anyinteraction with either HARP Manager or Postgres. The Consensus Layer itselfis the source of all truth from the Proxy’s perspective.

Colocation

The arrangement of the work units is such that their organization is requiredto follow these principles:

  1. The Manager and Postgres units must exist concomitantly within the same node.2. The contents of the Consensus Layer dictate the prescriptive role of all operational work units.

This delegates cluster Quorum responsibilities to the Consensus Layer itself, while HARP leverages it for critical role assignments and key/value storage. Neither storage or retrieval will succeed if the Consensus Layer is inoperable or unreachable, thus preventing rogue Postgres nodes from accepting connections.

As a result, the Consensus Layer should generally exist outside of HARP or HARP managed nodes for maximum safety. Our reference diagrams reflect this in orderto encourage such separation, though it is not required.

!!! Note * In order to operate and manage cluster state, BDR contains its own implementation of the Raft Consensus model. HARP may be configured to leverage this same layer to reduce reliance on external dependencies and to preserve server resources. However, there are certain drawbacks to this approach that are discussed in further depth in the section on the Consensus Layer.