BDR 4.1.0¶
This is a minor release of BDR 4 which includes new features as well as fixes for issues identified in previous versions.
Type |
Category |
Description |
|---|---|---|
Feature |
Reliability and operability |
Support in-place major upgrade of Postgres on a BDR node<p>This BDR release include as new command-line utility bdr_pg_upgrade which uses pg_upgrade to do a major version upgrade of Postgres on a BDR node.</p><p> This reduces the time and network bandwidth necessary to do major version upgrades of Postgres in a EDB Postgres Distributed cluster.</p> |
Feature |
Performance and scalability |
Replication Lag Control <p> Add configuration for a replication lag threshold after which the transaction commits get throttled. This allows limiting RPO without incurring the latency impact on every transaction that comes with synchronous replication.</p> |
Feature |
UX / Initial experience |
Distributed sequences by default <p> The default value of bdr.default_sequence_kind has been changed to ‘distributed’ which is new kind of sequence that uses SnowFlakeId for bigserial and Galloc sequences for serial column type.</p> |
Feature |
UX |
Simplified synchronous replication configuration <p> New syntax for specifying the synchronous replication options, with focus on BDR groups and SQL based management (as opposed to config file). </p><p> In future versions this will also replace the current Eager Replication and CAMO configuration options.</p> |
Feature |
High availability and disaster recovery |
Group Commit <p> The initial kind of synchronous commit that can be configured via the new configuration syntax.</p> |
Feature |
High availability and disaster recovery |
Allow a Raft request to be required for CAMO switching to Local Mode (RT78928)<p> Add a require_raft flag to the CAMO pairing configuration which controls the behavior of switching from CAMO protected to Local Mode, introducing the option to require a majority of nodes to be connected to allow to switch to Local Mode.</p> |
Feature |
High availability and disaster recovery |
Allow replication to continue on ALTER TABLE … DETACH PARTITION CONCURRENTLY of already detached partition (RT78362)<p> Similarly to how BDR 4 handles CREATE INDEX CONCURRENTLY when same index already exists, we now allow replication to continue when ALTER TABLE … DETACH PARTITION CONCURRENTLY is receiver for partition that has been already detached.</p> |
Feature |
User Experience |
Add additional filtering options to DDL filters.<p>DDL filters allow for replication of different DDL statements to different replication sets. Similar to how table membership in replication set allows DML on different tables to be replicated via different replication sets.</p><p>This release adds new controls that make it easier to use the DDL filters:<br /> - query_match - if defined query must match this regex<br /> - exclusive - if true, other matched filters are not taken into consideration (i.e. only the exclusive filter is applied), when multiple exclusive filters match, we throw error</p> |
Feature |
User Experience |
Add bdr.lock_table_locking configuration variable.<p>When enabled this changes behavior of LOCK TABLE command to take take a global DML lock</p> |
Feature |
Performance and scalability |
Implement buffered write for LCR segment file <p> This should reduce I/O and improve CPU usage of the Decoding Worker.</p> |
Feature |
User Experience |
Add support for partial unique index lookups for conflict detection (RT78368).<p>Indexes on expression are however still not supported for conflict detection.</p> |
Feature |
User Experience |
Add additional statistics to bdr.stat_subscription:<br /> - nstream_insert => the count of INSERTs on streamed transactions<br /> - nstream_update => the count of UPDATEs on streamed transactions<br /> - nstream_delete => the count of DELETEs on streamed transactions<br /> - nstream_truncate => the count of TRUNCATEs on streamed transactions<br /> - npre_commit_confirmations => the count pre-commit confirmations, when using CAMO<br /> - npre_commit => the count of pre-commits<br /> - ncommit_prepared => the count of prepared commits with 2PC<br /> - nabort_prepared => the count of aborts of prepared transactions with 2PC |
Feature |
User Experience |
Add execute_locally option to bdr.replicate_ddl_command (RT73533).<p>This allows optional queueing of ddl commands for replication to other groups without executing it locally.</p> |
Feature |
User Experience |
Add fast argument to bdr.alter_subscription_disable() (RT79798)<p>The argument only influences the behavior of immediate. When set to true (default) it will stop the workers without letting them finish the current work.</p> |
Feature |
User Experience |
Keep the bdr.worker_error records permanently for all types of workers.<p>BDR used to remove receiver and writer errors when those workers managed to replicate the LSN that was previously resulting in error. However this was inconsistent with how other workers behaved, as other worker errors were permanent and it also made the troubleshooting of past issues harder. So keep the last error record permanently for every worker type.</p> |
Feature |
User Experience |
Simplify bdr.{add,remove}_camo_pair functions to return void. |
Feature |
Initial Experience |
Add connectivity/lag check before taking global lock.<p>So that application or user does not have to wait for minutes to get lock timeout when there are obvious connectivity issues.</p><p>Can be set to DEBUG, LOG, WARNING (default) or ERROR.</p> |
Feature |
Initial Experience |
Only log conflicts to conflict log table by default. They are no longer logged to the server log file by default, but this can be overridden. |
Feature |
User Experience |
Improve reporting of remote errors during node join. |
Feature |
Reliability and operability |
Make autopartition worker’s max naptime configurable. |
Feature |
User Experience |
Add ability to request partitions upto the given upper bound with autopartition. |
Feature |
Initial Experience |
Don’t try replicate DDL run on subscribe-only node. It has nowhere to replicate so any attempt to do so will fail. This is same as how logical standbys behave. |
Feature |
User Experience |
Add bdr.accept_connections configuration variable. When false, walsender connections to replication slots using BDR output plugin will fail. This is useful primarily during restore of single node from backup. |
Bug fix |
Reliability and operability |
Keep the lock_timeout as configured on non-CAMO-partner BDR nodes<p> A CAMO partner uses a low lock_timeout when applying transactions from its origin node. This was inadvertently done for all BDR nodes rather than just the CAMO partner, which may have led to spurious lock_timeout errors on pglogical writer processes on normal BDR nodes.</p> |
Bug fix |
User Experience |
Show a proper wait event for CAMO / Eager confirmation waits (RT75900)<p>Show correct “BDR Prepare Phase”/”BDR Commit Phase” in bdr.stat_activity instead of the default “unknown wait event”.</p> |
Bug fix |
User Experience |
Reduce log for bdr.run_on_nodes (RT80973)<p> Don’t log when setting bdr.ddl_replication to off if it’s done with the “run_on_nodes” variants of function. This eliminates the flood of logs for monitoring functions.</p> |
Bug fix |
Reliability and operability |
Fix replication of arrays of composite types and arrays of builtin types that don’t support binary network encoding |
Bug fix |
Reliability and operability |
Fix replication of data types created during bootstrap |
Bug fix |
Performance and scalability |
Confirm end LSN of the running transactions record processed by WAL decoder so that the WAL decoder slot remains up to date and WAL sender get the candidate in timely manner. |
Bug fix |
Reliability and operability |
Don’t wait for autopartition tasks to complete on parting nodes |
Bug fix |
Reliability and operability |
Limit the bdr.standby_slot_names check when reporting flush position only to physical slots (RT77985, RT78290)<p> Otherwise flush progress is not reported in presence of disconnected nodes when using bdr.standby_slot_names.</p> |
Bug fix |
Reliability and operability |
Request feedback reply from walsender if we are close to wal_receiver_timeout |
Bug fix |
Reliability and operability |
Don’t record dependency of auto-paritioned table on BDR extension more than once.<p>This resulted in “ERROR: unexpected number of extension dependency records” errors from auto-partition and broken replication on conflicts when this happens.</p><p>Note that existing broken tables need to still be fixed manually by removing the double dependency from `pg_depend`</p> |
Bug fix |
Reliability and operability |
Improve keepalive handling in receiver.<p>Don’t update position based on keepalive when in middle of streaming transaction as we might lose data on crash if we do that.</p>There is also new flush and signalling logic that should improve latency in low TPS scenarios. |
Bug fix |
Reliability and operability |
Only do post CREATE commands processing when BDR node exists in the database. |
Bug fix |
Reliability and operability |
Don’t try to log ERROR conflicts to conflict history table. |
Bug fix |
Reliability and operability |
Fixed segfault where a conflict_slot was being used after it was released during multi-insert (COPY) (RT76439). |
Bug fix |
Reliability and operability |
Prevent walsender processes spinning when facing lagging standby slots (RT80295, RT78290).<p>Correct signaling to reset a latch so that a walsender process does consume 100% of a CPU in case one of the standby slots is lagging behind.</p> |
Bug fix |
Reliability and operability |
Fix handling of wal_sender_timeout when bdr.standby_slot_names are used (RT78290) |
Bug fix |
Reliability and operability |
Make ALTER TABLE lock the underlying relation only once (RT80204). |
Bug fix |
User Experience |
Fix reporting of disconnected slots in bdr.monitor_local_replslots. They could have been previously reported as missing instead of disconnected. |
Bug fix |
Reliability and operability |
Fix apply timestamp reporting for down subscriptions in bdr.get_subscription_progress() function and in the bdr.subscription_summary that uses that function. It would report garbage value before. |
Bug fix |
Reliability and operability |
Fix snapshot handling in various places in BDR workers. |
Bug fix |
User Experience |
Be more consistent about reporting timestamps and LSNs as NULLs in monitoring functions when there is no available value for those. |
Bug fix |
Reliability and operability |
Reduce log information when switching between writer processes. |
Bug fix |
Reliability and operability |
Don’t do superuser check when configuration parameter was specified on PG command-line. We can’t do transactions there yet and it’s guaranteed to be superuser changed at that stage. |
Bug fix |
Reliability and operability |
Use 64 bits for calculating lag size in bytes. To eliminate risk of overflow with large lag. |
Upgrades¶
This release supports upgrading from the following versions of BDR:
4.0.0 and higher
3.7.15
3.7.16
Please make sure you read and understand the process and limitations described in the Upgrade Guide before upgrading.