Get started#
Analytics Accelerator compatibility : Check supported PostgreSQL versions, operating systems, and other requirements.
Analytics Accelerator architecture : Understand the core architecture and how the vectorized engine works.
Concepts : Understand the fundamental principles of vectorized execution, data lake integration, and DirectScan.
Analytics Accelerator quickstart guide : Install PGAA, create a storage location and read table from our sample benchmark datasets.
Using PGAA#
Installing Analytics Accelerator : Step-by-step instructions for installing the extension and enabling the Seafowl background worker.
Reading tables in object storage : Connect directly to S3, GCS, or Azure Blob Storage to query Parquet, Delta, or Iceberg files using a PGFS storage location.
Integrating with Iceberg catalogs : Integrate with external Iceberg REST catalogs to manage table metadata.
Writing to object storage : Use
CREATE TABLE AS SELECT(CTAS) to export Postgres data into optimized lakehouse formats in your object store.
Replicating with PGD#
Implementing tiered tables Combine PGD AutoPartition and PGAA to create an automated data lifecycle. Move older partitions to object storage while keeping recent data in Postgres tables.
Replicating to analytics Convert standard heap tables into HTAP tables. Use continuous logical replication to maintain a real-time analytical copy of your transactional data in the data lake.
Offloading to analytics Perform surgical storage management by manually moving entire HTAP tables to the cold tier, truncating local data to reclaim disk space immediately.
Performance & optimization#
Accelerate with Spark : Offload massive datasets and complex distributed joins to a remote Spark cluster via Spark Connect. PGAA offers two integration modes depending on your performance requirements:
Distributed Spark execution Leverage a remote Spark cluster for high-concurrency analytical queries and distributed processing.
GPU-accelerated Spark with NVIDIA RAPIDS : Integrate with the NVIDIA RAPIDS Accelerator for Apache Spark to leverage GPU acceleration.
Monitoring and maintaining analytical tables : Audit storage utilization, monitor table health, and perform table maintenance tasks for PGAA-managed tables.
Optimizing query performance : Maximize query speeds by managing DirectScan execution, configuring compute pushdowns, and troubleshooting path fallbacks.
Reference#
Configuration parameters : The behavior of the PGAA extension is governed by Grand Unified Configuration (GUC) variables. These parameters allow you to switch executors, enable performance optimizations, and manage security credentials.
Functions : PGAA introduces a suite of SQL functions for administrative tasks, such as mapping new tables, monitoring storage health, and launching maintenance background jobs.
Table options : When mapping or creating analytical tables, specific options allow you to define how data is read from or written to your object store.
Data types and definitions : PGAA maps native Postgres data types to optimized columnar formats in the data lake.
Benchmark datasets : Access pre-configured schemas and data loading instructions for analytical datasets to baseline your performance.