EDB Postgres Lakehouse
======================

EDB Postgres® Lakehouse brings modern lakehouse architecture to
Postgres-based analytics.

It enables fast SQL-based analytics on data stored in object storage
(S3-compatible), using open table formats such as Apache Iceberg and
Delta Lake.

For implementation and management in Hybrid Manager (HM), see
`EDB Postgres Lakehouse <https://enterprisedb.com/docs/edb-postgres-ai/hybrid-manager/analytics/lakehouse>`_  .

What is the EDB Postgres Lakehouse
----------------------------------

EDB Postgres Lakehouse is an architecture pattern and set of
capabilities that:

- Integrate Postgres with modern lakehouse patterns

- Provide fast, scalable analytics on data in object storage

- Use vectorized query execution and columnar storage formats

- Leverage open table formats for interoperability and data governance

Related concept: :ref:`Data lakehouse <Data lakehouse>` 

Why Lakehouse matters for EDB analytics
---------------------------------------

Lakehouse enables Analytics Accelerator users to:

- Run fast SQL analytics on object storage — no data movement required

- Implement separation of storage and compute for cost efficiency

- Support interoperability with external tools (Spark, Trino, Flink,
  data science frameworks)

- Query PGD Tiered Tables offloaded to Iceberg seamlessly

- Enable unified OLTP + OLAP architectures with Postgres at the center

Related concept: :ref:`Analytics Accelerator concepts <Analytics Accelerator concepts>` 

How EDB implements Lakehouse architecture
-----------------------------------------

Core components:

- **EDB Postgres Lakehouse Nodes**:

- Stateless analytical compute nodes

- Provisioned and managed via Hybrid Manager (HM) or self-managed

- **Vectorized query engine**:

- Powered by `Apache DataFusion <https://datafusion.apache.org/>`_ 

- Processes data in Parquet and other columnar formats efficiently

- **PGAA**:

- Postgres extensions enabling Lakehouse behavior:

- External tables over Iceberg and Delta Lake

- Unified access to PGD hot data + offloaded cold data

- **PGFS**:

- Unified access layer to object storage

- Supports S3, GCS, MinIO, Azure Data Lake Storage, and compatible
  systems

- **Open table formats**:

- Apache Iceberg (full support including catalogs)

- Delta Lake (read-only support)

Related concepts:

- :ref:`Separation of storage and compute <Separation of storage and compute>` 

- :ref:`Vectorized query engines <Vectorized query engines>` 

- :ref:`Columnar storage formats <Columnar storage formats>` 

- :ref:`Open table formats <Open table formats>` 

Common use cases
----------------

.. csv-table::
  :header: Use case,EDB Postgres Lakehouse capability
  :widths: 10,30
  :align: left
  :class: longtable

  Business intelligence & reporting,"Run fast, scalable SQL on data in object storage"
  Historical analytics,Seamlessly query offloaded PGD Tiered Tables
  Data lake analytics,Query existing Iceberg and Delta Lake tables without ETL
  Data science pipelines,Provide efficient access to training data and features
  Hybrid architectures,Enable Postgres-centered OLTP + OLAP patterns

Role-based guidance
-------------------

Database administrators (DBAs)

- :ref:`Analytics Accelerator for your role: DBA <a persona-based guide>` 

Data scientists / analysts

- :ref:`Analytics Accelerator for your role: Data scientist / analyst <a persona-based guide>` 

DevOps / SRE

- :ref:`Analytics Accelerator for your role: DevOps / SRE <a persona-based guide>` 

Application developers

- :ref:`Analytics Accelerator for your role: Application developer <a persona-based guide>` 

Learning paths
--------------

- :ref:`Analytics Accelerator 101: Foundational concepts <Foundational concepts>` 

- :ref:`Analytics Accelerator 201: Practical application and core solutions <Practical application and core solutions>` 

- :ref:`Analytics Accelerator 301: Advanced techniques and optimization <Advanced techniques and optimization>` 

Related concepts
----------------

- :ref:`Data lakehouse <Data lakehouse>` 

- :ref:`Separation of storage and compute <Separation of storage and compute>` 

- :ref:`Vectorized query engines <Vectorized query engines>` 

- :ref:`Columnar storage formats <Columnar storage formats>` 

- :ref:`Open table formats <Open table formats>` 

Next steps
----------

For Hybrid Manager users

- `EDB Postgres Lakehouse <https://enterprisedb.com/docs/edb-postgres-ai/hybrid-manager/analytics/lakehouse>`_ 

How-To guides

- :ref:`Create a Lakehouse cluster <EDB Postgres Lakehouse>` 

- :ref:`Query existing Apache Iceberg tables <Apache Iceberg>` 

- `Query Delta Lake tables <https://enterprisedb.com/docs/edb-postgres-ai/hybrid-manager/analytics/learn/how-to/query-delta-lake-tables>`_ 

Explore more in the `Analytics Accelerator learning guide <https://enterprisedb.com/docs/edb-postgres-ai/hybrid-manager/analytics/learn/>`_  .
