アーキテクチャの概要

Hadoopは、分散ファイルシステムに大規模なデータセットを保存できるフレームワークです。

Hadoopデータラッパーは、HadoopファイルシステムとPostgresデータベース間のインターフェイスを提供します。Hadoopデータラッパーは、Postgresの SELECT ステートメントを、HiveQLまたはSparkSQLインターフェイスで認識されるクエリに変換します。

Using a Hadoop distributed file system with Postgres

PostgresでHadoop分散ファイルシステムを使用する

When possible, the Foreign Data Wrapper asks the Hive or Spark server to perform the actions associated with the WHERE clause of a SELECT statement. Pushing down the WHERE clause improves performance by decreasing the amount of data moving across the network.