How to Run the Connector Externally
To run the Ingestion via the UI you’ll need to use the OpenMetadata Ingestion Container, which comes shipped with custom Airflow plugins to handle the workflow deployment. If, instead, you want to manage your workflows externally on your preferred orchestrator, you can check the following docs to run the Ingestion Framework anywhere.External Schedulers
Get more information about running the Ingestion Framework Externally
Requirements
The connector reads metadata from the AWS Glue Data Catalog through the AWS Glue API. Grant the following AWS Identity and Access Management (IAM) permissions to the AWS identity that the connector uses:glue:GetDatabases: Lists the Glue databases in the catalog.glue:GetTables: Lists the tables and views in each Glue database.glue:GetTable: Reads the column details of each Iceberg table. For more information, see Iceberg Table Support.
<AWS_REGION> and <ACCOUNT_ID> with your values.
DESCRIBE permission on the databases and tables to ingest. Without that grant, Glue leaves those databases and tables out of its responses instead of returning an error.
Python Requirements
Use a Python version supported by theopenmetadata-ingestion package that matches your OpenMetadata server. To find the supported range for your release, check the Requires-Python metadata of that release’s ingestion package.
To run the Glue ingestion, you will need to install:
How Glue Maps to OpenMetadata
The connector doesn’t map a Glue database to an OpenMetadata database. The Glue Data Catalog becomes the OpenMetadata database, and each Glue database becomes a schema inside it.databaseName does more than rename the database. When you set it:
- The connector creates one database with that name and stops using
databaseFilterPatternto select catalogs, so leave that pattern empty. - Every Glue database the connector can see becomes a schema of that database, even when the Glue databases come from different catalogs. The ingestion run reports a warning when it sees more than one catalog.
- If two Glue databases from different catalogs share a name, one overwrites the other and its tables go missing.
databaseName unset. To choose which Glue databases to ingest, use schemaFilterPattern.
Filter Patterns
Each filter pattern insourceConfig matches the Glue name at its own level of this mapping. Patterns are case-insensitive regular expressions matched from the start of the name.
The following
sourceConfig applies these include patterns:
123456789012, wrap it in quotes. Otherwise, YAML reads it as a number, and the workflow rejects it because filter patterns must be strings.
If you set useFqnForFiltering: true, patterns match the fully qualified name instead. For example, the Glue database sales_raw in the catalog 123456789012 of the service local_glue has the fully qualified name local_glue.123456789012.sales_raw.
Metadata Ingestion
All connectors are defined as JSON Schemas. Here you can find the structure to create a connection to Glue. In order to create and run a Metadata Ingestion workflow, we will follow the steps to create a YAML configuration able to connect to the source, process the Entities if needed, and reach the OpenMetadata server. The workflow is modeled around the following JSON Schema1. Define the YAML Config
2. Run with the CLI
First, we will need to save the YAML file. Afterward, and with all requirements installed, we can run:dbt Integration
dbt Integration
Learn more about how to ingest dbt models