Skip to main content
This page is about running the Ingestion Framework externally!There are mainly 2 ways of running the ingestion:
  1. Internally, by managing the workflows from OpenMetadata.
  2. Externally, by using any other tool capable of running Python code.
If you are looking for how to manage the ingestion process from OpenMetadata, you can follow this doc.

Run the ingestion from GCP Composer

Requirements

OpenMetadata 2.0.3 ingestion requires Python 3.10 or later and SQLAlchemy 2.x. Use openmetadata-ingestion==2.0.3.0 with an OpenMetadata 2.0.3 server. For Composer environments whose managed dependencies conflict with these requirements, use the Kubernetes Pod Operator to run the ingestion image separately. The earlier Composer 2.5.4 / Airflow 2.6.3 example is not a verified compatibility combination for OpenMetadata 2.0.3.

Using the Python Operator

Use PythonOperator only when the Composer environment supports the ingestion package’s Python and dependency requirements. This approach installs the package directly into the managed environment. If Composer requires SQLAlchemy 1.x, use the Kubernetes Pod Operator approach instead.

Install the Requirements

In your environment, install openmetadata-ingestion[<plugins>]==x.y.z. Let the ingestion package resolve its SQLAlchemy dependency (>=2.0.0,<3). Do not pin SQLAlchemy to 1.4.27. That conflicts with OpenMetadata 2.0.3. Check the complete dependency set against your Composer environment before installing it. Replace x.y.z with the ingestion package version that matches your server. For an OpenMetadata 2.0.3 server, install ingestion package version 2.0.3.0. The plugin parameter lists the sources to ingest. For example, use openmetadata-ingestion[mysql,snowflake,s3]==2.0.3.0.

Prepare the DAG!

Note that this DAG is a usual connector DAG, just using the Airflow service with the Backend connection. As an example of a DAG pushing data to OpenMetadata under Google SSO, we could have:

Ingestion Workflow classes

We have different classes for different types of workflows. The logic is always the same, but you will need to change your import path. The rest of the method calls will remain the same. For example, for the Metadata workflow we’ll use:
The classes for each workflow type are:
  • Metadata: from metadata.workflow.metadata import MetadataWorkflow
  • Lineage: from metadata.workflow.metadata import MetadataWorkflow (same as metadata)
  • Usage: from metadata.workflow.usage import UsageWorkflow
  • dbt: from metadata.workflow.metadata import MetadataWorkflow
  • Profiler: from metadata.workflow.profiler import ProfilerWorkflow
  • Data Quality: from metadata.workflow.data_quality import TestSuiteWorkflow
  • Data Insights: from metadata.workflow.data_insight import DataInsightWorkflow
  • Elasticsearch Reindex: from metadata.workflow.metadata import MetadataWorkflow (same as metadata)

Using the Kubernetes Pod Operator

In this second approach we won’t need to install absolutely anything to the GCP Composer environment. Instead, we will rely on the KubernetesPodOperator to use the underlying k8s cluster of Composer. Then, the code won’t directly run using the hosts’ environment, but rather inside a container that we created with only the openmetadata-ingestion package. Note: This approach only has the openmetadata/ingestion-base ready from version 0.12.1 or higher!

Prepare the DAG!

Some remarks on this example code:

Kubernetes Pod Operator

You can name the task as you want (task_id and name). The important points here are the cmds, this should not be changed, and the env_vars. The main.py script that gets shipped within the image will load the env vars as they are shown, so only modify the content of the config YAML, but not this dictionary. Note that the example uses the image openmetadata/ingestion-base:2.0.3. The image version should be aligned with your OpenMetadata server version to avoid incompatibilities.
You can find more information about the KubernetesPodOperator and how to tune its configurations here. Note that depending on the kind of workflow you will be deploying, the YAML configuration will need to updated following the official OpenMetadata docs, and the value of the pipelineType configuration will need to hold one of the following values:
  • metadata
  • usage
  • lineage
  • profiler
  • TestSuite
Which are based on the PipelineType JSON Schema definitions