connectors

No menu items for this category

Home/Connectors/Metadata/Atlas

OpenMetadata Documentation

Atlas

PROD

Available In

Feature List

Metadata

Ingestion Deployment

To run the Ingestion via the UI you'll need to use the OpenMetadata Ingestion Container, which comes shipped with custom Airflow plugins to handle the workflow deployment. If you want to install it manually in an already existing Airflow host, you can follow this guide.

If you don't want to use the OpenMetadata Ingestion container to configure the workflows via the UI, then you can check the following docs to run the Ingestion Framework in any orchestrator externally.

Run Connectors from the OpenMetadata UI

Learn how to manage your deployment to run connectors from the UI

Run the Connector Externally

Get the YAML to run the ingestion externally

External Schedulers

Get more information about running the Ingestion Framework Externally

Requirements

Every table ingested will have a tag name AtlasMetadata.atlas_table. You can find all tags under Governance > Classification.

1. Create Database & Messaging Services

You need to create at least a Database Service before ingesting the metadata from Atlas. Make sure to note down the name, since we will use it to create Atlas Service.

For example, to create a Hive Service you can follow these steps:

1. Visit the Services Page

Click Settings in the side navigation bar and then Services.

The first step is to ingest the metadata from your sources. To do that, you first need to create a Service connection first.

This Service will be the bridge between OpenMetadata and your source system.

Once a Service is created, it can be used to configure your ingestion workflows.

Visit Services Page

Select your Service Type and Add a New Service

2. Create a New Service

Click on Add New Service to start the Service creation.

Create a new Service

Add a new Service from the Services page

3. Select the Service Type

Select Hive as the Service type and click Next.

Select Service

Select your Service from the list

4. Name and Describe your Service

Provide a name and description for your Service.

Service Name

OpenMetadata uniquely identifies Services by their Service Name. Provide a name that distinguishes your deployment from other Services, including the other Hive Services that you might be ingesting metadata from.

Note that when the name is set, it cannot be changed.

Add New Service

Provide a Name and description for your Service

5. Configure the Service Connection

In this step, we will configure the connection settings required for Hive.

Please follow the instructions below to properly configure the Service to read from your sources. You will also find helper documentation on the right-hand side panel in the UI.

Configure Service connection

Configure the Service connection by filling the form

2. Atlas Metadata Ingestion

Then, prepare the Atlas Service and configure the Ingestion:

1. Visit the Services Page

Click Settings in the side navigation bar and then Services.

The first step is to ingest the metadata from your sources. To do that, you first need to create a Service connection first.

This Service will be the bridge between OpenMetadata and your source system.

Once a Service is created, it can be used to configure your ingestion workflows.

Visit Services Page

Select your Service Type and Add a New Service

2. Create a New Service

Click on Add New Service to start the Service creation.

Create a new Service

Add a new Service from the Services page

3. Select the Service Type

Select Atlas as the Service type and click Next.

Select Service

Select your Service from the list

4. Name and Describe your Service

Provide a name and description for your Service.

Service Name

OpenMetadata uniquely identifies Services by their Service Name. Provide a name that distinguishes your deployment from other Services, including the other Atlas Services that you might be ingesting metadata from.

Note that when the name is set, it cannot be changed.

Add New Service

Provide a Name and description for your Service

5. Configure the Service Connection

In this step, we will configure the connection settings required for Atlas.

Please follow the instructions below to properly configure the Service to read from your sources. You will also find helper documentation on the right-hand side panel in the UI.

Configure Service connection

Configure the Service connection by filling the form

Connection Details

Host and Port: Host and port of the Atlas service.
Username: username to connect to the Atlas. This user should have privileges to read all the metadata in Atlas.
Password: password to connect to the Atlas.
databaseServiceName: source database of the data source. This is the service we created before: e.g., local_hive)
messagingServiceName: messaging service source of the data source.
Entity Type: Name of the entity type in Atlas.

6. Test the Connection

Once the credentials have been added, click on Test Connection and Save the changes.

Test Connection

Test the connection and save the Service

7. Configure Metadata Ingestion

In this step we will configure the metadata ingestion pipeline, Please follow the instructions below

Configure Metadata Ingestion

Configure Metadata Ingestion Page

8. Schedule the Ingestion and Deploy

Scheduling can be set up at an hourly, daily, weekly, or manual cadence. The timezone is in UTC. Select a Start Date to schedule for ingestion. It is optional to add an End Date.

Review your configuration settings. If they match what you intended, click Deploy to create the service and schedule metadata ingestion.

If something doesn't look right, click the Back button to return to the appropriate step and change the settings as needed.

After configuring the workflow, you can click on Deploy to create the pipeline.

Schedule the Workflow

Schedule the Ingestion Pipeline and Deploy

9. View the Ingestion Pipeline

Once the workflow has been successfully deployed, you can view the Ingestion Pipeline running from the Service Page.

View Ingestion Pipeline

View the Ingestion Pipeline from the Service Page