Skip to main content
Use this guide to configure the Salesforce Data 360 pipeline connector. Configure and run Salesforce Data 360 workflows externally with YAML. Metadata, lineage, and Pipeline Status require separate workflows:

How to Run the Connector Externally

To run the Ingestion via the UI you’ll need to use the OpenMetadata Ingestion Container, which comes shipped with custom Airflow plugins to handle the workflow deployment. If, instead, you want to manage your workflows externally on your preferred orchestrator, you can check the following docs to run the Ingestion Framework anywhere.

External Schedulers

Get more information about running the Ingestion Framework Externally

Requirements

OpenMetadata connects to Salesforce Data 360 through the Salesforce REST API using a connected app that authenticates with the OAuth 2.0 client credentials flow. Configure a connected app for the client credentials flow. The app provides the Consumer Key and Consumer Secret. Grant the app the Manage Data Cloud and Access Data Cloud APIs scopes. Give the integration user permission to view the Data Streams, Calculated Insights, and Data Transforms you want to ingest. To resolve lineage, ingest the Salesforce Data 360 database first and point data360DbServiceName to that service.

Python Requirements

Use a Python version supported by the openmetadata-ingestion package that matches your OpenMetadata server. To find the supported range for your release, check the Requires-Python metadata of that release’s ingestion package. To run the Salesforce Data 360 pipeline ingestion, install:

Metadata Ingestion

All connectors are defined as JSON Schemas. See the Data360 pipeline connection JSON Schema for the structure used to create a Salesforce Data 360 connection. To create and run a metadata ingestion workflow, create a YAML configuration. The configuration connects to the source, processes entities if needed, and reaches the OpenMetadata server. The workflow is modeled around the following JSON Schema.

1. Define the YAML Config

This sample config runs the Metadata workflow for Salesforce Data 360. It creates pipeline definitions and tags, but doesn’t ingest lineage or Pipeline Status.

2. Configure the Lineage Workflow

Lineage is not part of the Metadata workflow. Run a separate workflow with source type data360pipeline-lineage. Use the same service name and connection values as the Metadata workflow. data360DbServiceName is required for this workflow. The workflow fails if it is unset because the connector cannot resolve Data 360 objects to OpenMetadata tables. includeBulkLineage also runs only in this Lineage workflow.

3. Configure the Pipeline Status Workflow

Pipeline Status is not part of the Metadata workflow. Run a separate Usage workflow with source type data360pipeline-usage. Despite its workflow name in the UI, this agent ingests pipeline run status. It doesn’t ingest query usage.

2. Run with the CLI

First, we will need to save the YAML file. Afterward, and with all requirements installed, we can run:
Note that from connector to connector, this recipe will always be the same. By updating the YAML configuration, you will be able to extract metadata from different sources.