DataHub: A Generalized Metadata Search & Discovery Tool
📣 Next DataHub town hall meeting on March 20th, 9am-10am PST:
- Signup sheet & questions
- Details and recordings of past meetings can be found here
✨Mar 2020 Update:
- We're on Slack now! Ask questions and keep up with the latest announcement.
Introduction
DataHub is LinkedIn's generalized metadata search & discovery tool. To learn more about DataHub, check out our LinkedIn blog post and Strata presentation. You should also visit DataHub Architecture to get a better understanding of how DataHub is implemented and DataHub Onboarding Guide to understand how to extend DataHub for your own use case.
This repository contains the complete source code for both DataHub's frontend & backend. You can also read about how we sync the changes between our the internal fork and GitHub.
Quickstart
- Install docker and docker-compose. Make sure to configure Docker to allocate enough hardware resources for Docker engine. Tested & confirmed config: 4 CPUs, 8GB RAM, 2GB Swap area.
- Open Docker either from the command line or the Desktop app and ensure it is up and running.
- Clone this repo and cdinto the root directory for the cloned repository.
- Run below command to download and run all Docker containers in your local:
 This step takes long time and it might be hard to figure out when DataHub is fully up. You can refer to this guide to verify if DataHub is up and running.cd docker/quickstart && docker-compose pull && docker-compose up --build
- At this point, you should be able to start DataHubby opening http://localhost:9001 in your browser. You can sign in usingdatahubas both username and password. However, there is no data just yet.
- To ingest provided sample data to DataHub, switch to a new terminal, cdinto the cloneddatahubrepo, and run below command:
 After running this, you should be able to see sample data in DataHub.docker build -t ingestion -f docker/ingestion/Dockerfile . && cd docker/ingestion && docker-compose up
Refer to debugging guide if you have issues in any of the above steps.
Documents
- DataHub Architecture
- DataHub Onboarding Guide
- Docker Images
- Frontend
- Web App
- Generalized Metadata Service
- Metadata Ingestion
- Metadata Processing Jobs
Releases
See Releases page for more details.
Features & Roadmap
Check out DataHub's Features & Roadmap.
Contributing
We welcome contributions from the community. Please refer to the guidelines for more details. We also have a contrib directory for incubation.
Community
Join our slack channel for important discussions and announcements. You can also find out more about our past and upcoming town hall meetings.
FAQs
Frequently Asked Questions about DataHub can be found here.
Related Articles & Presentations
- DataHub: A Generalized Metadata Search & Discovery Tool
- Open sourcing DataHub: LinkedIn’s metadata search and discovery platform
- The evolution of metadata: LinkedIn’s story @ Strata Data Conference 2019
- Journey of metadata at LinkedIn @ Crunch Data Conference 2019
- Data Catalogue — Knowing your data
- How LinkedIn, Uber, Lyft, Airbnb and Netflix are Solving Data Management and Discovery for Machine Learning Solutions
- LinkedIn元数据之旅的最新进展—Data Hub
- 数据治理篇: 元数据之datahub-概述
- LinkedIn gibt die Datenplattform DataHub als Open Source frei
- Linkedin bringt Open-Source-Datahub
 
			