Data Science

Data Engineering and DataOps Foundation

Overview

This 2-day instructor-led course is designed to provide participants with a solid foundation in Data Engineering and DataOps principles and practices. Participants will gain hands-on experience with essential tools and techniques used in the field of data engineering and learn how to implement DataOps methodologies to streamline data pipelines and ensure data quality.

By the end of this course, participants will:

  • Gain a comprehensive understanding of data engineering and DataOps
  • Build practical experience in using relevant tools and techniques
  • Have the knowledge to apply these skills and tools in real-world data projects
  • Be equipped to contribute effectively to data-driven organizations

Duration

2 Days

Who Should Take This Course

Audience

This course is designed for developers, data engineers, IT professionals, analysts, data scientists, DevOps professionals, MLOps professionals, and technical practitioners interested in modern data engineering and DataOps practices.

Prerequisites

Participants should have a basic understanding of data concepts and familiarity with at least one programming language such as Python, JavaScript, Java, C, or C++. Comfort working with command line interfaces is recommended. Participants should also be able to access cloud based lab environments through SSH during the course.

Why You Should Take This Course

In Data Engineering and DataOps Foundation, a comprehensive 2-day course, participants will gain a deep understanding of critical data engineering and DataOps concepts. They’ll start by grasping the role of data engineering, differentiating it from data science, and exploring key data concepts and real-world use cases.

Participants will delve into techniques for data ingestion, data transformation, data warehousing and more with hands-on experience at every step in the data lifecycle. Learn about DevOps principles for data management, including metrics and KPIs, and continuous integration and continuous deployment (CI/CD) for data pipelines.

Upon completion, participants will be well-prepared to contribute effectively to data-driven organizations, utilizing their newfound expertise in data engineering and DataOps principles and tools.

Course Outline

Data Engineering and DataOps Foundation

Day 1 Master Data Engineering Essentials

1. Introduction to Data Engineering

  • The Role of Data Engineering in Modern Data Ecosystems
  • Key Concepts: Data Ingestion, Transformation, and Storage
  • Data Engineering vs. Data Science: Understanding the Distinctions
  • Industry Use Cases and Success Stories
  • Lab: Set Up Your Development Environment for Success

2. Data Ingestion and Storage

  • Batch vs. Real-time Data Ingestion
  • Data Ingestion Patterns: Push vs. Pull
  • Scalable Data Storage Solutions: AWS S3, HDFS, and More
  • Lab: Hands-On Experience Building a Data Ingestion Pipeline with Apache Kafka

3. Data Transformation and Processing

  • Data Transformation Techniques: Filtering, Aggregating, Joining
  • Introduction to Apache Spark for Data Processing
  • Optimizing Data Processing Pipelines for Performance
  • Lab: Dive Deep into Data Transformation with Hands-On Exercises Using Apache Spark

4. Data Storage and Warehousing

  • Data Storage Options: Relational vs. NoSQL
  • Data Warehousing Concepts and Benefits
  • Setting Up Data Warehouses with Amazon Redshift
  • Data Modeling for Effective Querying and Reporting
  • Lab: Practical Guide to Designing Data Warehouses with Amazon Redshift

Day 2 Elevate Your DataOps Competency

5. Introduction to DataOps

  • Understanding the DataOps Lifecycle
  • Agile and DevOps Principles Applied to Data
  • Benefits of Implementing DataOps
  • Key Metrics and KPIs in DataOps
  • Lab: Hands-On Setup of an Efficient DataOps Environment

6. Continuous Integration and Continuous Deployment (CI/CD) for Data

  • CI/CD Fundamentals and Their Adaptation for Data Pipelines
  • Tools for Building CI/CD Pipelines: Jenkins, Travis CI, GitLab CI/CD
  • Automated Testing and Validation in Data CI/CD
  • Lab: Build Your Own CI/CD Pipeline for Data Using Jenkins

7. Data Quality and Monitoring

  • The Importance of Data Quality in DataOps
  • Data Quality Frameworks and Best Practices
  • Implementing Data Quality Checks with Apache Airflow
  • Real-time Data Monitoring Strategies and Tools
  • Lab: Hands-On Data Quality Checks Implementation Using Apache Airflow

8. Data Governance and Compliance

  • Data Governance Essentials: Policies, Standards, and Data Owners
  • Compliance Regulations and Their Implications
  • Hands-On Data Governance with Apache Atlas
  • Auditing, Compliance Reporting, and Data Lineage
  • Lab: Practical Enforcement of Data Governance Policies with Apache Atlas
Search UMBC Training Centers