Get in Touch
 Duration 35 hours

Course Outline

Introduction, Objectives, and Migration Strategy

  • Defining course goals, aligning with participant profiles, and establishing success metrics
  • Reviewing high-level migration methodologies and assessing associated risks
  • Configuring workspaces, code repositories, and preparing lab datasets

Day 1 — Core Migration Concepts and Architecture

  • Exploring Lakehouse principles, Delta Lake fundamentals, and the Databricks ecosystem
  • Analyzing differences between SMP and MPP architectures and their impact on migration
  • Understanding the Medallion (Bronze→Silver→Gold) design pattern and an introduction to Unity Catalog

Day 1 Practical Exercise — Converting a Stored Procedure

  • Executing a hands-on migration of a sample stored procedure into a notebook
  • Converting temporary tables and cursor-based logic into DataFrame transformations
  • Validating results and comparing outputs against the original implementation

Day 2 — Advanced Delta Lake Features & Incremental Loads

  • Utilizing ACID transactions, commit logs, version control, and time travel capabilities
  • Implementing Auto Loader, MERGE INTO patterns, upserts, and managing schema evolution
  • Optimizing performance using OPTIMIZE, VACUUM, Z-ORDER, partitioning, and storage tuning

Day 2 Practical Exercise — Incremental Data Ingestion & Tuning

  • Setting up Auto Loader ingestion and establishing MERGE-based workflows
  • Applying OPTIMIZE, Z-ORDER, and VACUUM commands while verifying outcomes
  • Quantifying improvements in read and write performance

Day 3 — SQL in Databricks, Performance Analysis & Debugging

  • Leveraging analytical SQL features: window functions, higher-order functions, and JSON/array manipulation
  • Interpreting Spark UI metrics, including DAGs, shuffles, stages, tasks, and identifying bottlenecks
  • Applying query tuning techniques: broadcast joins, hints, caching, and minimizing spill

Day 3 Practical Exercise — SQL Refinement & Performance Optimization

  • Transforming resource-intensive SQL processes into optimized Spark SQL queries
  • Utilizing Spark UI traces to detect and resolve data skew and shuffle inefficiencies
  • Conducting before-and-after benchmarks and documenting tuning actions

Day 4 — Applied PySpark: Eliminating Procedural Logic

  • Understanding the Spark execution model: driver, executors, lazy evaluation, and partitioning strategies
  • Converting iterative loops and cursors into vectorized DataFrame operations
  • Implementing modularization, UDFs/pandas UDFs, UI widgets, and reusable libraries

Day 4 Practical Exercise — Refactoring Procedural Scripts

  • Rebuilding a procedural ETL script into modular PySpark notebooks
  • Incorporating parameters, unit-style testing, and creating reusable functions
  • Performing code reviews and applying industry best-practice checklists

Day 5 — Orchestration, Full Pipeline Integration & Best Practices

  • Configuring Databricks Workflows: job design, task dependencies, triggers, and error management
  • Designing incremental Medallion pipelines with integrated quality rules and schema validation
  • Integrating with Git (GitHub/Azure DevOps), establishing CI pipelines, and testing PySpark logic

Day 5 Practical Exercise — Constructing a Complete End-to-End Pipeline

  • Assembling a Bronze→Silver→Gold pipeline orchestrated via Workflows
  • Implementing logging, auditing, retry mechanisms, and automated validation checks
  • Executing the full pipeline, verifying outputs, and compiling deployment documentation

Operationalization, Governance, and Production Preparation

  • Best practices for Unity Catalog governance, data lineage, and access control
  • Managing costs, cluster sizing, autoscaling, and job concurrency patterns
  • Creating deployment checklists, defining rollback strategies, and developing runbooks

Final Assessment, Knowledge Sharing, and Future Directions

  • Participant presentations showcasing migration work and key takeaways
  • Conducting gap analysis, suggesting follow-up activities, and distributing training materials
  • Providing references, recommended learning paths, and support resources

Requirements

  • Foundational knowledge of data engineering principles
  • Proficiency with SQL and stored procedures (Synapse or SQL Server)
  • Working knowledge of ETL orchestration concepts (Azure Data Factory or equivalent tools)

Target Audience

  • Technology managers possessing a data engineering background
  • Data engineers seeking to shift procedural OLAP logic to Lakehouse patterns
  • Platform engineers tasked with overseeing Databricks adoption

Number of participants


Price per participant

Upcoming Courses

Related Categories