Course Outline
Introduction, Objectives, and Migration Strategy
- Defining course goals, aligning with participant profiles, and establishing success metrics
- Reviewing high-level migration methodologies and assessing associated risks
- Configuring workspaces, code repositories, and preparing lab datasets
Day 1 — Core Migration Concepts and Architecture
- Exploring Lakehouse principles, Delta Lake fundamentals, and the Databricks ecosystem
- Analyzing differences between SMP and MPP architectures and their impact on migration
- Understanding the Medallion (Bronze→Silver→Gold) design pattern and an introduction to Unity Catalog
Day 1 Practical Exercise — Converting a Stored Procedure
- Executing a hands-on migration of a sample stored procedure into a notebook
- Converting temporary tables and cursor-based logic into DataFrame transformations
- Validating results and comparing outputs against the original implementation
Day 2 — Advanced Delta Lake Features & Incremental Loads
- Utilizing ACID transactions, commit logs, version control, and time travel capabilities
- Implementing Auto Loader, MERGE INTO patterns, upserts, and managing schema evolution
- Optimizing performance using OPTIMIZE, VACUUM, Z-ORDER, partitioning, and storage tuning
Day 2 Practical Exercise — Incremental Data Ingestion & Tuning
- Setting up Auto Loader ingestion and establishing MERGE-based workflows
- Applying OPTIMIZE, Z-ORDER, and VACUUM commands while verifying outcomes
- Quantifying improvements in read and write performance
Day 3 — SQL in Databricks, Performance Analysis & Debugging
- Leveraging analytical SQL features: window functions, higher-order functions, and JSON/array manipulation
- Interpreting Spark UI metrics, including DAGs, shuffles, stages, tasks, and identifying bottlenecks
- Applying query tuning techniques: broadcast joins, hints, caching, and minimizing spill
Day 3 Practical Exercise — SQL Refinement & Performance Optimization
- Transforming resource-intensive SQL processes into optimized Spark SQL queries
- Utilizing Spark UI traces to detect and resolve data skew and shuffle inefficiencies
- Conducting before-and-after benchmarks and documenting tuning actions
Day 4 — Applied PySpark: Eliminating Procedural Logic
- Understanding the Spark execution model: driver, executors, lazy evaluation, and partitioning strategies
- Converting iterative loops and cursors into vectorized DataFrame operations
- Implementing modularization, UDFs/pandas UDFs, UI widgets, and reusable libraries
Day 4 Practical Exercise — Refactoring Procedural Scripts
- Rebuilding a procedural ETL script into modular PySpark notebooks
- Incorporating parameters, unit-style testing, and creating reusable functions
- Performing code reviews and applying industry best-practice checklists
Day 5 — Orchestration, Full Pipeline Integration & Best Practices
- Configuring Databricks Workflows: job design, task dependencies, triggers, and error management
- Designing incremental Medallion pipelines with integrated quality rules and schema validation
- Integrating with Git (GitHub/Azure DevOps), establishing CI pipelines, and testing PySpark logic
Day 5 Practical Exercise — Constructing a Complete End-to-End Pipeline
- Assembling a Bronze→Silver→Gold pipeline orchestrated via Workflows
- Implementing logging, auditing, retry mechanisms, and automated validation checks
- Executing the full pipeline, verifying outputs, and compiling deployment documentation
Operationalization, Governance, and Production Preparation
- Best practices for Unity Catalog governance, data lineage, and access control
- Managing costs, cluster sizing, autoscaling, and job concurrency patterns
- Creating deployment checklists, defining rollback strategies, and developing runbooks
Final Assessment, Knowledge Sharing, and Future Directions
- Participant presentations showcasing migration work and key takeaways
- Conducting gap analysis, suggesting follow-up activities, and distributing training materials
- Providing references, recommended learning paths, and support resources
Requirements
- Foundational knowledge of data engineering principles
- Proficiency with SQL and stored procedures (Synapse or SQL Server)
- Working knowledge of ETL orchestration concepts (Azure Data Factory or equivalent tools)
Target Audience
- Technology managers possessing a data engineering background
- Data engineers seeking to shift procedural OLAP logic to Lakehouse patterns
- Platform engineers tasked with overseeing Databricks adoption