Get in Touch

Course Outline

Introduction:

  • Apache Spark within the Hadoop ecosystem
  • Brief overview of Python and Scala

Foundational Concepts (Theory):

  • System architecture
  • RDDs
  • Transformations and Actions
  • Stages, Tasks, and Dependencies

Exploring fundamentals in a Databricks environment (Hands-on Workshop):

  • Practical exercises with the RDD API
  • Essential action and transformation functions
  • PairRDDs
  • Join operations
  • Caching strategies
  • Practical exercises with the DataFrame API
  • SparkSQL
  • DataFrame operations: select, filter, group, and sort
  • UDFs (User-Defined Functions)
  • An introduction to the Dataset API
  • Streaming capabilities

Understanding deployment in an AWS environment (Hands-on Workshop):

  • Foundations of AWS Glue
  • Key differences between AWS EMR and AWS Glue
  • Sample job implementations in both environments
  • Evaluation of advantages and trade-offs

Additional Topics:

  • Overview of Apache Airflow orchestration

Requirements

Programming proficiency (Python and Scala recommended)

Fundamental knowledge of SQL

 21 Hours

Number of participants


Price per participant

Testimonials (3)

Upcoming Courses

Related Categories