- Introduction to Batch Processing
- Introduction to Spark
- Installing Spark
- First Look at Spark/PySpark
- Spark DataFrames
- Preparing Yellow and Green Taxi Data
- SQL with Spark
- Anatomy of a Spark Cluster
- GroupBy in Spark
- Joins in Spark
- Operations on Spark RDDs
- Spark RDD mapPartition
- Connecting to Google Cloud Storage
- Creating a Local Spark Cluster
- Setting up a Dataproc Cluster
- Connecting Spark to BigQuery
Units 11 and 12 cover RDDs, which the module marks optional, as is the taxi-data preparation in unit 6 and the guided Linux install in unit 3.
Did you take notes? You can share them here
- Notes by Alvaro Navas
- Sandy's DE Learning Blog
- Notes by Alain Boisvert
- Alternative : Using docker-compose to launch spark by rafik
- Marcos Torregrosa's blog (spanish)
- Notes by Victor Padilha
- Notes by Oscar Garcia
- Notes by HongWei
- 2024 videos transcript by Maria Fisher
- 2025 Notes by Manuel Guerra
- 2025 Notes by Gabi Fonseca
- 2025 Notes on Installing Spark on MacOS (with Anaconda + brew) by Gabi Fonseca
- 2025 Notes by Daniel Lachner
- 2026 Notes by Ajay Katte
- 2026 Video PySpark installation on Windows (no pip install) by khanh
- Add your notes here (above this line)