Handle data at scale. Distributed storage, MapReduce, Spark, and parallel processing for datasets too big for one machine.
BeginnerAbout 4 hours12 lessons
What you will learn
What makes data big, the 5 Vs
Batch vs streaming
When you actually need Big Data tools
Distributed file systems and HDFS
Data lakes vs data warehouses
MapReduce, the core idea
Hadoop in a nutshell
Apache Spark basics
Curriculum
1
Big Data concepts
3 lessonsFree preview
What makes data big, the 5 VsPreview7m
Batch vs streamingPreview7m
When you actually need Big Data toolsPreview6m
2
Storing data at scale
2 lessons
Distributed file systems and HDFS9m
Data lakes vs data warehouses8m
3
Processing frameworks
4 lessons
MapReduce, the core idea9m
Hadoop in a nutshell8m
Apache Spark basics10m
Spark DataFrames, hands-on10m
4
Parallel processing
2 lessons
Parallelism and why it scales8m
MPI and MapReduce patterns8m
5
Hands-on checkpoint
1 lessons
Checkpoint: a MapReduce word count15m
Hands-on checkpoint
Every course ends with a Colab task you complete yourself. Our Gemini-assisted review gives feedback and points you to what to fix, it coaches you, it does not do it for you.