Learn Data Engineering
A free, structured path from SQL fundamentals to production Spark, Kafka, and lakehouse systems — built from real engineering experience.
Foundations
What data engineering is, the modern data stack, and how the pieces fit together.
SQL for Data Engineering
The language every data engineer lives in — joins, windows, and the query patterns that show up in every pipeline.
Python for Data Engineering
From pandas to your first ETL script, with the libraries and patterns you'll actually reach for.
- pandas Essentials for Data EngineersRead →
- Reading Files, APIs & JSON in PythonRead →
- Build Your First ETL Pipeline in PythonRead →PractiseSort by two keys in opposite directionsEasy · 8mFix the window that drops midnightMedium · 10mFix the partition key that trusts the stringMedium · 10mConvert currencies and total by countryMedium · 12mFix the load that doubles on retryMedium · 12mTop 3 from a stream you can read onceHard · 14m
Data Modeling & Warehousing
How to structure data so it stays fast, correct, and cheap to query.
Batch Processing With Spark
PySpark from fundamentals to the tuning that keeps production jobs fast.
Streaming With Kafka
Event streaming fundamentals and the failure modes that bite in production.
Lakehouse & Table Formats
Open table formats that bring warehouse reliability to the data lake.
Orchestration & Transformation
Scheduling, dependencies, and turning raw data into trusted models.
Cloud & Capstone
Run it all on the cloud and build a real end-to-end pipeline.
Comments
0Join the conversation. Sign in to leave a comment — we'd love to hear your thoughts.
