Retail medallion lakehouse
Raw sales files on ADLS Gen2 become clean silver tables and business ready gold tables, loaded incrementally.
PySpark, Delta Lake, medallion architecture and Unity Catalog on Azure.
from delta.tables import DeltaTablefrom pyspark.sql import functions as F bronze = spark.read.table("bronze.orders")silver = (bronze.dropDuplicates(["order_id"]) .withColumn("order_date", F.to_date("order_ts"))) (DeltaTable.forName(spark, "silver.orders").alias("t") .merge(silver.alias("s"), "t.order_id = s.order_id") .whenMatchedUpdateAll() .whenNotMatchedInsertAll() .execute())
Databricks is where large scale data processing happens on Azure. This course takes you from Spark basics to production lakehouse pipelines: PySpark transformations, Delta Lake, bronze, silver and gold layers, Auto Loader, Unity Catalog, Workflows and performance tuning. Everything is built on Azure with ADLS Gen2, the same setup I use at work.
You learn by doing: every week has hands-on exercises on Google Classroom, and every module ends with something you built yourself.
Drivers, executors, partitions, lazy evaluation and how jobs really run.
Joins, aggregations, window functions and clean reusable code.
MERGE, time travel, OPTIMIZE, VACUUM and schema evolution.
Bronze, silver and gold layers with incremental loads.
Catalogs, external locations, permissions and lineage.
Spark UI, shuffles, broadcast joins, Workflows and CI/CD.
A clear plan from the first session to the last. The pace adapts to you: if you already know a topic we move faster, and if something needs more time we take it.
Real projects you can put on GitHub and explain with confidence in interviews.
Raw sales files on ADLS Gen2 become clean silver tables and business ready gold tables, loaded incrementally.
Ingest device events continuously, handle late data and publish near real-time aggregates.
Take a slow, expensive job and make it faster and cheaper, with before and after numbers from the Spark UI.
The same tools used by data teams in real companies.
Every session is just you and me, never a batch. No recordings, so we can talk openly about your code.
Each session is 1 hour 15 minutes, on weekdays, weekends or both. We agree the schedule before starting.
Assignments, projects, quizzes, interview questions and homework, all organised in one place.
You build projects like the ones companies run, and get feedback on your code and design.
Common questions, live coding practice and mock interviews with honest feedback.
CV, LinkedIn, applications and interviews. I stay with you until you are hired, as long as you do the work.
You pay for your time with me, not per course. All prices are in US dollars.
Senior Data Engineer at Algo · Lahore, Pakistan
I build data platforms with Snowflake, Azure Databricks, Azure Data Factory and Microsoft Fabric every day, and I have been mentoring since 2021. I teach what I use at work, and I stay with you until you land the job.
Yes, a free or pay as you go Azure account is enough. I show you how to keep costs very low with small clusters, auto termination and job compute.
No. We build PySpark skills step by step, and most of it reads like SQL. Basic Python is enough to start.
Yes. Many teams run Databricks next to Snowflake or Fabric. I work with Databricks and Snowflake together every day.
200 USD per month whichever course you choose. You can also pay 100 USD every 15 days or 50 USD per week. All prices are in US dollars wherever you live.
If the training is not right for you after the first session, you get a full refund. After that, payments already used for completed sessions are not refunded, but unused prepaid time can be refunded.
Yes. As long as you attend your sessions and complete the assignments and projects, I keep helping you with CV, LinkedIn, mock interviews and job applications until you are hired.
Yes. Students join from many countries. We pick session times that work for your time zone before we start.
Python for data work: pandas, APIs, automation, testing and real ETL.
View course → 6 monthsThe complete path: SQL, Python, ADF, ADLS Gen2, Databricks, Synapse and Event Hubs.
View course → 3 monthsWarehousing, Snowpipe, streams and tasks, Snowpark and real cost control.
View course →Tell me your goals and your time zone. We plan your schedule together, and your first session can start this week.