Databricks and PySpark interview questions
The Databricks and PySpark questions interviewers ask most, with short answers you can explain in your own words. Tap a question to see the answer.
01What is the lakehouse?+
An architecture that combines data lake storage with warehouse features like ACID transactions and schema enforcement, using open formats like Delta.
02Driver vs executors?+
The driver plans the job and coordinates tasks. Executors run the tasks on partitions of data.
03What is lazy evaluation?+
Transformations build a plan and run only when an action like count, write or collect is called.
04Transformation vs action?+
Transformations (select, filter, join) return a new DataFrame lazily. Actions (count, show, write) trigger execution.
05Narrow vs wide transformations?+
Narrow ones work within a partition (filter, select). Wide ones need data from other partitions and cause a shuffle (groupBy, join).
06What is a shuffle and why is it expensive?+
Moving data between executors over the network and disk. It is often the slowest part of a job.
07What is a broadcast join?+
Sending a small table to every executor so the big table does not shuffle.
from pyspark.sql.functions import broadcast
df = big.join(broadcast(small), 'id')08repartition vs coalesce?+
repartition shuffles to any number of partitions. coalesce only reduces partitions without a full shuffle.
09What is data skew and how do you fix it?+
One key has far more rows, so one task runs much longer. Fix with AQE skew handling, filtering NULL keys, or salting.
10What is Adaptive Query Execution?+
Spark changes the plan at runtime: merges small partitions, switches join types and splits skewed partitions.
11What is Delta Lake?+
A storage layer on Parquet that adds ACID transactions, time travel, MERGE and schema enforcement.
12How do you upsert into Delta?+
With MERGE.
MERGE INTO silver.orders t USING updates s
ON t.order_id = s.order_id
WHEN MATCHED THEN UPDATE SET *
WHEN NOT MATCHED THEN INSERT *;13What do OPTIMIZE and VACUUM do?+
OPTIMIZE compacts small files. VACUUM removes old files no longer needed (after the retention period).
14What is the medallion architecture?+
Bronze (raw), silver (cleaned and conformed) and gold (business-ready) layers.
15What is Auto Loader?+
Incremental file ingestion that detects new files in cloud storage and processes each file once.
16What is Unity Catalog?+
Central governance for Databricks: catalogs, schemas, permissions, lineage and external locations across workspaces.
17Why avoid Python UDFs?+
Each row moves between the JVM and Python, which is slow. Built-in functions run inside Spark's engine.
18cache vs persist?+
cache keeps a DataFrame in memory (default level). persist lets you choose the storage level. Only cache data you reuse, and unpersist it.
19Job cluster vs all-purpose cluster?+
Job clusters start for a job and stop after, which is cheaper. All-purpose clusters are for interactive work.
20How do you debug a slow job?+
Open the Spark UI, find the longest stage, check spill and the gap between median and max task time, then fix skew, joins or file sizes.
More Databricks and PySpark interview questions
Free PDF downloads and premium packs with scenario questions and detailed model answers.
Coming soon
The premium Databricks and PySpark interview pack is being prepared.
🎤 Practise with a real mock interview
60 minutes live with Hikmat Ullah, plus written feedback. 30 USD, or 3 for 80 USD.
Azure Databricks Engineer
PySpark, Delta Lake, medallion architecture and Unity Catalog on Azure.
Learn it 1:1 → PROJECTSAzure Databricks Engineer projects
A free starter project, plus Small, Large and Enterprise projects.
See projects → MOREOther subjects
Free interview questions for all 12 subjects.
All interview questions →