FREE · 20 QUESTIONS WITH ANSWERS

Azure Data Engineering interview questions

The Azure Data Engineering questions interviewers ask most, with short answers you can explain in your own words. Tap a question to see the answer.

01What is Azure Data Factory?

A cloud service to orchestrate and move data with pipelines, activities, linked services and datasets.

02Linked service vs dataset?

A linked service is the connection (like a connection string). A dataset points to specific data on that connection, like a table or folder.

03What is an integration runtime?

The compute ADF uses: Azure IR for cloud data, self-hosted IR for on-premises or private networks.

04How do you build a metadata-driven pipeline?

Keep a control table of sources and targets, read it with Lookup, loop with ForEach and pass parameters to a generic Copy activity.

05How do you do incremental loads?

Store a watermark (last modified date or ID), load only newer rows, then update the watermark.

06Types of triggers in ADF?

Schedule, tumbling window (with dependencies and backfill), storage event and custom event triggers.

07What is ADLS Gen2?

Azure Blob Storage with a hierarchical namespace, built for analytics with folder-level ACLs.

08RBAC vs ACLs in ADLS?

RBAC grants access at account or container level. ACLs give fine-grained access on folders and files.

09How should a data lake be organised?

Zones like raw/bronze, curated/silver and serving/gold, partitioned by date where it helps, using Parquet or Delta.

10What is a managed identity?

An identity Azure manages for a service, so it can access other resources without storing passwords.

11Why use Key Vault?

To store secrets, keys and connection strings securely and reference them from ADF and Databricks.

12Synapse serverless vs dedicated SQL pool?

Serverless queries files in the lake and you pay per data scanned. Dedicated is a provisioned warehouse you pay for while it runs.

13What is Event Hubs?

A streaming ingestion service for millions of events per second, often read by Stream Analytics or Databricks.

14How do you run Databricks from ADF?

Use the Databricks Notebook or Job activity with a linked service, passing parameters and using job clusters.

15How do you handle failures in ADF?

Retries on activities, failure paths with alerts, logging run details to a table and Azure Monitor alerts.

16What is Slowly Changing Dimension Type 2?

Keeping history by adding a new row for each change, with valid_from, valid_to and an is_current flag.

17How does CI/CD work for ADF?

Git integration for development, publish creates ARM templates, and Azure DevOps or GitHub Actions deploy them to test and prod with parameters.

18Copy activity performance tips?

Use parallel copies, partition the source, choose the right integration runtime and avoid many tiny files.

19Parquet vs CSV?

Parquet is columnar, compressed and has a schema, so it is much faster and cheaper for analytics. CSV is simple but slow.

20Describe an end-to-end Azure pipeline.

ADF ingests from sources into ADLS bronze, Databricks cleans to silver and builds gold, Synapse or Fabric serves data, Power BI reports, with Key Vault, monitoring and CI/CD around it.

More Azure Data Engineering interview questions

Free PDF downloads and premium packs with scenario questions and detailed model answers.

Coming soon

The premium Azure Data Engineering interview pack is being prepared.

🎤 Practise with a real mock interview

60 minutes live with Hikmat Ullah, plus written feedback. 30 USD, or 3 for 80 USD.

Book a mock interview