dc dotCreds
Databricks Certified Data Engineer Associate Practice Test — A+ Source-Locked Hard Scenario Rebuild

Databricks Data Engineer Associate Practice Test

Start today’s free 10-question Databricks Data Engineer Associate set with source-backed explanations, local progress, and a fresh rotation every morning.

10 Free Daily Questions Source-backed Explanations 200 Verified Questions

Questions updated at Aug 23, 2026, 1:05 AM CDT

Go Pro - One Time Unlock

Unlock the full Databricks Certified Data Engineer Associate bank

200 verified questions Exam Mode Practice Mode Detailed explanations Weak-area review No subscription - one-time unlock

Get the complete source-backed bank with Interview Questions, the full Study Guide, full Course Notes, detailed explanations, weak-area review, and exam-style practice.

Interview Questions Full Study Guide Full Course Notes Exam Mode Practice Mode Guided Course Detailed explanations Weak-area review No subscription
$4.99 One-time payment
See bundle and PDF options

We will confirm your site email in one quick checkout step.

Why DotCreds?

Practice with explanations that teach.

Source links for every answer Every wrong answer explained Guided Course included Practice and Exam Mode Weak-area tracking Same verified bank across web practice

What you get with free practice

10 Free Questions Daily Fresh set every day from the live bank
Detailed Explanations Learn with clear source-backed answers
Track Your Progress Daily history and performance insights
Upgrade Anytime Unlock the full bank when you are ready
Today's 10 Databricks Data Engineer Associate questions

Use this Databricks Data Engineer Associate practice test to review Databricks Certified Data Engineer Associate. Questions rotate daily and each answer links back to the source used to write it.

Today’s Set
10 questions
Rotates at 10:00 AM local time
Progress
0/10
Answered on this page
Accuracy
0%
Loading countdown…

200 verified questions are in the live bank. Free daily questions are selected from a rotating sample set. Unlock Pro to access the full question bank.

Preparing today’s free questions... Ordering the final locked-bank set before showing the practice cards.
Question 1 of 10
Objective Use Unity Catalog ABAC policies to centrally control row filtering and column masking for sensitive data. Governance and Security

A team defines an ABAC policy based on a governed `sensitivity` tag but allows ordinary data editors to change that tag freely. Why is this dangerous?

Concept tested:
Question 2 of 10
Objective Understand Liquid Clustering and predictive optimization. Troubleshooting, Monitoring, and Optimization

A team changes the clustering columns on an existing liquid-clustered table and expects all old files to be physically reclustered immediately by the metadata change. Queries over historical data show little benefit. What is missing?

Concept tested:
Question 3 of 10
Objective Understand and build Gold-layer objects such as materialized views, views, streaming tables, and tables for BI and analytics in Unity Catalog. Data Transformation and Modeling

A clickstream source is append-oriented and should be incrementally processed into a continuously maintained Gold relation using streaming semantics. Which object is most aligned with that design?

Concept tested:
Question 4 of 10
Objective Configure Lakeflow Connect to reliably ingest data from diverse enterprise sources into Unity Catalog-governed tables. Data Ingestion and Loading

A managed connector supports the source, but the organization must apply a proprietary decryption library before records can be parsed. The managed connector exposes no hook for that library. Which design follows the Lakeflow Connect layering model?

Concept tested:
Question 5 of 10
Objective Configure notebook, SQL, dashboard, and pipeline tasks and their dependencies using the Lakeflow Jobs DAG. Working with Lakeflow Jobs

A team already has a managed Lakeflow pipeline that owns its source code, compute configuration, expectations, and target tables. They want a daily job to invoke one update and then continue to a downstream reconciliation task. What task type should they use?

Concept tested:
Question 6 of 10
Objective Deploy Declarative Automation Bundles to package, configure, and promote Lakeflow Jobs, pipelines, and workspace assets. Implementing CI/CD

Production deployments must enforce stronger safeguards than developer deployments, including branch validation and a non-human deployment identity. Which bundle capability should be considered?

Concept tested:
Question 7 of 10
Objective Understand Databricks Data Intelligence Platform compute services, including characteristics, limitations, cost models, and workload selection. Databricks Intelligence Platform

A platform owner is choosing between adding more manual cluster sizing controls and moving a supported workload to serverless. The business objective is to reduce operational burden, not to expose underlying nodes. Which characteristic should drive the decision?

Concept tested:
Question 8 of 10
Objective Implement data cleaning from bronze to silver with PySpark or SQL, including null handling and data type standardization. Data Transformation and Modeling

A bronze Delta table preserves raw point-of-sale records. In silver, `sale_ts` must be a timestamp, `amount` a decimal, and rows without `store_id` must not enter downstream finance metrics. A recent batch contains empty strings, malformed timestamps, and currency-formatted amount strings. Which transformation is the best silver-layer design?

Concept tested:
Question 9 of 10
Objective Differentiate managed and external tables and perform basic create, modify, delete, and conversion operations. Governance and Security

During conversion planning, one table is an external Delta table and another is a foreign table from a federated source. An engineer assumes both conversions use identical no-copy semantics. What must be corrected?

Concept tested:
Question 10 of 10
Objective Prioritize among Auto Loader, Lakeflow Connect, partner connectors, and other ingestion methods based on data volume, frequency, data type, and governance requirements. Data Ingestion and Loading

A monthly accounting feed contains a few hundred CSV files. Finance wants a SQL-centric load that is easy to retry and occasionally reprocess a selected subset of files. Which method is the better fit?

Concept tested:
Locked preview

You are viewing today’s free 10. Unlock 190 more questions.

Unlock full bank
Daily sample Rotating practice Free daily questions are selected from a rotating sample set.
Pro bank Full access Unlock Pro to access the full question bank, Exam Mode, Practice Mode, and random tests.
Databricks Certified Data Engineer Associate Pro $4.99 one-time

Unlock all 200 Databricks Data Engineer Associate questions, explanations, review tools, and exam-style practice.

50 Exam Practice Test $1.99 one-time

A 50-question Databricks Certified Data Engineer Associate PDF for short review sessions. Questions come first, then the answer review and explanations later in the file.

Linux / DevOps Access Bundle $6.99/month

Linux systems, Kubernetes, Terraform, data platform, and TensorFlow practice in one monthly unlock.

What’s includedCompTIA Linux+, CompTIA Cloud+, CompTIA CloudNetX, LFCS, CKA, Kubernetes CKS, Kubernetes CKAD, Terraform Associate, Databricks Data Engineer Associate, Databricks ML Associate, Databricks Generative AI Engineer Associate, TensorFlow Developer
AI / Machine Learning Access Bundle $6.99/month

AI, machine learning, MLOps, and generative AI practice in one monthly unlock.

What’s includedAWS AI Practitioner, AWS ML Engineer Associate, Google Generative AI Leader, Google ML Engineer, Databricks Data Engineer Associate, Databricks ML Associate, Databricks Generative AI Engineer Associate, IBM AI Engineering, NVIDIA GenAI LLM Associate, TensorFlow Developer, Stanford Machine Learning

Choose an unlock option to continue. We will confirm your site email in one quick checkout step.

Secure checkout powered by Stripe. Source-backed questions. Not brain dumps. Checkout stays on this page and unlocks the same Pro builder on this practice page.

Purchase options

Unlock the full Databricks Certified Data Engineer Associate bank.

Get the full bank, Exam Mode, Practice Mode, question sets, random tests, readiness tracking, saved box scores, and review tools for this exam.

The PDF versions keep questions first and move the answer review, explanations, and distractor notes to the back of the file.

200 verified exam-style questions Every choice explained Exam Mode and Practice Mode Question sets and random tests Readiness score and trends Previous test box scores

You've answered 0/10 questions in today's set.

Locked: 190 more questions in the full bank.

Locked: exam simulation mode, practice mode, readiness tracking, and saved review history.

Checkout stays on this page, so you can keep practicing, unlock the full bank, and start Exam Mode or Practice Mode when you are ready.

Cheat Sheets

7-day score keeper

Answer questions today and this will become a rolling 7-day scorecard.

Local history
Optional progress sync

Keep today’s practice moving

Guest progress saves automatically on this device. Add an email later when you want a magic link that keeps your daily Databricks Certified Data Engineer Associate practice in sync across browsers.

Guest progress saves on this device automatically

Guest progress is available without an account.

Source-backed answer review

The free daily Databricks Data Engineer Associate set includes crawlable question text, answer choices, correct answer labels, objective mapping, and source links. Only the first SEO card includes answer explanations and any extra learning features. Pro-only bank questions stay locked; this section mirrors only the 10 free daily questions already shown on this page.

Question 1 A team defines an ABAC policy based on a governed `sensitivity` tag but allows ordinary data editors to change that tag freely. Why is this dangerous?

Answer choices

  1. A. It is safe as long as editors lack MODIFY on table rows; tags cannot affect visibility.
  2. B. It is harmless because ABAC evaluates catalog names, not tags.
  3. C. It affects query performance because tags are used solely for clustering.
  4. D. Restrict and audit changes to governed tags that drive ABAC policy scope.

Correct answer

Restrict and audit changes to governed tags that drive ABAC policy scope.

ABAC derives policy application from governed attributes such as tags. If users can freely alter those attributes, they may change which security policy applies even without changing the data rows.

Wrong-answer review

  • A. It is safe as long as editors lack MODIFY on table rows; tags cannot affect visibility.: Changing policy-driving metadata can alter access independently of row-modification rights.
  • B. It is harmless because ABAC evaluates catalog names, not tags.: Governed tags are a core ABAC input, not decorative metadata.
  • C. It affects query performance because tags are used solely for clustering.: Tag-based ABAC controls security policy rather than physical data layout.

Extra learning features

Interview question

Q: You are using Unity Catalog ABAC to mask or filter sensitive data based on governed tags. How would you control who can change those tags, and why is that part of the security design rather than just metadata administration? Strong answer: The governed tags are inputs to ABAC policy evaluation, so changing a tag can change which policy applies to an object. I would restrict tag creation and assignment to authorized principals, use least privilege and separation of duties, and audit tag and policy changes. Ordinary data editors should not be able to remove or alter security-driving tags simply because they can edit data. Centralizing the policy at an appropriate catalog or schema scope also reduces per-table drift.

  • governed tags drive policy evaluation
  • restrict tag assignment
  • least privilege
  • separation of duties
  • audit tag changes
  • central policy scope

Caution: Do not describe ABAC policy code itself as preventing all tag changes. Security depends on permissions over governed tags and policy administration.

Objective/domain: Governance and Security

Source: Attribute-based access control in Unity Catalog

Question 2 A team changes the clustering columns on an existing liquid-clustered table and expects all old files to be physically reclustered immediately by the metadata change. Queries over historical data show little benefit. What is missing?

Answer choices

  1. A. Drop and recreate the Unity Catalog metastore because clustering metadata cannot be changed.
  2. B. Increase SQL warehouse concurrency because clustering changes are applied by query threads.
  3. C. Run `VACUUM` because VACUUM reclusters current data into the new key order.
  4. D. Run or allow `OPTIMIZE` to rewrite existing data under the new clustering layout

Correct answer

Run or allow `OPTIMIZE` to rewrite existing data under the new clustering layout

Objective/domain: Troubleshooting, Monitoring, and Optimization

Source: CLUSTER BY clause (TABLE)

Question 3 A clickstream source is append-oriented and should be incrementally processed into a continuously maintained Gold relation using streaming semantics. Which object is most aligned with that design?

Answer choices

  1. A. A materialized view chosen specifically because it uses batch rather than streaming semantics.
  2. B. A manually refreshed CSV export in workspace files.
  3. C. A streaming table maintained by a Lakeflow pipeline.
  4. D. A standard view whose query is recomputed from the entire source whenever queried.

Correct answer

A streaming table maintained by a Lakeflow pipeline.

Objective/domain: Data Transformation and Modeling

Source: Tables and views in Databricks

Question 4 A managed connector supports the source, but the organization must apply a proprietary decryption library before records can be parsed. The managed connector exposes no hook for that library. Which design follows the Lakeflow Connect layering model?

Answer choices

  1. A. Move to a more customizable standard connector or Structured Streaming pipeline where the required custom logic can run.
  2. B. Remain on the managed connector because managed connectors typically allow arbitrary runtime libraries.
  3. C. Use a larger SQL warehouse; compute size exposes managed-connector extension hooks.
  4. D. Disable Unity Catalog because governance prevents custom connector code.

Correct answer

Move to a more customizable standard connector or Structured Streaming pipeline where the required custom logic can run.

Objective/domain: Data Ingestion and Loading

Source: Standard connectors in Lakeflow Connect

Question 5 A team already has a managed Lakeflow pipeline that owns its source code, compute configuration, expectations, and target tables. They want a daily job to invoke one update and then continue to a downstream reconciliation task. What task type should they use?

Answer choices

  1. A. Use a Pipeline task that invokes the existing Lakeflow pipeline, then make reconciliation depend on it.
  2. B. Copy the pipeline's internal transformation code into a notebook task and maintain both versions.
  3. C. Use a SQL alert task as the pipeline executor.
  4. D. Use a Dashboard task because dashboards can initiate pipeline updates.

Correct answer

Use a Pipeline task that invokes the existing Lakeflow pipeline, then make reconciliation depend on it.

Objective/domain: Working with Lakeflow Jobs

Source: Pipeline task for jobs

Question 6 Production deployments must enforce stronger safeguards than developer deployments, including branch validation and a non-human deployment identity. Which bundle capability should be considered?

Answer choices

  1. A. Remove Git metadata from the bundle so branch state cannot affect deployment.
  2. B. Use a production-mode target with the appropriate safeguards and service-principal deployment identity.
  3. C. Run all production deployments from a developer's interactive notebook session.
  4. D. Use development mode in production because it relaxes validation and therefore prevents CI failures.

Correct answer

Use a production-mode target with the appropriate safeguards and service-principal deployment identity.

Objective/domain: Implementing CI/CD

Source: Declarative Automation Bundles deployment modes

Question 7 A platform owner is choosing between adding more manual cluster sizing controls and moving a supported workload to serverless. The business objective is to reduce operational burden, not to expose underlying nodes. Which characteristic should drive the decision?

Answer choices

  1. A. Serverless should be rejected because it cannot scale automatically.
  2. B. A SQL warehouse is mandatory for each workload that needs automatic scaling.
  3. C. Classic compute is preferable because serverless requires the team to choose worker instance types and autoscaling bounds.
  4. D. Serverless compute abstracts provisioning and scaling, so use it when the workload is supported

Correct answer

Serverless compute abstracts provisioning and scaling, so use it when the workload is supported

Objective/domain: Databricks Intelligence Platform

Source: Compute selection recommendations

Question 8 A bronze Delta table preserves raw point-of-sale records. In silver, `sale_ts` must be a timestamp, `amount` a decimal, and rows without `store_id` must not enter downstream finance metrics. A recent batch contains empty strings, malformed timestamps, and currency-formatted amount strings. Which transformation is the best silver-layer design?

Answer choices

  1. A. Load all columns as strings into silver and rely on Gold dashboards to interpret types at query time.
  2. B. Replace each null or failed cast with zero or the current timestamp so no rows are lost.
  3. C. Overwrite the bronze table after casting each column with Spark's inferred types so downstream jobs read one copy.
  4. D. Normalize empty strings to null, parse and cast the timestamp and amount with explicit rules, quarantine or reject rows missing required keys, and validate the resulting silver schema before promotion.

Correct answer

Normalize empty strings to null, parse and cast the timestamp and amount with explicit rules, quarantine or reject rows missing required keys, and validate the resulting silver schema before promotion.

Objective/domain: Data Transformation and Modeling

Source: PySpark functions

Question 9 During conversion planning, one table is an external Delta table and another is a foreign table from a federated source. An engineer assumes both conversions use identical no-copy semantics. What must be corrected?

Answer choices

  1. A. Neither table type can be converted to managed under any circumstances.
  2. B. Supported external-table conversion does not use MOVE/COPY, while foreign-table conversion requires the documented MOVE or COPY handling.
  3. C. Foreign-table conversion is metadata-only, while external Delta conversion typically deletes the source files.
  4. D. Both conversions typically require a full manual export to CSV first.

Correct answer

Supported external-table conversion does not use MOVE/COPY, while foreign-table conversion requires the documented MOVE or COPY handling.

Objective/domain: Governance and Security

Source: Convert external or foreign Delta Lake tables to Unity Catalog managed tables

Question 10 A monthly accounting feed contains a few hundred CSV files. Finance wants a SQL-centric load that is easy to retry and occasionally reprocess a selected subset of files. Which method is the better fit?

Answer choices

  1. A. COPY INTO.
  2. B. A managed SaaS connector for the file bucket.
  3. C. A continuous Kafka connector.
  4. D. Auto Loader is mandatory for any cloud file ingestion.

Correct answer

COPY INTO.

Objective/domain: Data Ingestion and Loading

Source: Ingest data from cloud object storage

Where to go after the daily web set

How are Databricks Data Engineer Associate questions generated?

dotCreds builds Databricks Data Engineer Associate practice questions from public exam objectives and Databricks exam and documentation references. The questions are written for realistic study practice, not copied from exam dumps.

How are explanations sourced?

Each question includes an explanation and, when available, a source link back to the provider documentation or reference used to validate the answer. That keeps the practice tied to study material you can actually review.

What score do I get?

The page tracks today's answered count and accuracy for the 10-question daily set, then saves a 7-day score history on this device so you can see your recent practice trend.

Why use this site?

The site is the fastest way to start Databricks Data Engineer Associate practice without installing anything. It is built for daily recall, quick weak-topic discovery, and source-backed explanations you can review immediately.