dc dotCreds
Databricks Certified Data Engineer Associate Practice Test — A+ Source-Locked Hard Scenario Rebuild

Databricks Data Engineer Associate Practice Test

Start today’s free 10-question Databricks Data Engineer Associate set with source-backed explanations, local progress, and a fresh rotation every morning.

10 Free Daily Questions Source-backed Explanations 200 Verified Questions

Questions updated at Aug 23, 2026, 8:12 PM CDT

Go Pro - One Time Unlock

Unlock the full Databricks Certified Data Engineer Associate bank

200 verified questions Exam Mode Practice Mode Detailed explanations Weak-area review No subscription - one-time unlock

Get the complete source-backed bank with Interview Questions, the full Study Guide, full Course Notes, detailed explanations, weak-area review, and exam-style practice.

Interview Questions Full Study Guide Full Course Notes Exam Mode Practice Mode Guided Course Detailed explanations Weak-area review No subscription
$4.99 One-time payment
See bundle and PDF options

We will confirm your site email in one quick checkout step.

Why DotCreds?

Practice with explanations that teach.

Source links for every answer Every wrong answer explained Guided Course included Practice and Exam Mode Weak-area tracking Same verified bank across web practice

What you get with free practice

10 Free Questions Daily Fresh set every day from the live bank
Detailed Explanations Learn with clear source-backed answers
Track Your Progress Daily history and performance insights
Upgrade Anytime Unlock the full bank when you are ready
Today's 10 Databricks Data Engineer Associate questions

Use this Databricks Data Engineer Associate practice test to review Databricks Certified Data Engineer Associate. Questions rotate daily and each answer links back to the source used to write it.

Today’s Set
10 questions
Rotates at 10:00 AM local time
Progress
0/10
Answered on this page
Accuracy
0%
Loading countdown…

200 verified questions are in the live bank. Free daily questions are selected from a rotating sample set. Unlock Pro to access the full question bank.

Preparing today’s free questions... Ordering the final locked-bank set before showing the practice cards.
Question 1 of 10
Objective Understand the core components of the Databricks Data Intelligence Platform, such as its architecture, Delta Lake, and Unity Catalog. Databricks Intelligence Platform

A regulated enterprise has three Databricks workspaces in the same AWS region. Security wants one governance view of tables and models across the workspaces, while engineering wants each workspace to retain its own development experience and compute choices. Which design best matches Databricks architecture?

Concept tested:
Question 2 of 10
Objective Ingest semi-structured and unstructured data into Unity Catalog-governed Delta tables using managed or Lakeflow ingestion patterns. Data Ingestion and Loading

A Lakeflow managed connector lands nested source records into Unity Catalog. Downstream Silver processing needs selected nested attributes as relational columns while preserving the original Bronze data for replay. Which layer responsibility is correct?

Concept tested:
Question 3 of 10
Objective Differentiate managed and external tables and perform basic create, modify, delete, and conversion operations. Governance and Security

A Unity Catalog external Delta table is no longer needed in Databricks, but an external application must continue using the files at the cloud storage path. What happens when the team drops the external table metadata?

Concept tested:
Question 4 of 10
Objective Deploy Declarative Automation Bundles to package, configure, and promote Lakeflow Jobs, pipelines, and workspace assets. Implementing CI/CD

A reviewer says Databricks Asset Bundles and Declarative Automation Bundles are unrelated products, so existing bundle-based CI/CD must be redesigned from scratch. What is the correct clarification?

Concept tested:
Question 5 of 10
Objective Configure notebook, SQL, dashboard, and pipeline tasks and their dependencies using the Lakeflow Jobs DAG. Working with Lakeflow Jobs

A finance validation consists entirely of a governed SQL file and must run on Databricks SQL after the transformation task. The team wants the Jobs UI to show it as a distinct unit of work. Which task type is the best fit?

Concept tested:
Question 6 of 10
Objective Understand the core components of the Databricks Data Intelligence Platform, such as its architecture, Delta Lake, and Unity Catalog. Databricks Intelligence Platform

Two streaming writers update a Delta table while analysts run long queries. The data team needs transactional consistency without maintaining separate batch and streaming copies. Which platform component provides the foundational storage behavior?

Concept tested:
Question 7 of 10
Objective Understand Liquid Clustering and predictive optimization. Troubleshooting, Monitoring, and Optimization

Analysts filter a large table mostly by `customer_id` and `event_date`, but an engineer proposes clustering by a random UUID because it has the highest cardinality. What is the better principle?

Concept tested:
Question 8 of 10
Objective Perform deduplication and aggregate operations on DataFrames, including count, approximate distinct count, mean, and summary. Data Transformation and Modeling

A revenue DataFrame has grain `order_line`. A developer groups only by `region` and sums `order_total`, but `order_total` is repeated on every line of an order. The reported revenue is inflated. What should be fixed first?

Concept tested:
Question 9 of 10
Objective Use Auto Loader with schema enforcement, schema evolution, and file discovery modes to land data into Unity Catalog-governed tables. Data Ingestion and Loading

Two Auto Loader streams read two different source directories but write to the same curated target. An engineer proposes sharing one streaming checkpoint to simplify operations. What is the correct response?

Concept tested:
Question 10 of 10
Objective Use Unity Catalog ABAC policies to centrally control row filtering and column masking for sensitive data. Governance and Security

A new table is tagged so an existing ABAC row-filter policy will apply. The data owner wants to deploy the tag directly in production without testing because “the policy code did not change.” What is the security concern?

Concept tested:
Locked preview

You are viewing today’s free 10. Unlock 190 more questions.

Unlock full bank
Daily sample Rotating practice Free daily questions are selected from a rotating sample set.
Pro bank Full access Unlock Pro to access the full question bank, Exam Mode, Practice Mode, and random tests.
Databricks Certified Data Engineer Associate Pro $4.99 one-time

50 Exam Practice Test $1.99 one-time

A 50-question Databricks Certified Data Engineer Associate PDF for short review sessions. Questions come first, then the answer review and explanations later in the file.

Databricks Bundle $9.99 one-time

Unlock all 3 active Databricks Bundle practice banks in one permanent purchase.

What’s includedData Engineer Associate, ML Associate, Generative AI Engineer Associate
All Access $6.99/month

Unlock every active practice exam, bundle and path experience, Pro course and study content, and included downloads.

What’s includedEvery current and future active practice exam, All active bundle and career-path practice content, Pro course lessons, study content, and supported paid downloads

Choose an unlock option to continue. We will confirm your site email in one quick checkout step.

Secure checkout powered by Stripe. Source-backed questions. Not brain dumps. Checkout stays on this page and unlocks the same Pro builder on this practice page.

Purchase options

Unlock the full Databricks Certified Data Engineer Associate bank.

Get the full bank, Exam Mode, Practice Mode, question sets, random tests, readiness tracking, saved box scores, and review tools for this exam.

The PDF versions keep questions first and move the answer review, explanations, and distractor notes to the back of the file.

200 verified exam-style questions Every choice explained Exam Mode and Practice Mode Question sets and random tests Readiness score and trends Previous test box scores

You've answered 0/10 questions in today's set.

Locked: 190 more questions in the full bank.

Locked: exam simulation mode, practice mode, readiness tracking, and saved review history.

Checkout stays on this page, so you can keep practicing, unlock the full bank, and start Exam Mode or Practice Mode when you are ready.

Cheat Sheets

7-day score keeper

Answer questions today and this will become a rolling 7-day scorecard.

Local history
Optional progress sync

Keep today’s practice moving

Guest progress saves automatically on this device. Add an email later when you want a magic link that keeps your daily Databricks Certified Data Engineer Associate practice in sync across browsers.

Guest progress saves on this device automatically

Guest progress is available without an account.

Source-backed answer review

The free daily Databricks Data Engineer Associate set includes crawlable question text, answer choices, correct answer labels, objective mapping, and source links. Only the first SEO card includes answer explanations and any extra learning features. Pro-only bank questions stay locked; this section mirrors only the 10 free daily questions already shown on this page.

Question 1 A regulated enterprise has three Databricks workspaces in the same AWS region. Security wants one governance view of tables and models across the workspaces, while engineering wants each workspace to retain its own development experience and compute choices. Which design best matches Databricks architecture?

Answer choices

  1. A. Create a separate metastore per notebook so lineage and privileges remain isolated at the code level, as the proposed databricks intelligence platform approach.
  2. B. Attach the workspaces to the same Unity Catalog metastore so governance is centralized while compute and workspace collaboration remain separate.
  3. C. Place all classic compute in the Databricks control plane so the metastore can see the data, within the documented operational, security, ownership, and validation requirements.
  4. D. Merge the three workspaces into one because a Unity Catalog metastore cannot govern more than one workspace, under the documented operational and governance requirements.

Correct answer

Attach the workspaces to the same Unity Catalog metastore so governance is centralized while compute and workspace collaboration remain separate.

Unity Catalog metastores are the central governance system and can be attached to multiple workspaces in the same region. Workspaces remain collaboration and execution environments, while governance can be shared.

Wrong-answer review

  • A. Create a separate metastore per notebook so lineage and privileges remain isolated at the code level, as the proposed databricks intelligence platform approach.: Governance is organized at metastore/catalog/schema/object scope, not per notebook.
  • C. Place all classic compute in the Databricks control plane so the metastore can see the data, within the documented operational, security, ownership, and validation requirements.: Classic compute runs in the customer's AWS account, not in the Databricks control plane.
  • D. Merge the three workspaces into one because a Unity Catalog metastore cannot govern more than one workspace, under the documented operational and governance requirements.: A single metastore can be linked to multiple workspaces in the same region, so merging the workspaces is not required.

Extra learning features

Why candidates miss this

The distractors 'Create a separate metastore per notebook' and 'Place all classic compute in the Databricks control plane' suggest isolated environments, which would negate the benefits of centralized governance. The decisive clue is the emphasis on a 'single governance view,' highlighting the need for a unified metastore to manage access and lineage across workspaces. Likely wrong answer: Create a separate metastore per notebook so lineage and privileges remain isolated at the code level. Review focus: High-level architecture

Interview question

Q: You are designing governance for several Databricks workspaces in the same cloud region. How would you centralize governance without forcing every engineering team into one workspace, and what would you use if a catalog must be visible only from selected workspaces? Strong answer: I would attach the workspaces in that region to the same Unity Catalog metastore so they share governed metadata, permissions, and securable objects while retaining separate workspace development and compute choices. If a catalog must be isolated to specific workspaces, I would use workspace-catalog bindings rather than creating separate metastores just for workspace separation.

  • one metastore per region
  • multiple workspaces per metastore
  • centralized Unity Catalog governance
  • workspace autonomy
  • workspace-catalog bindings

Caution: A strong answer should separate regional metastore scope from workspace-level isolation and should not propose one metastore per notebook or team.

Objective/domain: Databricks Intelligence Platform

Source: High-level architecture

Question 2 A Lakeflow managed connector lands nested source records into Unity Catalog. Downstream Silver processing needs selected nested attributes as relational columns while preserving the original Bronze data for replay. Which layer responsibility is correct?

Answer choices

  1. A. Use a column mask to transform nested JSON into relational columns, for the described technical objective and its associated operational control requirements, as proposed.
  2. B. Discard the raw nested payload as soon as the first parse succeeds, for the described technical objective and its associated operational control requirements, for evaluation.
  3. C. Make Gold the raw landing layer and Bronze the business aggregate layer, for the required business outcome.
  4. D. Keep the raw/nested representation in Bronze and perform deterministic parsing/normalization into Silver, for the stated data ingestion and loading requirement.

Correct answer

Keep the raw/nested representation in Bronze and perform deterministic parsing/normalization into Silver, for the stated data ingestion and loading requirement.

Objective/domain: Data Ingestion and Loading

Source: Databricks Certified Data Engineer Associate Exam Guide - May 4, 2026

Question 3 A Unity Catalog external Delta table is no longer needed in Databricks, but an external application must continue using the files at the cloud storage path. What happens when the team drops the external table metadata?

Answer choices

  1. A. The table is automatically converted to a managed table before the metadata is removed, for the stated security, delivery, and accountability requirements.
  2. B. Databricks deletes the external data files because each Unity Catalog DROP owns the underlying lifecycle, as the organization’s selected response.
  3. C. The cloud storage path is moved into Databricks-controlled managed storage, as the recommended response to this scenario.
  4. D. Unity Catalog removes the table metadata, while the externally managed data files remain at the storage location, in practice.

Correct answer

Unity Catalog removes the table metadata, while the externally managed data files remain at the storage location, in practice.

Objective/domain: Governance and Security

Source: Managed versus external assets in Unity Catalog

Question 4 A reviewer says Databricks Asset Bundles and Declarative Automation Bundles are unrelated products, so existing bundle-based CI/CD must be redesigned from scratch. What is the correct clarification?

Answer choices

  1. A. Asset Bundles now refer to Unity Catalog tables and cannot define jobs, for the described technical objective and its associated operational control requirements, for evaluation.
  2. B. Declarative Automation Bundles are the current name for the capability formerly called Databricks Asset Bundles, within the stated policy framework.
  3. C. Declarative Automation Bundles are a Git provider and do not deploy Databricks resources, under the organization’s defined implementation and exception-management process.
  4. D. The reviewer is correct; Asset Bundles were replaced by Dashboard tasks, within cross-functional operational-accountability boundaries.

Correct answer

Declarative Automation Bundles are the current name for the capability formerly called Databricks Asset Bundles, within the stated policy framework.

Objective/domain: Implementing CI/CD

Source: What are Declarative Automation Bundles?

Question 5 A finance validation consists entirely of a governed SQL file and must run on Databricks SQL after the transformation task. The team wants the Jobs UI to show it as a distinct unit of work. Which task type is the best fit?

Answer choices

  1. A. Use a `For each` task with the SQL text as the input array, for the described technical objective and its associated operational control requirements, under this approach.
  2. B. Use a Dashboard task because SQL files and dashboards use the same task semantics, under the described working with lakeflow jobs criteria.
  3. C. Use a SQL task configured to run the SQL file on a supported SQL warehouse and depend on the transformation task, within the stated policy framework.
  4. D. Store the SQL in a notebook solely so a Notebook task can execute it, even though a SQL task directly supports files, within the defined security and accountability boundaries.

Correct answer

Use a SQL task configured to run the SQL file on a supported SQL warehouse and depend on the transformation task, within the stated policy framework.

Objective/domain: Working with Lakeflow Jobs

Source: SQL task for jobs

Question 6 Two streaming writers update a Delta table while analysts run long queries. The data team needs transactional consistency without maintaining separate batch and streaming copies. Which platform component provides the foundational storage behavior?

Answer choices

  1. A. A Git folder, because source control prevents concurrent data modifications, for the described technical objective and its associated operational control requirements, within this context.
  2. B. Unity Catalog, because lineage alone serializes all writes to the physical files, for the described technical objective and its associated operational control requirements.
  3. C. A SQL warehouse, because query compute is what creates table transaction history, as the proposed design for the complete governed operational workflow.
  4. D. Delta Lake, because its transaction log provides ACID semantics and supports both batch and Structured Streaming against the same table.

Correct answer

Delta Lake, because its transaction log provides ACID semantics and supports both batch and Structured Streaming against the same table.

Objective/domain: Databricks Intelligence Platform

Source: What is Delta Lake in Databricks?

Question 7 Analysts filter a large table mostly by `customer_id` and `event_date`, but an engineer proposes clustering by a random UUID because it has the highest cardinality. What is the better principle?

Answer choices

  1. A. Choose clustering behavior based on observed filter/query patterns and measured benefit, within the described context.
  2. B. Typically cluster on each table column so all possible queries are covered, for the specified implementation requirement.
  3. C. Choose keys alphabetically because liquid clustering ignores query access patterns, for the stated scenario.
  4. D. Typically cluster on the highest-cardinality column regardless of query predicates, as described.

Correct answer

Choose clustering behavior based on observed filter/query patterns and measured benefit, within the described context.

Objective/domain: Troubleshooting, Monitoring, and Optimization

Source: CLUSTER BY clause (TABLE)

Question 8 A revenue DataFrame has grain `order_line`. A developer groups only by `region` and sums `order_total`, but `order_total` is repeated on every line of an order. The reported revenue is inflated. What should be fixed first?

Answer choices

  1. A. Deduplicate or aggregate to one row per order before summing order totals by region, as the primary proposed approach.
  2. B. Use `approx_count_distinct` on order total and multiply it by region count, within this context.
  3. C. Increase the shuffle partition count because inflated totals indicate reducer skew, within the proposed design.
  4. D. Convert `order_total` to string before summing so repeated values compare exactly, within the documented operational, security, ownership, and validation requirements.

Correct answer

Deduplicate or aggregate to one row per order before summing order totals by region, as the primary proposed approach.

Objective/domain: Data Transformation and Modeling

Source: PySpark DataFrame API

Question 9 Two Auto Loader streams read two different source directories but write to the same curated target. An engineer proposes sharing one streaming checkpoint to simplify operations. What is the correct response?

Answer choices

  1. A. Use one checkpoint if both sources are JSON, for the required data ingestion and loading outcome.
  2. B. Do not share it; each source ingestion workload requires a separate streaming checkpoint, in the described situation.
  3. C. Share the checkpoint because a common target table makes the streams one logical source, within the stated policy framework.
  4. D. Checkpoints are unnecessary when the target is Delta, for the described technical objective and its associated operational control requirements, when applied.

Correct answer

Do not share it; each source ingestion workload requires a separate streaming checkpoint, in the described situation.

Objective/domain: Data Ingestion and Loading

Source: Configure schema inference and evolution in Auto Loader

Question 10 A new table is tagged so an existing ABAC row-filter policy will apply. The data owner wants to deploy the tag directly in production without testing because “the policy code did not change.” What is the security concern?

Answer choices

  1. A. Review and test a policy-driving tag assignment as a security change, within the defined security and accountability boundaries.
  2. B. The primary risk is increased Spark shuffle because tags add a partition column, within the proposed design.
  3. C. No concern; tag assignments does not affect whether an ABAC policy applies, within this design.
  4. D. Tag changes require a cluster restart but cannot change authorization, for this requirement.

Correct answer

Review and test a policy-driving tag assignment as a security change, within the defined security and accountability boundaries.

Objective/domain: Governance and Security

Source: Attribute-based access control in Unity Catalog

Where to go after the daily web set

How are Databricks Data Engineer Associate questions generated?

dotCreds builds Databricks Data Engineer Associate practice questions from public exam objectives and Databricks exam and documentation references. The questions are written for realistic study practice, not copied from exam dumps.

How are explanations sourced?

Each question includes an explanation and, when available, a source link back to the provider documentation or reference used to validate the answer. That keeps the practice tied to study material you can actually review.

What score do I get?

The page tracks today's answered count and accuracy for the 10-question daily set, then saves a 7-day score history on this device so you can see your recent practice trend.

Why use this site?

The site is the fastest way to start Databricks Data Engineer Associate practice without installing anything. It is built for daily recall, quick weak-topic discovery, and source-backed explanations you can review immediately.