Big Data Engineer Interview: 15 In-Depth Questions
Covers real-time data warehouses, Spark skew remediation, Flink Exactly-Once, Lakehouse table formats, and governance.
Questions reflect common real-world prompts. The three answer layers are illustrative examples, not real interview transcripts.
① Common plain answer
"I inspect the Spark UI to identify slow-running tasks, increase shuffle partition numbers, or allocate additional worker memory."
Increasing partitions or memory blindly fails when specific keys concentrate disproportionate records, lacking two-stage salting and broadcast joins.
② Interviewer follow-up logic
③ Quantified high-score answer
Diagnosing Apache Spark data skew requires inspecting Spark UI stage telemetry to identify extreme percentile duration gaps where the 95th-percentile task consumes orders of magnitude more Shuffle Read bytes than median tasks. Remediation depends directly on the skewed operational boundary: GroupBy aggregations versus Join transformations. For skewed GroupBy operations, we implement two-stage aggregation: prefixing skewed partition keys with randomized integer salts (e.g., 0 to 19) to distribute records across twenty parallel executors for local pre-aggregation, followed by stripping the salt prefix during a second pass for final global reduction. For Join skew where the dimension table remains under 8GB, we force Broadcast Hash Joins via broadcast hints to eliminate shuffle phases entirely. In our behavioral tracking pipeline processing 1.4 billion events daily, severe user-activity skew repeatedly stalled stage execution for 48 minutes on single straggler tasks. Introducing key-salting combined with broadcast hints reduced total stage duration down to 4.2 minutes, stabilizing cluster compute utilization while completely eliminating out-of-memory executor crashes.
① Common plain answer
"I check the Flink Web UI to see which operator turns red with high backpressure, then increase TaskManager parallelism or scale workers."
Scaling workers blindly can exhaust downstream connection pools, lacking bottom-up operator metric traversal, thread flame graphs, and I/O profiling.
② Interviewer follow-up logic
③ Quantified high-score answer
Tracing backpressure requires bottom-up traversal along the operator topology: locating the downstream-most operator exhibiting High backpressure while its successor shows none. We inspect its on-CPU flame graph: if CPU-bound, we isolate costly deserialization, excessive regex evaluations, or un-cached external synchronous calls. If I/O-bound, we investigate downstream sink saturation (e.g., HBase or Elasticsearch) and enable asynchronous batched writes. If JVM garbage collection pauses stall processing, we tune off-heap allocations and RocksDB state-backend memory budgets to restore equilibrium.
① Common plain answer
"We use timestamp attributes from raw records as event time, set a watermark delay of a few seconds, and drop late arriving records."
Dropping late data violates financial data integrity, lacking understanding of multi-partition minimum watermark alignment and side-output reconciliations.
② Interviewer follow-up logic
③ Quantified high-score answer
Watermarks represent monotonically advancing timestamps tracking event-time progress. We deploy bounded-out-of-orderness watermark generators, accommodating anticipated 5-second network transmission delays. When parallel input streams merge, operators advance based on the minimum watermark among all active input channels, preventing premature window closure. For extreme outlier records exceeding the delay allowance, we prevent data loss by routing late records via side outputs (sideOutputLateData) into dedicated dead-letter topics for batch reconciliation.
① Common plain answer
"We turn on Flink Checkpointing, configure Kafka producer retries, and set isolation.level to read_committed on consumer clients."
Lacks rigorous explanation of the distributed Two-Phase Commit (2PC) state machine, barrier alignment mechanics, and transactional boundary timeouts.
② Interviewer follow-up logic
③ Quantified high-score answer
End-to-End Exactly-Once requires coupling distributed snapshots with two-phase commit protocols (2PC). Flink triggers periodic checkpoints using Chandy-Lamport barriers: source operators snapshot Kafka offsets, while stateful operators serialize internal state into durable storage. Concurrently, the FlinkKafkaProducer coordinates 2PC: initiating a pre-commit transaction writing speculative data to Kafka transaction partitions while barriers pass. Once the JobManager confirms all operator checkpoints succeeded globally, the coordinator broadcasts a formal commit, exposing records to downstream consumers.
① Common plain answer
"We break flat wide tables into fact and dimension tables, placing numerical metrics in facts and descriptive attributes in dimensions."
Ignores the three fundamental fact table patterns, their distinct lifecycles, and granularity requirements across enterprise business processes.
② Interviewer follow-up logic
③ Quantified high-score answer
Dimensional modeling selects fact table architectures based on business process lifecycles: Transaction Fact Tables capture atomic, discrete events at maximum granularity (individual payment transactions), offering maximum flexibility but requiring aggregations across periods. Periodic Snapshot Fact Tables record point-in-time status at fixed intervals (monthly bank balances, daily inventory snapshots). Accumulating Snapshot Fact Tables model workflows spanning multiple milestone dates (an order progressing through checkout, payment, fulfillment, and delivery) with dynamic in-place updates.
① Common plain answer
"Data lakes store raw files on HDFS or cloud object storage, allowing SQL queries to execute ACID updates faster than legacy Hive tables."
Superficial understanding misses hierarchical metadata structures (Manifest Lists), hidden partitioning, and Merge-on-Read versus Copy-on-Write tradeoffs.
② Interviewer follow-up logic
③ Quantified high-score answer
Modern table formats eliminate the operational limitations of legacy Hive directories. Apache Iceberg’s core strength is its directory-agnostic hierarchical metadata architecture: organizing snapshots into Manifest Lists and Manifest Files, eliminating O(N) object storage directory listings while enabling hidden partitioning and zero-copy branching. Apache Hudi specializes in high-frequency streaming ingestion, leveraging Merge-on-Read (MOR) with bloom filter indexes for sub-minute incremental updates. Delta Lake delivers deep Spark engine coupling with transaction logs.
① Common plain answer
"Previously we ran Hive for batch and Kafka+Flink for real-time, which was hard to maintain, so we replaced everything with a single lake table."
Fails to analyze dual-codebase maintenance overhead, metric divergence, and data reconciliation nightmares, lacking unified storage and compute designs.
② Interviewer follow-up logic
③ Quantified high-score answer
Lambda architectures impose high engineering friction: duplicate business logic, doubled compute costs, and perpetual data discrepancies. We engineered a phased migration to an Iceberg Lakehouse: In the storage tier, Iceberg provides a single unified storage layer across cloud object storage, supporting ACID transactions and near-real-time streaming ingestion. In the compute tier, Flink handles streaming ingestion while Spark handles large-scale batch transformations against identical underlying tables. This eliminated metric drift and reduced cloud storage costs by 45%.
① Common plain answer
"We write periodic cron jobs using HDFS commands to merge files, or tell developers to append coalesce(10) to the end of Spark jobs."
Ad-hoc scripts are brittle stopgaps, failing to analyze how micro-batch streaming and uncontrolled dynamic partitioning exhaust NameNode memory.
② Interviewer follow-up logic
③ Quantified high-score answer
Small files degrade distributed file systems by saturating NameNode RAM and spawning millions of tiny executor tasks. We enforce end-to-end multi-tiered governance: At the ingestion tier, streaming jobs enforce commit thresholds, flushing files only when reaching 128MB or 15-minute intervals. At the computation tier, Spark jobs repartition data by partition keys before writing, preventing cross-node fragmentation. At the lakehouse tier, we automate asynchronous compaction pipelines in Iceberg, rewriting small files into optimized columnar Parquet files, cutting file counts by 90%.
① Common plain answer
"We draw table dependency diagrams on internal company wikis, and have analysts manually verify operational reporting tables each morning."
Manual documentation becomes instantly obsolete, lacking automated SQL AST parsing, runtime metadata lineage, and automated pipeline circuit breakers.
② Interviewer follow-up logic
③ Quantified high-score answer
Data governance requires automated lineage extraction coupled with active pipeline circuit breaking. For lineage, we parse SQL queries via ANTLR4 and OpenLineage hooks, extracting table-level and column-level dependency graphs across Hive, Spark, and Flink into a centralized graph database. For Data Quality Control (DQC), critical pipeline stages execute automated assertion rules: validating primary key uniqueness, foreign key integrity, and day-over-day metric fluctuations. If revenue metrics deviate beyond 30%, the engine trips circuit breakers, halting downstream updates.
① Common plain answer
"Use ClickHouse if you need raw speed on flat wide tables; use StarRocks or Doris if your analytics require multi-table relational joins."
Superficial rules of thumb ignore vectorized execution architectures, Cost-Based Optimizers (CBO), and decoupled storage-compute models.
② Interviewer follow-up logic
③ Quantified high-score answer
OLAP selection hinges on query topology and operational complexity. ClickHouse delivers unmatched single-table scan speeds and compression via hand-tuned SIMD vectorization, making it optimal for denormalized flat log analytics; however, its distributed join engine is fragile and re-sharding involves high operational overhead. StarRocks employs a vectorized CBO optimizer and pipeline execution engine, excelling at multi-table shuffle joins across star schemas while natively supporting real-time upserts via its Primary Key model. We deploy ClickHouse for flat logging, and StarRocks for ad-hoc relational joins.
① Common plain answer
"We mask user phone numbers during ETL ingestion and restrict access permissions to sensitive financial database tables."
Ad-hoc masking leaves exposure to differential privacy re-identification, lacking centralized fine-grained row/column filtering and auditing.
② Interviewer follow-up logic
③ Quantified high-score answer
Data lakehouse security requires centralized policy governance and dynamic runtime masking. We deploy Apache Ranger across Spark, Trino, and Hive: First, data catalogs tag sensitive PII columns across all storage tiers. Second, Ranger enforces attribute-based and role-based access control (ABAC/RBAC); when non-authorized analysts query sensitive tables, Ranger injects query rewrites into execution plans, masking phone numbers and national IDs with dynamic regex hashing. Third, row-level filters restrict data visibility by department. All queries are audited centrally for regulatory compliance.
① Common plain answer
"I check which pipeline task crashed, rerun the failed job, and notify leadership in message channels that reporting will be delayed."
Blindly rerunning failed pipelines risks compounding cluster resource saturation, lacking critical-path analysis, priority preemption, and structured triage.
② Interviewer follow-up logic
③ Quantified high-score answer
Morning dashboard pipeline failures require disciplined critical-path triage: Step one isolates the bottleneck: I inspect scheduler Gantt charts to identify the exact task blocking the critical delivery path. Step two preempts cluster capacity: we suspend low-priority ad-hoc analytical queues, reallocating 100% of YARN compute capacity to the blocking ETL job. If dependencies cannot complete before the executive meeting, we activate our degradation plan: generating top-line revenue indicators first while deferring granular dimensional slices. Concurrently, a designated liaison provides objective timeline updates to leadership.
① Common plain answer
"I let company leadership decide which definition to use, or create two separate reporting tables in the data warehouse calculating both."
Building duplicate diverging tables is an abdication of architectural responsibility, deepening data silos and undermining corporate decision-making.
② Interviewer follow-up logic
③ Quantified high-score answer
Metric disputes stem from divergent business process viewpoints rather than technical flaws. I refuse to maintain competing tables. Instead, I establish an alignment working group: understanding that business operations optimizes for operational velocity (e.g., Gross Bookings at checkout) while finance optimizes for GAAP compliance (Net Revenue post-refund settlement). We resolve conflict through semantic decomposition: establishing explicit metric definitions in our corporate Metrics Store ("Order GMV" vs "Settled Net Revenue"), deprecating ambiguous generic terms, and publishing an enterprise Metric Whitepaper ratified by leadership.
① Common plain answer
"I configure automated scripts to kill any query running longer than ten minutes, and restrict analysts to querying only the first 100 rows."
Aggressive query killing paralyzes legitimate business intelligence, lacking resource pool isolation, query cost estimation, and proactive enablement.
② Interviewer follow-up logic
③ Quantified high-score answer
Governing ad-hoc analytics requires hardware isolation, gateway validation, and active developer education. First, we enforce physical resource boundaries: splitting compute queues into guaranteed production queues and weighted ad-hoc queues, preventing runaway queries from impacting production pipelines. Second, our query gateway validates incoming SQL, intercepting full-table scans lacking partition predicates or un-bounded Cartesian joins. Third, we establish weekly query optimization clinics: reviewing expensive queries with analysts to demonstrate partition pruning, fostering collaboration.
① Common plain answer
"Big data technologies are becoming legacy, so data engineers should quickly transition toward model training and generative AI algorithms."
Ignoring the reality that frontier AI pre-training, multi-modal data curation, and real-time RAG retrieval rely entirely on distributed big data foundations.
② Interviewer follow-up logic
③ Quantified high-score answer
Generative AI does not diminish data engineering; it elevates it into deeper systems domains. High-performance data engineering is the foundational fuel for AI: First, pre-training trillion-token models demands distributed data processing frameworks like Spark and Ray to execute petabyte-scale MinHash deduplication, text extraction, and filtering. Second, production enterprise RAG is fundamentally an exercise in real-time streaming pipelines combined with vector indexing. Third, data flywheels rely on lineage and feature stores. My enduring focus is bridging distributed data foundations with AI serving architectures.
Keep practicing in another role
After Big Data Engineer, these are the adjacent roles to practice next
Backend Interview: 15 In-Depth Questions
Core Tech · Distributed Systems · BQ
View bank
Same functionData Analyst Interview: 15 In-Depth Questions
Metrics · Attribution · BQ
View bank
Skill extensionLLM Inference & Infra Interview: 15 In-Depth Questions
Inference Serving · Memory Optimization · High Concurrency · BQ
View bank
Don't see your role? Browse all 25 roles →
Finished the breakdown? Try a realistic mock interview
Start a round without signing up. Experience in-depth follow-up questions and surface your real project highlights.
No credit card required · Free 600 credits on signup