Skip to main content
End-to-End Platform Development

Development Arena

Build production data engineering pipelines from the ground up across modern lakehouse and warehouse architectures. Architect Bronze, Silver, and Gold layers, configure Delta tables, integrate external APIs, design dimensional data marts, and implement CDC pipelines.

High Priority
200 XPDEV-001

Build Customer Lakehouse Pipeline

Build a production-ready customer data pipeline from source ingestion through Bronze, Silver, and Gold layers for analytics consumption.

DatabricksPySparkDelta LakeSnowflake
Award: 🏗️ Lakehouse Builder
Difficulty:Intermediate
Start Building →
High Priority
250 XPDEV-002

Incremental Sales Pipeline

Build a reliable incremental sales pipeline that processes cloud storage transaction files, handles duplicates and late data, and delivers trusted analytics datasets.

DatabricksPySparkDelta LakeSnowflakeCloud Storage
Award: ⚡ Incremental Engineer
Difficulty:Intermediate
Start Building →
High Priority
300 XPDEV-003

REST API Integration Pipeline

Build an ingestion pipeline consuming third-party REST API customer activity with pagination, transient error handling, rate limiting, and Silver/Gold curation.

DatabricksPySparkREST APIJSONDelta LakeSnowflake
Award: 🔌 Integration Engineer
Difficulty:Intermediate
Start Building →
High Priority
400 XPDEV-004

Enterprise Sales Reporting Data Mart

Design and implement a dimensional reporting data mart with explicit grain definition, fact/dimension relationships, SCD handling, and Snowflake reconciliation.

SnowflakeDatabricksPySparkSQLDelta Lake
Award: 🏛️ Data Mart Architect
Difficulty:Intermediate → Advanced
Start Building →
Critical Priority
500 XPDEV-005

CDC-Based Customer Data Pipeline

Design a Change Data Capture (CDC) processing pipeline handling INSERT, UPDATE, and DELETE event streams with out-of-order sequencing, replay, and current-state materialization.

DatabricksPySparkDelta LakeSnowflakeCDC Concepts
Award: 🔄 Change Data Engineer
Difficulty:Advanced
Start Building →
Critical Priority
600 XPDEV-006

Real-Time Sales Streaming Pipeline

Build a low-latency streaming pipeline continuously ingesting real-time sales events, handling watermarking, late data, stateful micro-batch aggregations, and Snowflake synchronization.

DatabricksPySparkStructured StreamingDelta LakeSnowflake
Award: 🌊 Streaming Engineer
Difficulty:Advanced
Start Building →
High Priority
700 XPDEV-007

Customer ML Feature Engineering Pipeline

Design a production-grade ML feature engineering pipeline transforming customer and transaction history into reusable feature stores with strict point-in-time leakage prevention.

DatabricksPySparkDelta LakeMLflowSnowflake
Award: 🧠 Feature Engineer
Difficulty:Advanced
Start Building →
High Priority
800 XPDEV-008

Production CI/CD Data Platform

Implement automated GitHub Actions CI/CD workflows for multi-environment data pipelines with linting, unit/integration testing, secret management, and rollback mechanisms.

GitHub ActionsDatabricksPySparkSnowflakeCloud CI/CD
Award: 🚀 Production Data Engineer
Difficulty:Expert
Start Building →
Critical Priority
900 XPDEV-009

Multi-Source Enterprise Data Platform

Build a unified enterprise platform integrating Customer DB, Sales Files, Product REST API, and Store Reference Data with entity resolution, SLA freshness, and cross-source reconciliation.

DatabricksPySparkSnowflakeREST APICloud StorageDelta Lake
Award: 🏛️ Data Platform Architect
Difficulty:Expert
Start Building →
Critical Priority
1000 XPDEV-010

Enterprise Data Engineering Capstone

Architect and build an end-to-end enterprise lakehouse platform combining batch ingestion, real-time streaming, CDC, REST APIs, ML feature stores, CI/CD, and Snowflake data marts.

DatabricksPySparkStructured StreamingCDCMLflowGitHub ActionsSnowflake
Award: 🏆 Enterprise Data Engineer
Difficulty:Expert / Capstone
Start Building →