Job Requirements
- 4+ years of experience as a Data Engineer, including 3+ years of hands-on experience with Databricks, Apache Spark, and Delta Lake in a production environment.
- Design and implement a greenfield Databricks Lakehouse for the CRM domain by integrating Salesforce and other critical data sources while establishing scalable data pipelines and ML infrastructure to support customer forecasting and strategic business decisions.
- Serve as the technical design authority for end-to-end ETL/ELT pipelines, transforming raw data from Salesforce, Redshift, and data lakes into a high-quality Medallion architecture.
- Build and manage Databricks infrastructure using Infrastructure as Code (Terraform), including workspace provisioning, configuration, security, and Unity Catalog implementation to ensure GDPR and PII compliance.
- Lead the implementation of data monitoring dashboards and BI solutions using Databricks SQL and MicroStrategy, while architecting bi-directional data integration (Reverse ETL) to enable CRM activation and advanced customer analytics.
- Provide technical leadership by mentoring engineers, conducting code reviews, and establishing engineering best practices for operational excellence, including monitoring, alerting, and performance optimization.
- Demonstrated expertise in designing and implementing enterprise-scale data platforms with cross-functional integrations, particularly involving Salesforce and other enterprise CRM systems.
- Strong proficiency in Terraform, advanced Python programming, and CI/CD practices for building scalable ETL/ELT pipelines and machine learning infrastructure using MLflow and Databricks Genie.
- Deep understanding of data governance, security, compliance, and data quality frameworks, including Unity Catalog, sensitive data handling, and PII protection standards.