Databricks Engineer
Location: Washington, DC (REMOTE)
Duration: 7+ months
Seeking for a Databricks Data Engineer with experience building scalable data pipelines using the Databricks Lakehouse Platform, PySpark, and SQL. This role involves designing and optimizing data solutions, implementing the Medallion Architecture, managing data governance through Unity Catalog, and supporting CI/CD processes for reliable data deployments.
Key Responsibilities:
- Design, develop, and maintain scalable ETL/ELT data pipelines using the Databricks Lakehouse Platform.
- Develop data transformation processes using PySpark and SQL.
- Implement and maintain the Medallion Architecture (Bronze, Silver, and Gold layers) for efficient data processing.
- Optimize Spark jobs, SQL queries, and Databricks workloads for performance and scalability.
- Configure and manage Unity Catalog, including row-level security, column-level security, data lineage, and secure data sharing.
- Develop, test, and deploy Databricks solutions using CI/CD pipelines and version control tools.
- Monitor and troubleshoot production data pipelines to ensure high availability and reliability.
- Collaborate with data architects, analysts, and business stakeholders to deliver data solutions that meet business requirements.
- Implement data quality checks, validation processes, and governance best practices.
- Document technical designs, workflows, and operational procedures.
Required Qualifications
- Bachelor's degree in Computer Science, Information Technology, Engineering, or a related field.
- Strong experience with the Databricks Lakehouse Platform.
- Proficiency in PySpark and SQL.
- Experience implementing Medallion Architecture.
- Hands-on experience with Unity Catalog, including data governance, row-level security, column-level security, and data lineage.
- Experience with ETL/ELT pipeline development and data transformation.
- Knowledge of performance tuning and optimization techniques in Databricks.
- Experience with CI/CD tools and version control systems such as Azure DevOps, Git, or GitHub.
- Strong analytical, troubleshooting, and problem-solving skills.
- Excellent verbal and written communication skills.
Preferred Qualifications
- Experience with SAS.
- Experience with Unix/Linux environments.
- Advanced SQL development and optimization skills.
- Experience working with Delta Lake and modern cloud-based data platforms.
- Familiarity with Agile development methodologies.
Required Skills
- Databricks Lakehouse Platform
- PySpark
- SQL
- Medallion Architecture
- Unity Catalog
- Delta Lake
- CI/CD
- Git / Azure DevOps
- Data Pipeline Development