Senior Databricks Engineer

Diverse Lynx, LLC

  • Malvern, PA
  • 6 days ago
  • $110,000–$130,000 Per Year

Highlights

Delta Lake & Lakehouse: Building medallion architectures (Bronze, Silver, Gold layers) and using optimizations like Z-Ordering and ACID transactions. Pipelines & Orchestration: Developing with Delta Live Tables (DLT) and Databricks Workflows to design resilient, idempotent ELT/ETL pipelines.

Numbers & Facts

LocationMalvern, PA
Salary$110,000–$130,000 Per Year

Description

Role: Senior Databricks Engineer

Location- Malvern, PA (Onsite)

Job Type: Full Time

Salary Range: $110000 to $130000/Annum + Full Time Benefits

Experience Required 10-15 years of experience.

Job Description

Must Have Technical/Functional Skills

  • Apache Spark: Mastery of Spark SQL and the DataFrame API for distributed data processing, performance tuning, and optimizing execution plans.
  • Delta Lake & Lakehouse: Building medallion architectures (Bronze, Silver, Gold layers) and using optimizations like Z-Ordering and ACID transactions
  • Pipelines & Orchestration: Developing with Delta Live Tables (DLT) and Databricks Workflows to design resilient, idempotent ELT/ETL pipelines.
  • Unity Catalog: Configuring data governance, row/column-level security, and data lineage
  • Languages: Advanced proficiency in Python (PySpark) and SQL
  • Storage & Data Lakes: Expert in reading/writing to Amazon S3 and integrating with AWS data warehouses (e.g., Amazon Redshift).
  • Security & IAM: Implementing cross-account IAM roles, S3 bucket policies, and Customer-Managed Keys (CMK) via AWS KMS
  • .Compute Management: Managing cluster policies, instance profiles, and utilizing Spot Instances to optimize Databricks compute costs on AWS.
  • Ecosystem Integration: Familiarity with integrating Databricks jobs alongside native services like AWS Glue, Amazon Kinesis, and AWS Step Functions
  • CI/CD & Version Control: Git-based workflows using Databricks Git Folders and managing deployments.
  • AI & MLOps: Using MLflow for experiment tracking and model registries.
  • Familiarity with LLMs, Vector Search, and GenAI integrations.

Roles & Responsibilities

  • Data Pipeline Development: Extract, transform, and load (ETL) data from multiple sources, building batch and streaming pipelines.Spark & Code Optimization: Write highly efficient PySpark, Scala, or SQL code. Troubleshoot distributed processing bottlenecks and optimize job performance.
  • Lakehouse Management: Work with the Medallion architecture to transition data through Bronze (raw), Silver (cleansed), and Gold (aggregated) layers in Delta Lake.
  • Orchestration & Automation: Schedule and automate workflows using Databricks Jobs, Workflows, or Delta Live Tables (DLT). Configure automated tests, retries, and error handling.
  • Governance & Security: Implement data privacy and access controls utilizing the Databricks Unity Catalog, ensuring compliance and secure data sharing.
  • Collaboration: Work alongside Data Scientists, Data Analysts, and ML Engineers to prepare data features and support BI reporting, ML training, and generative AI initiatives.

Diverse Lynx LLC is an Equal Employment Opportunity employer. All qualified applicants will receive due consideration for employment without any discrimination. All applicants will be evaluated solely on the basis of their ability, competence and their proven capability to perform the functions outlined in the corresponding role. We promote and support a diverse workforce across all levels in the company.

Similar Jobs

See more jobs