Senior Databricks Engineer

Diverse Lynx, LLC

  • Malvern, PA
  • 5 days ago

    Highlights

    Delta Lake & Lakehouse: Building medallion architectures (Bronze, Silver, Gold layers) and using optimizations like Z-Ordering and ACID transactions. Pipelines & Orchestration: Developing with Delta Live Tables (DLT) and Databricks Workflows to design resilient, idempotent ELT/ETL pipelines.

    Numbers & Facts

    LocationMalvern, PA

    Description

    Senior Databricks Engineer

    Malvern PA - Onsite

    Fulltime

    $140K

    strong Databricks Platform engineering experience with AWS cloud expertise

    Job Description

    Must Have Technical/Functional Skills

    • Apache Spark: Mastery of Spark SQL and the Data Frame API for distributed data processing, performance tuning, and optimizing execution plans.
    • Delta Lake & Lakehouse: Building medallion architectures (Bronze, Silver, Gold layers) and using optimizations like Z-Ordering and ACID transactions
    • Pipelines & Orchestration: Developing with Delta Live Tables (DLT) and Databricks Workflows to design resilient, idempotent ELT/ETL pipelines.
    • Unity Catalog: Configuring data governance, row/column-level security, and data lineage
    • Languages: Advanced proficiency in Python (PySpark) and SQL
    • Storage & Data Lakes: Expert in reading/writing to Amazon S3 and integrating with AWS data warehouses (e.g., Amazon Redshift).
    • Security & IAM: Implementing cross-account IAM roles, S3 bucket policies, and Customer-Managed Keys (CMK) via AWS KMS
    • Compute Management: Managing cluster policies, instance profiles, and utilizing Spot Instances to optimize Databricks compute costs on AWS.
    • Ecosystem Integration: Familiarity with integrating Databricks jobs alongside native services like AWS Glue, Amazon Kinesis, and AWS Step Functions
    • CI/CD & Version Control: Git-based workflows using Databricks Git Folders and managing deployments.
    • AI & MLOps: Using MLflow for experiment tracking and model registries.
    • Familiarity with LLMs, Vector Search, and GenAI integrations.

    Roles & Responsibilities

    • Data Pipeline Development: Extract, transform, and load (ETL) data from multiple sources, building batch and streaming pipelines.
    • Spark & Code Optimization: Write highly efficient PySpark, Scala, or SQL code. Troubleshoot distributed processing bottlenecks and optimize job performance.
    • Lakehouse Management: Work with the Medallion architecture to transition data through Bronze (raw), Silver (cleansed), and Gold (aggregated) layers in Delta Lake.
    • Orchestration & Automation: Schedule and automate workflows using Databricks Jobs, Workflows, or Delta Live Tables (DLT). Configure automated tests, retries, and error handling.
    • Governance & Security: Implement data privacy and access controls utilizing the Databricks Unity Catalog, ensuring compliance and secure data sharing.
    • Collaboration: Work alongside Data Scientists, Data Analysts, and ML Engineers to prepare data features and support BI reporting, ML training, and generative AI initiatives.

    Diverse Lynx LLC is an Equal Employment Opportunity employer. All qualified applicants will receive due consideration for employment without any discrimination. All applicants will be evaluated solely on the basis of their ability, competence and their proven capability to perform the functions outlined in the corresponding role. We promote and support a diverse workforce across all levels in the company.

    Similar Jobs

    See more jobs