Senior Data Engineer – Identity Resolution & Large-Scale Data Engineering

Sumeru Solutions

  • Bellevue, WA
  • 3 days ago

    Highlights

    The ideal candidate will have proven experience delivering data solutions at billion-record scale, with strong expertise in Azure Databricks, PySpark, Python, Azure Data Factory, and distributed data processing. The role requires strong production-scale engineering judgment-the ability to identify technical and scalability constraints early, evaluate alternatives quickly, and deliver tested, scalable, and production-ready solutions.

    Numbers & Facts

    LocationBellevue, WA

    Description

    Role: Senior Data Engineer Identity Resolution & Large-Scale Data Engineering

    Location: Bellevue, WA

    Overview:

    We are seeking an experienced Senior Data Engineer to design, build, and optimize production-grade data solutions supporting Identity Resolution (IDR) and large-scale customer data initiatives.

    The ideal candidate will have proven experience delivering data solutions at billion-record scale, with strong expertise in Azure Databricks, PySpark, Python, Azure Data Factory, and distributed data processing.

    The role requires strong production-scale engineering judgment-the ability to identify technical and scalability constraints early, evaluate alternatives quickly, and deliver tested, scalable, and production-ready solutions.

    Experience with Identity Resolution or comparable high-throughput matching systems is highly preferred.

    Key Responsibilities

    Design and optimize "large-scale data processing solutions capable of operating at 1B+ record scale".

    Develop scalable data pipelines using Azure Databricks, PySpark, Python, and ADF".

    Optimize Spark workloads including large-scale joins, partitioning, shuffling, data skew, memory, and compute utilization.

    Design and improve solutions for Identity Resolution, Entity Resolution, Record Linkage, Deduplication, or similar high-volume matching problems.

    Evaluate technical approaches against production data volumes and infrastructure constraints.

    Identify and communicate scalability and environment constraints early and recommend viable alternatives.

    Perform performance and scalability testing using production-like workloads.

    Build reliable data pipelines with appropriate validation, error handling, monitoring, and recovery mechanisms.

    Collaborate with architects, engineers, data scientists, and stakeholders to deliver solutions from design through production.

    Required Qualifications

    7+ years of Data Engineering experience.

    Proven experience working with very large-scale datasets, preferably 1B+ records.

    Strong hands-on experience with Azure Databricks and PySpark.

    Strong Python and SQL skills.

    Experience with Azure Data Factory (ADF) and Azure data services.

    Strong understanding of distributed processing and Spark performance optimization.

    Experience delivering production-grade, scalable data solutions.

    Ability to identify technical constraints early and make timely technical decisions.

    Strong problem-solving and communication skills.

    Preferred Qualifications

    Experience with Identity Resolution / Entity Resolution / Record Linkage / Deduplication / Fuzzy Matching.

    Experience with Cosmos DB and low-latency APIs.

    Experience with Azure Functions or Azure Web Apps.

    Experience with Event Hub, Kafka, or Spark Structured Streaming.

    ML experience related to matching or customer data.

    Exposure to Generative AI, LLMs, RAG, or Agentic AI.

    Telecommunications or large-scale customer data experience.

    Similar Jobs

    See more jobs