Introduction
We are seeking a highly skilled individual to help build and expand our next-generation data processing platform. This role will focus on designing and developing distributed data processing solutions using Apache Spark while leveraging an existing enterprise Kafka ecosystem. The successful candidate will play a critical role in scaling our data processing capabilities, enabling cloud adoption, and supporting strategic migration initiatives involving Databricks and AWS. This position offers the opportunity to work on high-volume, mission-critical data platforms processing large-scale real-time and batch workloads.
Required Skills & Qualifications
- 5 years of hands-on development experience with Apache Spark.
- Strong expertise in Spark SQL, Structured Streaming, DataFrames/Datasets, performance tuning, and optimization.
- Experience integrating Spark with Apache Kafka.
- Strong development skills in Java and Python.
- Experience designing and supporting large-scale distributed systems.
- Solid understanding of software engineering principles, object-oriented design, and testing practices.
- Experience working with relational and distributed data platforms.
- Bachelor’s Degree
- Prior work experience at client or in client's Industry
Applicants must be able to work directly for Artech on W2
Preferred Skills & Qualifications
- Databricks experience (SQL).
- AWS experience, including services such as S3, EMR, Glue, Lambda, ECS/EKS.
- Experience supporting large-scale cloud migration initiatives.
- Experience with lakehouse architectures and modern data platform patterns.
- Experience handling multi-terabyte or petabyte-scale datasets.
- Financial services or regulated industry experience.
Day-to-Day Responsibilities
- Design, develop, and optimize Spark-based applications for large-scale data processing.
- Build streaming and batch data solutions leveraging Apache Kafka and Apache Spark.
- Develop reusable frameworks and components supporting enterprise data engineering initiatives.
- Support migration efforts from on-premises data platforms to Databricks and AWS.
- Collaborate with architects and platform engineers to implement cloud-native data solutions.
- Optimize Spark jobs for performance, scalability, reliability, and operational efficiency.
- Participate in system design, code reviews, testing, and production support activities.
- Implement CI/CD, monitoring, and observability practices for distributed data applications.
- Troubleshoot and resolve complex issues across Spark, Kafka, Databricks, and cloud environments and with other enterprise messaging systems.
For immediate consideration please click APPLY to begin the screening process with Alex.