Senior Data Software Engineer with AWS and Terraform

EPAM Systems Inc

  • Georgia, GA
  • 4 days ago

    Highlights

    The role involves ingesting and transforming large datasets with PySpark on AWS Glue and delivering curated, validated data into Snowflake, with a core focus on data quality, validation, and reconciliation for downstream analytics. We are seeking a Senior Data Software Engineer to join a client-facing delivery team building and hardening cloud-native data pipelines on AWS as part of a data platform modernization program.

    Numbers & Facts

    LocationGeorgia, GA

    Description

    Back to Search

    Senior Data Software Engineer with AWS and Terraform

    Remote in Georgia, & 4 others

    Data Software Engineering

    apply

    FacebookLinkedInSend via email

    Looking for something else?

    Find a vacancy that works for you. Send us your CV to receive a personalized offer.

    Find me a job

    Location-specific conditions & benefits*

    Choose an option

    We are seeking a Senior Data Software Engineer to join a client-facing delivery team building and hardening cloud-native data pipelines on AWS as part of a data platform modernization program. The role involves ingesting and transforming large datasets with PySpark on AWS Glue and delivering curated, validated data into Snowflake, with a core focus on data quality, validation, and reconciliation for downstream analytics. This position is delivered at a Senior Consultant level with high autonomy and direct client stakeholder communication.

    Responsibilities

    • Design, build, and optimize scalable batch and incremental ETL/ELT pipelines using PySpark on AWS Glue

    • Configure Glue jobs, crawlers, triggers, connections, bookmarks, workflows, and the Glue Data Catalog

    • Tune workers, partitioning, and shuffle behavior for cost and performance optimization

    • Model and load curated datasets into Snowflake with staging, transformation, and publishing layers

    • Implement automated data quality and validation frameworks, including schema/contract enforcement and null/uniqueness/referential checks

    • Develop row-count and financial reconciliation processes, anomaly detection, and quarantine/reject handling

    • Configure and extend Glue Data Quality (DQDL) rules per requirements

    • Write clean, modular, testable Python with unit/integration tests and reusable libraries

    • Integrate pipelines with AWS services such as S3, IAM, Lambda, Athena, CloudWatch, Step Functions, and Secrets Manager

    • Instrument observability through logging, metrics, alerting, and pipeline SLA monitoring

    • Participate in code reviews, CI/CD automation, and documentation

    • Engage directly with client stakeholders in requirements refinement, design walkthroughs, status reporting, and act as technical advisor within the workstream

    Requirements

    • 3+ years of experience with Python for production-level data engineering, including OOP and functional patterns

    • Expertise in PySpark for distributed data processing and the DataFrame API

    • Advanced proficiency in Snowflake, including data warehousing and staging/transformation layers

    • Skills in AWS Glue, including job configuration, crawlers, Data Catalog, and DQDL

    • Background in data quality engineering, including validation frameworks and reconciliation

    • Proficiency in AWS services including S3, IAM, Lambda, Athena, and CloudWatch

    • English proficiency at B2 level or higher

    Nice to have

    • Familiarity with Generative AI / LLM concepts

    • Knowledge of Airflow / Step Functions orchestration

    • Familiarity with Great Expectations or similar data quality frameworks

    • Knowledge of Terraform / CloudFormation

    Similar Jobs

    See more jobs