Big Data Lead (AWS. Python, Pyspark)

Damco Solutions Inc

  • Atlanta, GA
  • 1 day ago

    Highlights

    Deep experience with AWS services including Glue, EMR, Redshift, RDS, Athena, Lambda, Step Functions, SNS, SQS, S3, CloudWatch, and CloudTrail. - Architect and implement scalable batch and near-real-time data pipelines using Python, PySpark, Glue, EMR, Lambda, and Step Functions.

    Numbers & Facts

    LocationAtlanta, GA

    Description

    Role: Principal AWS Data Engineer
    Job Title: Big Data Lead
    Job Skills: AWS. Python, Pyspark

    Job Description
    - Sr Dev - AWS Data Engineer with 12+ years of software development experience
    - Architect and implement scalable batch and near-real-time data pipelines using Python, PySpark, Glue, EMR, Lambda, and Step Functions.
    - Define data storage, transformation, and access patterns across S3, Redshift, Athena, RDS, and DynamoDB.
    - Design event-driven and integration solutions using SNS, SQS, Kinesis, and APIs.
    - Establish standards for data quality, observability, monitoring, logging, and failure recovery.
    - Optimize data processing, storage, and query performance across AWS platforms.
    - Lead troubleshooting of production issues and drive root cause analysis and corrective actions.
    - Support cloud migration, data migration, and modernization initiatives.
    - Review solution designs, code quality, and production readiness across projects.
    - Collaborate with architects, platform teams, and business stakeholders to align technical solutions with business needs.
    - Drive best practices for CI/CD, deployment automation, and documentation.

    Required Skills and Qualifications
    - Solid IT background and experience.
    - Experience as an application developer for projects similar in scope and responsibility.
    - Advanced knowledge of AWS services and architecture.
    - Strong experience with Python object-oriented programming, PySpark, SQL, and ETL frameworks.
    - Deep experience with AWS services including Glue, EMR, Redshift, RDS, Athena, Lambda, Step Functions, SNS, SQS, S3, CloudWatch, and CloudTrail.
    - Strong understanding of enterprise data architecture, migration, and support models.
    - Experience with GitLab and Terraform and release/deployment practices.

    Preferred Skills
    - Experience with Kinesis, ECS, Batch, Beanstalk, and DynamoDB.
    - Exposure to infrastructure-as-code and reusable platform patterns.
    - Have a good understanding of performance engineering of code pipelines and near real time systems
    - Good understanding on Agents and MCP


    Required Qualifications

    • Bachelor's or Master's degree in Computer Science, Information Technology, Engineering, or related field.
    • 12+ years of software development and data engineering experience.
    • Extensive experience designing and implementing large-scale cloud-native data solutions.
    • Advanced expertise in Python, object-oriented programming, PySpark, SQL, and ETL frameworks.
    • Deep hands-on experience with AWS services including:
      • AWS Glue
      • EMR
      • Lambda
      • Step Functions
      • S3
      • Redshift
      • Athena
      • RDS
      • SNS
      • SQS
      • CloudWatch
      • CloudTrail
    • Strong understanding of enterprise data architecture, integration patterns, and data migration strategies.
    • Experience implementing event-driven and serverless architectures in AWS.
    • Strong understanding of data governance, security, monitoring, and operational support models.
    • Hands-on experience with GitLab, CI/CD pipelines, release management, and deployment automation.
    • Experience with Infrastructure as Code using Terraform.
    • Excellent problem-solving, communication, and stakeholder management skills.

    Similar Jobs