Data Engineer

Expert In Recruitment Solutions

  • malvern, PA
  • 11 days ago

    Highlights

    Pipeline Management: Experience building end-to-end data pipelines, managing SFTP connections, and integrating with diverse data sources (Data Lakes, flat files, etc.). Data Modeling: Proficiency in designing data models and creating multi-layered data structures to feed executive and operational reports.

    Numbers & Facts

    Locationmalvern, PA

    Description

    Data Engineer
    hybird - malvern pa
    2 openings
    assessment required


    Job Description

    Technical Stack & Requirements:
    • Core Competencies: AWS, Glue, S3, IAM, Redshift, Python, Spark.
    • Development Environment: Experience with Amazon SageMaker is mandatory, as this is our primary environment for code development before production deployment.
    • Pipeline Management: Experience building end-to-end data pipelines, managing SFTP connections, and integrating with diverse data sources (Data Lakes, flat files, etc.).
    • Infrastructure & Automation: Experience with SNS notifications and CI/CD tools (Jenkins) to ensure process efficiency.
    • Nice-to-Have Skills: Experience with Databricks or Snowflake is highly desirable as we transition toward these platforms.
    • Data Modeling: Proficiency in designing data models and creating multi-layered data structures to feed executive and operational reports.

    Key Responsibilities:
    • Manage priorities across multiple workstreams and handle production support as needed.
    • Maintain robust pipelines with automated checks, particularly for event-driven jobs.
    • Communicate effectively with external vendors and internal business stakeholders to gather requirements and provide solutions.
    • Collaborate with senior engineering staff to ensure the security and integrity of our AWS development environments.
    Required Qualifications
    • 5+ years of Data Engineering or IT experience.
    • 4+ years of AWS cloud development experience.
    • Expert Python programming skills.
    • Experience with Sagemaker.
    • Robust hands-on experience with Apache Spark / PySpark.
    • Experience building cloud-native ETL pipelines.
    • Robust understanding of distributed computing.
    • Experience designing scalable data models.
    • Experience with Git and CI/CD pipelines.
    • Robust analytical and problem-solving skills.
    • Build scalable and auto recoverable pipelines.
    • Bachelor's degree in Computer Science, Engineering, Information Systems, or related discipline.
    • Working directly with business SME's and stakeholders.

    Good to have
    • Experience with Databricks tool.

    Similar Jobs

    See more jobs