Role Technical Lead AWS PySpark
Location Onsite South SFO / Remote
Client Genentech
Mandatory Skills: Managing Data Ingestion for Unstructured data
Position Overview
We are seeking a highly skilled and technical AWS PySpark Lead to spearhead our data engineering initiatives. In this role, you will lead a team of data engineers to design, build, and optimize robust, scalable data pipelines on the AWS cloud. The ideal candidate brings a strong architectural mindset, deep expertise in distributed data processing, and a proven track record of optimizing PySpark, SQL workloads, and native AWS ETL tools. If you have a background as an AWS Data Engineer or AWS Solution Architect, excel in Agile environments, and are passionate about maintaining high data quality standards, we want you on our team.
Key Responsibilities
Pipeline Architecture & ETL: Design and implement large-scale, high-performance data pipelines using AWS EMR, PySpark, AWS Glue, and AWS DataBrew.
Unstructured Data & GenAI Pipelines: Architect end-to-end processing pipelines for unstructured data, including optimal ingestion processes, metadata extraction, and integration with Large Language Models (LLMs) to extract key metadata and content.
Vector DB Integration: Manage, optimize, and effectively store embeddings and data across enterprise Vector Databases to support advanced search and AI workloads.
Technical Leadership & Agile Delivery: Mentor and manage a team of data engineers. Drive Agile Sprints using JIRA, planning and actively managing the backlog and leveraging JIRA reporting/dashboards to track team velocity, bottlenecks, and project health.
Reporting: Daily, Weekly, and Monthly reporting to ensure all stakeholders are fully informed and engaged.
Data Governance & Quality: Architect solutions that guarantee high Data Quality and implement comprehensive Data Lineage tracking across all enterprise data workflows.
Performance Optimization: Take ownership of system performance by applying advanced PySpark optimization techniques (handling data skewness, memory management, partitioning, and broadcasting) and rigorous SQL optimization.
Serverless Data Processing: Architect and deploy event-driven data workflows and microservices utilizing AWS Lambda.
Data Warehousing: Model, manage, and optimize data storage and querying within AWS Redshift for analytical reporting and BI consumption.
Accelerated Delivery: Utilize AI coding assistants (such as GitHub Copilot and Claude Code) to streamline development, improve code quality, and accelerate project delivery timelines.
Cloud Architecture: Apply AWS Solution Architect principles to ensure data infrastructure is secure, highly available, cost-efficient, and scalable.
Required Qualifications & Skills (Must-Have)
Experience: Minimum of 6 years of hands-on experience in Data Engineering, Big Data, or Cloud Architecture.
Unstructured Data Processing: Hands-on experience in processing unstructured data, including designing optimal ingestion processes and metadata extraction pipelines.
LLM & GenAI Data Engineering: Proven experience leveraging LLMs to parse unstructured data and perform automated extraction of both metadata and core content.
Vector Database Expertise: Deep expertise in Vector Databases (e.g., Pinecone, Milvus, Qdrant, OpenSearch Vector Engine, Pgvector) with demonstrated experience in effectively indexing, managing, and storing vector data.
AWS Expertise: Proven experience operating as an AWS Data Engineer or AWS Solutions Architect. (Active AWS certifications are highly preferred).
PySpark Mastery: Exceptional proficiency in Python and Apache Spark. Must have a deep understanding of Spark's internal workings and hands-on experience optimizing heavy PySpark workloads.
AWS ETL Ecosystem: Strong, demonstrated experience with core AWS data services including AWS EMR, AWS Glue, AWS DataBrew, AWS Lambda, AWS Redshift, AWS S3, IAM, and AWS Step Functions.
Data Governance: Deep understanding of and practical experience with implementing automated Data Quality checks and establishing Data Lineage from source to destination.
SQL & Database Skills: Advanced SQL proficiency with a track record of tuning complex queries for performance across distributed databases.
Agile & JIRA: Demonstrated experience driving Agile methodologies. Must be well-versed in JIRA, specifically in configuring workflows and generating detailed JIRA reports for stakeholders.
DevOps & CI/CD: Strong proficiency in version control and DevOps practices using Git and GitHub Actions for automated CI/CD pipelines.
AI-Assisted Development: Demonstrated experience successfully integrating AI tools (Copilot, Claude Code) into daily engineering workflows to boost productivity.
Nice to Have
Experience with workflow orchestration tools like Apache Airflow.
Knowledge of Infrastructure as Code (IaC) tools such as Terraform or AWS CloudFormation.
Education
Bachelors or Masters in Information Technology, Computer Science or relevant field.
Work Environment
This job operates in a professional office environment. This role routinely uses standard office equipment, including but not limited to, computers, phones, and photocopiers.
Physical Demands
This position requires the frequent and repetitive use of a computer, keyboard, and mouse. Hand and finger dexterity is required.
Other Duties
Please note this job description is not designed to cover or contain a comprehensive listing of activities, duties or responsibilities that are required of the employee for this job. Duties, responsibilities, and activities may change at any time with or without notice.
EEO
Saama provides equal employment opportunities to all employees and applicants for employment and prohibits discrimination and harassment of any type without regard to race, color, religion, age, sex, national origin, disability status, genetics, protected veteran status, sexual orientation, gender identity or expression, or any other characteristic protected by federal, state or local laws.
This policy applies to all terms and conditions of employment, including recruiting, hiring, placement, promotion, termination, layoff, recall, transfer, leaves of absence, compensation, and training.