Note: Candidate should be working from Customer Denver, CO office from day one
Description:
The Data Developer works within the ETL and Operations team to design, develop, and maintain data pipelines on AWS that drive analytic solutions from diverse, disparate sources (network telemetry, WiFi platform data, billing, and provisioning), routinely processing billions of records per day. The ideal candidate is a skilled data engineer who solves problems in a highly technical, cross-functional environment, with oversight and support from Lead Developers.
Major Duties and Responsibilities:- Coordinate, build, and manage new data ingests from network, WiFi, billing, and provisioning sources.
- Implement updates, fixes, and optimizations across a large suite of production ETL jobs.
- Build and tune Spark SQL and PySpark transformations on AWS EMR at billions-of-records-per-day scale.
- Write and optimize Athena (Trino/Presto) queries for adhoc analysis, validation, and stakeholder support.
- Contribute to the migration from legacy orchestration (Step Functions) to Airflow / MWAA DAGs and a centralized Python Job Framework.
- Participate in GitLab merge request reviews and follow CI/CD deployment workflows (staging to production).
- Monitor production pipelines, respond to data quality alerts and job failures, and partner with analysts and data scientists on aggregation processes.
Required Qualifications- Strong SQL - Spark SQL and Athena (Trino/Presto), including comfort moving logic between both dialects.
- AWS & orchestration - EMR, S3, Glue Data Catalog, Athena, Lambda, Step Functions, and Airflow / MWAA.
- Large-scale data - billions of records per day on partitioned data lakes (Parquet with ZSTD on S3).
- Python scripting - PySpark and orchestration / control scripts.
- Pipelines & optimization - building and tuning pipelines end to end (partition pruning, join strategy, shuffle and AQE tuning).
- Version control & shell - Git / GitLab (branching, MRs, peer review) and Bash for automation on Linux / EMR.
- Collaboration & learning - strong communication with analysts, scientists, and architects; ability to pick up new technologies quickly.
Education & experience:- Bachelor's degree in Computer Science, Engineering, Information Systems, or a related field, or equivalent experience; 2+ years of data engineering / ETL experience.
Preferred Qualifications:- Apache Iceberg - table format, MERGE patterns, and Spark-managed DDL (team is actively adopting).
- Data modeling - dimensional modeling, slowly changing dimensions, and medallion (bronze/silver/gold) layering.
- AI-assisted development - comfort using approved AI coding assistants (Amazon Q, Kiro, GitLab Duo) to accelerate development, review, and troubleshooting; AI / ML pipeline exposure a plus.
- Hadoop / Hive - familiarity with the Hadoop ecosystem and HiveQL for legacy table definitions and SQL-on-Hadoop concepts.
Please add the below to the comment section:
Legal Name:
Current Location: (City, State & Zip Code):
Home location:
Relocate:
Bill Rate:
CTH After 3 Months:
Travelling Availability:
Availability to Start:
Phone/Mobile Number:
Email Address:
Visa Type:
Visa Expiration Date:
Hiring Status:
Are you working directly with the contractor's visa holder:
If not indicate # of layers and names of the company:
Indicate if the Candidate has worked in CG before and where:
Ex-*** Employee:
LinkedIn Account: (If available)
Vendor Point of Contact for Interview Scheduling:
Contractor approved to share its resume to client: