The consultant will be responsible for understanding end-to-end data flows, developing and troubleshooting data pipelines, optimizing distributed processing workloads, performing SQL-based data reconciliation, and supporting production data platforms. We are looking for a hands-on Data Architect / Data Engineering Consultant with strong expertise in Python, PySpark, Spark optimization, distributed data processing, Hive/Impala, SQL, and production ETL troubleshooting .
Numbers & Facts
Location
Irving, TX
Description
We are looking for a hands-on Data Architect / Data Engineering Consultant with strong expertise in Python, PySpark, Spark optimization, distributed data processing, Hive/Impala, SQL, and production ETL troubleshooting.
The consultant will be responsible for understanding end-to-end data flows, developing and troubleshooting data pipelines, optimizing distributed processing workloads, performing SQL-based data reconciliation, and supporting production data platforms. Required Qualifications
12-15 years in data architecture, data engineering, or enterprise architecture.
7 or more Hands-on development background.
Experience with Databricks and Snowflake. And/or-
Strong expertise in data warehousing, Lakehouse architecture, ETL/ELT, data modelling, and event-driven integration.
Experience in regulated financial services.
Key Responsibilities
Design, develop, enhance, and troubleshoot ETL/ELT data pipelines.
Develop and maintain data processing solutions using Python and PySpark.
Work extensively with Apache Spark, including performance tuning and optimization.
Analyze Spark jobs to identify performance bottlenecks related to partitions, shuffles, joins, data skew, caching, serialization, and resource utilization.
Work with distributed data processing concepts and large-volume datasets.
Develop complex SQL queries for data transformation, validation, reconciliation, and troubleshooting.
Perform source-to-target reconciliation and investigate data discrepancies.
Work with Hive and Impala for querying and processing large datasets.
Troubleshoot production ETL failures, data quality issues, performance problems, and batch-processing failures.
Perform root-cause analysis and implement permanent fixes for recurring production issues.
Understand and troubleshoot end-to-end data flows, from source systems through ETL processing to downstream consumers.
Work with Linux environments, shell commands, batch processing, and job scheduling.
Collaborate with engineering, application, and business teams to resolve complex data issues.
Participate in technical design discussions and provide recommendations for scalable and maintainable data solutions.
Document technical designs, data flows, troubleshooting procedures, and production resolutions