ABOUT THE ROLE
REPAY is looking for a Data Engineer to join our growing team. The Data Engineer is responsible for designing, building, optimizing, and maintaining scalable cloud-based data infrastructure and data pipelines that enable reliable data processing, analytics, reporting, and business intelligence capabilities. This role focuses on developing production-grade data pipelines, data models, and ETL/ELT processes using modern data engineering tools and platforms, including AWS, Databricks, PySpark, SQL, and Python. The Data Engineer partners with BI, Product, Engineering, and client-facing teams to ensure high-quality, well-documented, and performance-optimized data solutions that support business insights and operational decision-making.
ROLES & RESPONSIBILITIES
- Design, build, and maintain scalable, reliable cloud-based data pipelines and data infrastructure.
- Deliver high-quality data models and curated datasets that support analytics, reporting, and data-driven decision-making.
- Optimize Spark, PySpark, and SQL workloads to improve performance, reliability, cost efficiency, and scalability.
- Support production data pipelines through monitoring, troubleshooting, incident resolution, and continuous improvement.
- Implement data engineering standards, CI/CD practices, automated deployment processes, unit testing, and code quality expectations.
- Partner with BI, Product, Engineering, and client-facing teams to translate business and reporting requirements into scalable data solutions.
- Document technical solutions, data flows, pipeline logic, and operational processes to support knowledge sharing and long-term maintainability.
- Design, build, maintain, and optimize data pipelines using Python, SQL, PySpark, Databricks, and AWS-based data services.
- Develop ETL/ELT processes that support data warehousing, analytics, reporting, and business intelligence use cases.
- Build and optimize Spark jobs, with a focus on performance, scalability, reliability, and efficient resource utilization.
- Design and implement data models for structured, semi-structured, and NoSQL data where applicable.
- Implement CI/CD practices, automated deployments, unit tests, and code quality standards for data engineering workflows.
- Monitor, troubleshoot, and support production data pipelines, resolving issues and recommending improvements.
- Collaborate with BI Analysts, Product, Engineering, Data, and client-facing teams to understand requirements and support reporting needs.
- Document technical solutions, data flows, pipeline logic, and operational processes.
- Share technical knowledge through documentation, mentorship, and team knowledge-sharing sessions.
- Stay current with advancements in data engineering, cloud platforms, Spark, Databricks, data warehousing, and analytics technologies.
- Participate in client-facing design sessions, technical presentations, workshops, or training as needed.
- Other duties as assigned.
QUALIFICATIONS
Required
- Undergraduate or Masters' degree in Computer Science, Statistics, or Analytics.
- Minimum of 3-5 years of experience in Data Engineering, preferably working with AWS-based cloud data platforms.
- Hands-on experience building, maintaining, and supporting cloud-based data pipelines.
- Strong knowledge of PySpark, preferably on the Databricks platform.
- Hands-on experience with Databricks.
- Strong proficiency in SQL, including query optimization.
- Strong proficiency in Python.
- Strong knowledge of data modeling, data warehousing, ETL/ELT, and analytics concepts.
- Experience with CI/CD practices, automated deployment processes, unit testing, and code quality standards.
- Experience troubleshooting, monitoring, and supporting production data pipelines.
- Experience documenting technical solutions, data flows, and pipeline logic.
- Strong analytical and problem-solving skills, with the ability to translate business requirements into scalable data solutions.
- Excellent written and verbal communication skills, including the ability to explain technical concepts to technical and non-technical stakeholders.
- Ability to collaborate effectively across BI, Product, Engineering, Data, and client-facing teams.
- Strong organizational skills and ability to manage multiple priorities in a fast-paced environment.
- Proactive, ownership-oriented mindset with the ability to work independently and drive solutions from design through production support.
- Professionalism and composure when supporting production issues or participating in client-fac