| Location | Durham, NC |
The role centers on data modeling, ETL and ELT development, and the ongoing ownership of pipelines, databases, dashboards, reports, and alerting systems. You will be responsible for the full lifecycle of these solutions, including requirements gathering, architecture and design, development, deployment, monitoring, support, and optimization.
This position requires strong communication and cross functional collaboration. You will work with both internal and external stakeholders to ensure data pipelines and automation solutions are reliable, scalable, well documented, and easy to use.Responsibilities• Build and maintain data pipelines and processing frameworks supporting dataset creation, annotation workflows, and large scale data movement across batch and incremental processing patterns.
• Design and maintain data models and pipelines for large scale multimodal datasets including text, image, video, and audio.
• Design and support high volume logging and telemetry systems with a focus on operational efficiency, monitoring, and troubleshooting.
• Develop and maintain dashboards, reports, metrics, and alerting solutions that track throughput, quality, cost, and overall program health.
• Design and maintain data models, schemas, databases, lakehouse environments, dimensional models, and reporting structures.
• Build, optimize, and support Databricks workloads including transformations, orchestration, governance, and performance tuning.
• Implement data quality validation, monitoring, and alerting mechanisms to improve reliability and trust in data assets.
• Develop and maintain workflows that move data into and out of annotation platforms including preprocessing, post processing, validation, and delivery of annotated datasets.
• Support engineering needs related to annotation project user interfaces, testing activities, and quality assurance initiatives.
• Develop backend services, automation tools, and production software using Python and established engineering best practices.
• Troubleshoot complex technical issues spanning multiple systems and identify root causes to implement sustainable solutions.
• Partner with business, engineering, and operational teams to gather requirements, communicate progress, and document systems and procedures.
• Independently evaluate ambiguous requests, design solutions, manage priorities, and execute cross functional initiatives.
• Integrate pipelines and applications with cloud platforms, storage systems, and data annotation technologies while ensuring strong observability and reliability.QualificationsRequired Qualifications
• Bachelor's Degree in Computer Science, Engineering, or a related field, or equivalent professional experience.
• Four or more years of professional experience in Data Engineering or a related technical discipline.
• Proven experience designing, building, and operating production ETL and ELT pipelines supporting large and complex datasets.
• Strong proficiency in SQL and Python.
• Experience with version control systems, testing frameworks, and continuous integration and continuous deployment practices.
• Experience with workflow orchestration tools such as Airflow, Dagster, Databricks Jobs, or similar platforms.
• Experience designing and maintaining data models, databases, warehouses, or lakehouse architectures.
• Experience supporting high volume logging or telemetry solutions.
• Experience building dashboards, reports, metrics, and alerting solutions.
• Hands on experience with Databricks, Spark, Delta Lake, or similar large scale data platforms.
• Experience working with AWS, Azure, or Google Cloud Platform.
• Demonstrated proficiency using AI tools for coding and data related workflows while applying strong engineering fundamentals and technical judgment.
• Strong problem solving and troubleshooting skills.
• Excellent written and verbal communication skills.
• Ability to work independently, manage competing priorities, and proactively communicate status and risks.
Preferred Qualifications
• Eight or more years of Data Engineering or Analytics Engineering experience.
• Experience supporting human in the loop annotation programs and data labeling workflows.
• Direct experience with SuperAnnotate.
• Experience supporting machine learning, Large Language Model, or Vision Language Model training data pipelines.
• Experience working with large scale multimodal datasets including text, image, video, and audio.
• Experience with JavaScript, TypeScript, HTML, or full stack application development supporting internal tools and workflow applications.Tools and Technologies• Python
• SQL
• Databricks
• Apache Spark
• Delta Lake
• Airflow
• Dagster
• Databricks Jobs
• Git
• Continuous Integration and Continuous Deployment
• AWS
• Azure
• Google Cloud Platform
• Data Warehouses
• Lakehouse Platforms