Job Mission
Join a pioneering research organization developing next-generation lithography light source technologies. Our laser-produced plasma (LPP) system combine high-power lasers, advanced optics, plasma-based EUV generation, sensing, controls, and physics-based modeling to address some of the most complex challenges in semiconductor manufacturing.
As Source Research continues to expand its use of data-driven engineering, machine learning, simulation and physics-based modeling, a scalable and well-governed data ecosystem has become essential. High-quality, accessible, and connected data enables faster technology development, deeper system understanding, more effective trade studies, and better-informed technology and roadmap decisions.
In this role, you will help shape the data foundation that supports research and development activities across Source Research. Working closely with lab owners and experimental, modeling, and ML scientists, you will build and improve data pipelines, integrate diverse data sources, and enable reliable access to research data at scale. You will also help establish practical architecture standards and best practices that ensure our data platform remains scalable, secure, maintainable, and aligned with the broader data landscape.
his role combines hands-on development with technical leadership in shaping the data foundation for Source Research. You will build, operate, and continuously improve data pipelines, integrating new data sources, improving reliability, and enabling scientists and engineers to use high-quality data at scale. You will also define practical architecture standards that keep the platform consistent, secure, future-ready, and aligned with clients data landscape.
Qualifications
• Bachelor’s or Master’s degree in Computer Science, Statistics, Math, Data Science, or a related field.
• 10+ years of relevant experience in data engineering, data architecture, or scientific/engineering data platforms.
• Strong hands-on development experience in Python and modern data engineering tooling.
• Proven experience building and operating scalable big data pipelines, analytics platforms, and data products that support data-intensive scientific and engineering workflows.
• Experience with cloud and distributed data platforms such as Azure, Azure Databricks, Apache Spark, Kubernetes, and data lake architectures.
• Solid understanding of data modeling, metadata management, data lineage, data quality, governance, security and access control.
• Experience supporting scientific or engineering workflows (for example simulation, HPC, instrumentation, or time-series sensor data).
• Familiarity with AI/ML workflows and MLOps practices is a plus.
• Strong communication and collaboration skills, with the ability to translate technical details into clear and actionable guidance.