Fantom Corporation is a mission-focused organization supporting critical programs across the defense and intelligence community. We partner with our customers to deliver high-impact technical solutions while fostering a culture built on trust, expertise, and long-term career growth.
We are seeking a Data & Software Engineer to join a small, agile team developing advanced data solutions for a custom application supporting mission-critical operations. This role is focused on designing, building, and optimizing scalable data pipelines and ETL workflows while leveraging modern cloud technologies, distributed data processing frameworks, and AI/ML capabilities.
The ideal candidate has strong Python programming skills, experience building production-grade data pipelines, and expertise with Apache Spark, AWS cloud services, relational databases, and modern data engineering best practices.
Responsibilities
- Design, develop, and maintain scalable, end-to-end data pipelines and ETL/ELT workflows supporting enterprise applications
- Build and optimize distributed data processing solutions using Apache Spark and PySpark
- Develop robust Python applications and automation scripts to support data engineering initiatives
- Deploy and manage data pipelines using workflow orchestration platforms such as AWS Step Functions or Apache Airflow
- Containerize and deploy data applications within AWS cloud environments using Docker, Podman, or similar technologies
- Design, optimize, and maintain PostgreSQL and MySQL databases, including schema design, indexing, and query optimization for analytical workloads
- Collaborate with stakeholders to gather requirements, assess technical feasibility, and design scalable data solutions
- Troubleshoot and resolve data quality issues, pipeline failures, and performance bottlenecks
- Implement and support data governance, security, privacy, and compliance best practices
- Develop and maintain technical documentation, architecture diagrams, and engineering standards
- Support large-scale data migration and platform modernization initiatives
- Integrate AI/ML services, machine learning models, and advanced analytics capabilities into enterprise data platforms
- Utilize Git and CI/CD pipelines to support automated testing, deployment, and version control
Required Qualifications
- Must be fully cleared with a recent polygraph
- Must be willing and able to work fully onsite at the location listed in this posting
- 5+ years of experience as a Data Engineer, Software Engineer, or similar technical role
- Demonstrated experience building production-scale data pipelines and ETL/ELT workflows
- Advanced programming experience with Python, including libraries such as Pandas and NumPy
- Strong experience with Apache Spark and PySpark for distributed data processing
- Experience with workflow orchestration tools such as Apache Airflow or AWS Step Functions
- Experience deploying containerized applications using Docker, Podman, or similar technologies
- Experience developing cloud-native applications within AWS environments
- Hands-on experience with AWS services including Amazon S3, AWS Lambda, and AWS Step Functions
- Strong experience with PostgreSQL and MySQL, including schema design, performance tuning, and query optimization
- Advanced SQL skills supporting large-scale analytical workloads
- Experience with Git and CI/CD practices for data engineering and software deployment
- Understanding of data governance, privacy, security, and compliance principles
- Strong analytical, troubleshooting, and problem-solving skills
- Experience collaborating directly with stakeholders to gather requirements and deliver technical solutions with minimal oversight
- #CJ
Desired Qualification
- Experience designing and implementing Lakehouse architectures using Apache Iceberg
- Experience configuring and supporting enterprise data platform technologies, including:
- Apache Ranger
- Trino
- Apache Polaris or Unity Catalog OSS
- Apache Superset
- Experience with Infrastructure as Code (Terraform or AWS CloudFormation)
- Proficiency with Bash scripting for automation and operational support
- Experience tracking data lineage using OpenLineage or similar tools
- Working knowledge of Java
- Experience implementing data quality frameworks, automated testing, and validation strategies
- Experience supporting enterprise data modernization and large-scale migration initiatives
- Experience integrating AI/ML services, including OCR, NLP, speech-to-text, language detection, translation services, topic modeling, Large Language Models (LLMs), and Retrieval-Augmented Generation (RAG) pipelines
- Experience processing geospatial data using technologies such as H3, PostGIS, or similar frameworks
- Experience with NoSQL databases such as DynamoDB
- Experience developing engineering standards, documentation, and reusable design patterns