Data Scientist / Machine Learning Engineer

SMX USA

  • Chantilly, VA
  • 2 days ago
  • Remote
  • $72–$77.35 Per Hour

Highlights

The selected Data Scientist will play an important role in developing predictive models, optimizing sampling and analytical processes, and establishing reproducible machine-learning workflows that integrate with NASS's broader Lakehouse and cloud environment. SMX Services & Consulting is seeking an experienced Senior Data Scientist / Machine Learning Engineer to support a major modernization initiative for USDA's National Agricultural Statistics Service.

Numbers & Facts

LocationChantilly, VA (
Remote
)
Salary$72–$77.35 Per Hour

Description

Senior Data Scientist / Machine Learning Engineer

Agency/Client: U.S. Department of Agriculture (USDA), National Agricultural Statistics Service (NASS)
Program: Sampling Program Enhancement / GENESIS Modernization
Location: Remote — supporting USDA NASS, Washington, DC
Duration: 24-month engagement
Pay Rate:$77.35/hour
Openings: 1
Minimum Experience: 5 years
Previous USDA/NASS Experience: Strong plus

Position Overview

SMX Services & Consulting is seeking an experienced Senior Data Scientist / Machine Learning Engineer to support a major modernization initiative for USDA's National Agricultural Statistics Service.

NASS relies on the Generalized Enhanced Sampling Information System (GENESIS) for sample frame analysis, survey population creation, probability sample selection, Sample Master generation, and downstream survey administration. The modernization effort is transitioning GENESIS away from legacy technologies toward a scalable, cloud-ready, service-oriented architecture supporting modern statistical and AI/ML capabilities.

The selected Data Scientist will play an important role in developing predictive models, optimizing sampling and analytical processes, and establishing reproducible machine-learning workflows that integrate with NASS's broader Lakehouse and cloud environment.

Key Responsibilities

The successful candidate will:

  • Develop predictive models supporting survey optimization.
  • Perform feature engineering and algorithm optimization.
  • Design and tune supervised-learning models.
  • Validate model performance using A/B testing and other appropriate validation techniques.
  • Develop reproducible ML and analytical pipelines.
  • Support integration of AI/ML capabilities with NASS's Lakehouse architecture.
  • Ensure models and analytical processes are reproducible.
  • Support FAIR principles: Findable, Accessible, Interoperable, and Reusable.
  • Incorporate appropriate data-governance practices into model-development workflows.
  • Identify and address potential model bias.
  • Work with developers, architects, statisticians, data specialists, and NASS technical SMEs.
  • Document model design, methodology, assumptions, testing, and results.
  • Support implementation of automated, monitored, and documented data pipelines.
  • Contribute to the modernization team's testing, validation, and knowledge-transfer activities.

The SOW's acceptance criteria specifically require integration with the NASS Lakehouse architecture, automated and monitored pipelines, reproducible AI models, FAIR tagging/metadata, and tiered access governance.

Required Technical Qualifications

Candidates should demonstrate:

  • Minimum 5 years of AI/ML model development and deployment experience.
  • Strong Python skills.
  • Strong SQL skills.
  • Experience with one or more major ML frameworks, including:
    • TensorFlow
    • PyTorch
    • Scikit-learn
  • Supervised machine-learning experience.
  • Feature engineering.
  • Algorithm optimization and tuning.
  • Model testing and validation.
  • Experience designing reproducible pipelines.
  • Hands-on Lakehouse architecture experience.
  • Cloud-platform experience, with Microsoft Azure preferred.
  • Understanding of data governance.
  • Understanding of model bias detection.
  • Experience supporting reproducible analytical environments.

Preferred Background

Particularly attractive candidates will have previous experience supporting:

  • USDA or USDA NASS
  • Federal statistical or data programs
  • Federal government technology programs
  • Large-scale survey or sampling systems
  • Azure data/AI environments
  • Databricks or similar Lakehouse platforms
  • Statistical workflows using Python, R, and/or SAS
  • Highly regulated or sensitive-data environments

Security & Eligibility

The work is unclassified, but the system processes highly sensitive PII. Candidates must be U.S. Citizens or Lawful Permanent Residents and must be able to successfully complete USDA-required fingerprinting and background investigation requirements.

The SOW identifies NACI as the minimum Personnel Security Investigation, while reserving USDA's ability to require a higher investigation and/or clearance depending upon the position sensitivity designation.

Similar Jobs

See more jobs