Software Engineer, Autonomy

Artech LLC

  • Foster City, CA
  • 30+ days ago
  • $75–$97.80 Per Hour

Highlights

Design, build, and maintain Airflow DAGs that orchestrate end-to-end dataset creation and refresh pipelines, including embedding generation, distributed k-means clustering, FAISS index building, embedding cache construction, and dataset statistics computation. Experience designing, building, and optimizing distributed data processing pipelines at scale using technologies like Spark, Databricks, AWS EMR, AWS Batch, or Ray Core/Data.

Numbers & Facts

LocationFoster City, CA
Salary$75–$97.80 Per Hour

Description

Introduction

The Autonomy Behavior ML Data Optimization team is seeking a Software Engineer with strong data processing and pipeline engineering skills. The role involves building, scaling, and optimizing ScenarioScout, a scenario discovery platform that transforms large-scale driving data into actionable insights. This tool empowers teams across the autonomous vehicle development lifecycle by enabling fast, semantic similarity search over millions of driving scenario embeddings.

Required Skills & Qualifications

  • 3 years of professional software engineering experience with a focus on data processing and pipeline engineering.
  • Experience designing, building, and optimizing distributed data processing pipelines at scale using technologies like Spark, Databricks, AWS EMR, AWS Batch, or Ray Core/Data.
  • Strong proficiency in Python with experience building production data pipelines and web services (FastAPI, Uvicorn, or similar async frameworks).
  • Experience building and maintaining data visualization dashboards, with proficiency in SQL, and familiarity with PySpark/Scala for large-scale data manipulation.
  • Experience with workflow orchestration tools (Airflow, Prefect, Dagster, or similar) for managing complex multi-step data processing pipelines.
  • Applicants must be able to work directly for Artech on W2.

Preferred Skills & Qualifications

  • Machine Learning concepts, particularly embeddings, clustering (k-means), and similarity search.
  • Building or operating ML serving infrastructure (e.g., TensorFlow Serving, TorchServe, Triton, or custom model serving like RayServe).
  • Full-stack development with experience owning applications end-to-end from frontend to infrastructure.

Day-to-Day Responsibilities

  • Build and maintain the Python/FastAPI backend powering high-throughput embedding search, bulk execution APIs, and dataset management endpoints.
  • Design, build, and maintain Airflow DAGs that orchestrate end-to-end dataset creation and refresh pipelines, including embedding generation, distributed k-means clustering, FAISS index building, embedding cache construction, and dataset statistics computation.
  • Collaborate with ML researchers on integrating new embedding types, improving embedding quality, and exploring LLM-powered natural language querying capabilities.

For immediate consideration please click APPLY to begin the screening process with Alex.

Similar Jobs

See more jobs