Senior Software Engineer, Data Platform

Rheaction

  • Los Angeles, California
  • 13 days ago
  • Remote

    Highlights

    This is a highly autonomous, zero-to-one role focused on turning large volumes of complex, real-world data into clean, consistent, and reliable datasets for machine learning and product teams. Experience with AWS or another major cloud platform and deploying production data pipelines.

    Numbers & Facts

    LocationLos Angeles, California (
    Remote
    )

    Description

    Location: Remote — Los Angeles area preferred
    Employment: Full-time
    Experience: 5–8 years
    Compensation: $200,000–$240,000 + 0.25%–0.5% equity
    Work Authorization: U.S. citizenship required

    About the Role

    We're looking for a Senior Software Engineer, Data Platform to build and own a data platform from the ground up.

    This is a highly autonomous, zero-to-one role focused on turning large volumes of complex, real-world data into clean, consistent, and reliable datasets for machine learning and product teams.

    You'll architect high-throughput pipelines, establish data quality and reproducibility standards, and build the infrastructure and tooling needed to make data easy to discover, process, label, and use.

    What You'll Do

    • Architect and build high-throughput data pipelines from scratch.
    • Develop pipelines for processing, aligning, transforming, and calibrating complex datasets.
    • Build backfill and reprocessing frameworks for historical data.
    • Establish dataset lineage, versioning, and reproducibility.
    • Build APIs and tooling for dataset discovery and access.
    • Develop data-quality metrics, monitoring, and alerting.
    • Optimize storage, compression, partitioning, and data access patterns.
    • Build internal tools and workflows for ML and data-labeling teams.
    • Partner closely with ML engineers to understand and support their data requirements.

    What We're Looking For

    • 5–8 years of software or data engineering experience.
    • Proven ability to architect a data platform or high-throughput pipeline from a blank page.
    • Strong Python and data tooling experience, such as PyArrow, Polars, Pandas, NumPy, or SciPy.
    • Experience with AWS or another major cloud platform and deploying production data pipelines.
    • Experience with a services language such as Go, Rust, or TypeScript.
    • Strong understanding of data lineage, versioning, reproducibility, and data quality.
    • Experience building backfill and reprocessing systems.
    • High ownership and ability to operate independently in a startup environment.
    • U.S. citizenship required.

    Strong Plus

    • Founding or early data-platform engineering experience.
    • Experience with sensor, audio, video, telemetry, or geospatial data.
    • Airflow, Prefect, MLflow, W&B, or similar tooling.
    • Experience building labeling tools or ML data workflows.
    • Strong full-stack/generalist capabilities.
    • Experience treating datasets as a product, with strong documentation and discoverability.

    Similar Jobs