ML Infrastructure Engineer

Clera

  • San Mateo, California
  • 7 days ago

    Highlights

    This is a hands-on infrastructure engineering role at an early-stage enterprise AI company building a context layer that makes AI agents reliable, accurate, and secure for mission-critical business operations. Hands-on experience designing and scaling inference-serving infrastructure using tools such as TensorFlow Serving, TorchServe, Triton, KServe, or equivalent custom systems.

    Numbers & Facts

    LocationSan Mateo, California

    Description

    About the Role

    This is a hands-on infrastructure engineering role at an early-stage enterprise AI company building a context layer that makes AI agents reliable, accurate, and secure for mission-critical business operations. You'll own the systems that keep those agents running fast and reliably in production — from design through deployment — working closely with ML and infrastructure teams to scale inference at increasing concurrency.

    What You'll Do

    • Own inference and model-serving infrastructure end to end, from architecture design through production deployment.

    • Build and scale systems that enable AI agents to run reliably and efficiently under high concurrency in production environments.

    • Collaborate with ML and infrastructure teams to ensure seamless integration and drive performance optimization.

    • Identify infrastructure bottlenecks and lead the engineering effort to resolve them.

    What We're Looking For

    • 5+ years of experience building and operating machine learning inference systems, model-serving platforms, or ML infrastructure in production environments.

    • Hands-on experience designing and scaling inference-serving infrastructure using tools such as TensorFlow Serving, TorchServe, Triton, KServe, or equivalent custom systems.

    • Demonstrated ability to optimize production ML systems for latency, throughput, and reliability at scale.

    • Strong proficiency with containerization and orchestration technologies — Docker and Kubernetes — for deploying ML workloads.

    • Experience building or maintaining distributed systems that handle concurrent requests and manage resource allocation under load.

    • Solid command of monitoring, observability, and debugging tooling for production systems (e.g., Prometheus, Grafana, ELK, distributed tracing).

    • Experience deploying and managing ML systems on cloud platforms such as AWS, GCP, or Azure.

    • Proficiency in at least one systems or backend language: Python, Go, Rust, C++, or Java.

    • Experience with knowledge graphs, semantic search, or graph databases (e.g., Neo4j, Amazon Neptune) is a plus.

    • Familiarity with real-time or low-latency inference systems, agentic AI pipelines, or enterprise data infrastructure is a plus.

    Location

    On-site in San Mateo, California, United States. Visa sponsorship is not available for this role.

    Similar Jobs

    See more jobs