Job Title: AI Lab Tech Engineer
Location: Fremont, CA (Hybrid: 2 days Onsite / 3 days Remote)
Job Type: Long-Term Contract
Role Overview
We are seeking an Infrastructure Engineer to own the execution layer beneath our Reinforcement Learning (RL) environments: the core systems that enable AI agents to operate coherently inside realistic, multi-tool worlds for extended periods.
This position focuses on hard systems engineering applied to AI infrastructure. As agent task horizons expand, our training environments must maintain stability, speed, and cost-efficiency. You will build and scale the platform that provides secure sandboxing, high-performance execution, and state-management capabilities (snapshotting, restoring, inspecting, and branching running environments) for production workloads. You will collaborate with research and data teams, frontier labs, and enterprise clients to translate complex environment requirements into resilient systems.
Key Responsibilities
Environment Execution & Sandboxing: Design and maintain the isolation and execution layer for RL environments. Build systems to snapshot and restore environment state (disk, process, memory, and accelerator state) to enable pausing, resuming, inspecting, and branching agent rollouts.
Fault Detection & Recovery: Implement automated machinery to identify failure modes early (infra faults, reward hacks, fairness issues), revert to known-good states, patch, and resume execution.
Long-Horizon & Multi-Node Scaling: Extend execution capabilities to handle complex, multi-node environments where agents interact across multiple tools and services over hours or days.
Performance & Cost Optimization: Drive throughput, latency optimization, and cost-per-rollout efficiency. Manage scheduling and resource utilization to maximize environment rollouts per dollar without compromising system reliability.
Observability & Profiling: Profile bottlenecks from container initialization to teardown. Develop deep observability into thousands of concurrent, long-running agent rollouts.
Environment Platform Engineering: Build frameworks and tooling for specifying, packaging, and deploying RL environments used by internal researchers and agents. Provide debugging tooling to trace failures across complex agent runs.
Model Deployment: Deploy and manage small and large AI models on on-premises hardware.
Production Standards: Scale prototypes into production-grade systems with high testing, validation, and documentation standards.
Required Qualifications & Skills
Systems & Infrastructure: Proven background in building production distributed systems, execution engines, or container/sandboxing infrastructure at scale.
Low-Level Systems Knowledge: Deep technical experience with containerization and isolation tech (namespaces, cgroups, VMs, gVisor, Firecracker), filesystems, and process/state management.
Languages & Tools: High proficiency in Python. Experience with systems languages (Rust, Go, or C++) and modern development tools (e.g., Claude Code) is strongly preferred.
Cloud & On-Premises: Practical experience with major cloud platforms (AWS, GCP) alongside hands-on experience deploying models on on-premises hardware.
Performance Engineering: Demonstrated track record in profiling, workload scheduling, resource management, and cost optimization for compute-heavy workloads.
Soft Skills: Strong communication skills to effectively translate research requirements into production infrastructure and engage directly with technical stakeholders and enterprise clients.