Member of Technical Staff, Inference Systems

Transparent Search Group

  • Palo Alto, California
  • 16 days ago
  • $230,000–$350,000 Per Year

Highlights

BEST-FIT CANDIDATE: 2+ yrs; LLM inference internals; vLLM/SGLang/TRT-LLM; Rust/C++ systems; visa: transfers + new H-1B/TN; location: Palo Alto 5 days. TARGET COMPANIES (suggested (vLLM/SGLang/TRT-LLM contributors)): Together AI, Fireworks AI, Baseten, Modal, NVIDIA, Anyscale.

Numbers & Facts

LocationPalo Alto, California
Salary$230,000–$350,000 Per Year

Description

Member of Technical Staff, Inference Systems

Company: Photon
Location: Palo Alto, CA (on-site 5 days per week)
Compensation: $230,000 - $350,000 + competitive equity
Employment Type: Full-time
Visa Sponsorship: Visa transfers and new sponsorship (H-1B, TN)

About Photon

Photon is building a next-generation AI inference platform from the ground up with a relentless focus on performance. It was founded by Stanford alumni with deep AI infrastructure experience, including early work at Together AI, and is engineering the entire inference stack with Rust at its core.

Photon has raised a $10M seed from notable investors and is currently in stealth ahead of announcing its fundraise and product.

The Role

Photon is hiring Members of Technical Staff (2+ years) to build a high-performance inference platform from scratch. You are a systems engineer who knows inference internals (attention, KV cache, batching, scheduling) and wants to own the whole stack rather than a narrow slice. You will join a small team building a new inference system in Rust, where you shape every design decision.

What You Will Do

  • Build a new inference runtime in Rust, owning batching, scheduling, request routing and the serving stack.
  • Design KV cache management, prefix caching and optimizations that reduce latency and cost per token.
  • Scale serving across GPUs and nodes.
  • Profile, benchmark and ship performance improvements across the inference pipeline.
  • Make core architecture decisions with the founding team.

What You Bring

  • 2+ years of systems engineering experience
  • Deep knowledge of inference internals: attention, KV cache, batching and scheduling
  • Experience inside inference engines such as vLLM, SGLang or TensorRT-LLM
  • Strong systems programming (Rust, C++ or similar)
  • Ability to work on-site in Palo Alto 5 days a week

Tech Stack

Rust, Python, PyTorch, C++, Go, vLLM, SGLang, TensorRT-LLM, CUDA, Triton, NCCL


REVENUE: 14% of first-year salary. Est. fee per hire $32K-$49K; 10 seat(s) = up to $406K if all filled.

TARGET COMPANIES (suggested (vLLM/SGLang/TRT-LLM contributors)): Together AI, Fireworks AI, Baseten, Modal, NVIDIA, Anyscale.

BEST-FIT CANDIDATE: 2+ yrs; LLM inference internals; vLLM/SGLang/TRT-LLM; Rust/C++ systems; visa: transfers + new H-1B/TN; location: Palo Alto 5 days.

Similar Jobs

See more jobs