You will bridge the gap between GPU kernels and production serving, focusing on distributed inference orchestration, hardware-software co-optimization, and the elimination of system-level bottlenecks to ensure maximum throughput and minimum latency. Profiling & Benchmarking: Systematically identify AI bottlenecks (NVLink, PCIe, HBM bandwidth) and establish rigorous TPS/TTFT benchmarking suites.