As well, person should be able to monitor the lab/data-center for any issues such as power outages, network outages, liquid cooling leaks, etc… Person should have the ability to physically install and replace hardware, debug FW and OS issues especially related to GPUs. Person will be responsible for going through a ticketing system to check tickets that come in during the individuals working hours: debugging SW and HW problems root causing and fixing said issue.