Description: What you do at AMD changes everything
At AMD, we push the boundaries of what is possible. We believe in changing the world for the better by driving innovation in high-performance computing, graphics, and visualization technologies building blocks for gaming, immersive platforms, and the data center.
Developing great technology takes more than talent: it takes amazing people who understand collaboration, respect, and who will go the extra mile to achieve unthinkable results. It takes people who have the passion and desire to disrupt the status quo, push boundaries, deliver innovation, and change the world. If you have this type of passion, come join our team.
JOB TITLE HERE: Emulation Datacenter Engineer
THE ROLE
In this role the individual will work with external vendors and internal AMD teams across business units to deliver industry leading Hardware Emulation and Prototyping datacenter solutions to allow the company to meet challenges in producing state of the art compute, graphics and adaptive silicon technologies.
THE PERSON
Looking for a detail-oriented individual with excellent communication skills to support and develop hardware emulation and prototyping datacenters and infrastructure. This position is part of AMD s Virtual Bring-Up (VBU) central methodology team.
KEY RESPONSIBILITIES
"LSF/Slurm administration and maintenance, analyzing jobs and job data, handling user requests and incidents using current tools and recommending new ones. Supporting multiple/distributed clusters on-prem and cloud. Managing server platform hardware upgrades, Datacenter migrations, and system firmware/software patching.
"Interface with IT Teams to deploy new hardware and resolve issues with existing infrastructure software toolsets.
"Work with external vendors to plan, configure and deploy new hardware / software solutions within AMD's engineering environment.
"Work closely with internal customer teams to develop and deploy new use scenarios and triage and root-cause issues in current production configurations and use-models.
"Responsible for installing, maintaining and repairing hardware, configuring vendors software tools to manage operating system mass deployments & supporting data center infrastructure to accommodate growth.
"Execution of audits of data centers to ensure up time of facility by software reporting tools, reporting on power and cooling usages.
"Responsible for inventory management of data center operational materials.
"Maintain audits to support space planning forecasting. Management to build out related racks and rack infrastructure to plan for hardware installations and moves.
"Responsible for the asset documentation into a Data center infrastructure management tool (DCIM).
"Providing hardware support for IT assets following incident management/change management workflows.
PREFERRED EXPERIENCE
"LSF/Slurm Administration/Use experience.
"Fluent in Unix administration, configuration, driver installation, and general system configuration.
"Enterprise level emulation or prototyping system management and maintenance.
"Physical lab or datacenter space management & system organization in a limited space.
"Focus on continuous improvement to keep the hardware resources running as efficiently as possible.
"Hardware support for datacenter systems and infrastructure.
"AI workflow integration / augmentation.
PREFERRED SOFTWARE EXPERIENCE
"RHEL8 / RHEL9
"Bash / Perl / Python
"Ansible / Puppet / cfengine / pxeboot
"Device42 / Nagios / PowerBI
ACADEMIC CREDENTIALS
"Preferred BS in Computer Science/Engineering or Technical IT Certificate
Custom Fields:
Name: Please indicate the exact site or location
Value: None
Name: Is travel required for this position?
Value: Yes
Name: Work Site Type
Value: Onsite/Hybrid
Name: Is this a returning FTE or Contractor (re-engage)?
Value: No
Name: Time Occasional
Value: No
Name: Is this position related to Veriscale?
Value: No
Name: Is this an Engineering or IT related position?
Value: Yes