Location: Onsite 5 days per week - Santa Clara, CA
Note: Residents must be onsite
5 days per week. No exceptions
Description:
Updated JD :
Dell will provide PE XE training – approximately 50 hours of required training – Tech Direct – Training is paid and completed during background check process
- PE XE Resident Program // Typically 48 weeks duration –
- Limited XE experience – looking for candidates w/ strong server and linux experience then leverage training to transfer and upskill
- Dedicated full time – 40 hours per week for a set time embedded with customer – day 2 activities and support – bring them in to observe day 1 and move into day 2 – no deployment, not break-fix – help operationalize the deployed equipment
- Technical Expertise – Customer Engagement – Operational Ownership
- PE Rack/Tower experience
- RHEL/Ubuntu – admin experience
- Experience in large environments and customer facing
- Independent and professional communication
- Linux admin skills
- Data Center Operations and Enterprise infrastructure experience
- Runbooks and reporting
PowerEdge XE Resident ResponsibilitiesDay-to-Day ActivitiesData Center Operations- Support daily operational activities.
- Proactively identify and communicate issues, including:
- Amber lights
- Environmental concerns
- Hot aisles/doors
- Escort Dell field engineers
Infrastructure Administration- BIOS lifecycle management
- User administration
- Patching and updates
- Version maintenance
Documentation & Asset Management- Elevation diagrams/rack elevations
- Runbooks
- Serial number tracking
- LOIS parts management (when applicable)
Customer & Dell Coordination- Coordinate maintenance activities
- Schedule break/fix activities (Dell and third-party)
- Perform customer-directed infrastructure updates
Must have skills Dell XE Server, understands Dell Support process and troubleshooting
Nice to have skills (if any) Performance evaluation
- NVIDIA thermal issue.... Taxing the GPU over a year or 2 at a high peg rate. (known issue to Dell under NDA between NVIDIA and Dell)
- Then tanks the GPU at 50% to spread the workload - Ends up degrating
- Zoom has 15 down for 5 weeks (inc. this week)
- Our Support has been very slow - "Massive backlog because of known issue"
- CoreWeave / Meta are front runner for supply - Causing issues for others than need parts
- Taking hours to get logs from the data centers - Using a lot of Zoom FTE times
- Dell then takes 2 weeks to get parts
Goal:
- Put a body in DC to pull logs for Zoom; might need to help to swap parts
- Extra admin"
- Need for 5 days but not continuous - Might need a day or 2 here or there
- Santa Clara - Ai group
- Put together runbooks and documentation