System One logo

Senior Production Support Engineer - Apigee

System One

  • Lafayette, LA
  • 3 days ago
  • $85,000 Per Year

Highlights

This role is responsible for troubleshooting complex Kubernetes workloads, API gateway and runtime issues, routing and backend connectivity, certificates/TLS, DNS, HTTP, and other platform related production problems. The engineer will lead technical triage, coordinate resolution across application, infrastructure, network, cloud, and Level 3 engineering teams, and support releases, platform changes, failover activities, and production recovery.
System One

Numbers & Facts

LocationLafayette, LA
IndustryStaffing/Employment Agencies
Salary$85,000 Per Year
Company Size2,500 to 4,999 employees
Websitehttps://systemone.com

Description


Job Title: Senior Production Support Engineer - Apigee
Duration: Full Time / Permanent Position
Location: Lafayette, LA, Knoxville, TN, Columbia, SC, Birmingham, AL
Work Mode:
5 Days Onsite

Position Description
Systemone is seeking a Senior Production Support Engineer to provide advanced operational support for a highly available enterprise API platform running on Kubernetes and Apigee Hybrid within an AWS environment. Candidates must have deep Kubernetes and Apigee experience is required rather than general AWS support alone.

This role is responsible for troubleshooting complex Kubernetes workloads, API gateway and runtime issues, routing and backend connectivity, certificates/TLS, DNS, HTTP, and other platform related production problems. The engineer will lead technical triage, coordinate resolution across application, infrastructure, network, cloud, and Level 3 engineering teams, and support releases, platform changes, failover activities, and production recovery.

The position is primarily focused on production support and operational engineering, with targeted scripting and automation rather than application feature development.

Your future duties and responsibilities

  • Provide senior level production support for Kubernetes and Apigee Hybrid environments.
  • Troubleshoot Kubernetes pods, deployments, replica sets, services, configurations, namespaces, and runtime issues.
  • Use kubectl to investigate workload health, application failures, and configuration problems.
  • Troubleshoot Apigee API Gateway, including API proxies, routing, target endpoints, runtime components, and backend integrations.
  • Investigate certificates, TLS, DNS, HTTP, and application/network connectivity issues.
  • Use curl, ping, traceroute, and related utilities to diagnose API and network problems.
  • Use AWS, Linux, Splunk, and CloudWatch to investigate complex production issues.
  • Lead technical triage and coordinate incident resolution across engineering and support teams.
  • Communicate technical findings, business impact, risks, mitigation plans, and recommended actions.
  • Support planned releases, platform changes, failover activities, and production recovery.
  • Maintain technical documentation and contribute to operational, automation, and resiliency improvements.
  • Participate in rotating 24x7 primary on call coverage.

Required qualifications to be successful in this role
  • 8+ years of hands on Kubernetes administration and production troubleshooting experience.
  • Hands on Apigee API Gateway experience; Apigee Hybrid experience is strongly preferred.
  • Advanced Linux troubleshooting skills and strong hands on experience with kubectl.
  • Strong knowledge of Kubernetes pods, deployments, replica sets, services, configurations, and namespaces.
  • Experience troubleshooting API proxies, routing, target endpoints, certificates, TLS, and backend integrations.
  • Working knowledge of TCP/IP, DNS, and HTTP, along with curl, ping, and traceroute.
  • Experience supporting applications and platforms in an AWS production environment.
  • Experience with technical incident management, production troubleshooting, and complex incident triage.
  • Ability to use logs, metrics, dashboards, and alerts to isolate production issues.
  • Strong communication and stakeholder management skills, particularly during production incidents.
  • Ability to work independently on complex investigations and coordinate effectively with Level 3 and engineering teams.

Desired Skillset:
?Preferred: AWS Associate certification, Splunk, CloudWatch, GitLab, CI/CD, and experience supporting API platforms within a regulated enterprise.


Educational Requirements:
Bachelor's degree in Computer Science, Information Systems, Engineering, or a related technical field.



Ref: #404-IT Pittsburgh


About Company

Every day, System One focuses on services and solutions that require a high degree of specialization, in-demand technical skills, and large-scale operational expertise. We are essential partners to those on the front lines of our nation’s most critical infrastructure, technology, and life sciences initiatives. 

Founded more than 40 years ago as a staffing partner to the engineering industry, today System One is a diversified organization operating in over 50 locations and putting more than 9,000 people to work in the United States, Canada, and the United Kingdom.

Similar Jobs