Amazon Web Services (AWS), Application Programming Interface (API), Atlassian JIRA, Command Line, Configuration Management, DNS (Domain Name System), Engineering, HTTP (HyperText Transport Protocol), Linux Operating System, Load Balancing, MCP - Microsoft Certified Professional, Machine Tool, Multiplatform/Cross-Platform, Network Operations Center, On Call, Operational Support, Operations Management, Pattern Matching, Product Engineering, Production Support, Reporting Dashboards, SSL-TLS (Secure Socket Layer - Transport Layer Security), Slack, Software as a Service (SaaS), Splunk, Standard Operating Procedures (SOP), Technical Support, Traffic Shaping, Writing Skills
LOCATION
Dallas, TX
POSTED
1 day ago
Role :: L1 Support Engineer
Location: Dallas, TX / Austin, TX (Onsite Only)
Rate : C2C for Junior position (2-5 Years experience)
C2C for Mid Level position (5-8 Years' experience)
for Senior (8+ Years)
About the Role
This is a hands-on L1 support role - you will be the first responder for support tickets and PagerDuty pages during US business hours, triaging and validating alerts, driving comprehensive first-level diagnosis, and resolving what can be resolved at L1. When a problem is genuinely beyond L1 scope, you escalate to L2 with a well-evidenced, well-documented handoff - after a thorough first-level investigation, not instead of one.
This is a support role, not a design role - but it is heavy on critical thinking. SOPs and runbooks are your starting point, not a substitute for judgment: the platform generates significant alert volume, including false alerts, and your value is in reasoning through what a page actually means, separating real incidents from noise, and investigating comprehensively before anything reaches L2. It is a strong role for engineers with traffic-domain expertise who want structured production experience on a Tier 0 platform at real enterprise scale.
What You'll Do
Serve as the primary first-responder for support tickets, PagerDuty pages, and Slack help requests during US business hours
Validate alerts before acting or escalating - distinguish real incidents from false alerts, and flag recurring false positives so alert quality improves over time
Apply established SOPs and runbooks - with judgment - to triage and resolve incidents across the platform's Gateway and Mesh components, where the bulk of alert volume lands
Execute standard operational tasks defined in runbooks - cert renewal steps, config change requests, rate-limit adjustments, service onboarding checklists
Escalate to L2 only after comprehensive first-level assessment - with a clean, well-documented handoff (what happened, what you tried, what evidence you gathered, what you think is going on)
Own ticket bridge and communications for Sev-3 tickets during your window; loop L2 in for Sev-1 and Sev-2 per escalation matrix
Maintain accurate ticket updates in Jira - clear status, next steps, blockers
Contribute to SOP and runbook improvements when you notice a step is unclear, outdated, or missing
Perform comprehensive shift hand-off to the IDC team at end of your window - signed handover with open tickets and in-flight work
What We're Looking For
2 5 years of hands-on production support / operations engineering experience
Working domain expertise in both API gateways and service mesh - API gateways (Apigee, Kong, AWS API Gateway, or comparable) and service mesh (Istio, Envoy, Linkerd) are core to this role, not optional; exposure to configuration management systems, MCP Gateway or protocol gateways, and load balancing / traffic routing is a strong plus
Strong critical thinking under pressure - you can form and test hypotheses about unfamiliar failures rather than only pattern-matching against a runbook, and you know when something's off before the dashboard says so
Comfortable running Linux command line under time pressure - reading logs, running basic network diagnostics (curl, dig, netstat), navigating kubectl for pod / service inspection
Practical experience with PagerDuty, Jira, and Slack in an operational context
Reading-level proficiency with observability dashboards - Splunk, Wavefront, Datadog, Grafana, or comparable - enough to spot anomalies and pull evidence for an escalation
Strong written communication for incident notes and escalation handoffs - you can write a clear, factual timeline that another engineer can pick up
Discipline to apply the SOP accurately, judgment to investigate beyond it when it doesn't apply, and initiative to update it when it needs a fix
US work authorization; able to work Austin business hours with occasional secondary on-call rotation
Nice to Have
Prior experience in a 24/7 managed-service operations environment (NOC, SOC, or platform support team)
Familiarity with kubernetes / container operations - enough to check pod status, read logs, restart a pod when the runbook says to
Understanding of network fundamentals - TLS, mTLS, DNS, load balancers, rate limiting, HTTP status codes
Prior work in fintech, SaaS, or large-scale consumer product engineering environments where uptime matters
Shift Model
Primary window: Austin business hours - 8:00 AM 6:00 PM Central Time (Monday through Friday)
On-call rotation: 1-in-4 secondary on-call rotation for Sev-3 escalations during US hours
Coverage overlap: 60 minutes of shift-handoff overlap with the IDC (India / Bangalore) team at end of each business day
Weekend / holiday coverage: Delivered by IDC team; US L1 team is Monday-Friday