L1 Support Engineer :: Dallas, TX / Austin, TX (Onsite Only)

Talent Movers

Dallas, TX

JOB DETAILS
SKILLS
Amazon Web Services (AWS), Application Programming Interface (API), Atlassian JIRA, Command Line, Configuration Management, DNS (Domain Name System), Engineering, HTTP (HyperText Transport Protocol), Linux Operating System, Load Balancing, MCP - Microsoft Certified Professional, Machine Tool, Multiplatform/Cross-Platform, Network Operations Center, On Call, Operational Support, Operations Management, Pattern Matching, Product Engineering, Production Support, Reporting Dashboards, SSL-TLS (Secure Socket Layer - Transport Layer Security), Slack, Software as a Service (SaaS), Splunk, Standard Operating Procedures (SOP), Technical Support, Traffic Shaping, Writing Skills
LOCATION
Dallas, TX
POSTED
1 day ago

Role :: L1 Support Engineer

Location: Dallas, TX / Austin, TX (Onsite Only)

Rate : C2C for Junior position (2-5 Years experience)

C2C for Mid Level position (5-8 Years' experience)

for Senior (8+ Years)

About the Role

This is a hands-on L1 support role - you will be the first responder for support tickets and PagerDuty pages during US business hours, triaging and validating alerts, driving comprehensive first-level diagnosis, and resolving what can be resolved at L1. When a problem is genuinely beyond L1 scope, you escalate to L2 with a well-evidenced, well-documented handoff - after a thorough first-level investigation, not instead of one.

This is a support role, not a design role - but it is heavy on critical thinking. SOPs and runbooks are your starting point, not a substitute for judgment: the platform generates significant alert volume, including false alerts, and your value is in reasoning through what a page actually means, separating real incidents from noise, and investigating comprehensively before anything reaches L2. It is a strong role for engineers with traffic-domain expertise who want structured production experience on a Tier 0 platform at real enterprise scale.

What You'll Do
  • Serve as the primary first-responder for support tickets, PagerDuty pages, and Slack help requests during US business hours
  • Validate alerts before acting or escalating - distinguish real incidents from false alerts, and flag recurring false positives so alert quality improves over time
  • Apply established SOPs and runbooks - with judgment - to triage and resolve incidents across the platform's Gateway and Mesh components, where the bulk of alert volume lands
  • Perform thorough first-level diagnosis - log correlation, error code lookups (4xx/5xx investigation), certificate expiry checks, traffic pattern review - using established tooling
  • Execute standard operational tasks defined in runbooks - cert renewal steps, config change requests, rate-limit adjustments, service onboarding checklists
  • Escalate to L2 only after comprehensive first-level assessment - with a clean, well-documented handoff (what happened, what you tried, what evidence you gathered, what you think is going on)
  • Own ticket bridge and communications for Sev-3 tickets during your window; loop L2 in for Sev-1 and Sev-2 per escalation matrix
  • Maintain accurate ticket updates in Jira - clear status, next steps, blockers
  • Contribute to SOP and runbook improvements when you notice a step is unclear, outdated, or missing
  • Perform comprehensive shift hand-off to the IDC team at end of your window - signed handover with open tickets and in-flight work
What We're Looking For
  • 2 5 years of hands-on production support / operations engineering experience
  • Working domain expertise in both API gateways and service mesh - API gateways (Apigee, Kong, AWS API Gateway, or comparable) and service mesh (Istio, Envoy, Linkerd) are core to this role, not optional; exposure to configuration management systems, MCP Gateway or protocol gateways, and load balancing / traffic routing is a strong plus
  • Strong critical thinking under pressure - you can form and test hypotheses about unfamiliar failures rather than only pattern-matching against a runbook, and you know when something's off before the dashboard says so
  • Comfortable running Linux command line under time pressure - reading logs, running basic network diagnostics (curl, dig, netstat), navigating kubectl for pod / service inspection
  • Practical experience with PagerDuty, Jira, and Slack in an operational context
  • Reading-level proficiency with observability dashboards - Splunk, Wavefront, Datadog, Grafana, or comparable - enough to spot anomalies and pull evidence for an escalation
  • Strong written communication for incident notes and escalation handoffs - you can write a clear, factual timeline that another engineer can pick up
  • Discipline to apply the SOP accurately, judgment to investigate beyond it when it doesn't apply, and initiative to update it when it needs a fix
  • US work authorization; able to work Austin business hours with occasional secondary on-call rotation
Nice to Have
  • Prior experience in a 24/7 managed-service operations environment (NOC, SOC, or platform support team)
  • Familiarity with kubernetes / container operations - enough to check pod status, read logs, restart a pod when the runbook says to
  • Understanding of network fundamentals - TLS, mTLS, DNS, load balancers, rate limiting, HTTP status codes
  • Prior work in fintech, SaaS, or large-scale consumer product engineering environments where uptime matters
Shift Model
  • Primary window: Austin business hours - 8:00 AM 6:00 PM Central Time (Monday through Friday)
  • On-call rotation: 1-in-4 secondary on-call rotation for Sev-3 escalations during US hours
  • Coverage overlap: 60 minutes of shift-handoff overlap with the IDC (India / Bangalore) team at end of each business day
  • Weekend / holiday coverage: Delivered by IDC team; US L1 team is Monday-Friday

About the Company

T

Talent Movers