Senior Site Reliability Engineer- San Francisco, CA, The US

KodyPay Ltd

  • San Francisco, CA
  • 30+ days ago

    Highlights

    You will own production observability, incident response, service-level management, and cloud infrastructure reliability across mission-critical payment processing systems operating in Europe, Asia, and North America. Diagnose, triage, mitigate, and coordinate resolution of production incidents across payment services, Kubernetes platforms, databases, messaging systems, and cloud infrastructure.

    Numbers & Facts

    LocationSan Francisco, CA

    Description

    Senior Site Reliability Engineer (Payments Infrastructure)

    Kody is seeking a Senior Site Reliability Engineer to ensure the reliability, availability, scalability, and operational excellence of our global payment platform. You will own production observability, incident response, service-level management, and cloud infrastructure reliability across mission-critical payment processing systems operating in Europe, Asia, and North America.

    Responsibilities

    • Participate in a follow-the-sun production on-call rotation as a primary incident responder.
    • Diagnose, triage, mitigate, and coordinate resolution of production incidents across payment services, Kubernetes platforms, databases, messaging systems, and cloud infrastructure.
    • Define and maintain SLOs, SLIs, error budgets, alerting standards, and operational readiness processes.
    • Drive reliability improvements through automation, observability, capacity planning, performance optimization, and post-incident reviews.
    • Partner with engineering teams to improve resilience, security, and operational maturity in PCI-DSS-regulated environments.
    • Lead incident management during SEV1/SEV2 events and improve response effectiveness and MTTR.

    Similar Jobs

    See more jobs