Senior Site Reliability Engineer (Payments Infrastructure)Kody is seeking a Senior Site Reliability Engineer to ensure the reliability, availability, scalability, and operational excellence of our global payment platform. You will own production observability, incident response, service-level management, and cloud infrastructure reliability across mission‑critical payment processing systems operating in Europe, Asia, and North America.ResponsibilitiesParticipate in a follow-the-sun production on‑call rotation as a primary incident responderDiagnose, triage, mitigate, and coordinate resolution of production incidents across payment services, Kubernetes platforms, databases, messaging systems, and cloud infrastructureDefine and maintain SLOs, SLIs, error budgets, alerting standards, and operational readiness processesDrive reliability improvements through automation, observability, capacity planning, performance optimization, and post‑incident reviewsPartner with engineering teams to improve resilience, security, and operational maturity in PCI‑DSS‑regulated environmentsLead incident management during SEV1/SEV2 events and improve response effectiveness and MTTRCross‑Border Collaboration: Act as a key technical bridge between our US operations and international engineering hubs, leveraging bilingual communication to streamline complex technical alignmentRequirements5+ years of experience in Site Reliability Engineering, Platform Engineering, DevOps, or Cloud Infrastructure roles supporting mission‑critical production systemsStrong hands‑on experience with AWS, Kubernetes (EKS), Terraform, PostgreSQL, Redis, Kafka, Linux, networking, and modern observability platformsDeep understanding of distributed systems, cloud‑native architectures, high availability, disaster recovery, capacity planning, and performance optimizationProven experience operating payment, banking, fintech, or other highly regulated systems with stringent security, compliance, and uptime requirementsStrong knowledge of SRE principles, including SLOs, SLIs, error budgets, incident management, alert governance, and operational excellenceLeadership & Operational ExcellenceDemonstrates strong ownership and accountability, taking end‑to‑end responsibility for service reliability and customer impactPossesses a strong sense of urgency during production incidents while maintaining sound judgment and structured decision‑making under pressureApplies a systematic and methodical approach to troubleshooting, root‑cause analysis, and incident resolution in complex distributed environmentsData‑driven mindset with the ability to leverage metrics, telemetry, trends, and service‑level indicators to prioritize reliability investments and operational improvementsContinuously drives engineering excellence through iterative improvement, automation, standardization, and elimination of operational toilProven ability to lead cross‑functional incident response efforts, coordinate stakeholders, and communicate effectively during high‑severity production eventsChampions a culture of operational readiness, continuous learning, post‑incident improvement, and blameless accountabilityDemonstrates strong mentoring and technical leadership skills, influencing engineering teams to build reliable, scalable, and resilient systems by designBenefitsCompetitive packages aligned with California market standardsLead a dynamic and innovative team in a very rapidly growing companyCollaborative, inclusive environment where your contributions are recognized and valued#J-18808-Ljbffr