Site Reliability Engineer

Artech LLC

  • Dearborn, MI
  • 6 days ago

    Highlights

    You'll work hands-on with cloud infrastructure, BigQuery workloads, CI/CD pipelines, and enterprise monitoring tools to keep critical systems healthy, performant, and reliable at scale. Participate in building advanced tooling for system access monitoring, log session recording, administration of reliability across multiple geographically distributed data centers.

    Numbers & Facts

    LocationDearborn, MI

    Description

    Introduction

    We are seeking a skilled professional to join our team, focused on observability, monitoring, and technical consulting across our GCP-based data platforms. You'll work hands-on with cloud infrastructure, BigQuery workloads, CI/CD pipelines, and enterprise monitoring tools to keep critical systems healthy, performant, and reliable at scale.

    Required Skills & Qualifications

    • Proficiency with Google Cloud Platform (GCP)
    • Experience with monitoring/observability tools, ideally Dynatrace or comparable tools like Datadog, New Relic
    • Familiarity with ITSM tools such as ServiceNow (incident, problem, change management)
    • Practitioner experience in at least one coding language or framework
    • 4 years in IT and 3 years in development
    • Prior work experience at client or in client's Industry
    • Applicants must be able to work directly for Artech on W2

    Preferred Skills & Qualifications

    • Familiarity with the use of AI tools – agents, skills, LLMs, copilot
    • Experience defining and tracking SLAs/SLOs/SLIs
    • Experience with GCP Cloud Run and Python

    Day-to-Day Responsibilities

    • Collaborate with Infrastructure teams in implementing critical solutions by automating routine tasks
    • Monitor and manage production environments, proactively identifying and resolving issues
    • Participate in building advanced tooling for system access monitoring, log session recording, administration of reliability across multiple geographically distributed data centers
    • Engage with engineering teams to improve on-call efficiencies, drive incident management and post-mortem analysis
    • Perform capacity planning and optimization to support growing demands and traffic patterns
    • Maintain, monitor, and alert systems for proactive system health checks
    • Continuously improve system performance, stability, and security through data-driven analysis and optimization
    • Facilitate knowledge sharing by creating and maintaining comprehensive documentation & diagrams

    For immediate consideration please click APPLY to begin the screening process with Alex.

    Similar Jobs

    HTC Global Services Inc

    CDN Platform Engineer

    • Dearborn, MI
    28 days ago
    See more jobs