Production Support Engineer

ASM Tech Solutions LLC

  • Lake mary, FL
  • Today

    Highlights

    Key Responsibilities Monitor, troubleshoot, and resolve Level 2 production incidents across AI platforms, cloud infrastructure, data pipelines, model-serving environments, and associated services. Build and maintain monitoring, alerting, and observability capabilities for infrastructure, applications, data pipelines, model operations, and distributed compute workloads.

    Numbers & Facts

    LocationLake mary, FL

    Description

    As of now regular shift during US hours. Hybrid work ( 3 days from Office is Mandatory and 2 days remote)
    Location: Lake Mary or 240 NY Office

    Key Responsibilities
    1. Monitor, troubleshoot, and resolve Level 2 production incidents across AI platforms, cloud infrastructure, data pipelines, model-serving environments, and associated services.
    2. Provide hands-on support for deployment, orchestration, and operational management of AI/ML workloads across cloud-native environments.
    3. Build and maintain monitoring, alerting, and observability capabilities for infrastructure, applications, data pipelines, model operations, and distributed compute workloads.
    4. Perform root cause analysis for production issues, implement permanent fixes, and drive problem-management activities to improve service reliability.
    5. Collaborate with DevOps, MLOps, data engineering, platform engineering, and application teams to maintain and enhance CI/CD and deployment automation.
    6. Develop and support automation, self-service tooling, recovery mechanisms, and self-healing controls to reduce manual operational effort.
    7. Monitor and troubleshoot batch processes, workflow orchestration, data ingestion, model training, and model deployment failures.
    8. Apply ITIL-based incident, problem, change, and release management processes to support stable production operations.
    9. Analyse support-ticket and incident trends; recommend and implement operational improvements, including AI-driven automation where appropriate.

    Qualifications & Skills
    Mandatory:
    1. 3–5 years of experience in Level 2 application, platform, DevOps, or production support roles.
    2. Strong hands-on experience with UNIX/Linux, SQL, and shell or Python scripting.
    3. Experience troubleshooting cloud-native applications, distributed systems, containers, and Kubernetes-based environments.
    4. Working knowledge of CI/CD pipelines, deployment automation, and source-control platforms such as GitLab.
    5. Experience with monitoring, logging, and observability tools such as Splunk, Grafana, AppDynamics, Prometheus, or similar tools.
    6. Understanding of incident management, root cause analysis, problem management, and ITIL support processes.
    7. Strong analytical and problem-solving skills, with a client-service mindset.
    8. Ability to troubleshoot data-pipeline, workflow, API, and production deployment issues.

    Good-to-Have:
    1. Exposure to MLOps practices, including model deployment, model monitoring, feature/data pipelines, and AI workload orchestration.
    2. Experience with cloud platforms such as AWS, Azure, or GCP.
    3. Experience with infrastructure automation and configuration-management tools such as Ansible, Terraform, or Ansible Tower.
    4. Familiarity with workflow and job-scheduling tools, such as Contro-M or equivalent enterprise schedulers.
    5. Knowledge of Docker, Kubernetes, and distributed compute/data-processing technologies.
    6. Exposure to Kafka, MQ, or other messaging and event-streaming platforms.
    7. Experience implementing self-healing, auto-remediation, resilience, backup, or disaster-recovery mechanisms.
    8. Familiarity with AI-driven operational tooling, ticket-trend analysis, or automation agents.

    Similar Jobs

    See more jobs