Apply now »

Specialist - System Management

Job Req Id:  1482101

We are currently looking for a Site Reliability Engineer (SRE) to join our team supporting mission-critical production environments.

 

Role Overview

As a Site Reliability Engineer, you will be responsible for ensuring the reliability, scalability, and performance of enterprise platforms and applications. You will work closely with development and operations teams to automate processes, improve observability, and maintain highly available production systems.

 

📍 Location: Toluca, Mexico

🏠 Work Model: Remote

Key Responsibilities

Manage and support production environments with 24x7 operational readiness and on-call participation.

Build automation solutions using programming languages such as Python, Java, or Go.

Deploy, manage, and optimize Kubernetes-based environments.

Implement and maintain monitoring and observability solutions using Grafana, Prometheus, and Alert Manager.

Support CI/CD pipelines and infrastructure automation using Jenkins, Ansible, and Terraform.

Troubleshoot performance, reliability, and availability issues across distributed systems.

Collaborate with development, infrastructure, and platform teams to improve system resilience and operational efficiency.

 

Required Qualifications

Experience coding in at least one programming language: Python, Java, or Go.

Hands-on experience with Kubernetes administration and operations.

Strong experience with monitoring and observability tools such as Grafana, Prometheus, and Alert Manager.

Experience supporting 24x7 production environments and participating in on-call rotations.

Hands-on experience with CI/CD and automation tools including Jenkins, Ansible, and Terraform.

 

Preferred Qualifications

Experience with AIOps tools and practices.

Exposure to cloud platforms such as AWS or Google Cloud Platform (GCP).

Min Salary: 
Max Salary: 


Job Segment: Developer, Java, Manager, Technology, Management

Apply now »