Talent.com
TransUnion
Senior Site Reliability & Observability Engineer (SRE)TransUnion • Miguel Hidalgo, Mexico City, Mexico
Senior Site Reliability & Observability Engineer (SRE)

Senior Site Reliability & Observability Engineer (SRE)

TransUnion • Miguel Hidalgo, Mexico City, Mexico
Hace 4 días
Descripción del trabajo

TransUnions Job Applicant Privacy Notice

Team Overview

The Reliability Engineering team ensures the stability availability and performance of Buró de Créditos mission-critical platforms and services. Through observability automation and incident management practices the team drives operational excellence and continuous improvement. Working closely with Infrastructure Development Database and Security teams they help maintain resilient systems that support critical business operations.

This job is assigned as On-Site Essential and requires in- person work at an assigned TU office location as a condition of employment.

Role Overview And Core Responsibilities

  • Ensure the reliability stability and availability of mission-critical services through Site Reliability Engineering (SRE) best practices.

  • Operate and continuously improve the organizations observability platform including metrics logs traces and alerting capabilities.

  • Monitor critical systems proactively and respond to operational incidents to minimize service disruption and business impact.

  • Lead major incident response activities coordinating recovery efforts and stakeholder communication during service outages.

  • Conduct root cause analysis (RCA) and facilitate blameless post-mortems to identify systemic improvements and prevent recurrence.

  • Define monitor and improve Service Level Indicators (SLIs) Service Level Objectives (SLOs) and error budgets.

  • Automate operational processes and reduce manual effort (toil) through scripting infrastructure automation and platform engineering practices.

  • Support and troubleshoot distributed environments including Apache Cassandra Kafka Kubernetes relational databases and NoSQL platforms.

  • Collaborate with development and infrastructure teams to improve deployment reliability operational readiness and platform performance.

  • Drive continuous improvement initiatives focused on monitoring availability scalability resilience and operational efficiency.


Required Knowledge And Experiences

  • Bachelors degree in Computer Science Systems Engineering Telecommunications Engineering or a related technical discipline providing the foundation required to manage complex distributed environments.

  • Proven experience in Site Reliability Engineering (SRE) Production Operations Platform Engineering Infrastructure Engineering or Reliability-focused roles.

  • Strong knowledge of distributed systems observability practices incident management and troubleshooting methodologies for mission-critical environments.

  • Experience managing major incidents root cause analysis processes service restoration activities and operational excellence initiatives.

  • Understanding of reliability frameworks including SLOs SLIs error budgets continuous improvement and service management best practices.

#LI-SG4


TransUnion Overview:

At TransUnion we encourage and are committed to creating a real positive impact and shared sense of purpose within our Workforce for Good which empowers our people to grow innovate and contribute to a better future for our communities and customers. We strive to build an environment where our associates are in the drivers seat of their professional development while having access to help along the way. We recognize that success comes when our associates thrive both professionally and personally; thats why we prioritize work/life flexibility and offer resources for our teams across the globe to collaborate and drive excellence.

Be a part of our Workforce for Good youll work with great people pioneering products and cutting-edge technology.


TransUnion Job Title



Consultant IT Support


Required Experience:

Senior IC


Employment Type : Full-Time
Experience: years
Vacancy: 1

Crear una alerta de empleo para esta búsqueda

Senior Site Reliability & Observability Engineer (SRE) • Miguel Hidalgo, Mexico City, Mexico

Ofertas similares

Technical Project Manager Data Center (Remote)

RM Staffing B.V.Huixquilucan, SLP, MX

Clients need a single point of contact who actually understands hardware, not just a relationship manager who has to relay every technical question.Projects span hardware deployment, structured cab... Mostrar más

Customer Support & Operations Lead (AI-First) 🤖 - LATAM - Remote

Atomic HRNaucalpan de Juárez, Mexico, Mexico

Our client operates a two-sided marketplace and software platform in the outdoor hospitality space.On one side, travelers use the platform to discover and book campgrounds and RV parks.On the other... Mostrar más

Sales Development Specialist Data Center Services (APAC timezone)

RM Staffing B.V.Cuajimalpa de Morelos, SLP, MX

Reboot Monkey is a global datacenter services provider headquartered in Haarlem, Netherlands, operating.We deliver colocation, IP transit, smart hands, remote hands, and managed datacenter services... Mostrar más

LM438: HRIS Business Senior Advisor

FedExCuajimalpa de Morelos, Ciudad de Mexico, MX

Englis: Develop process improvements and upgrades of the Human Resources (HR) technology systems and Human Resources Information System (HRIS) including, planning, system design, technical design, ... Mostrar más