About the role
Job Title
Service Reliability Engineer LOCATION: BOGOTA
SRE's (All Levels)
Summary of the role
We are looking for Site Reliability Engineers (SREs) to join our global teams supporting mission-critical airline platforms and systems.
In this role, you will focus on ensuring system reliability, availability, scalability, and performance, while driving automation and operational excellence across distributed environments.
You will collaborate with globally distributed teams in a follow-the-sun model, supporting production systems and continuously improving reliability and operational processes.
Key Responsibilities:
Reliability & Production Operations
Ensure high availability, scalability, and resilience of production systems
Perform incident, problem, and change management following ITIL practices
Conduct root cause analysis (RCA) and drive resolution of production issues
Support systems across multiple environments (test, staging, production)
Participate in on-call / follow-the-sun rotations
Automation & Reliability Engineering
Automate repetitive operational tasks, deployments, and recovery processes
Improve system reliability through engineering solutions, not manual fixes
Contribute to continuous improvement of operational processes and efficiency
Support implementation of deployment strategies (e.g., blue/green, canary)
Observability & Performance
Build and enhance monitoring, alerting, and observability frameworks
Improve visibility across metrics, logs, and traces
Track and improve SLOs/SLAs and system performance
Perform proactive reliability analysis and capacity planning
Infrastructure & Platform Reliability
Operate and support systems across cloud and on-prem environments
Work with containerized and distributed systems (where applicable)
Support Linux and/or Windows-based environments depending on the platform
Contribute to system architecture improvements, resilience, and scalability
Application Reliability (Multi-stack)
Troubleshoot distributed systems, microservices, and APIs
Diagnose issues using logs, monitoring tools, and profiling techniques
Support applications across different stacks, such as:.NET / C# applications (…