Skip to main content
← Back to Jobs

Senior Site Reliability Engineer

OmiliaPhilippines🌍 Remote
Full-timeSenior
👁️ 0 views📝 0 applicationsPosted 9/5/2026Expires 11/4/2026
Tailor Resume for This JobCheck ATS Score

Get alerts for roles like this

More Senior Site Reliability Engineer roles in Philippines — straight to your inbox. No account needed.

Applying to this role? Tailor your résumé to this job description in one click, then download it clean — no watermark, no subscription.

Job Description

We are looking for a Senior Site Reliability Engineer with Cloud platform experience.

This individual will be part of a team responsible for operating and maintaining production clusters and developing our observability solutions; they will collaborate with team members to develop automation strategies, monitoring & alerting, and ensuring overall platform reliability.

Your goal will be to become an integral part of the team, making every challenge of the platform – your own challenge, and solving them accordingly.

Responsibilities

Ensure platform reliability and availability across production and pre-production environments through proactive monitoring, alerting, and automation. First response for incidents, contribute to problem management and root cause analysis.

Supporting the development team's effort towards reliability, creating a solid reliability culture within the development lifecycle. Develop troubleshooting documentation for production support resources.

Collaborate with Engineering teams to develop optimised and productive runbooks, operational documentation and automation of operational tasks. Collaborate with development and cloud engineering teams to embed reliability and performance into the software delivery lifecycle.

Design, implement, and evolve observability solutions (metrics, logs, traces, dashboards) using tools such as Prometheus, Grafana, and ELK. Participate in on-call rotations and continuously improve alert quality and response processes.

Champion a culture of reliability, performance, and continuous improvement across teams.

Requirements

Bachelor's Degree or MS in Engineering or equivalent. Experience in operating at least one container orchestration cluster (Kubernetes, Docker Swarm). Experience developing or maintaining software for production services at scale. Experience with ELK. Experience with AWS. Experience with Grafana/Prometheus stack. Strong scripting skills (Bash, Python or Go). Excellent communication skills. Thinking out

Required Skills

KubernetesDocker SwarmELKAWSGrafanaPrometheusBashPythonGoCloud platform experienceProduction supportAutomation strategiesMonitoringAlertingObservability solutionsTroubleshooting documentationOperational documentationRunbooksIncident managementRoot cause analysisReliability cultureContinuous improvement

The best-paying roles in your field. Every week. Free.

Join 10,000+ professionals getting job alerts and salary insights in their inbox

We respect your privacy. Unsubscribe anytime with one click.