Skip to main content
← Back to Jobs

Senior Site Reliability Engineer (SRE) Engineer

Umanist Staffing LLCIndia
Full-time7-15
₹20L - ₹25L
per year
👁️ 0 views📝 0 applicationsPosted 9/5/2026Expires 10/5/2026
Tailor Resume for This JobCheck ATS Score

Get alerts for roles like this

More Senior Site Reliability Engineer (SRE) Engineer roles in India — straight to your inbox. No account needed.

Applying to this role? Tailor your résumé to this job description in one click, then download it clean — no watermark, no subscription.

Job Description

Senior Site Reliability Engineer (SRE) Engineer Location: Viman Nagar, Pune – Work From Office Experience Overall(must have): 8 Years CTC: Up to ₹25 LPA Notice Period: Immediate Joiners Only within 15d or (if serving max 30days) Working Hours: 3:00 PM – 12:00 AM, Monday to Friday On-Call: 24/7 Production Support – On-Call Rotation Required Employment Type: Full-Time About the Role We are looking for an experienced Senior Site Reliability Engineer (SRE) / DevOps Engineer to manage and improve the reliability, scalability, performance, security, and observability of mission-critical production environments. The role requires strong hands-on expertise in Cloud, Kubernetes, DevOps automation, Monitoring & Observability, Incident Management, and SRE practices . The ideal candidate should be comfortable handling production incidents while also driving long-term initiatives around reliability, automation, scalability, and reduction of operational toil. Must-Have Skills & Experience1. SRE & Production Operations • Relevant 7+ years of relevant experience in SRE / DevOps / Cloud Infrastructure / Production Engineering . • Hands-on experience with 24/7 production support and on-call operations . • Strong experience in incident management, troubleshooting, RCA, and post-mortems . • Good understanding of SLI, SLO, SLA, Error Budgets, MTTR, and reliability engineering . • Experience with toil reduction, capacity planning, high availability, disaster recovery, and failover strategies . • Ability to improve system availability, performance, scalability, and operational reliability. 2. Cloud & Infrastructure • Strong hands-on experience with Microsoft Azure, AWS, and/or GCP . • Strong understanding of cloud infrastructure, networking, IAM, storage, compute, and cloud-native services. • Hands-on experience with at least one major cloud platform and good exposure to multi-cloud environments. • Experience with: • Azure: VMs, Networking, Storage, IAM, Azure Monitor, AKS • AWS: EC2, S3, RDS, IAM, VPC, CloudWatch, EKS • GCP: Compute Engine, Cloud Storage, IAM, VPC, GKE, Cloud Monitoring 3. Kubernetes & Containerization • Strong hands-on experience with Kubernetes and containerized workloads. • Experience with AKS / EKS / GKE or equivalent Kubernetes environments. • Hands-on experience with Helm deployments. • Understanding of Kubernetes troubleshooting, scaling, networking, and workload management. 4. Infrastructure as Code & DevOps • Hands-on experience with Terraform / Infrastructure as Code (IaC) . • Experience with Git-based workflows using GitHub, GitLab, or Azure Repos . • Strong DevOps automation and CI/CD understanding. • Strong scripting skills in Python and/or Bash . 5. Monitoring & Observability • Strong hands-on experience with OpenTelemetry . • Experience with monitoring and observability tools such as: • Prometheus • Grafana • Datadog • Azure Monitor • AWS CloudWatch • GCP Cloud Monitoring • Strong understanding of metrics, logs, distributed tracing, and alerting . • Experience implementing monitoring based on Golden Signals : • Latency • Traffic • Errors • Saturation • Ability to develop symptom-based, user-impact-focused alerting. 6. Linux & Networking • Strong knowledge of Linux system administration . • Strong understanding of: • DNS • TCP/IP • Load Balancing • SSL/TLS • Networking fundamentals • Experience supporting highly available production environments. 7. Incident & Reliability Engineering • Ability to rapidly diagnose and resolve high-severity production incidents . • Experience driving MTTR reduction . • Strong debugging and analytical problem-solving skills. • Ability to identify recurring issues and implement permanent corrective/preventive solutions. Good-to-Have Skills • Experience working across Azure + AWS + GCP in a multi-cloud environment. • Knowledge of Go (Golang) . • Experience with OpenSearch / ELK Stack . • Experience supporting AI/ML workloads in production. • Exposure to Azure AI Services and Azure AI Foundry . • Experience supporting RAG (Retrieval-Augmented Generation) workloads. • Experience designing infrastructure for AI/ML platforms. • Experience building enterprise-wide OpenTelemetry observability frameworks . • Strong understanding of distributed systems architecture. • Exposure to advanced cloud-native architectures and reliability patterns. • Experience with security, compliance, vulnerability remediation, secrets management, and network segmentation. Key ResponsibilitiesProduction & Incident Management • Participate in the 24/7 on-call rotation . • Diagnose, mitigate, and resolve production incidents. • Lead RCA and post-incident reviews. • Implement corrective and preventive actions. • Continuously improve MTTR and production stability. Reliability Engineering • Define and improve SLIs, SLOs, SLAs, and Error Budgets . • Identify and eliminate operational toil. • Conduct reliability and capacity reviews. • Improve redundancy, failover, disaster recovery, and system resilience. Cloud & Infrastructure • Manage and optimize cloud infrastructure across Azure, AWS, and/or GCP. • Manage Kubernetes clusters and containerized applications. • Implement and maintain Infrastructure as Code using Terraform. • Support CI/CD and Git-based development workflows. Observability & Performance • Build and improve monitoring, logging, metrics, and tracing. • Implement OpenTelemetry and distributed tracing . • Establish Golden Signals-based monitoring and alerting. • Identify and resolve infrastructure and application performance bottlenecks. Security • Implement cloud security best practices around IAM, network segmentation, and secrets management . • Support vulnerability remediation and compliance initiatives. • Collaborate with Development, Security, and Infrastructure teams. Ideal Candidate We are looking for someone with: • Strong SRE mindset and production ownership . • Excellent troubleshooting and incident-management skills. • Hands-on expertise in Cloud + Kubernetes + Terraform + Observability . • Strong understanding of OpenTelemetry and Golden Signals . • Experience working in highly available, production-critical environments. • Ability to remain calm and make effective decisions during critical incidents. • Strong communication and cross-functional collaboration skills. • Passion for automation, scalability, reliability, and continuous improvement . Important Hiring Criteria Must be: • 7+ years relevant experience • Immediate joiner • Willing to work from office in Viman Nagar, Pune • Comfortable with 3:00 PM – 12:00 AM shift • Comfortable with 24/7 on-call rotation • Strong hands-on SRE/DevOps experience • Strong Cloud + Kubernetes + Observability experience • Strong production incident management experience Good to have: • Multi-cloud: Azure + AWS + GCP • OpenTelemetry • AI/ML or RAG production workloads • Azure AI / AI Foundry • Go • OpenSearch / ELK • Distributed systems

Required Skills

AWSAzureCI/CDCross-functional CollaborationDevOpsGCPGitGoIntellectual PropertyKubernetesLinuxMachine LearningProblem SolvingPythonTerraform

The best-paying roles in your field. Every week. Free.

Join 10,000+ professionals getting job alerts and salary insights in their inbox

We respect your privacy. Unsubscribe anytime with one click.