Senior Site Reliability Engineer (SRE) Engineer
Umanist Staffing LLC•India
Full-time7-15
₹20L - ₹25L
per year
👁️ 0 views•📝 0 applications•Posted 9/5/2026•Expires 10/5/2026
Get alerts for roles like this
More Senior Site Reliability Engineer (SRE) Engineer roles in India — straight to your inbox. No account needed.
Applying to this role? Tailor your résumé to this job description in one click, then download it clean — no watermark, no subscription.
Job Description
Senior Site Reliability Engineer (SRE) Engineer
Location: Viman Nagar, Pune – Work From Office
Experience Overall(must have): 8 Years
CTC: Up to ₹25 LPA
Notice Period: Immediate Joiners Only within 15d or (if serving max 30days)
Working Hours: 3:00 PM – 12:00 AM, Monday to Friday
On-Call: 24/7 Production Support – On-Call Rotation Required
Employment Type: Full-Time
About the Role
We are looking for an experienced Senior Site Reliability Engineer (SRE) / DevOps Engineer to manage and improve the reliability, scalability, performance, security, and observability of mission-critical production environments.
The role requires strong hands-on expertise in Cloud, Kubernetes, DevOps automation, Monitoring & Observability, Incident Management, and SRE practices . The ideal candidate should be comfortable handling production incidents while also driving long-term initiatives around reliability, automation, scalability, and reduction of operational toil.
Must-Have Skills & Experience1. SRE & Production Operations
•
Relevant 7+ years of relevant experience in SRE / DevOps / Cloud Infrastructure / Production Engineering .
•
Hands-on experience with 24/7 production support and on-call operations .
•
Strong experience in incident management, troubleshooting, RCA, and post-mortems .
•
Good understanding of SLI, SLO, SLA, Error Budgets, MTTR, and reliability engineering .
•
Experience with toil reduction, capacity planning, high availability, disaster recovery, and failover strategies .
•
Ability to improve system availability, performance, scalability, and operational reliability.
2. Cloud & Infrastructure
•
Strong hands-on experience with Microsoft Azure, AWS, and/or GCP .
•
Strong understanding of cloud infrastructure, networking, IAM, storage, compute, and cloud-native services.
•
Hands-on experience with at least one major cloud platform and good exposure to multi-cloud environments.
•
Experience with:
•
Azure: VMs, Networking, Storage, IAM, Azure Monitor, AKS
•
AWS: EC2, S3, RDS, IAM, VPC, CloudWatch, EKS
•
GCP: Compute Engine, Cloud Storage, IAM, VPC, GKE, Cloud Monitoring
3. Kubernetes & Containerization
•
Strong hands-on experience with Kubernetes and containerized workloads.
•
Experience with AKS / EKS / GKE or equivalent Kubernetes environments.
•
Hands-on experience with Helm deployments.
•
Understanding of Kubernetes troubleshooting, scaling, networking, and workload management.
4. Infrastructure as Code & DevOps
•
Hands-on experience with Terraform / Infrastructure as Code (IaC) .
•
Experience with Git-based workflows using GitHub, GitLab, or Azure Repos .
•
Strong DevOps automation and CI/CD understanding.
•
Strong scripting skills in Python and/or Bash .
5. Monitoring & Observability
•
Strong hands-on experience with OpenTelemetry .
•
Experience with monitoring and observability tools such as:
•
Prometheus
•
Grafana
•
Datadog
•
Azure Monitor
•
AWS CloudWatch
•
GCP Cloud Monitoring
•
Strong understanding of metrics, logs, distributed tracing, and alerting .
•
Experience implementing monitoring based on Golden Signals :
•
Latency
•
Traffic
•
Errors
•
Saturation
•
Ability to develop symptom-based, user-impact-focused alerting.
6. Linux & Networking
•
Strong knowledge of Linux system administration .
•
Strong understanding of:
•
DNS
•
TCP/IP
•
Load Balancing
•
SSL/TLS
•
Networking fundamentals
•
Experience supporting highly available production environments.
7. Incident & Reliability Engineering
•
Ability to rapidly diagnose and resolve high-severity production incidents .
•
Experience driving MTTR reduction .
•
Strong debugging and analytical problem-solving skills.
•
Ability to identify recurring issues and implement permanent corrective/preventive solutions.
Good-to-Have Skills
•
Experience working across Azure + AWS + GCP in a multi-cloud environment.
•
Knowledge of Go (Golang) .
•
Experience with OpenSearch / ELK Stack .
•
Experience supporting AI/ML workloads in production.
•
Exposure to Azure AI Services and Azure AI Foundry .
•
Experience supporting RAG (Retrieval-Augmented Generation) workloads.
•
Experience designing infrastructure for AI/ML platforms.
•
Experience building enterprise-wide OpenTelemetry observability frameworks .
•
Strong understanding of distributed systems architecture.
•
Exposure to advanced cloud-native architectures and reliability patterns.
•
Experience with security, compliance, vulnerability remediation, secrets management, and network segmentation.
Key ResponsibilitiesProduction & Incident Management
•
Participate in the 24/7 on-call rotation .
•
Diagnose, mitigate, and resolve production incidents.
•
Lead RCA and post-incident reviews.
•
Implement corrective and preventive actions.
•
Continuously improve MTTR and production stability.
Reliability Engineering
•
Define and improve SLIs, SLOs, SLAs, and Error Budgets .
•
Identify and eliminate operational toil.
•
Conduct reliability and capacity reviews.
•
Improve redundancy, failover, disaster recovery, and system resilience.
Cloud & Infrastructure
•
Manage and optimize cloud infrastructure across Azure, AWS, and/or GCP.
•
Manage Kubernetes clusters and containerized applications.
•
Implement and maintain Infrastructure as Code using Terraform.
•
Support CI/CD and Git-based development workflows.
Observability & Performance
•
Build and improve monitoring, logging, metrics, and tracing.
•
Implement OpenTelemetry and distributed tracing .
•
Establish Golden Signals-based monitoring and alerting.
•
Identify and resolve infrastructure and application performance bottlenecks.
Security
•
Implement cloud security best practices around IAM, network segmentation, and secrets management .
•
Support vulnerability remediation and compliance initiatives.
•
Collaborate with Development, Security, and Infrastructure teams.
Ideal Candidate
We are looking for someone with:
•
Strong SRE mindset and production ownership .
•
Excellent troubleshooting and incident-management skills.
•
Hands-on expertise in Cloud + Kubernetes + Terraform + Observability .
•
Strong understanding of OpenTelemetry and Golden Signals .
•
Experience working in highly available, production-critical environments.
•
Ability to remain calm and make effective decisions during critical incidents.
•
Strong communication and cross-functional collaboration skills.
•
Passion for automation, scalability, reliability, and continuous improvement .
Important Hiring Criteria
Must be:
•
7+ years relevant experience
•
Immediate joiner
•
Willing to work from office in Viman Nagar, Pune
•
Comfortable with 3:00 PM – 12:00 AM shift
•
Comfortable with 24/7 on-call rotation
•
Strong hands-on SRE/DevOps experience
•
Strong Cloud + Kubernetes + Observability experience
•
Strong production incident management experience
Good to have:
•
Multi-cloud: Azure + AWS + GCP
•
OpenTelemetry
•
AI/ML or RAG production workloads
•
Azure AI / AI Foundry
•
Go
•
OpenSearch / ELK
•
Distributed systems
Required Skills
AWSAzureCI/CDCross-functional CollaborationDevOpsGCPGitGoIntellectual PropertyKubernetesLinuxMachine LearningProblem SolvingPythonTerraform
Prepare to Win This Role
Everything you need to ace the interview and negotiate top-of-band compensation.