A marketplace of instructor-created courses across technology, business and creative skills.
Site Reliability Engineers (SRE)
Get alerts for roles like this
More Site Reliability Engineers (SRE) roles in Remote — straight to your inbox. No account needed.
Applying to this role? Tailor your résumé to this job description in one click, then download it clean — no watermark, no subscription.
Job Description
The Role: You will join a team working with Observability, Escalations, Post-mortems, Correction of Errors, and other practices that will contribute to the company's goal of cloud resiliency.
You will be responsible for driving processes around reliability, best practices, cultural change, and enforcement of these practices.
The main responsibilities of the position include: Honor and practice the Resiliency pillar of the Well Architected Framework in all tasks and responsibilities Conduct Chaos Engineering experiments and relevant exercises to improve resiliency and fault-tolerance Research workloads for migrating to the cloud with minimal disruption and impact Monitor cloud migration projects to ensure seamless transitions Design, consult, re-platform, and re-factor the observability of current cloud infrastructure Coordinate with other IT departments and teams regarding observability for both individual and organizational needs Regularly assess cloud deployments for compliance with the company’s standards and best practices Investigate and correct areas where observability is lagging Stay up to date and provide training on new and current technologies, services, tools, methodologies, and practices Occasionally participate in service capacity planning, software performance analysis, and system tuning Mentor colleagues in technical skills and knowledge Analyze, oversee, and remediate the company’s resiliency Participate in on-call support 24/7 based on a rotation schedule Main requirements: BSc/MSc degree in Computer Science or related field 5+ years of cloud services experience, with at least 3 years on AWS cloud 3+ years of experience in SRE or a similar role Experience with monitoring, APM, logging, and notification tools Familiarity with incident, problem and change management procedures and practices Advanced knowledge of SRE practices and methods Understanding and practice of Service Levels Strong troubleshooting skills and the ability to mentor others
Required Skills
Partner picks for Site Reliability Engineers (SRE) in Remote
Matched to the skills this page calls for and the candidate's location.
- UdemyVerified partnerPartner course providercovers aws
- Partner course provider
Structured programmes for software engineers, data science and DevOps.
covers awsBengaluruVisit partner → - Partner course provider
Certification training in cloud, data, cyber security, project management and digital marketing.
covers awsBengaluruVisit partner →
Partners are ResumeKart affiliates or institutes it works with; ResumeKart may earn a commission when a candidate enrols. Placement is decided by relevance, not payment. How ResumeKart earns
Similar Jobs
Other open roles matched to this job's skills and location.
Prepare to Win This Role
Everything you need to ace the interview and negotiate top-of-band compensation.