Structured programmes for software engineers, data science and DevOps.
Staff Engineer, Site Reliability Engineering
Get alerts for roles like this
More Staff Engineer, Site Reliability Engineering roles in Markham, York region — straight to your inbox. No account needed.
Applying to this role? Tailor your résumé to this job description in one click, then download it clean — no watermark, no subscription.
Job Description
Hybrid Markham, Ontario, Canada Full time JR-202618328 **Job Description** **Vacancy Status:** **Yes -** This posting is for an existing vacancy within the organization and is open to new applications.
(Backfill) **AI Disclosure:** **As part of the application process, Artificial Intelligence will be used in the hiring process for this role** **Work Arrangement:** Hybrid: This role is categorized as hybrid. This means the successful candidate is expected to report to Markham office three times per week, at minimum.
**About the role** General Motors is transforming the automotive landscape through its next-generation Software-Defined Vehicle platform. Data is central to that transformation, powering safety, personalization, energy optimization, operational decision-making, and connected customer experiences.
We are seeking a Staff Engineer to help make GM's data platforms reliable, observable, operable, and scalable. This is a senior technical leadership role for someone who can move comfortably between system-level design, production operations, incident response, automation, and customer partnership.
You will help define and spread the engineering patterns that make services easier to operate. You will work with SRE, data engineering, infrastructure, developer experience, application, and product teams to improve reliability from design through production and continuously improve how the organization operates.
**What** **you'll** **do** + Lead the design and implementation of scalable, fault-tolerant, and observable infrastructure supporting vehicle telemetry, data ingestion, and platform operations.
+ Lead production readiness efforts across multiple teams-engaging directly in code, shaping reliability standards, guiding architectural improvements, and ensuring applications launch with resilient deployments, strong observability, and predictable operations.
+ Design, implement, and improve CI/CD delivery pipelines that make releases repeatable, safe, observable, and fast. Establishappropriate qualitygates, artifact promotion, deployment verification, progressive delivery, and rollback practices.
+ Partner across SRE, product, and application teams to design and implement meaningful SLOs, SLIs, observability,monitoringand alerting, runbooks, and operational best practices. + Build and improve reusable AI workflows, skills, and evaluations.
Applyappropriate validationtechniques, including regression testing, structured evaluations, and LLM-as-a-judge approacheswhereuseful. + Automate operational work, including incident intake, triage, diagnostics, remediation, evidence collection, service requests, and customer-facing status workflows.
+ Participate in a weekly on-call rotation with 12-hour shifts; the rotation cycles every eight weeks. Lead incident response, communicate clearly under pressure, and coordinate effective mitigation and recovery.
+ Participate in post-incident reviews and drive durable, system-level fixes that prevent recurrence rather than relying on short-term patches or repeated manual workarounds.
+ Partner directly with internal customers to understand their needs, explain technical trade-offs, and improve service outcomes with tact, empathy, and clear communication.
+ Influence technical direction across teams, mentor engineers, and raise engineering standards through design reviews, code reviews, documentation, and hands-on leadership with cross-functional engineering projects.
+ Balance reliability, performance, security, delivery speed, and cost when making technical decisions-especially under pressure. **What you bring** + 8+ years in SRE, DevOps, or systems engineering, including experience managing or mentoring high-impact teams.
+ Track recordofdesigning and buildingandmaintaininghigh-scale, cloud-native systems in production (preferably Azure, AWS, or GCP).
+ Hands-on experience architecting observability patterns, including standardized instrumentation, OTELcollector configuration, SLO/SLIdefinitions, anddeploying observability resources like monitors, alerts, and dashboards.
+ Strong understanding of production readiness, service ownership, SLOs, incident management, post-incident learning, and continuous reliability improvement. + Experienceparticipatingin an on-call rotation and leading technical response to production incidents.
+ Experience designing,operating, and improving CI/CD pipelines. Understanding ofGitOps, release strategies, quality gates, deployment verification, progressive delivery, and safe rollback is expected.
+ Strong programming ability in Python, Go, Java, ora comparablelanguage, with disciplined code review, version control, testing, and maintainability practices. + Familiarity with AI-assisted software development, LLM application practices, agentic workflows, and evaluation techniques.
+ Ability to work effectively with internal customers, including in difficult or high-pressure situations, with professionalism, tact, and empathy. + Excellent ownership attitude and the ability tooperatewith pace, judgment, and accountability in a high-velocity environment.
+ Ability to influence without relying on formal authority and to increase adoption of shared engineering patterns. + Strong written and verbal communication skills for both technical and non-technical audiences.
+ BS / MS / PhD in computer science, engineering, physics, mathematics, or another relevant, technical field **Preferred experience** + Azure Databricks + Azure Event Hubs + Azure Kubernetes Service (AKS) + Kubernetes configuration management with Helm andKustomize + Infrastructure as code, especially Terraform + GitHub Actions, Argo CD, andGitOps-based deployment models + Prometheus, Grafana, Datadog,OpenTelemetry, or comparable observability platforms and tools + LLM application developmentand testing withPromptfoo, agentic workflowsandreusable AI skillswith CoPilot + Experienceoperatinglarge-scale data ingestion, processing, and delivery systems such asFivetran, Apache Flink, Kafka, and Pulsar + Experience with vehicle telemetry, connected-vehicle platforms, or other high-volume event-driven systems **Why Join Us?
** This is more than an engineering role - it's an opportunity to shape the **future of mobility** . At GM, you'll join a team committed to **cutting-edge technology, sustainability** , and **inclusive innovation** .
With meaningful projects, a collaborative culture, and a global mission, your impact will be tangible and far-reaching. **Compensation:** The salary range for this role is $147,000 to $196,600.
The actual base salary a successful candidate will be offered within this range will vary based on factors relevant to the position. GM DOES NOT PROVIDE IMMIGRATION-RELATED SPONSORSHIP FOR THIS ROLE. DO NOT APPLY FOR THIS ROLE IF YOU WILL NEED GM IMMIGRATION SPONSORSHIP NOW OR IN THE FUTURE.
**Benefits Overview** The goal of the General Motors of Canada total rewards program is to support the health and well-being of you and your family.
Our comprehensive compensation plan currently includes the following benefits, in addition to many others: + Paid time off including vacation days, holidays, and supplemental benefits for pregnancy, parental and adoption leave; + Healthcare, dental, and vision benefits; + Life insurance plans to cover you and your family; + Company and matching contributions to a Defined Contribution Pension plan to help you save for retirement; + GM Vehicle Purchase Plan for you and your family.
**About GM** Our vision is a world with Zero Crashes, Zero Emissions and Zero Congestion and we embrace the responsibility to lead the change that will make our world better, safer and more equitable for all.
**Why Join Us** We believe we all must make a choice every day - individually and collectively - to drive meaningful change through our words, our deeds and our culture. Every day, we want every employee to feel they belong to one General Motors team. **Non-Discrimination and Equal Employment Opportuniti
Required Skills
Upskill for This Role
Courses from Udemy and edX matched to this role's skills.

AWS Certified Cloud Practitioner CLF-C02 Practice Tests 2026

AWS Certified CloudOps Engineer Associate Practice Exams

AWS Certified Developer Associate DVA-C02 Practice Test 2026

AWS Certified Security - Specialty (SCS-C03) Exam 2026

Frontend Deployment on AWS

300 Questions For 2026 AWS CLF-C02 Exam w/Explanations
ResumeKart may earn a commission from these links at no extra cost to you.
Partner picks for Staff Engineer, Site Reliability Engineering in Markham
Matched to the skills this page calls for and the candidate's location.
- Partner course providercovers javacovers pythoncovers devopscovers awsBengaluruVisit partner →
- Partner course provider
Certification training in cloud, data, cyber security, project management and digital marketing.
covers awscovers azurecovers devopsBengaluruVisit partner → - edXVerified partnerPartner course provider
Courses and programmes from universities and institutions worldwide.
covers pythoncovers leadership
Partners are ResumeKart affiliates or institutes it works with; ResumeKart may earn a commission when a candidate enrols. Placement is decided by relevance, not payment. How ResumeKart earns
Similar Jobs
Other open roles matched to this job's skills and location.
Prepare to Win This Role
Everything you need to ace the interview and negotiate top-of-band compensation.