Skip to main content
ResumeKart
← Back to Jobs

Senior Lead Engineer, Platform Operations & Observability/Ingénieur principal, Opérations de pl[]

McKesson•Canada
Full-time7-15
👁️ 0 views•📝 0 applications•Posted 9/20/2026•Expires 10/21/2026
Tailor Resume for This JobCheck ATS Score

Get alerts for roles like this

More Senior Lead Engineer, Platform Operations & Observability/Ingénieur principal, Opérations de pl[] roles in Canada — straight to your inbox. No account needed.

Applying to this role? Tailor your résumé to this job description in one click, then download it clean — no watermark, no subscription.

Job Description

McKesson, l’une des 10 premières entreprises du classement Fortune Global 500, touche à pratiquement tous les aspects des soins de santé et s’emploie à faire une réelle différence. Nous sommes reconnus pour notre capacité à offrir un savoir, des produits et des services qui rendent les soins de qualité plus accessibles et plus abordables. Chez nous, la santé, le bonheur et le bien‑être de nos gens et des personnes que nous desservons sont prioritaires—et nous tiennent à cœur. Ce que tu fais chez McKesson a de l’importance. Nous favorisons une culture où tu peux t’épanouir et avoir un impact, et où tu es encouragé à proposer de nouvelles idées. Ensemble, nous façonnons l’avenir de la santé pour nos patients, nos communautés et nos équipes. Si tu souhaites dès aujourd’hui contribuer à la santé de demain, nous aimerions avoir de tes nouvelles. McKesson is an impact-driven, Fortune 10 company that touches virtually every aspect of healthcare. We are known for delivering insights, products, and services that make quality care more accessible and affordable. Here, we focus on the health, happiness, and well‑being of you and those we serve – we care. What you do at McKesson matters. We foster a culture where you can grow, make an impact, and are empowered to bring new ideas. Together, we thrive as we shape the future of health for patients, our communities, and our people. If you want to be part of tomorrow’s health today, we want to hear from you. About The Role Team/Project: Canada B2C Digital Solution. Main application in the portfolio is a B2C Platform for pharmacy patients. Team of around 20 people. McKesson is seeking a Senior Lead Engineer, Platform Operations & Observability to lead the reliability, observability, and operational excellence of enterprise healthcare technology platforms. In this role, in accordance with Application Monitoring & Observability Lead, you will drive monitoring strategies, incident management practices, change governance, and root cause analysis initiatives while helping build scalable, secure, and resilient systems. You will collaborate with software engineering, platform, security, and operations teams to improve service reliability, automate operational processes, and establish best practices for production readiness. This position also provides technical leadership and mentorship to engineering teams while influencing reliability standards and long‑term platform strategy. What You’ll Do • Lead monitoring and observability strategies across enterprise applications, platforms, and services. • Design, implement, and optimize dashboards, alerts, telemetry, logging, and performance monitoring solutions. • Drive incident management processes, major incident response, escalation coordination, service restoration activities and postmortem incident. • Conduct root cause analysis (RCA) investigations and lead corrective and preventive action planning. • Partner with engineering teams to improve platform reliability, resiliency, scalability, and operational readiness. • Lead change management reviews and promote safe deployment and release practices. • Ability to execute regression and validation test plans following each production deployment. • Provide technical leadership, coaching, and mentoring to engineers while establishing engineering best practices. • Influence architecture, automation, CI/CD, and operational excellence initiatives supporting enterprise platforms. Basic Requirements • 7+ years of professional experience in Software Engineering, Site Reliability Engineering, Platform Engineering, DevOps, or related technical roles. • Bachelor's degree in Computer Science, Engineering, Information Technology, or equivalent experience. • Experience supporting large‑scale production environments and enterprise applications. • Hands‑on experience with monitoring, observability, logging, alerting, and application performance monitoring tools. • Proven experience leading incident management and production support activities. • Experience performing root cause analysis and implementing preventive solutions. • Experience with CI/CD, automation, DevOps practices, and software delivery pipelines. • Experience with microservices, APIs, distributed systems, and cloud‑based architectures. • Lead production readiness reviews and operational acceptance activities prior to major releases. Preferred Skills / Experience • Experience with tools such as Dynatrace, Prometheus, Dotcom Monitor or similar observability platforms. • Experience with Kubernetes, containers, and cloud platforms such as Azure. • Knowledge of ITIL‑aligned incidents, problems, and change management practices. • Experience defining SLAs, MTTR and service reliability metrics. • Demonstrated technical leadership and mentorship of engineering teams. • Experience operating in regulated or highly compliant environments. • Experience driving platform modernization and operational excellence initiatives. • Strong analytical, troubleshooting, continuous improvement and stakeholder communication skills. • Fair understanding and mastery of AI tools (Copilot, Rovo) and AI agents. Travel / Work Environment / Physical Requirements • Hybrid, two mandatory days at the Dobrin Office (Usually on Monday and Wednesday). • Ability to work at a computer for extended periods and participate in virtual collaboration activities. • Participation in on‑call support and prod deployment rotations out of business hours may be required based on organizational needs. À propos du poste Équipe / Projet : Solution numérique B2C Canada. La principale application du portefeuille est une plateforme B2C destinée aux patients en pharmacie. L’équipe est composée d’environ 20 personnes. McKesson est à la recherche d’un Ingénieur principal responsable, Opérations de plateforme et observabilité pour diriger la fiabilité, l’observabilité et l’excellence opérationnelle des plateformes technologiques de santé de l’entreprise. Dans ce rôle, en collaboration avec le responsable de la surveillance applicative et de l’observabilité, vous dirigerez les stratégies de surveillance, les pratiques de gestion des incidents, la gouvernance des changements et les initiatives d’analyse des causes fondamentales, tout en contribuant à la création de systèmes évolutifs, sécurisés et résilients. Vous collaborerez avec les équipes d’ingénierie logicielle, de plateforme, de sécurité et d’exploitation afin d’améliorer la fiabilité des services, d’automatiser les processus opérationnels et d’établir les meilleures pratiques en matière de préparation à la production. Ce poste offre également un leadership technique et du mentorat aux équipes d’ingénierie tout en influençant les normes de fiabilité et la stratégie à long terme des plateformes. Ce que vous ferez • Diriger les stratégies de surveillance et d’observabilité à travers les applications, plateformes et services de l’entreprise. • Concevoir, mettre en œuvre et optimiser les tableaux de bord, alertes, solutions de télémétrie, de journalisation et de surveillance de la performance. • Piloter les processus de gestion des incidents, la réponse aux incidents majeurs, la coordination des escalades, les activités de rétablissement des services et les analyses post‑incident. • Réaliser des analyses des causes fondamentales (RCA) et diriger la planification des actions correctives et préventives. • Collaborer avec les équipes d’ingénierie afin d’améliorer la fiabilité, la résilience, l’évolutivité et la préparation opérationnelle des plateformes. • Diriger les revues de gestion des changements et promouvoir des pratiques sécuritaires de déploiement et de mise en production. • Être en mesure d’exécuter des plans de tests de régression et de validation après chaque déploiement en production. • Fournir un leadership technique, de l’accompagnement et du mentorat aux ingénieurs tout en établissant les meilleures pratiques d’ingénie

Required Skills

Software EngineeringSite Reliability EngineeringPlatform EngineeringDevOpsmonitoringobservabilityloggingalertingapplication performance monitoringincident managementroot cause analysisCI/CDautomationmicroservicesAPIsdistributed systemscloud-based architecturesDynatracePrometheusDotcom MonitorKubernetescontainersAzureITILSLAsMTTRtechnical leadershipmentorshipregulated environmentsplatform modernizationanalytical skillsAI tools

Partner picks for Senior Lead Engineer, Platform Operations & Observability/Ingénieur principal, Opérations de pl[] in Canada

Matched to the skills this page calls for and the candidate's location.

Partner
  • Partner course provider

    Certification training in cloud, data, cyber security, project management and digital marketing.

    covers azurecovers devopscovers itil
  • Partner course provider

    Structured programmes for software engineers, data science and DevOps.

    covers devopscovers kubernetes
  • edXVerified partner
    Partner course provider

    Courses and programmes from universities and institutions worldwide.

Partners are ResumeKart affiliates or institutes it works with; ResumeKart may earn a commission when a candidate enrols. Placement is decided by relevance, not payment. How ResumeKart earns

The best-paying roles in your field. Every week. Free.

Join 10,000+ professionals getting job alerts and salary insights in their inbox

We respect your privacy. Unsubscribe anytime with one click.