Job Description
SRE Technical Project Manager Location: Toronto, ON
Work Arrangement: Onsite
Employment Type: Full-Time FTE Job Summary: We are seeking an experienced SRE / DevOps Engineer Technical Project Manager focused on reliability engineering, automation, observability, and cloud operations. The ideal candidate will have strong hands-on expertise with Dynatrace, AWS, Azure, Ansible, Terraform, CI/CD, and Kubernetes , along with the ability to coordinate technical initiatives and drive reliability improvements. Required Skills & Qualifications
• Strong expertise in Dynatrace , including APM, Davis AI, RUM, and infrastructure monitoring.
• Experience using Dynatrace Davis AI for root cause analysis, anomaly detection, predictive insights, and alert optimization.
• Hands-on experience with OneAgent, Smartscape, distributed tracing, SLIs/SLOs, dashboards, and alert management .
• Strong automation experience using Ansible for deployments, provisioning, patching, and remediation.
• Hands-on cloud experience with AWS (primary) and Azure , including serverless and cloud-native architectures.
• Strong Infrastructure as Code (IaC) experience with Terraform, CloudFormation, or AWS CDK .
• Experience implementing CI/CD pipelines using Jenkins, GitHub Actions, and GitLab CI .
• Experience integrating monitoring and observability tools with CI/CD pipelines.
• Strong knowledge of Docker, Kubernetes, ECS, and AKS .
• Experience with High Availability (HA), Disaster Recovery (DR), incident response, and reliability engineering .
• Strong programming/scripting skills in Python (boto3) and Bash .
• Experience with AWS CloudWatch and Azure Monitor .
• Exposure to Prometheus, Grafana, and ELK is an advantage. Key Responsibilities
• Design, implement, and maintain reliable, scalable, and highly available cloud infrastructure.
• Lead observability initiatives using Dynatrace across applications, infrastructure, and user experience.
• Configure and optimize Dynatrace OneAgent, Smartscape, distributed tracing, dashboards, alerts, SLIs, and SLOs.
• Leverage Davis AI for automated anomaly detection, root cause analysis, predictive insights, and alert optimization.
• Develop and maintain Ansible playbooks for deployment, provisioning, patching, and automated remediation.
• Automate infrastructure provisioning and configuration using Terraform, CloudFormation, or CDK .
• Build and maintain CI/CD pipelines and integrate observability and monitoring capabilities into deployment workflows.
• Support containerized workloads using Docker, Kubernetes, ECS, and AKS .
• Implement and maintain HA, DR, monitoring, alerting, and incident response processes.
• Troubleshoot complex application, infrastructure, cloud, and network reliability issues.
• Collaborate with development, infrastructure, security, cloud, and business teams to improve system reliability and operational efficiency.
• Drive automation and continuous improvement across SRE and DevOps processes.
• Provide technical leadership and coordinate delivery of reliability, observability, and automation initiatives. Preferred Qualifications
• Experience in Site Reliability Engineering (SRE), DevOps, Cloud Engineering, or Technical Project Management .
• Strong understanding of cloud-native architecture and enterprise observability.
• Excellent communication, stakeholder management, problem-solving, and technical leadership skills.
• Experience managing multiple technical initiatives in a fast-paced enterprise environment.