Skip to main content
← Back to Jobs

KnowledgeWorks Global - Senior DevOps Engineer - Docker/Kubernetes

Knowledgeworks GlobalMumbai
Full-time7-15
👁️ 5 views📝 0 applicationsPosted 8/15/2026Expires 9/26/2026
Tailor Resume for This JobCheck ATS Score

Get alerts for roles like this

More KnowledgeWorks Global - Senior DevOps Engineer - Docker/Kubernetes roles in Mumbai — straight to your inbox. No account needed.

Applying to this role? Tailor your résumé to this job description in one click, then download it clean — no watermark, no subscription.

Job Description

Sr. DevOps Engineer Observability Strategy (Web & AI Applications) : - Design and implement an end-to-end observability & alert stack covering both traditional web applications and AI/ML services. - Develop and maintain tooling such as Prometheus, Grafana, ELK/OpenSearch, Datadog, or equivalent. - Build AI-specific observability : model latency/throughput tracking, GPU utilization, token usage, drift detection, and inference quality signals. Infrastructure Scaling & Reliability : - Design and manage infrastructure capable of scaling across on-premise data centers and public cloud. - Own capacity planning, load testing, auto-scaling, and cost-optimization initiatives across compute, storage, and networking. - Implement Infrastructure as Code (Terraform, Ansible, or equivalent) to ensure environments are reproducible, version-controlled, and auditable. - Lead disaster recovery, backup, and high-availability strategy for critical systems. DevOps & CI/CD Delivery : - Partner closely with Solution Architects to translate project and system designs into concrete DevOps execution plans. - Design, build, and maintain CI/CD pipelines (Jenkins, GitHub Actions or ArgoCD/Flux for GitOps) across multiple projects and teams. - Containerize and orchestrate applications using Docker and Kubernetes, including Helm chart and manifest management. - Embed security and compliance checks (SAST/DAST, secrets scanning, image scanning) directly into the delivery pipeline (DevSecOps). GPU Infrastructure & Model Deployment : - Provision, configure, and manage GPU infrastructure (on-prem clusters and cloud GPU instances) for model training and inference. - Deploy, scale, and monitor ML/LLM models in production using tools such as Triton Inference Server, vLLM, or similar. - Optimize GPU utilization, cost, and throughput across multi-tenant workloads; manage CUDA/driver/toolkit versions. - Collaborate with data science/ML engineering teams on MLOps pipelines model versioning, experiment tracking, and model registry. Requirements : - 5-10 years of hands-on DevOps/SRE/Infrastructure engineering experience, including at least 23 years in a senior or lead capacity. - Deep expertise in Git-based workflows and repository management at scale (GitHub/GitLab/Bitbucket). - Proven experience designing observability stacks (Prometheus, Grafana, ELK/EFK, Datadog, New Relic, OpenTelemetry). - Strong background in cloud platforms (AWS, Azure, and/or GCP) and on-premise/hybrid infrastructure. - Expert-level skills with Infrastructure as Code (Terraform, Ansible, CloudFormation, or Pulumi). - Strong Kubernetes and Docker experience, including multi-cluster and multi-environment management. - Hands-on experience building and maintaining CI/CD pipelines end to end. - Working knowledge of GPU infrastructure (NVIDIA CUDA, drivers, NCCL) and experience deploying ML/AI models to production. - Proficiency in scripting/automation languages : Python, Bash, and/or Go. - Solid understanding of networking, load balancing, DNS, and security fundamentals in distributed systems. - Experience partnering with architects and engineering leads to translate designs into infrastructure and delivery plans. - Hands-on with SAST, DAST, and SCA tooling (e.g., SonarQube, Snyk, Checkmarx, OWASP DependencyCheck) integrated directly into CI/CD pipelines. - Container and image security : vulnerability scanning (Trivy, Grype, Clair), minimal/hardened base images, and signed/verified image provenance (Cosign/Sigstore). - Familiarity with Secrets management and credential hygiene using tools such as HashiCorp Vault, AWS Secrets Manager, or Azure Key Vault. - Excellent communication skills and comfort operating cross-functionally with development, data science, and product teams.

Required Skills

DockerKubernetesPrometheusGrafanaELK/OpenSearchDatadogTerraformAnsibleCloudFormationPulumiPythonBashGoNVIDIA CUDANCCLTrivyGrypeClairCosignSigstore

The best-paying roles in your field. Every week. Free.

Join 10,000+ professionals getting job alerts and salary insights in their inbox

We respect your privacy. Unsubscribe anytime with one click.