Job Description
We're looking for a
Senior DevOps Engineer
This role is Office Based, Hyderabad Office
Senior DevOps / Cloud Platform Engineer – ML & AI Infrastructure
Job Summary
We are looking for a Senior DevOps / Cloud Platform Engineer with strong experience in AWS, Kubernetes, CI/CD, infrastructure automation, and ML/AI infrastructure to design, deploy, manage, and optimize cloud infrastructure supporting our machine learning and AI services.
The ideal candidate will have hands-on experience with AWS EKS, SageMaker, Bedrock, Docker, Kubernetes, Terraform, Helm, GitHub Actions, Databricks, Elasticsearch, and self-hosted LLM deployments . This role will work closely with Data Engineering, Machine Learning, and Software Engineering teams to build reliable, scalable, secure, and cost-efficient platforms for ML services across development and production environments.
In this role you will...
Key Responsibilities
AWS & Kubernetes Infrastructure
• Design, deploy, and administer AWS infrastructure supporting ML and AI workloads.
• Manage Amazon EKS clusters , including cluster provisioning, upgrades, scaling, networking, and troubleshooting.
• Work with AWS SageMaker, AWS Bedrock, EKS, ECS, and related AWS services .
• Configure and manage Kubernetes Ingress controllers such as NGINX and AWS ALB .
• Manage Cloudflare Tunnels, DNS, Cloudflare configuration, networking, and security .
• Troubleshoot application, networking, compute, and infrastructure issues across AWS and Kubernetes environments.
• Implement best practices for security, reliability, availability, and scalability.
CI/CD & Azure-to-AWS Migration
• Build and maintain CI/CD pipelines for ML and AI services.
• Develop and manage GitHub Actions and Azure DevOps pipelines using YAML.
• Migrate repositories and CI/CD workflows from Azure DevOps to GitHub/AWS .
• Automate build, test, containerization, deployment, and release processes.
• Establish deployment strategies across development, staging, and production environments.
ML Service Deployment
• Deploy and manage ML services across AWS EKS/ECS and SageMaker .
• Build and maintain Docker containers and Kubernetes deployments.
• Manage environment segregation and configuration across Dev, QA, and Production.
• Develop and maintain Kubernetes manifests and Helm charts .
• Troubleshoot ML service deployment, networking, scaling, and runtime issues.
App Runner to EKS Migration
• Lead migration of existing services from AWS App Runner to Amazon EKS .
• Containerize applications and develop Kubernetes manifests/Helm charts.
• Design appropriate Kubernetes architecture, networking, ingress, scaling, and deployment strategies.
• Ensure minimal service disruption during migration and establish operational best practices on EKS.
Self-Hosted LLM & AI Infrastructure
• Deploy and manage self-hosted Large Language Models and inference services.
• Work with model serving frameworks such as vLLM .
• Design containerized infrastructure for GPU-based model serving and inference.
• Manage model versions, deployments, configurations, and rollback strategies.
• Support migration of ML services from managed APIs/services to self-hosted models .
• Work with engineering teams on API integration and inference infrastructure.
Databricks Administration
• Administer Databricks workspaces, clusters, permissions, and access controls .
• Manage cluster configuration, policies, and resource utilization.
• Support LMI Insights and related ML/AI workloads.
• Troubleshoot Databricks infrastructure and connectivity issues.
• Implement appropriate security and access-control practices.
Elasticsearch Infrastructure
• Design, deploy, and manage Elasticsearch clusters .
• Perform cluster sizing, scaling, configuration, and performance optimization.
• Manage indices, mappings, retention, and data lifecycle requirements.
• Support Kibana configuration, dashboards, and troubleshooting.
• Monitor Elasticsearch health, capacity, and performance.
Monitoring, Reliability & Auto-Scaling
• Implement monitoring and observability for Kubernetes, AWS, and ML services.
• Use Prometheus, Grafana, and AWS CloudWatch for monitoring and alerting.
• Configure Kubernetes HPA/VPA and other auto-scaling mechanisms.
• Establish proactive alerting for infrastructure and application health.
• Perform capacity planning and resource optimization.
• Identify opportunities for AWS infrastructure and compute cost optimization .
Infrastructure as Code & Automation
• Build and maintain infrastructure using Terraform .
• Develop reusable Terraform modules for AWS and Kubernetes infrastructure.
• Manage Kubernetes deployments using Helm charts .
• Automate infrastructure provisioning, configuration, deployments, and operational tasks.
• Maintain infrastructure documentation and deployment standards.
Cross-Team Collaboration
• Partner closely with Data Engineering, ML Engineering, Data Science, and Software Engineering teams.
• Understand data pipelines, SQL, APIs, and ML service architecture sufficiently to troubleshoot end-to-end workflows.
• Coordinate infrastructure requirements for new ML models and services.
• Participate in production incident resolution, root-cause analysis, and continuous improvement.
• Establish engineering standards around deployment, monitoring, security, and operational ownership.
You've got what it takes if you have...
Required Skills & Experience
• 5+ years of experience in DevOps, Cloud Infrastructure, SRE, or Platform Engineering.
• Strong hands-on experience with AWS .
• Strong experience administering Amazon EKS and Kubernetes in production.
• Hands-on experience with:
• AWS EKS
• AWS SageMaker
• AWS Bedrock
• AWS ECS
• AWS App Runner
• Kubernetes
• Docker
• NGINX / AWS ALB Ingress
• Cloudflare / Cloudflare Tunnels
• Strong experience with Terraform and Helm .
• Strong experience developing CI/CD pipelines using GitHub Actions and/or Azure DevOps.
• Strong YAML scripting and Git experience.
• Experience migrating CI/CD pipelines and repositories from Azure to AWS/GitHub .
• Experience deploying and operating ML/AI services.
• Experience with self-hosted LLM/model serving , preferably vLLM .
• Experience with GPU-based workloads is highly desirable.
• Experience with Databricks administration .
• Experience managing Elasticsearch and Kibana .
• Experience with Prometheus, Grafana, and CloudWatch .
• Strong understanding of Kubernetes HPA/VPA, networking, ingress, DNS, and service discovery .
• Strong understanding of cloud networking fundamentals.
• Experience with production troubleshooting, monitoring, capacity planning, and cost optimization.
• Strong understanding of security, IAM, secrets management, and access control.
Preferred / Nice-to-Have Skills
• Experience supporting Generative AI / LLM platforms .
• Experience with GPU infrastructure and NVIDIA/CUDA environments.
• Experience with model lifecycle and model version management.
• Experience migrating workloads between managed cloud services and Kubernetes.
• Experience with AWS networking such as VPC, load balancers, security groups, and Route 53.
• Experience with API gateways and microservice architectures.
• Experience with Python or shell scripting for infrastructure automation.
• Experience working with Data Engineering and ML teams in a production environment.
What You'll Own
• AWS ML/AI infrastructure
• EKS cluster administration and upgrades
• ML service deployment and production operations
• CI/CD automation
• App Runner → EKS migration
• Self-hosted LLM infrastructure and vLLM
• Databricks platform administration
• Elasticsearch infrastructure
• Monitoring and auto-scaling
• Terraform and Helm-based infrastructure automation
• Cloud cost, reliability, and performance optimization
Ideal Candidate
The ideal candidate is a hands-on infrastructure engineer who can independently take an ML/AI service from containerization → CI/CD