Skip to main content
ResumeKart
← Back to Jobs

Network Engineer — GPU Infrastructure

Nava•Bangalore, Karnataka
Full-time1-3
₹40L - ₹60L
per year
👁️ 0 views•📝 0 applications•Posted 9/13/2026•Expires 10/13/2026
Tailor Resume for This JobCheck ATS Score

Get alerts for roles like this

More Network Engineer — GPU Infrastructure roles in Bangalore, Karnataka — straight to your inbox. No account needed.

Applying to this role? Tailor your résumé to this job description in one click, then download it clean — no watermark, no subscription.

Job Description

About Nava Nava is a neocloud company purpose-built for the AI era. We design, deploy, and operate large-scale GPU infrastructure and deliver inference-as-a-service to teams building the next generation of AI products. Our platform runs on NVIDIA GPU systems, high-performance RDMA fabrics, and a fully automated, software-defined operations model. Every engineer at Nava works close to the bare metal, on infrastructure built to keep GPUs available and models serving. About the team The Compute team owns Nava's GPU infrastructure end-to-end: from bare-metal provisioning of NVIDIA GPU systems through cluster software, RDMA integration, and production operation. We keep thousands of GPUs healthy and highly utilized so training and inference workloads run fast and reliably. Responsibilities · Provision and configure GPU nodes: OS, drivers, CUDA, GPU Operator, and RDMA connectivity. · Write and maintain automation using Python, Ansible, and Terraform. · Support cluster deployments, run health checks and benchmarks, and document results. · Monitor GPU infrastructure, respond to alerts, and resolve or escalate issues. · Maintain runbooks and operational documentation. · Collaborate with cross-functional engineering teams to understand business and technical requirements and deliver software solutions. Required qualifications · Strong hands-on experience with Kubernetes administration and architecture. · Bare Metal GPU cluster deployment and management experience. · Kubernetes networking, storage, and security. · Infrastructure automation (Ansible, Terraform, etc.). · Strong Linux systems expertise and working knowledge of the NVIDIA GPU stack (CUDA, drivers, GPU Operator) and RDMA (RoCEv2/InfiniBand). · 2–4 years in systems/infrastructure engineering with strong Linux fundamentals and a learning-oriented mindset. Preferred qualifications · NVIDIA GPU ecosystem exposure. · AI/ML infrastructure experience. · Monitoring tools such as Prometheus, Grafana, Dynatrace, Datadog, or Zabbix. · Hybrid cloud and datacenter infrastructure experience. · NCCL and NVLink tuning, Slurm, and distributed training/inference workloads. Technology environment NVIDIA GPU systems (DGX/HGX-class, Blackwell), NVLink, CUDA, GPU Operator; RoCEv2 / InfiniBand RDMA fabrics; Kubernetes (networking, storage, security), Slurm; Ansible, Terraform, Python; monitoring with Prometheus, Grafana, Dynatrace, Datadog, or Zabbix; Linux at scale.

Required Skills

KubernetesGPUAnsibleTerraformLinuxCUDAGPU OperatorRDMARoCEv2InfiniBandNVIDIA GPUAIMLPrometheusGrafanaDynatraceDatadogZabbixHybrid clouddatacenterNCCLNVLinkSlurm

Partner picks for Network Engineer — GPU Infrastructure in Bangalore

Matched to the skills this page calls for and the candidate's location.

Partner
  • Partner course provider

    Structured programmes for software engineers, data science and DevOps.

    covers kubernetes
  • Partner competition

    Hackathon platform used by student and community hackathons across India.

Partners are ResumeKart affiliates or institutes it works with; ResumeKart may earn a commission when a candidate enrols. Placement is decided by relevance, not payment. How ResumeKart earns

The best-paying roles in your field. Every week. Free.

Join 10,000+ professionals getting job alerts and salary insights in their inbox

We respect your privacy. Unsubscribe anytime with one click.