Skip to main content
ResumeKart
← Back to Jobs

Senior Network Engineer — AI Datacenter

Nava•Bangalore, Karnataka
Full-timeSenior
₹40L - ₹60L
per year
👁️ 0 views•📝 0 applications•Posted 9/13/2026•Expires 10/13/2026
Tailor Resume for This JobCheck ATS Score

Get alerts for roles like this

More Senior Network Engineer — AI Datacenter roles in Bangalore, Karnataka — straight to your inbox. No account needed.

Applying to this role? Tailor your résumé to this job description in one click, then download it clean — no watermark, no subscription.

Job Description

About Nava Nava is a neocloud company purpose-built for the AI era. We design, deploy, and operate large-scale GPU infrastructure and deliver inference-as-a-service to teams building the next generation of AI products. Our platform runs on NVIDIA GPU systems, high-performance RDMA fabrics, and a fully automated, software-defined operations model. Every engineer at Nava works close to the metal, on infrastructure built to keep GPUs saturated and models serving. About the team The Network Engineering team owns the full lifecycle of Nava's AI datacenter fabric, from architecture and design through deployment, automation, and production operation. We build the lossless, high-throughput networks that carry GPU-to-GPU traffic for distributed training and low-latency inference. This is a hands-on team where the network is treated as code. Responsibilities · Design and deliver leaf-spine fabrics using BGP and EVPN-VXLAN, and implement RoCEv2 and InfiniBand lossless networking for GPU backend traffic. · Build and extend network automation in Python, Ansible, and Terraform to provision, validate, and operate the fabric. · Lead deployment and turn-up of new GPU clusters, including acceptance testing and performance validation against line-rate targets. · Triage and resolve production network incidents; tune congestion control (PFC/ECN) to keep RDMA traffic lossless under real load. · Partner with Compute and Storage engineers to integrate the network end-to-end and eliminate bottlenecks. · Contribute to standards and mentor Network Engineers. Required qualifications · 5–8 years in datacenter or production network engineering with strong hands-on BGP and EVPN-VXLAN experience. · Strong hands-on command of core networking protocols: BGP, OSPF, IS-IS, TCP/IP, IPv4 and IPv6, DNS, DHCP, and MPLS. · Experience with networking protocols such as TCP/IP, VPN, DNS, DHCP, and SSL/TLS. · Solid understanding of datacenter fabric concepts (leaf-spine, EVPN-VXLAN) and RDMA fabrics (RoCEv2 and InfiniBand), including lossless Ethernet. · Automation skills: Python plus Ansible and/or Terraform. · Comfortable operating in a fast-paced, on-call production environment. Preferred qualifications · Experience with NVIDIA networking (Spectrum-X, Quantum InfiniBand, BlueField DPUs) and NCCL traffic patterns. · Experience operating GPU clusters for large-scale distributed training or inference. · Familiarity with network telemetry, streaming analytics, and closed-loop automation. · Relevant certifications (e.g., CCNP or vendor equivalents). · Experience with Cisco and/or Arista platforms is a strong plus. Technology environment At Nava, our network infrastructure is deeply integrated with modern AI workloads—designed for scale, performance, and automation. Below is an overview of the key technologies and platforms you’ll work with daily: • Hardware & Interconnects : NVIDIA DGX/HGX systems (including Blackwell architecture), NVLink, Spectrum and Quantum-based switches, BlueField DPUs, and InfiniBand fabrics. • RDMA & Lossless Networking : RoCEv2 and InfiniBand RDMA, with deep familiarity in congestion control mechanisms (PFC, ECN) and lossless Ethernet configuration. • Control & Data Plane Protocols : BGP (including MP-BGP for EVPN), OSPF, IS-IS, EVPN-VXLAN for fabric virtualization, MPLS, and full IPv4/IPv6 stack support. • Automation & Tooling : Python for custom tooling, Ansible for configuration management, Terraform for infrastructure-as-code, and Linux-based toolchains (e.g., iproute2, netlink). • Observability & Operations : Telemetry via streaming telemetry (gNMI/sFlow), integration with monitoring stacks (Prometheus/Grafana), and experience with closed-loop automation for anomaly detection and remediation. You’ll operate in a unified stack where networking, compute, and storage are co-designed—making deep technical understanding and automation-first thinking essential to success.

Required Skills

BGPEVPN-VXLANRoCEv2InfiniBandPythonAnsibleTerraformOSPFIS-ISTCP/IPIPv4IPv6DNSDHCPMPLSlossless EthernetNVIDIA networkingSpectrum-XQuantum InfiniBandBlueField DPUsNCCLnetwork telemetrystreaming analyticsclosed-loop automationCCNPCiscoArista platforms

Partner picks for Senior Network Engineer — AI Datacenter in Bangalore

Matched to the skills this page calls for and the candidate's location.

Partner
  • Partner course provider

    Structured programmes for software engineers, data science and DevOps.

    covers python
  • edXVerified partner
    Partner course provider

    Courses and programmes from universities and institutions worldwide.

    covers python

Partners are ResumeKart affiliates or institutes it works with; ResumeKart may earn a commission when a candidate enrols. Placement is decided by relevance, not payment. How ResumeKart earns

The best-paying roles in your field. Every week. Free.

Join 10,000+ professionals getting job alerts and salary insights in their inbox

We respect your privacy. Unsubscribe anytime with one click.