Structured programmes for software engineers, data science and DevOps.
Senior Software Engineer, SRE and Production Engineering - DGX Cloud
Get alerts for roles like this
More Senior Software Engineer, SRE and Production Engineering - DGX Cloud roles in United States — straight to your inbox. No account needed.
Applying to this role? Tailor your résumé to this job description in one click, then download it clean — no watermark, no subscription.
Job Description
NVIDIA DGX Cloud builds and operates large-scale GPU infrastructure for AI workloads. We are looking for Software Engineers with SRE or Production Engineering experience who have worked hands-on with bare-metal NVIDIA systems.
This team builds the software and operational tooling that moves GPU capacity from installed hardware to production service supporting an IaaS production environment of BMaaS, VMaaS.
What you’ll be doing Build automation for bare-metal provisioning, hardware validation, firmware and software upgrades, repair, and cluster lifecycle management. Develop tools that interact with BMC and Redfish interfaces to monitor hardware health, manage server state, and assist recovery workflows.
Handle and advance NVIDIA NVL72 systems and BlueField-3 or later DPUs throughout cloud partner and on-premises environments. Diagnose failures across servers, DPUs, GPU systems, CPU systems, networking, Linux, and Kubernetes; turn recurring issues into automated detection and repair.
Define validation and handoff criteria so new capacity enters production safely and consistently. Take part in on-call duties, incident response, root-cause analysis, and follow-up to implement permanent solutions.
Work with hardware, networking, platform, data center operations, and partner teams to resolve issues across ownership boundaries. What we need to see: 5+ years building software for or operating production infrastructure, including substantial hands-on bare-metal experience.
Strong Go or Python skills, with a record of delivering production automation and services. Direct experience with BMC and Redfish in server provisioning, health inspection, power management, or fault diagnosis.
Practical experience working directly with NVIDIA GPU hardware, such as NVL72 systems, and BlueField-3 or newer DPUs. Experience with Linux, firmware and driver management, network boot, and the server lifecycle from initial provisioning through repair. Experience managing production reliability th
Required Skills
Upskill for This Role
Courses from Udemy and edX matched to this role's skills.

【2026年更新】KCNA-JP 認定Kubernetesクラウドネイティブアソシエイト 模擬問題集

Architecting with Google Kubernetes Engine: Foundations

Deploy Microservices with Azure Kubernetes Service

MLflow for Kubernetes: Deploy and Manage ML Models at Scale

Certified Kubernetes CKAD Questions and Answers
![Kubernetes and Cloud Native Associate (KCNA) [Exams 2026]](https://i.udemycdn.com/course/480x270/5619202_4dda_4.jpg)
Kubernetes and Cloud Native Associate (KCNA) [Exams 2026]
ResumeKart may earn a commission from these links at no extra cost to you.
Partner picks for Senior Software Engineer, SRE and Production Engineering - DGX Cloud in United States
Matched to the skills this page calls for and the candidate's location.
- Partner course providercovers pythoncovers kubernetesBengaluruVisit partner →
- edXVerified partnerPartner course provider
Courses and programmes from universities and institutions worldwide.
covers python - UdemyVerified partnerPartner course provider
A marketplace of instructor-created courses across technology, business and creative skills.
covers python
Partners are ResumeKart affiliates or institutes it works with; ResumeKart may earn a commission when a candidate enrols. Placement is decided by relevance, not payment. How ResumeKart earns
Prepare to Win This Role
Everything you need to ace the interview and negotiate top-of-band compensation.