Structured programmes for software engineers, data science and DevOps.
Senior HPC Cluster Administrator - Deep Learning Frameworks Infrastructure
Get alerts for roles like this
More Senior HPC Cluster Administrator - Deep Learning Frameworks Infrastructure roles in Poland — straight to your inbox. No account needed.
Applying to this role? Tailor your résumé to this job description in one click, then download it clean — no watermark, no subscription.
Job Description
NVIDIA 's Deep Learning Frameworks (DLFW) Infrastructure team is looking for a deeply technical Senior HPC Cluster Administrator to lead the design, deployment, and reliability of our large-scale GPU compute clusters.
These systems run the most demanding deep learning training, inference, and high-performance computing workloads in the industry — from DGX/HGX platforms to ground-breaking Grace Blackwell systems.
You will drive architectural decisions across compute, networking, and storage, and partner closely with software, research, and product teams to keep our infrastructure ahead of the workloads it supports.
What you'll be doing: Own the full lifecycle of GPU compute clusters — procurement, provisioning, configuration management, monitoring, and deprecation — across heterogeneous Linux environments (DGX, HGX, embedded systems) Design and scale storage solutions (NFS, Lustre, WekaFS, or equivalent) with a clear roadmap for capacity and performance growth Lead automation of infrastructure using modern IaC tools (Ansible, Terraform) and CI/CD pipelines (GitLab) Manage and optimize job scheduling via Slurm, including fair-share policies, reservation management, and MIG/GPU partitioning strategies Maintain and improve observability stacks (Prometheus, Grafana, DCGM) and drive proactive resolution of hardware and software incidents Collaborate with ML engineers and software teams to tune cluster configuration for large-scale distributed training workloads Evaluate and introduce new technologies — networking fabrics (InfiniBand, NVLink, EFA/RDMA), storage tiers, container runtimes — to improve performance and reliability Mentor junior engineers and contribute to team-wide engineering standards What we need to see: BS/MS in CS, EE, CE, or equivalent hands-on experience 5+ years of experience deploying and administering large-scale HPC or ML training clusters Deep expertise in Linux systems administration at scale Strong scripting and automation skills in Python and/or
Required Skills
Upskill for This Role
Courses from Udemy and edX matched to this role's skills.

Dive Into Ansible - Beginner to Expert in Ansible - DevOps

Ansible: Beginner to Pro

Ansible Mastery: Complete Guide to Infrastructure Automation

RedHat Specialist in Developing Ansible Automation 5 exames!

Red Hat RHCE EX294 & EX374 Ansible: 1500 Exam Questions

Ansible 11.0 for Beginners with Examples - DevOps
ResumeKart may earn a commission from these links at no extra cost to you.
Partner picks for Senior HPC Cluster Administrator - Deep Learning Frameworks Infrastructure in Poland
Matched to the skills this page calls for and the candidate's location.
- Partner course providercovers pythoncovers machine learningBengaluruVisit partner →
- edXVerified partnerPartner course provider
Courses and programmes from universities and institutions worldwide.
covers pythoncovers machine learning - Partner course provider
University and industry courses, professional certificates and online degrees.
covers pythoncovers machine learning
Partners are ResumeKart affiliates or institutes it works with; ResumeKart may earn a commission when a candidate enrols. Placement is decided by relevance, not payment. How ResumeKart earns
Prepare to Win This Role
Everything you need to ace the interview and negotiate top-of-band compensation.