Skip to main content
ResumeKart
← Back to Jobs

Senior HPC Cluster Administrator - Deep Learning Frameworks Infrastructure

NVIDIA•Poland•🌍 Remote
Full-timeSenior
zł293k - zł507k
per year
👁️ 0 views•📝 0 applications•Posted 9/24/2026•Expires 10/23/2026
Tailor Resume for This JobCheck ATS Score

Get alerts for roles like this

More Senior HPC Cluster Administrator - Deep Learning Frameworks Infrastructure roles in Poland — straight to your inbox. No account needed.

Applying to this role? Tailor your résumé to this job description in one click, then download it clean — no watermark, no subscription.

Job Description

NVIDIA 's Deep Learning Frameworks (DLFW) Infrastructure team is looking for a deeply technical Senior HPC Cluster Administrator to lead the design, deployment, and reliability of our large-scale GPU compute clusters.

These systems run the most demanding deep learning training, inference, and high-performance computing workloads in the industry — from DGX/HGX platforms to ground-breaking Grace Blackwell systems.

You will drive architectural decisions across compute, networking, and storage, and partner closely with software, research, and product teams to keep our infrastructure ahead of the workloads it supports.

What you'll be doing: Own the full lifecycle of GPU compute clusters — procurement, provisioning, configuration management, monitoring, and deprecation — across heterogeneous Linux environments (DGX, HGX, embedded systems) Design and scale storage solutions (NFS, Lustre, WekaFS, or equivalent) with a clear roadmap for capacity and performance growth Lead automation of infrastructure using modern IaC tools (Ansible, Terraform) and CI/CD pipelines (GitLab) Manage and optimize job scheduling via Slurm, including fair-share policies, reservation management, and MIG/GPU partitioning strategies Maintain and improve observability stacks (Prometheus, Grafana, DCGM) and drive proactive resolution of hardware and software incidents Collaborate with ML engineers and software teams to tune cluster configuration for large-scale distributed training workloads Evaluate and introduce new technologies — networking fabrics (InfiniBand, NVLink, EFA/RDMA), storage tiers, container runtimes — to improve performance and reliability Mentor junior engineers and contribute to team-wide engineering standards What we need to see: BS/MS in CS, EE, CE, or equivalent hands-on experience 5+ years of experience deploying and administering large-scale HPC or ML training clusters Deep expertise in Linux systems administration at scale Strong scripting and automation skills in Python and/or

Required Skills

AnsibleCI/CDDeep LearningLinuxMachine LearningProcurementPythonTerraform

Partner picks for Senior HPC Cluster Administrator - Deep Learning Frameworks Infrastructure in Poland

Matched to the skills this page calls for and the candidate's location.

Partner
  • Partner course provider

    Structured programmes for software engineers, data science and DevOps.

    covers pythoncovers machine learning
  • edXVerified partner
    Partner course provider

    Courses and programmes from universities and institutions worldwide.

    covers pythoncovers machine learning
  • Partner course provider

    University and industry courses, professional certificates and online degrees.

    covers pythoncovers machine learning

Partners are ResumeKart affiliates or institutes it works with; ResumeKart may earn a commission when a candidate enrols. Placement is decided by relevance, not payment. How ResumeKart earns

The best-paying roles in your field. Every week. Free.

Join 10,000+ professionals getting job alerts and salary insights in their inbox

We respect your privacy. Unsubscribe anytime with one click.