Skip to main content
ResumeKart
← Back to Jobs

Software Engineer - AI Research Clusters

NVIDIAUnited States🌍 Remote
Full-time1-3
$124k - $196k
per year
👁️ 0 views📝 0 applicationsPosted 9/10/2026Expires 11/9/2026
Tailor Resume for This JobCheck ATS Score

Get alerts for roles like this

More Software Engineer - AI Research Clusters roles in United States — straight to your inbox. No account needed.

Applying to this role? Tailor your résumé to this job description in one click, then download it clean — no watermark, no subscription.

Job Description

NVIDIA is at the forefront of innovations in Artificial Intelligence, High-Performance Computing, and Visualization. Our invention—the GPU—functions as the visual cortex of modern computing and is central to groundbreaking applications from generative AI to autonomous vehicles.

We are now looking for a Software Engineer to help accelerate the next era of machine learning innovation.

In this role, you will propose and implement engineering solutions to ensure delivery of functional, reliable, secure, and performance-optimal GPU clusters to internal researchers, enable them to focus on training and development by reducing operational disruption and overhead, empower them for self-service continuous improvement on reliability, operational excellence & performance.

Your work will empower scientists and engineers to train, fine-tune, and deploy the most advanced ML models on some of the world’s most powerful GPU systems.

What You'll Be Doing: In this position, you will work with coworkers across the AI Platform organization to understand the pain points of validating, monitoring and operating GPU clusters at scale. Then you will design, develop and maintain engineering solutions to solve those pain points systematically.

You will also research in traditional AIOps and the emerging Agentic AI, and leverage it to further reduce the operation toil. You will participate in on-call support for systems, platforms built and owned by the team. What We Need To See: BS/MS in Computer Science, Engineering, or equivalent experience.

2+ years in software/platform engineering, including 1 year in ML infrastructure or distributed systems. Experience in software development lifecycle on Linux-based platforms. Strong coding skills in languages such as Python, C++ or Rust. Experience with Docker, Kubernetes, GitLab CI, automated deployments.

Experience with AIOps or Agentic AI and apply it successfully in production environment. Ways To Stand Out From The Crowd: Proficiency with full-stac

Required Skills

Computer ScienceEngineeringsoftware/platform engineeringML infrastructuredistributed systemsLinuxPythonC++RustDockerKubernetesGitLab CIautomated deploymentsAIOpsAgentic AI

Partner picks for Software Engineer - AI Research Clusters in United States

Matched to the skills this page calls for and the candidate's location.

Partner
  • edXVerified partner
    Partner course provider

    Courses and programmes from universities and institutions worldwide.

    covers pythoncovers computer science
  • UdemyVerified partner
    Partner course provider

    A marketplace of instructor-created courses across technology, business and creative skills.

    covers pythoncovers docker

Partners are ResumeKart affiliates or institutes it works with; ResumeKart may earn a commission when a candidate enrols. Placement is decided by relevance, not payment. How ResumeKart earns

The best-paying roles in your field. Every week. Free.

Join 10,000+ professionals getting job alerts and salary insights in their inbox

We respect your privacy. Unsubscribe anytime with one click.