← Back to Jobs

Senior AI Tools Engineer, SRE Operations - GeForce NOW

NVIDIACanada, United States🌍 Remote
Full-timeSenior
$144k - $230k
per year
👁️ 0 views📝 0 applicationsPosted 8/6/2026Expires 10/5/2026

Get alerts for roles like this

More Senior AI Tools Engineer, SRE Operations - GeForce NOW roles in Canada, United States — straight to your inbox. No account needed.

Applying to this role? Tailor your résumé to this job description in one click, then download it clean — no watermark, no subscription.

Job Description

NVIDIA has been transforming computer graphics, PC gaming, and accelerated computing for more than 25 years. It’s a unique legacy of innovation that’s fueled by great technology—and amazing people. Today, we’re tapping into the unlimited potential of AI to define the next era of computing.

An era in which our GPU acts as the brains of computers, robots, and self-driving cars that can understand the world. Doing what’s never been done before takes vision, innovation, and the world’s best talent.

As an NVIDIA N, you’ll be immersed in a diverse, supportive environment where everyone is inspired to do their best work. Come join the team and see how you can make a lasting impact on the world. We are seeking a passionate AI Tools Engineer to join the Site Reliability Engineering (SRE) Data Team.

Applicants with SRE or equivalent experience are encouraged. What you will be doing: You will build and deploy sophisticated AI-powered tools and products. These tools support the operation and optimization of a critical production global Geforce Now service.

This role is critical for transforming extensive production data streams—such as signals, metrics, and logs—into actionable intelligence. The intelligence automates root cause analysis for incidents and predicts future service trends and patterns.

Build and implement robust AI/ML tools capable of analyzing production data to identify root causes for complex incidents and identify future operational trends. Lead the development of brand-new LLM- and Agent-based systems to improve operational efficiency.

Establish and maintain excellent data management practices, including building pipelines to transform and handle large-scale data sources vital for model development. Take charge of and enhance LLM-based pipelines while integrating a strong grasp of LLM progress into product development.

Act as a resident authority on AI Frameworks, recommending the best platforms, toolsets, and architectural approaches to ensure the long-term