Skip to main content
← Back to Jobs

Senior Solutions Architect, AI Factory Observability and Visualization - NVIS

NVIDIACanada, United States🌍 Remote
Full-timeSenior
$184k - $357k
per year
👁️ 0 views📝 0 applicationsPosted 7/13/2026Expires 9/11/2026
Tailor Resume for This JobCheck ATS Score

Get alerts for roles like this

More Senior Solutions Architect, AI Factory Observability and Visualization - NVIS roles in Canada, United States — straight to your inbox. No account needed.

Applying to this role? Tailor your résumé to this job description in one click, then download it clean — no watermark, no subscription.

Job Description

NVIDIA 's Infrastructure Specialists team is hiring a Senior Solutions Architect - AI Factory Observability & Visualization!

This remote role develops full-spectrum visibility that supports the smooth functioning of HPC systems and AI factories, transforming intricate telemetry across network and compute into straightforward, actionable perspectives.

The role has a complete, end-to-end understanding of the HPC/AI system, running and interpreting microbenchmarks and workloads to confirm system readiness, then establishing the observability that maintains this state.

The work involves collaborating across NVIDIA teams to help partners see, understand, and respond to HPC system and AI factory performance, from hardware to workload.

What You Will be Doing: Run AI factory validation tools, microbenchmarks, and workloads provided by the team, and interpret results to assess system health and performance. Gain a comprehensive understanding of the system from start to finish, including network topology, interconnects, and compute.

Establish what "healthy" represents across the stack — the metrics, logs, and signals that confirm a system is functioning well, and the thresholds that show it isn't. Build and extend the telemetry surface across hardware, fabric, and workload, crafting how data is collected, transformed, stored, and surfaced.

Serve as the observability expert, investigating gaps in visibility to ensure it reflects true system behavior. Develop automation (Python, Shell) for collecting, transforming, and presenting system and network data.

Recommend improvements to system visibility, data sources, and reporting that give teams clearer insight.

Collaborate with hardware, software, networking, datacenter, and product groups to ready HPC systems and AI factories for customer deployment, contributing documentation and readiness materials throughout the process.

What We Need to See: Bachelor's degree or equivalent experience in Computer Science, Mathematics, Engineeri

Required Skills

DocumentationPython

The best-paying roles in your field. Every week. Free.

Join 10,000+ professionals getting job alerts and salary insights in their inbox

We respect your privacy. Unsubscribe anytime with one click.