Skip to main content
ResumeKart
← Back to Jobs

ML Infrastructure Engineer

Bright Vision Technologies•United States•🌍 Remote
Full-time7-15
$100k - $150k
per year
👁️ 0 views•📝 0 applications•Posted 8/16/2026•Expires 10/15/2026
Tailor Resume for This JobCheck ATS Score

Get alerts for roles like this

More ML Infrastructure Engineer roles in United States — straight to your inbox. No account needed.

Applying to this role? Tailor your résumé to this job description in one click, then download it clean — no watermark, no subscription.

Job Description

ML Infrastructure Engineer - Remote Bright Vision Technologies is a technology consulting and software development company delivering cloud, AI, data, and enterprise solutions across the United States.

This is a fantastic opportunity to join an established and well-respected organization offering tremendous career growth potential. Job Title: ML Infrastructure Engineer Location: 100% Remote (U. S.)

Position Type: Full-time, Direct W2 Salary Range: $100,000–$150,000 Annually Experience Required: 6+ years Sponsorship: U. S. Citizens, Green Card Holders, EAD Holders, and H-1B transfer candidates are encouraged to apply. We are unable to sponsor new H-1B visa petitions for this position.

Job Summary

We are seeking an AI Infrastructure Engineer to design, build, and operate the platform layer that powers large-scale AI training and inference workloads.

The role focuses on GPU clusters, distributed training frameworks, scheduling, storage performance, and developer experience for ML engineers and researchers, with strong emphasis on reliability, efficiency, and cost control.

The ideal candidate has built or operated production AI infrastructure at scale, understands the interaction between hardware, kernel, scheduler, and ML framework, and brings strong software engineering discipline to platform work.

Key Responsibilities

Design and operate GPU and accelerator infrastructure for training and inference, spanning on-prem clusters, cloud-managed services, and hybrid configurations. Build scheduling, queueing, and resource-sharing systems that maximize accelerator utilization across many teams.

Integrate frameworks such as PyTorch, JAX, DeepSpeed, FSDP, Megatron-LM, and Ray Train into a unified platform offering. Operate high-performance storage systems and data pipelines that keep accelerators fed with training data at near-line-rate.

Design networking architectures supporting RDMA, InfiniBand, NCCL, and high-bandwidth collective communication. Build observability for AI work

Required Skills

GPU clustersdistributed training frameworksschedulingstorage performancePyTorchJAXDeepSpeedFSDPMegatron-LMRay TrainRDMAInfiniBandNCCLhigh-bandwidth collective communicationcloud-managed serviceshybrid configurationsobservability

Partner picks for ML Infrastructure Engineer in United States

Matched to the skills this page calls for and the candidate's location.

Partner
  • edXVerified partner
    Partner course provider

    Courses and programmes from universities and institutions worldwide.

  • UdemyVerified partner
    Partner course provider

    A marketplace of instructor-created courses across technology, business and creative skills.

  • Partner competition

    Hackathon platform used by student and community hackathons across India.

Partners are ResumeKart affiliates or institutes it works with; ResumeKart may earn a commission when a candidate enrols. Placement is decided by relevance, not payment. How ResumeKart earns

Prepare to Win This Role

Everything you need to ace the interview and negotiate top-of-band compensation.

The best-paying roles in your field. Every week. Free.

Join 10,000+ professionals getting job alerts and salary insights in their inbox

We respect your privacy. Unsubscribe anytime with one click.