Machine Learning Data Engineer
Get alerts for roles like this
More Machine Learning Data Engineer roles in United States — straight to your inbox. No account needed.
Applying to this role? Tailor your résumé to this job description in one click, then download it clean — no watermark, no subscription.
Job Description
Machine Learning Data Engineer – Remote Bright Vision Technologies is a technology consulting and software development company delivering cloud, AI, data, and enterprise solutions across the United States.
This is a fantastic opportunity to join an established and well-respected organization offering tremendous career growth potential. Job Title: Machine Learning Data Engineer Location: 100% Remote (U. S.)
Position Type: Full-time, Direct W2 Salary Range: $80,000–$100,000 Annually Experience Required: 6+ years Sponsorship: U. S. Citizens, Green Card Holders, EAD Holders, and H-1B transfer candidates are encouraged to apply. We are unable to sponsor new H-1B visa petitions for this position.
Job Summary
We are seeking an Machine Learning Data Engineer to build and operate the large-scale data systems that power modern AI training and evaluation pipelines.
The role combines deep data engineering expertise with a strong understanding of AI workloads, focusing on ingestion, transformation, quality assurance, lineage, and high-throughput delivery of data to training jobs across diverse modalities.
The ideal candidate has experience operating petabyte-scale data systems, strong software engineering fundamentals, and clear understanding of how data infrastructure choices propagate into model quality and training efficiency.
Required Qualifications Bachelor’s or Master’s degree in Computer Science or a related field. Six or more years of data engineering experience, with significant work supporting ML or AI workloads. Strong proficiency in Python and at least one JVM or systems language.
Deep experience with modern data processing frameworks such as Spark, Ray, or Beam. Hands-on experience operating petabyte-scale storage and pipeline systems. Strong understanding of distributed systems, data modeling, and storage formats.
Experience with dataset versioning, lineage, and reproducibility for ML workflows. Familiarity with high-throughput data loading for accelerator-based trai
Similar Jobs
Other open roles matched to this job's skills and location.