Technical Program Manager, Platform
Scale AI•San Francisco, CA; New York, NY
Full-timeMid Level
👁️ 13 views•📝 0 applications•Posted 6/5/2026•Expires 9/15/2026
Get alerts for roles like this
More Technical Program Manager, Platform roles in San Francisco, CA; New York, NY — straight to your inbox. No account needed.
Applying to this role? Tailor your résumé to this job description in one click, then download it clean — no watermark, no subscription.
Job Description
As a Technical Program Manager for the Platform team, you will partner with engineering teams to directly accelerate the development and maturity of the Scale Generative AI Platform (SGP) . We are looking for a TPM who has actively built and shipped products in the past and understands how to deliver robust, scalable developer tooling and distributed systems. In this role, you will own the strategic alignment and end-to-end execution of our most critical infrastructure initiatives—from initial scoping to measurable, company-wide and customer-ready adoption. You will serve as the core communication backbone and connective tissue between platform engineering, product teams, and executive leadership. Operating in a hyper-growth, demanding AI environment, you will translate SGP’s architectural complexities into clear execution strategies , unblock engineering bottlenecks, proactively mitigate deployment risks, and ensure our foundational platforms deliver reliable, performant, and secure systems capable of global-scale deployment. Key Responsibilities Lifecycle & Platform Delivery: Lead strategic planning and high-velocity execution for SGP core capabilities (orchestration layers, model serving, APIs). Manage features from technical scoping and architecture design through production launch. Cross-Functional GenAI Alignment: Drive execution and manage complex technical dependencies across systems engineering, Core ML, Research, and Product teams to deliver unified SGP capabilities with architectural consistency. Technical Translation & Requirements: Translate complex infrastructure metrics (LLM inference optimization, GPU utilization, compute orchestration) into actionable roadmaps. Map demands like multi-tenancy, data privacy, and isolation into platform features. Risk & Dependency Mitigat
Required Skills
AgileAWSCross-functional CollaborationData PrivacyGCPJIRAKubernetesLeadershipLearning & DevelopmentMachine LearningProgram ManagementProject ManagementRecruitmentStrategic PlanningTPM