Site Reliability Engineer, Apple Data Platform / Big Data Platform
Apple•Mc Neil, Travis County
Full-timeMid Level
$140k - $140k
per year
👁️ 0 views•📝 0 applications•Posted 8/3/2026•Expires 9/2/2026
Get alerts for roles like this
More Site Reliability Engineer, Apple Data Platform / Big Data Platform roles in Mc Neil, Travis County — straight to your inbox. No account needed.
Applying to this role? Tailor your résumé to this job description in one click, then download it clean — no watermark, no subscription.
Job Description
The Apple Services Engineering team (ASE) is one of the most exciting examples of Apple"s long-held passion for combining art and technology. These are the people who power the App Store, Apple TV, Apple Music, Apple Podcasts, and Apple Books — at extensive scale, meeting high expectations to deliver a huge variety of entertainment in over 35 languages to more than 150 countries.\\n\\nWithin ASE, the Apple Data Platform SRE team keeps a massive, multi-cloud platform running for thousands of internal engineers building the next generation of data and AI products at Apple. We sit at the intersection of infrastructure, automation, and customer success — running incident response, providing hands-on support to internal teams, and partnering with developers to make cutting-edge services like Spark, Flink, Airflow, Trino, Notebooks, and LLM-based agent platforms reliable at scale.\\n
This is a rare opportunity to build deep expertise across one of the most technically diverse platforms at Apple — while specialising in the big data engines and catalog/governance layers that power analytics and data engineering across the company. As an SRE on Apple Data Platform, you"ll operate and support the team"s full portfolio, from ML/AI platform services to multi-cloud infrastructure, and grow into the team"s go-to expert for big data platform services — including Spark, Flink, Airflow, Trino, Notebooks, REST Catalog services (such as Glue Catalog), and data governance. Just as importantly, you"ll be a first point of contact for the internal customers who rely on these services daily — someone who can translate a confusing error or a vague support request into a clear diagnosis and a fast resolution.\\n\\nWe"re looking for a self-motivated engineer who thrives on ownership — someone who wants a set of services to call their own, the autonomy to drive their reliability roadmap, and the collaborative instinct to keep that work aligned with the team"s broader direction. If you love solving hard operational problems, take genuine satisfaction in helping frustrated customers get unblocked, and want a front-row seat to how Apple"s data engineering platform scales, this role offers real room to grow your scope and impact over time.\\n
Operate, monitor, and triage production and non-production environments across the ADP portfolio — data processing, ML/AI, and multi-cloud infrastructure.\\nParticipate in a rotating on-call schedule across supported services, including occasional weekday and weekend coverage.\\nOwn the operational health of big data platform services as SME — driving reliability, support, and customer guidance for Spark, Flink, Airflow, Trino, Notebooks, REST Catalog, and governance tooling.\\nServe as a primary point of contact for internal customers via Slack — clearly communicating status, root cause, and next steps during active issues.\\nScreen, triage, and resolve customer-reported service issues and support tickets, prioritizing based on customer impact and urgency.\\nPartner with dev teams across time zones to onboard new services — understanding architecture, then designing monitoring, alerting, and dashboards (Prometheus, Grafana, Splunk).\\nBuild automation and self-healing tooling that reduces manual toil and scales the team"s operational capacity.\\nIdentify, escalate, and resolve production issues to protect platform reliability and customer experience.\\nChampion customer success by helping internal teams understand platform capabilities and adopt tools effectively.\\nCollaborate with SRE and dev partner teams, engineering, and program management to align execution with team and org goals.
Bachelor"s Degree in Computer Science, an engineering-related field, or equivalent related experience.\\n1-4 years in a Site Reliability Engineering, DevOps, or Infrastructure-focused role.\\nProficient in Python; working knowledge of Golang a plus.\\nDeep understanding of one or more Big Data technologies (Spark, Flink, Airflow, Trino, Notebooks).\\nExperience with Kubernetes and at least one major cloud provider (AWS or GCP).\\nExcellent written and verbal communication skills, with the ability to explain technical issues clearly to non-expert customers.\\nSolid grounding in SRE principles, with prior on-call, production-support, or customer-facing support role experience.
Experience with REST Catalog services (e.g., Glue Catalog) and data governance frameworks.\\nPrior experience in a customer-facing or technical support role, with a demonstrated passion for customer success.\\nFamiliarity with observability tooling: Prometheus, Grafana, Splunk, PagerDuty.\\nWorking knowledge of CI/CD pipelines and deployment workflows.\\nExperience with S3 and cloud storage/networking fundamentals.\\nFamiliarity with data pipeline orchestration and workflow scheduling patterns.\\nA track record of automating manual operations through scripting or tooling.\\nIntellectual curiosity and a drive to keep learning — for yourself, your team, and the org.\\n