Job Description
Salary: £45,000 - 80,000 per year
Requirements:
• We require expert-level experience deploying and operating production infrastructure or platform services in a Linux-based environment using Kubernetes, containers, cloud or private-cloud platforms, infrastructure-as-code, and CI/CD/GitOps tooling.
• We require very strong understanding of security fundamentals for production platforms, including identity, secrets, access control, network segmentation, vulnerability management, and audit logging.
• We value familiarity with AI platform concepts such as model routing, MCP servers, agentic workflows, RAG systems, or LLM observability.
• We require very strong automation and scripting skills, for example Terraform, Go, Python, or similar.
• We require experience with incident management, problem management, demand forecasting, and production readiness practices.
• We value experience operating internal developer platforms, AI platforms, model gateways, MCP infrastructure, or other shared engineering platforms.
• We value experience with service mesh, policy-as-code, workload identity, sandboxing, secure runtime environments, or multi-tenant platform designs.
• We value experience with regulated or security-sensitive engineering environments.
• We value working knowledge of Open AI tools, products, and APIs. Responsibilities:
• We will build, deploy, and operate the infrastructure for centrally hosted AI platform services, including MCP server infrastructure, model gateway services, and supporting control-plane components.
• We will design runtime patterns for isolation, scalability, secure execution, capacity management, and cost-aware operation.
• We will automate provisioning, configuration, upgrades, and lifecycle management using infrastructure-as-code and GitOps patterns.
• We will define and implement service-level indicators, service-level objectives, alerting, dashboards, runbooks, and support workflows.
• We will handle incident response, post-incident review, and vendor outage management for AI services embedded in engineering workflows.
• We will build telemetry that helps us understand AI platform health, usage, performance, cost, and operational risk.
• We will ensure platform components meet production readiness, security, and compliance expectations.
• We will help make secure AI usage the default by providing reliable paved paths rather than manual or fragmented infrastructure.
• We will provide technical leadership as a recognised expert and lead engineer on our AI platform technology stack.
• We will own reliability, scalability, monitoring, alerting, incident response, runbooks, and operational readiness for AI platform components.
• We will develop automation for provisioning, deployment, configuration, backup, recovery, patching, upgrades, and lifecycle management.
• We will implement secure runtime patterns, including workload isolation, secrets management, identity integration, network controls, and auditability.
• We will contribute to platform roadmap planning and prioritisation.
• We will participate in production support and our paid on-call rota for high-impact incident response. Technologies:
• AI
• CI/CD
• Cloud
• Embedded
• GitOps
• Incident Management
• Support
• Kubernetes
• LLM
• Linux
• MCP
• Network
• Python
• RAG
• Security
• Terraform
• ARM
More:
We are building a central AI control plane to support the safe, scalable, and efficient use of AI across our engineering teams. This hands-on, technical leadership, platform engineering role helps direct and own delivery of the runtime platforms that make AI services reliable, secure, observable, and supportable at our scale. You will work across Kubernetes, cloud, identity, secrets, networking, telemetry, incident management, and automation to provide the production foundation for our AI platform. We offer an encouraging and inclusive environment that supports continuous learning and professional growth, along with a hybrid working approach designed to balance high performance and personal wellbeing. We are committed to equal opportunities and a respectful workplace.
last updated 36 week of 2026