Skip to main content
← Back to Jobs

Senior Site Reliability Engineer (Hiring Globally)

COGNATIVUnited Kingdom🌍 Remote
Full-time3-7
👁️ 0 views📝 0 applicationsPosted 8/30/2026Expires 9/29/2026
Tailor Resume for This JobCheck ATS Score

Get alerts for roles like this

More Senior Site Reliability Engineer (Hiring Globally) roles in United Kingdom — straight to your inbox. No account needed.

Applying to this role? Tailor your résumé to this job description in one click, then download it clean — no watermark, no subscription.

Job Description

About the role: We run a distributed, camera-based video monitoring and AI alerting platform.

The system spans the full spectrum of modern and legacy infrastructure: an AWS-hosted fleet of Java microservices and workers, a GPU-backed computer-vision inference pipeline, a real-time streaming and presence layer, and thousands of on-premise edge "media boxes" deployed in the field that ingest camera feeds, serve HLS video, and stream events back to the cloud.

This is a reliability-first role. Your primary job is to keep a large, mixed operational estate healthy at scale: meaningful service objectives, trustworthy alerting, sound capacity, tested disaster recovery, and fast, calm incident response.

Delivery pipelines matter, but here they exist in service of reliability, not the other way around. We are not hiring a pipeline-and-self-service DevOps engineer who treats operations as a side concern.

We are hiring an SRE who owns uptime and operational quality, and who can write the software to make that uptime measurable and automatic. If you think in SLOs, error budgets, and blameless postmortems, and you are happiest when a noisy, fragile system becomes quiet and predictable on your watch, this is your role.

What you'll keep reliableYou will own the operational health of all of the following: Computer-vision / AI models: frame-based inference services running on GPU EC2 (g4dn-class, AWS Deep Learning AMIs), fed by camera frames from S3 and a Kafka (Amazon MSK) event bus, with Redis (ElastiCache) for state.

Outputs flow through an alerts pipeline (OutgoingInferenceMessage to SNS/IoT to notification workers). Python services: the AI/alerts inference tier and supporting tooling. Legacy Java services: ~140 Java 8 services and libraries (REST APIs, SQS/SNS workers, Lambda functions) running on Jetty 9.

4, deployed to Elastic Beanstalk, ECS, and Lambda. Edge appliances ("media boxes"): Ubuntu 22. 04 / Docker Compose appliances managed remotely over AWS IoT Core secure tunnellin

Required Skills

AWSJavaPythonGPUKafkaRedisElastiCacheS3SNSIoTDockerUbuntuElastic BeanstalkECSLambdaJettyAIComputer VisionMicroservicesSQSEvent BusDisaster RecoveryIncident ResponseBlameless PostmortemsSLOsError BudgetsStreamingReal-timeMonitoringAlertingCapacity Planning

The best-paying roles in your field. Every week. Free.

Join 10,000+ professionals getting job alerts and salary insights in their inbox

We respect your privacy. Unsubscribe anytime with one click.