Skip to main content
← Back to Jobs

Senior Software / Site Reliability Lead Engineer

General Dynamics Mission SystemsUnited States🌍 Remote
Full-time7-15
$143k - $158k
per year
👁️ 0 views📝 0 applicationsPosted 8/11/2026Expires 10/10/2026
Tailor Resume for This JobCheck ATS Score

Get alerts for roles like this

More Senior Software / Site Reliability Lead Engineer roles in United States — straight to your inbox. No account needed.

Applying to this role? Tailor your résumé to this job description in one click, then download it clean — no watermark, no subscription.

Job Description

Basic Qualifications Bachelor's degree in Software Engineering, or related Science, Technology, Engineering or Mathematics field, plus a minimum of 8 years of relevant experience; or Master's degree, plus 6 years relevant experience.

CLEARANCE REQUIREMENTS: Ability to obtain a Department of Defense Secret security clearance is required at time of hire. Applicants selected will be subject to a U. S. Government security investigation and must meet eligibility requirements for access to classified information.

Due to the nature of work performed within our facilities, U. S. citizenship is required.

Responsibilities

for this Position What You Will Own Cross-pod reliability standards . Set the reliability bar and ensure it is met consistently across applications. Collaborate with Functional SREs to connect technical reliability metrics to business-side outcomes.

You own the engineering signal; together you tell the full reliability story. SLOs and reliability metrics . Own definitions of service level objectives for every AI service that goes to production. Establish error budgets and use them to drive engineering decisions — not just measure uptime.

Monitoring and observability . Implement and maintain the full observability stack — logging, metrics, tracing, and dashboards. You will know when something is degrading before users do. Design and manage alerting infrastructure that tells you what's wrong, not just that something is wrong.

Alerts you build catch real problems; they don't cry wolf. Incident response. Own on-call procedures, escalation paths, and incident management end-to-end. Lead post-incident reviews and maintain the reliability improvement backlog.

When something breaks, you coordinate the response and ensure it doesn't break the same way again. Production Readiness. Define and enforce the criteria that determine whether an AI service is ready for production. You are the gate between "it works in dev" and "it's ready to ship." Toil elimination. Identify and

Required Skills

Software Engineering

The best-paying roles in your field. Every week. Free.

Join 10,000+ professionals getting job alerts and salary insights in their inbox

We respect your privacy. Unsubscribe anytime with one click.