Skip to main content
← Back to Jobs

Site Reliability Engineer (SRE)

Bright Vision TechnologiesUnited States🌍 Remote
Full-time7-15
$100k - $180k
per year
👁️ 0 views📝 0 applicationsPosted 8/14/2026Expires 10/13/2026
Tailor Resume for This JobCheck ATS Score

Get alerts for roles like this

More Site Reliability Engineer (SRE) roles in United States — straight to your inbox. No account needed.

Applying to this role? Tailor your résumé to this job description in one click, then download it clean — no watermark, no subscription.

Job Description

Site Reliability Engineer (SRE) - Remote Bright Vision Technologies is a technology consulting and software development company delivering cloud, AI, data, and enterprise solutions across the United States.

This is a fantastic opportunity to join an established and well-respected organization offering tremendous career growth potential. Job Title: Site Reliability Engineer (SRE) Location: 100% Remote (U. S.)

Position Type: Full-time, Direct W2 Salary Range: $100,000–$180,000 Annually Experience Required: 10+ years Sponsorship: U. S. Citizens, Green Card Holders, EAD Holders, and H-1B transfer candidates are encouraged to apply. We are unable to sponsor new H-1B visa petitions for this position.

Job Summary

We are seeking an experienced Site Reliability Engineer to ensure the availability, performance, and operational excellence of large-scale distributed systems in production.

As an SRE you will live at the boundary between development and operations, applying strong software engineering principles to infrastructure and operations problems, and continually pushing the platform toward higher reliability with lower operational toil.

The ideal candidate will combine deep systems knowledge with strong programming skills, a measurement-driven mindset, and the discipline to design, automate, and operate complex services so that reliability becomes a first-class engineering deliverable rather than a reactive concern.

Key Responsibilities

Define, instrument, and continually refine service-level objectives (SLOs), service-level indicators (SLIs), and error budgets for critical services, and use those measures to drive concrete engineering and prioritization decisions.

Lead incident response and resolution for production issues, acting as a calm and effective incident commander when needed, and ensuring high-quality post-incident reviews that drive lasting improvements. Design and implement comprehensive monitoring, logging, and tracing strategies using Prometheus, Grafana, Op

Required Skills

Site Reliability EngineeringDistributed SystemsSoftware EngineeringInfrastructure AutomationSLOsSLIsError BudgetsIncident ResponseMonitoringLoggingTracingPrometheusGrafana

The best-paying roles in your field. Every week. Free.

Join 10,000+ professionals getting job alerts and salary insights in their inbox

We respect your privacy. Unsubscribe anytime with one click.