Site Reliability Engineer

Full-Time
United States (Remote)
Posted 1 month ago

Company Overview

We are a dynamic and rapidly expanding enterprise technology organization operating in a cloud-first environment where system reliability, service availability, infrastructure performance, automation, monitoring, incident response, and operational resilience directly influence business performance. Our organization supports enterprise customers, product teams, software engineering groups, data teams, DevOps partners, security teams, infrastructure teams, customer success teams, and executive stakeholders who depend on stable digital platforms and highly available services.

As companies scale across cloud-native applications, distributed systems, AI-enabled products, and enterprise SaaS platforms, site reliability engineering has become one of the most important foundations for customer trust and operational excellence. Modern organizations need SRE professionals who can do more than respond to alerts or maintain infrastructure. They need engineers who can build reliable systems, improve observability, automate operational workflows, reduce incident frequency, support scalable architecture, improve deployment safety, and help engineering teams design services that are resilient from the start.

Our company is investing heavily in cloud infrastructure modernization, Kubernetes platforms, observability, service-level objectives, incident response maturity, reliability automation, infrastructure as code, performance engineering, disaster recovery, DevSecOps practices, AI-enabled operations, and scalable platform engineering. We are seeking a highly capable Site Reliability Engineer who can improve system reliability, strengthen production operations, and support the engineering practices that keep enterprise services secure, performant, and available.

This is not a traditional operations role focused only on ticket handling or reactive production support. The Site Reliability Engineer will serve as a hands-on technical contributor responsible for improving reliability across applications, infrastructure, deployment pipelines, monitoring systems, and incident response workflows. This role requires strong engineering judgment, automation discipline, systems thinking, and the ability to partner with software teams to reduce operational risk and improve service quality.

The selected candidate will work closely with Software Engineering, Platform Engineering, DevOps, Cloud Infrastructure, Security, Product, Data Engineering, QA, Customer Support, Compliance, and leadership teams to ensure systems are reliable, scalable, observable, secure, cost-aware, and aligned with business needs.

This is a strong opportunity for a reliability-focused engineer who wants to work with modern cloud systems, influence technical operations, improve production resilience, and contribute to a remote-first organization where reliability is treated as a strategic driver of customer trust, engineering speed, and long-term growth.

Your Role at the Company

As Site Reliability Engineer, you will support the reliability, availability, scalability, performance, and operational maturity of cloud-based applications and infrastructure.

You will be responsible for improving monitoring, alerting, automation, deployment safety, incident response, production readiness, infrastructure resilience, and service health across modern technology platforms.

You will serve as a trusted technical partner to engineering, DevOps, platform, security, and product teams while maintaining accountability for production reliability, automation quality, operational documentation, root cause analysis, and continuous improvement.

The ideal candidate combines strong systems engineering experience, cloud infrastructure knowledge, scripting ability, observability expertise, incident response discipline, and practical communication to support reliable enterprise technology operations.

What You’ll Do

Monitor, maintain, and improve the reliability, availability, scalability, and performance of production systems, cloud services, APIs, databases, and platform infrastructure.

Build and improve observability through monitoring, logging, tracing, service health checks, dashboards, alert tuning, synthetic monitoring, and operational reporting.

Support incident response by investigating service issues, coordinating technical response, communicating status, restoring service, and participating in post-incident reviews.

Develop automation to reduce manual operational work, improve reliability, standardize recovery procedures, and support repeatable infrastructure and deployment workflows.

Partner with Software Engineering teams to improve application reliability, production readiness, service ownership, dependency management, and operational design.

Partner with DevOps and Platform teams to improve CI/CD pipelines, deployment validation, rollback procedures, environment consistency, and release safety.

Support cloud infrastructure across AWS, Azure, Google Cloud, or hybrid environments with a focus on resilience, performance, security, and cost efficiency.

Support containerized and orchestration environments such as Docker, Kubernetes, EKS, AKS, GKE, ECS, OpenShift, Helm, or similar platforms.

Define and support service-level indicators, service-level objectives, error budgets, uptime goals, alert thresholds, reliability metrics, and operational scorecards.

Troubleshoot production incidents involving latency, service errors, failed deployments, database issues, network problems, scaling constraints, infrastructure failures, or customer-impacting outages.

Improve disaster recovery readiness, backup validation, failover procedures, capacity planning, dependency mapping, and business continuity practices.

Partner with Security teams to support secure infrastructure practices, access controls, secrets management, vulnerability remediation, compliance checks, and incident readiness.

Create and maintain runbooks, architecture documentation, incident response procedures, monitoring guides, operational standards, and reliability improvement plans.

Identify recurring incidents, technical debt, manual processes, performance bottlenecks, noisy alerts, platform risks, and opportunities to improve reliability engineering maturity.

What You’ll Bring

5+ years of experience in site reliability engineering, DevOps, cloud infrastructure, platform engineering, systems engineering, production operations, or software engineering.

Experience working within enterprise SaaS, AI-enabled technology, cloud platforms, cybersecurity, fintech, healthcare technology, ecommerce, digital products, or large-scale technology environments preferred.

Strong experience with cloud platforms such as AWS, Azure, or Google Cloud, including compute, networking, storage, IAM, managed databases, load balancing, monitoring, and security services.

Experience with Linux systems, networking fundamentals, DNS, load balancing, firewalls, HTTP, TLS, service discovery, distributed systems, and production troubleshooting.

Experience with observability tools such as Datadog, New Relic, Prometheus, Grafana, Splunk, CloudWatch, Azure Monitor, OpenTelemetry, ELK, Honeycomb, or similar platforms.

Experience with infrastructure as code and automation tools such as Terraform, CloudFormation, Pulumi, Ansible, Chef, Puppet, Helm, Bash, Python, Go, or PowerShell.

Experience with CI/CD tools such as GitHub Actions, GitLab CI, Jenkins, CircleCI, Azure DevOps, Argo CD, Harness, or similar platforms.

Experience with container platforms such as Docker, Kubernetes, EKS, AKS, GKE, ECS, OpenShift, service mesh, container registries, or orchestration workflows preferred.

Strong understanding of incident response, root cause analysis, service-level objectives, error budgets, monitoring strategy, change management, and production readiness.

Familiarity with databases, caching systems, message queues, API services, microservices, event-driven architecture, and cloud-native application patterns preferred.

Ability to work with Software Engineering, DevOps, Platform Engineering, Security, Product, Data Engineering, QA, Customer Support, Compliance, and leadership teams.

Strong troubleshooting, documentation, scripting, analytical thinking, communication, and follow-through skills.

Bachelor’s degree in Computer Science, Information Technology, Software Engineering, Systems Engineering, Computer Engineering, or a related technical field preferred.

AWS, Azure, Google Cloud, Kubernetes, Terraform, DevOps, SRE, Linux, security, or networking certification preferred.

Benefits

Competitive compensation package.

Comprehensive medical, dental, and vision healthcare coverage.

Flexible remote-first work environment.

Performance bonus eligibility.

Reliability and platform performance incentive opportunities.

Service availability and incident reduction incentive opportunities.

Long-term incentive opportunities where applicable.

Retirement savings plan with company contribution.

Professional development and continuing education reimbursement.

Site reliability engineering, cloud infrastructure, Kubernetes, observability, incident response, DevOps, infrastructure as code, AI-enabled operations, and distributed systems training resources.

Wellness and mental health support programs.

Paid time off and company holidays.

Opportunity to support enterprise-level platform reliability, cloud modernization, incident response maturity, and scalable digital transformation initiatives.

Access to modern cloud platforms, observability systems, CI/CD tools, automation frameworks, security platforms, collaboration tools, and AI-enabled engineering operations resources.

Personal Capabilities and Qualifications

Reliability-focused engineer with strong ownership of production health, system performance, service availability, and operational excellence.

Strong problem-solver with the ability to investigate complex issues across applications, infrastructure, databases, networks, pipelines, and cloud services.

Automation-minded and improvement-oriented, with the ability to reduce manual work, improve consistency, and build repeatable reliability practices.

Calm and effective under pressure, especially during production incidents, customer-impacting outages, deployment failures, latency spikes, or urgent escalations.

Systems thinker with the ability to understand how services, infrastructure, dependencies, data flows, security controls, and business operations connect.

Security-conscious and responsible, with the ability to follow access control, secrets management, compliance, and infrastructure hardening practices.

Clear communicator who can explain reliability risks, incident findings, operational improvements, service health, and technical recommendations to technical and non-technical audiences.

Collaborative and able to work effectively with software engineers, DevOps partners, platform teams, security teams, product managers, customer support teams, and leadership.

Metrics-driven and practical, with the ability to use monitoring data, incident trends, deployment metrics, latency reports, and capacity signals to guide improvements.

High integrity and discretion when handling production systems, customer data, cloud credentials, security-sensitive information, architecture details, incident records, and internal documentation.

Strategic Support

Support engineering leadership with service reliability, platform health, incident response maturity, cloud operations, and operational improvement.

Help improve customer trust by reducing downtime, improving system performance, strengthening monitoring, and improving response to production issues.

Support software delivery by improving deployment safety, rollback readiness, release validation, environment consistency, and production readiness practices.

Align SRE work with business priorities, customer expectations, product roadmaps, security standards, compliance needs, and enterprise growth goals.

Strengthen observability maturity by improving dashboards, alerts, traces, logs, service health checks, SLOs, and operational reporting.

Partner with Software Engineering teams to improve service ownership, resilience patterns, dependency awareness, error handling, and scalable application design.

Partner with Platform and DevOps teams to improve infrastructure automation, Kubernetes operations, CI/CD reliability, cloud architecture, and operational tooling.

Partner with Security teams to support secure operations, vulnerability remediation, access governance, secrets management, compliance automation, and incident readiness.

Support cost and capacity planning by reviewing usage trends, scaling behavior, resource efficiency, storage growth, traffic patterns, and cloud utilization.

Help turn site reliability engineering into a strategic capability that improves uptime, engineering speed, customer confidence, operational resilience, and long-term enterprise value.

Working Conditions

Remote-first site reliability engineering role.

Periodic travel may be required for engineering offsites, reliability workshops, incident review sessions, architecture planning, company gatherings, cloud vendor meetings, or strategic technology sessions.

High-visibility technical role supporting production reliability, cloud infrastructure, observability, incident response, and platform operations.

Fast-paced environment focused on uptime, security, automation, performance, scalability, documentation, and continuous improvement.

Regular collaboration with Software Engineering, Platform Engineering, DevOps, Cloud Infrastructure, Security, Product, Data Engineering, QA, Customer Support, Compliance, and leadership teams.

Opportunity to influence reliability engineering maturity, service health, incident response, observability, cloud architecture, deployment safety, and operational resilience.

Requires flexibility during production incidents, release windows, urgent escalations, monitoring issues, platform migrations, customer-impacting events, and security reviews.

Role requires handling confidential cloud architecture details, customer data, system credentials, access records, production incident information, security-sensitive materials, and internal documentation with discretion.

Job Function

Site Reliability Engineering.

Production Reliability.

Cloud Infrastructure Support.

Platform Reliability.

Observability Engineering.

Incident Response.

SRE Automation.

Infrastructure as Code.

Cloud Operations.

CI/CD Reliability.

Kubernetes Operations.

Monitoring and Alerting.

Disaster Recovery Support.

Performance Engineering.

Remote Reliability Engineering.

Compensation & Benefits

Compensation Package: $240,000 – $360,000

Base Salary: $240,000 – $360,000.

Annual Performance Bonus.

Service Reliability Incentives.

Platform Performance and Incident Reduction Incentives.

Cloud Operations and Automation Incentives.

Long-Term Incentive Eligibility where applicable.

Additional benefits may include:

• Comprehensive healthcare coverage.

• Retirement savings plan with company contribution.

• Wellness and mental health programs.

• Professional education assistance.

• Flexible remote work environment.

• Paid time off and company holidays.

• Site reliability engineering, cloud infrastructure, Kubernetes, Terraform, observability, DevOps, incident response, AI-enabled operations, and distributed systems training.

• Access to modern cloud platforms, observability tools, CI/CD systems, automation frameworks, security platforms, collaboration systems, and AI-enabled engineering resources.

Equal Opportunity Statement

We are committed to fostering a workplace where innovation, diversity, and inclusion drive meaningful business outcomes. We provide equal employment opportunities to all applicants and employees regardless of race, religion, gender, age, disability, veteran status, sexual orientation, or any other protected characteristic under applicable law.

Why Join Us

Support the reliability foundation behind enterprise platforms, customer-facing applications, cloud services, and scalable digital products.

Work in a remote-first organization using modern cloud platforms, observability systems, automation frameworks, and AI-enabled engineering resources.

Partner with engineering, DevOps, platform, security, product, and customer support teams on meaningful reliability and production operations work.

Contribute to a company investing heavily in cloud modernization, SRE practices, DevOps maturity, automation, observability, and operational resilience.

Improve uptime, deployment confidence, incident response, monitoring quality, platform performance, and customer trust through meaningful technical work.

Build a strong SRE career platform by turning systems expertise, automation discipline, reliability ownership, and operational judgment into measurable enterprise value.

Job Features

Job CategoryIncident Response, Systems Engineering & Digital Transformation

Apply For This Job

A valid phone number is required.

Connecting great people
with great companies, faster.