Senior Site Reliability Engineer

AI Matched Verified Fresh Posting

Senior Site Reliability Engineer

Vivaldi

July 12, 2025

Location: Remote
Salary: $180,000 - $200,000
Job Type: Full-Time
Seniority Level: Senior-Level

Job Description

We’re looking for a Staff or Senior SRE who has helped build and run modern infrastructure in early-stage, high-growth startups. If you've taken systems from fragile to resilient, know how to move fast without burning things down, and care deeply about helping engineers ship reliably at scale—this role was made for you.

This is an opportunity to join a scaling company during a foundational growth phase. You'll work closely with product engineers and platform teams to design, build, and refine the systems that power everything from daily development workflows to production uptime and observability. We're a remote-first team building tools used by millions of users globally, and we need someone who thrives in ambiguity, takes ownership, and knows how to bring structure to fast-moving environments.

What You’ll Be Responsible For

  • Designing and improving the reliability and performance of distributed systems across cloud environments.
  • Building observability into everything—from telemetry pipelines and alerting to incident response automation.
  • Creating infrastructure that can handle scale and complexity without compromising developer speed.
  • Defining SLIs, SLOs, and risk thresholds that help engineering teams ship confidently.
  • Leading incident triage, postmortems, and operational playbook development.
  • Introducing and evolving CI/CD pipelines, infrastructure as code, and deployment workflows.
  • Collaborating closely with developers to ensure production systems support experimentation, safety, and velocity.

What We’re Looking For

  • 10+ years of experience in SRE, DevOps, or software/platform engineering, ideally with increasing scope and ownership.
  • Proven history of working in startup environments—especially during stages of rapid growth or platform scale-up.
  • Strong in at least two modern programming languages (e.g., Python, Go, TypeScript, JavaScript).
  • Deep knowledge of cloud infrastructure (especially AWS: ECS, EKS, CloudRun) and IaC tools like Terraform or CDK.
  • Experience operating and scaling observability platforms such as Prometheus, Datadog, OpenTelemetry, or ELK.
  • Familiarity with Kubernetes, container orchestration, and cloud networking.
  • Skilled in designing resilient systems, debugging distributed environments, and improving deployment pipelines.
  • Comfortable with security, networking, Linux systems, and performance tuning at the infrastructure level.
  • Exposure to database optimization and tuning for data-intensive applications is a plus.
  • Formal degree in Computer Science (or related technical field) or equivalent industry experience.

You’ll Succeed Here If You:

  • Have a history of building in environments with high ownership and low process overhead.
  • Approach infrastructure with a product mindset—always thinking about usability and impact.
  • Are a clear communicator who knows how to collaborate across teams and functions.
  • Stay calm under pressure and believe every incident is a chance to strengthen the system.
  • Enjoy diving into complex technical problems and working toward long-term improvements, not quick fixes.
  • Care about empowering developers to move fast safely.

Why This Role Is Unique

  • You’ll shape systems from the ground up during a period of rapid growth and platform evolution.
  • Your work will support products used globally across scientific, technical, and research communities.
  • You’ll join a high-caliber team that values pragmatic solutions, humility, and real-world impact.
  • You’ll have the freedom to work remotely while being part of a supportive, mission-aligned team.
Canada or United States, Only.

Requirements & Responsibilities

  • AWS, GO, Python, Typescript, Datadog, Prometheus