Remote
Remote
Senior
Full Time
17 days ago
SRERelease EngineeringProduction OperationsRemoteAWSKubernetesInfrastructure as CodeOn-callIncident Response
Requirements
- •5+ years in SRE, production operations, platform engineering, or release engineering
- •Experience operating production systems at scale and carrying on-call
- •Fluency in SLAs, SLOs, error budgets, DORA metrics, and observability tooling (Prometheus, Grafana, Alertmanager, or similar)
- •Experience leading incident response with tools like incident.io, PagerDuty, or Opsgenie
- •Experience running blameless postmortems and reducing MTTD/MTTR
- •Confidence operating on AWS (multiple accounts, IAM, VPC) in production
- •Comfortable with infrastructure-as-code tools (Pulumi, Terraform) and Kubernetes
- •Ability to script and automate to eliminate toil
- •Clear communication skills with infrastructure specialists and product engineers
- •Ability to thrive in async, globally distributed teams
- •Comfortable navigating ambiguity and iterating toward better systems
What You'll Do
- •Own the reliability of deployment and release systems and control plane against SLOs and error budgets
- •Standardize and instrument deployment workflows
- •Drive disaster-recovery readiness and reproducible deployments
- •Build and operate health and SLO monitoring using synthetic testing
- •Reduce mean-time-to-detect and mean-time-to-recover for deploy-related incidents
- •Participate in on-call, lead blameless postmortems, and create runbooks and automation
- •Improve deployment observability and auditability
- •Document operational procedures and reliability knowledge
- •Define and track SLAs, SLOs, error budgets, and DORA delivery metrics with alerting
- •Ensure deployments fail fast and safely when health checks degrade
- •Harden access and break-glass workflows for incident response
- •Partner with product engineering and platform teams to align release practices with reliability targets
Benefits
- •Fully remote work with global hiring and WeWork membership or co-working allowance
- •ESOP (equity ownership) for all team members
- •Tech allowance for work environment setup
- •Health insurance coverage 100% for employees and 80% for dependents globally
- •Annual company off-sites for connection and collaboration
- •Flexible asynchronous work schedule
- •Annual professional development education allowance
