LATAM•Brazil•Uruguay•Peru•Colombia•Argentina•Paraguay•Canada•Chile
Remote
Mid Level
Full Time
about 16 hours ago
Site Reliability EngineerTelemetryPrometheusTerraformKubernetesObservabilityDistributed Systems
Requirements
- •3+ years of experience as a Site Reliability Engineer, Platform/Infrastructure Engineer, Observability Engineer, or similar production engineering role.
- •Comfortable managing production systems at scale that collect, process, store, and serve telemetry such as metrics, logs, traces, or profiles.
- •Experience with Prometheus or a Prometheus-compatible monitoring stack, including metrics collection, querying, and alerting.
- •Experience troubleshooting distributed production systems, including availability, latency, data flow, and capacity issues.
- •Experience with Infrastructure as Code, particularly Terraform, and CI/CD.
- •Experience operating containerised workloads with Nomad, Kubernetes, or similar platforms.
- •Solid scripting/programming ability and comfort using AI tools and agents (e.g., Claude) to accelerate delivery.
- •Strong incident response, documentation, and collaboration skills.
What You'll Do
- •Operate and improve the shared platform for metrics, logs, traces, alerting, dashboards, and profiling.
- •Maintain metrics collection, long-term storage, querying, dashboards, and alerting using Prometheus-compatible systems, VictoriaMetrics, Grafana, and modern alerting tools.
- •Operate log pipelines using Vector, Splunk, and Loki, including reliability, throughput, and troubleshooting.
- •Operate distributed tracing and profiling capabilities using Grafana Alloy, Tempo, OpenTelemetry, and Pyroscope.
- •Deploy and manage telemetry services using Terraform, Terragrunt, and container orchestration across multiple environments.
- •Troubleshoot missing data, slow queries, broken alerts, pipeline backpressure, and capacity issues.
- •Build reusable configuration and automation that helps teams manage dashboards, alerts, and telemetry integrations safely.
- •Participate in incident response and on-call, write runbooks, and improve the platform using what we learn from incidents.
Nice to Have
- •Experience with VictoriaMetrics, Grafana, Tempo, Loki, Vector, Splunk, Alertmanager, or OpenTelemetry.
- •Experience with PromQL or LogQL, dashboards or alerts as code.
- •Experience with maintaining operators and their related CRDs in Kubernetes.
- •Experience with Consul, Vault, AWS, and on-premises or datacentre infrastructure.
- •Experience operating high-volume logging, streaming, or data pipelines.
- •Experience making practical trade-offs between observability data volume, performance, and cost.
- •Background in highly regulated or financial services environments where change management and audit trails are critical.
