Kraken logo
    K

    Site Reliability Engineer - Telemetry

    Kraken
    LATAMBrazilUruguayPeruColombiaArgentinaParaguayCanadaChile
    Remote
    Mid Level
    Full Time
    about 16 hours ago
    Site Reliability EngineerTelemetryPrometheusTerraformKubernetesObservabilityDistributed Systems

    Requirements

    • 3+ years of experience as a Site Reliability Engineer, Platform/Infrastructure Engineer, Observability Engineer, or similar production engineering role.
    • Comfortable managing production systems at scale that collect, process, store, and serve telemetry such as metrics, logs, traces, or profiles.
    • Experience with Prometheus or a Prometheus-compatible monitoring stack, including metrics collection, querying, and alerting.
    • Experience troubleshooting distributed production systems, including availability, latency, data flow, and capacity issues.
    • Experience with Infrastructure as Code, particularly Terraform, and CI/CD.
    • Experience operating containerised workloads with Nomad, Kubernetes, or similar platforms.
    • Solid scripting/programming ability and comfort using AI tools and agents (e.g., Claude) to accelerate delivery.
    • Strong incident response, documentation, and collaboration skills.

    What You'll Do

    • Operate and improve the shared platform for metrics, logs, traces, alerting, dashboards, and profiling.
    • Maintain metrics collection, long-term storage, querying, dashboards, and alerting using Prometheus-compatible systems, VictoriaMetrics, Grafana, and modern alerting tools.
    • Operate log pipelines using Vector, Splunk, and Loki, including reliability, throughput, and troubleshooting.
    • Operate distributed tracing and profiling capabilities using Grafana Alloy, Tempo, OpenTelemetry, and Pyroscope.
    • Deploy and manage telemetry services using Terraform, Terragrunt, and container orchestration across multiple environments.
    • Troubleshoot missing data, slow queries, broken alerts, pipeline backpressure, and capacity issues.
    • Build reusable configuration and automation that helps teams manage dashboards, alerts, and telemetry integrations safely.
    • Participate in incident response and on-call, write runbooks, and improve the platform using what we learn from incidents.

    Nice to Have

    • Experience with VictoriaMetrics, Grafana, Tempo, Loki, Vector, Splunk, Alertmanager, or OpenTelemetry.
    • Experience with PromQL or LogQL, dashboards or alerts as code.
    • Experience with maintaining operators and their related CRDs in Kubernetes.
    • Experience with Consul, Vault, AWS, and on-premises or datacentre infrastructure.
    • Experience operating high-volume logging, streaming, or data pipelines.
    • Experience making practical trade-offs between observability data volume, performance, and cost.
    • Background in highly regulated or financial services environments where change management and audit trails are critical.

    About Kraken

    Kraken is a cryptocurrency exchange platform that provides parachain auctions, staking, and index services.

    San Francisco, CA
    1000 - 5000
    Finance