NAMER
Remote
Mid Level
Full Time
about 1 month ago
💰$ 119,000 - $ 178,500
incident managementAI-powered workflowsremotetechnical operationsautomationincident response
Requirements
- •Experience in incident response, technical operations, or a reliability-adjacent role
- •Hands-on familiarity with incident tooling such as incident.io, PagerDuty, and Slack-based workflow automation
- •Comfortable building with SQL and reporting tools like Databricks, Looker, Grafana
- •Able to diagnose operational issues using logs, system context, and integration debugging
- •Can build lightweight automations, configure APIs, and prototype AI workflows without needing an engineer
- •Demonstrable daily use of AI for drafting, summarizing, building workflows, and triaging
- •Experience building repeatable AI-powered workflows
- •Applies verification and judgment to AI outputs especially in high-pressure incident contexts
- •Can articulate how AI usage has improved speed, quality, or operational capacity
- •Ability to independently drive problems to resolution with strong attention to detail in configuration, data quality, and documentation
- •Works visibly with proactive updates and no need to be chased
- •Proven track record of respectfully disagreeing and committing to decisions
- •Clear written communication across engineering, support, and operations audiences
- •Proactive in surfacing risks, blockers, and improvement opportunities
- •Comfortable pushing back on off-program requests and escalating trade-offs
- •Able to communicate highly technical incident information to non-technical audiences clearly and without jargon
What You'll Do
- •Own incident tooling operations including maintaining reliability and configuration of incident.io, PagerDuty, Slack workflows, on-call rotations, and escalation paths
- •Build and maintain AI-powered workflows such as incident thread summarization, postmortem draft generation, follow-up triage, severity classification, and data hygiene workflows
- •Build and sustain the Incident Commander community of practice, run regular touchpoints, share learnings, and coach responders
- •Operate data and reporting systems including building, maintaining, and troubleshooting dashboards and reports using Databricks, Grafana, Looker
- •Maintain documentation and enablement assets like playbooks, templates, and incident guides
- •Drive continuous improvement by participating in incidents and postmortem reviews, identifying patterns, and implementing improvements
- •Partner across the organization coordinating with Engineering, Support, GTM, Legal, and Finance teams
- •Support the incident program at scale by onboarding and enabling new stakeholder groups and building infrastructure for self-sufficient operation
