San Jose, California, United States•San Jose, United States (US)
Remote
Senior
Full Time
21 days ago
💰$178,000 - $321,000
AIcloudAWSGCPPythonTypeScriptGoagentic runtimemulti-agentdata engineering
Requirements
- •7+ years building and operating resilient backend or platform systems in production including on-call ownership
- •Proven brownfield migrations from prototype to production grade while in daily use
- •Strong engineering fundamentals in data structures and algorithms
- •Fluent Python programming skills
- •Strong SQL and data modeling skills
- •Proficiency in an additional systems language such as TypeScript/Node or Go
- •Experience building agentic runtime and harness including model-agnostic agent orchestration, model routing and evaluation, agent SDKs, MCP servers, skill- and hook-based agent tooling, and evaluation/red-team harnesses with responsible-AI controls
- •Data engineering with provenance and lineage
- •Security and data-protection engineering including encryption, identity and access management, secrets, and retention
- •Cloud experience with AWS or GCP including infrastructure-as-code (Terraform), CI/CD, high availability, disaster recovery, service-level objectives, and observability
- •Experience shipping inside locked-down enterprise environments with security guardrails
- •Ability to translate audit needs into systems and explain technical risk to non-engineers
What You'll Do
- •Inherit, operate, and progressively migrate the working prototype estate to the target platform without interrupting daily and board-cycle workflows
- •Re-architect Hive Mind into a resilient AWS or GCP platform with high availability, disaster recovery, defined service-level objectives, and full observability, and own it in production
- •Build the agentic runtime and harness including orchestration, multi-model routing, evaluation, red-team, and regression harnesses with responsible-AI controls
- •Build the Responsible-AI and model-governance layer including hallucination, bias, drift controls, output validation, guardrails, and complete logging
- •Build data infrastructure with provenance and lineage including immutable audit trails, versioned evidence, reproducible pipelines, and traceability
- •Engineer data protection including encryption, key and secrets management, least-privilege access, sensitive-data handling, residency, and defensible retention
- •Stand up the computer-assisted audit technique (CAAT) and continuous-monitoring data foundation with analytics over full populations and real-time exceptions streaming
- •Own the cloud foundation including infrastructure-as-code, CI/CD, identity, networking, observability, and cost controls, making the platform examinable
Nice to Have
- •Experience with multi-agent orchestration frameworks, MCP servers, and tooling/plugin development across Anthropic, OpenAI, and Google model ecosystems
- •Experience with model-risk or AI-governance programs and frameworks such as the NIST AI Risk Management Framework
- •Experience with continuous-auditing or continuous-controls-monitoring platforms, streaming and real-time data at scale, statistical anomaly detection, applied machine learning beyond LLMs, vector stores and retrieval, LLM cost engineering
- •Experience passing external audit, SOC 2, or SOX
- •Crypto and blockchain literacy; regulated financial-services, fintech, or crypto experience
Benefits
- •Competitive total compensation package
- •Learning and Development programs and Education subsidy for employees' growth and development
- •Various team building programs and company events
- •Wellness and meal allowances
- •Comprehensive healthcare schemes for employees and dependants
