Site Reliability Engineer
Strobe is building the world's largest power plant: not a single site, but a distributed fleet of buildings, batteries, EVs, and generators that buy and sell power in real time across wholesale energy markets.
Keep a fleet of power plants online. Own reliability across the cloud platform, the control room, and the hardware in the field, for systems where downtime means missed grid events and real dollars.
What you'd own
Uptime and latency of the dispatch path, from the cloud decision to the command on site
Observability: metrics, logs, and alerts that page on real impact and nothing else
On-call, incident response, and postmortems for systems that control physical equipment
AWS infrastructure as code (Terraform), CI/CD, and safe deploys
Edge fleet reliability: connectivity, OTA updates, and remote diagnostics (balena, containerized runtimes)
Network links to utilities and grid operators (VPNs, tunnels, DNP3) and how they fail
Capacity, cost, and resilience of the data and control platforms
Strong fit if you have
Production SRE or infrastructure experience on AWS
Terraform at scale (modules, remote state, multi-workspace patterns)
Strong Linux, networking, and debugging fundamentals
Alerting and on-call that people trust, because you built it
Python, Rust, or Go for tooling
Bonus points
Edge or IoT fleets, industrial protocols, or OT networks
Site connectivity: SD-WAN, VPNs, cellular failover
Energy or other safety-critical, real-time systems
Stack: AWS / Terraform / Python / Rust / balena / Grafana / DNP3 / Modbus. AI-agent-native monorepo with deep investment in agent-enabled engineering efficiency.
Small team, high ownership, no layers.