Strobe
←All roles

Site Reliability Engineer

San Francisco, CA
Full-time
On-site

Strobe is building the world's largest power plant: not a single site, but a distributed fleet of buildings, batteries, EVs, and generators that buy and sell power in real time across wholesale energy markets.

Keep a fleet of power plants online. Own reliability across the cloud platform, the control room, and the hardware in the field, for systems where downtime means missed grid events and real dollars.

What you'd own

Uptime and latency of the dispatch path, from the cloud decision to the command on site

Observability: metrics, logs, and alerts that page on real impact and nothing else

On-call, incident response, and postmortems for systems that control physical equipment

AWS infrastructure as code (Terraform), CI/CD, and safe deploys

Edge fleet reliability: connectivity, OTA updates, and remote diagnostics (balena, containerized runtimes)

Network links to utilities and grid operators (VPNs, tunnels, DNP3) and how they fail

Capacity, cost, and resilience of the data and control platforms

Strong fit if you have

Production SRE or infrastructure experience on AWS

Terraform at scale (modules, remote state, multi-workspace patterns)

Strong Linux, networking, and debugging fundamentals

Alerting and on-call that people trust, because you built it

Python, Rust, or Go for tooling

Bonus points

Edge or IoT fleets, industrial protocols, or OT networks

Site connectivity: SD-WAN, VPNs, cellular failover

Energy or other safety-critical, real-time systems

Stack: AWS / Terraform / Python / Rust / balena / Grafana / DNP3 / Modbus. AI-agent-native monorepo with deep investment in agent-enabled engineering efficiency.

Small team, high ownership, no layers.