Senior Software Engineer, Grid Communications
Gridware is a San Francisco-based technology company dedicated to protecting and enhancing the electrical grid. We pioneered a groundbreaking new class of grid management called active grid response (AGR), focused on monitoring the electrical, physical, and environmental aspects of the grid that affect reliability and safety. Gridware’s advanced Active Grid Response platform uses high-precision sensors to detect potential issues early, enabling proactive maintenance and fault mitigation. This comprehensive approach helps improve safety, reduce outages, and ensure the grid operates efficiently. The company is backed by climate-tech and Silicon Valley investors.
About the role
As a member of the Grid Communications and Platform team, you will develop a control system for our distributed fleet of devices. This system will enable performant communications with our devices, including the ingestion of millions of events per day, operations such as sensor data retrievals and device configuration updates, and supporting the work of other teams reliant on our device's data. You will own, test, operate, and plan the future of our services end-to-end. Along the way, you’ll partner closely with other software teams such as Data Engineering and DevOps, as well as cross-functional teams such as Firmware, Operations, and Data Science. As an early member of the GCAP team, your technical and non-technical decisions will influence the long-term architecture and trade-offs of our system.
Responsibilities
- Design, build, test, and operate software that is scalable, observable, secure, and fault-tolerant.
- Create low-latency, event-driven pipelines for high-volume device telemetry and command processing, and make key architectural decisions from protocol design to infrastructure.
- Own our simulated and hardware-in-the-loop testing software in collaboration with the Firmware team.
- Collaborate with Data Science and Operations teams to ensure your designs account for a variety of environmental, connectivity, and operating conditions.
- Lead cross-functional projects to optimize system performance end-to-end.
- Own observability, monitoring, and incident response capabilities to support reliable production operations.
Requirements
- 5+ years of hands-on experience developing distributed, event-driven systems on cloud-native platforms using Python or Go.
- Experience with Kafka, Kinesis, Redpanda, or Apache Pulsar.
- Experience with using observability, monitoring, and logging tools such as Grafana, Prometheus, Loki, or similar.
- Strong communication skills and the ability to navigate ambiguity.
Skills
Bonus Skills:
- Production experience optimizing transport layer protocols (TCP/UDP/QUIC).
- Deep understanding of networking, DNS, TLS, and DTLS.
- Experience developing software for distributed physical, IoT, or sensing systems.
- Experience with Apollo Router / GraphQL federation gateways.
- Expertise with Kubernetes, GitOps workflows, and Infrastructure as Code.
- Experience in high-growth startup environments where you must wear many hats.
This describes the ideal candidate; many of us have picked up this expertise along the way. Even if you meet only part of this list, we encourage you to apply!
Benefits
- Health, Dental & Vision (Gold and Platinum plans with some providers fully covered).
- Paid parental leave.
- Alternating day off (every other Monday).
- “Off the Grid”, a two-week per year paid break for all employees.
- Commuter allowance.
- Company-paid training.