GridCARE logo
GridCARE

Senior Site Reliability Engineer

Redwood City, USAHybridPosted 2 weeks ago

Apply opens GridCARE's site. When you're back, we'll ask whether you applied.

Job type
Full-time
Work mode
Hybrid
Level
Senior
Department
Information Technology
Experience
Not listed
Posted
Sep 14, 2026

About the role

JOB DESCRIPTION

We're looking for a Senior SRE to own the reliability, scalability, and observability of our production systems. You'll work closely with platform and data engineering to keep high-throughput, data-intensive services running at the availability our customers (utilities, data center operators) require.

RESPONSIBILITIES

  • Design and operate infrastructure on AWS using Terraform and Kubernetes
  • Build monitoring, alerting, and observability (Prometheus, Grafana, Datadog, or similar) with meaningful SLOs/SLIs
  • Automate away toil — deployment pipelines, capacity management, self-healing systems
  • Partner with engineering on architecture reviews to catch reliability and scalability risks before they ship
  • Manage database and data pipeline reliability for large-scale, real-time grid data processing
  • Drive security and compliance best practices across infrastructure

QUALIFICATIONS

Required

  • 5+ years in SRE, DevOps, or infrastructure engineering roles
  • Deep experience with Kubernetes, Terraform/IaC, and cloud platforms (AWS Preferred)
  • Strong scripting/programming ability (Python, Bash)
  • Observability Experience (Prometheus, Grafana, Datadog)
  • Track record of running on-call for production systems and leading incident response
  • Experience with CI/CD pipelines (Github Actions) and infrastructure automation
  • Experience with Gitops concepts and tooling (ArgoCD/Flux)
  • Solid understanding of networking, distributed systems, and database reliability
  • Comfortable operating in a fast-moving startup environment with ambiguity

Preferred

  • Experience with data-intensive or real-time processing systems
  • Background in energy, climate tech, or critical infrastructure
  • Experience scaling infrastructure through hypergrowth
  • On-Prem Kubernetes Deployment Experience
  • Windows Server Administration Experience

WHAT WE OFFER

  • Competitive salary, performance bonus, and equity.
  • Comprehensive health, dental, and vision coverage.
  • Lunch provided three days a week in office.
  • Hybrid schedule for local employees: 3 days in office for collaboration, 2 days remote for focused work.
  • Access to leading academic, industry, and government partners in the AI-energy ecosystem.
  • A mission-driven team focused on shaping the future of the energy transition.

Join us in tackling one of the most important infrastructure challenges of our time — enabling the energy foundation for the age of AI.