- Job type
- Full-time
- Work mode
- Remote
- Level
- Principal
- Department
- Software Development
- Experience
- 12+ years experience
- Posted
- Sep 29, 2026
About the role
ABOUT POSITION:
Are you passionate about automation, DevOps, public cloud services, Kubernetes, and observability? As a site reliability engineer at Legion, you will be responsible for building the tools, infrastructure, and services to create a secure, highly scalable, and cost-effective AWS/Kubernetes-based cloud platform. You will be working with everything from infrastructure tooling, automation, build and deployment pipelines, monitoring and logging, containerization, and more! Simply put, the SRE team keeps Legion running for our customers.
ROLES AND RESPONSIBILITIES:
- Support Legion’s public cloud platform, utilizing multiple cloud services and containerization
- Develop infrastructure automation leveraging Terraform, Chef, Jenkins, and Golang.
- Create and manage production alerts. Respond to alerts and conduct a root-cause investigation
- Develop automated operational runbooks
- Support the deployment of Legion’s WFM solution during and off regular office hours
- This role involves participation in on-call rotation
- Build and operate internal AI agents that reduce SRE toil — incident triage, runbook execution, alert enrichment, capacity and cost analysis — and own their guardrails, rollback, and reliability
- Own the platform that lets other engineers build agents safely: sandboxed execution, scoped credentials, tool/MCP integrations, cost controls, and agent observability
BASIC QUALIFICATION:
- 12+ years experience in SRE, DevOps, or other SaaS operations.
- 3+ years of hands-on experience with managing AWS services and cloud infrastructure in AWS utilizing Terraform. AWS Certification preferred
- 3+ years experience with containerized cloud solutions utilizing Docker, Kubernetes, or AWS EKS. Familiarity with HELM charts
- 5+ years experience with at least one programming language like Go, Python, Bash, Perl
- 5+ years experience with one or more Infrastructure automation (Terraform, Ansible, etc.), CI/CD pipelines (GIT, Jenkins, etc.), and configuration management tools (Ansible, Chef, etc).
- Hands-on experience with one or more Linux/Unix platforms like RedHat/CentOS/Ubuntu/Amazon Linux.
- Hands-on experience building agentic AI systems in production — tool integrations, evaluation, and failure handling — beyond use of AI coding assistants
- Bachelors or Masters degree in Computer Science/Engineering or related field
OTHER QUALIFICATION:
- Track record of introducing AI into an engineering org's SDLC with measurable results, and the judgment to name where it should not be used.
- 3+ years of experience with observability tools: Splunk, Nagios, Elasticsearch, Kibana, CloudWatch, and Logstash and ways to scale these systems.
- 3+ years of experience in AWS RDS or Aurora MySQL
- Demonstrated ability to work with remote teams