Migaru AI builds an AI-native CRM that automatically files emails, calls, and meetings into the right account, contact, and deal records, keeping the pipeline current without manual data entry. We're looking for a Site Reliability Engineer to help us run a reliable, scalable, and secure platform as it processes a growing volume of customer conversations and data in real time.
What you'll do
- Own the reliability, availability, and performance of production systems, including setting and tracking SLOs/SLIs
- Build and maintain monitoring, alerting, and incident response tooling to detect and resolve issues before customers notice
- Lead or support incident response, write blameless postmortems, and drive follow-up fixes to prevent recurrence
- Design and implement infrastructure-as-code, CI/CD pipelines, and deployment automation
- Partner with engineering teams to improve system architecture for scalability, resilience, and cost efficiency
- Plan and run capacity planning, load testing, and disaster recovery exercises
What we're looking for
- Experience operating production infrastructure in a cloud environment (AWS, GCP, or Azure)
- Strong scripting or programming skills (Python, Go, or similar) for automation and tooling
- Hands-on experience with container orchestration (Kubernetes or similar) and infrastructure-as-code (Terraform or similar)
- Solid understanding of networking, Linux systems, and distributed systems fundamentals
- Experience with observability tooling (metrics, logging, tracing) and on-call incident response
- Clear communication skills and comfort working cross-functionally with product and engineering teams
Nice to have
- Experience operating systems that process sensitive customer data, including security and compliance considerations
- Familiarity with database reliability and performance tuning
- Experience supporting a SaaS product through periods of rapid growth
