DoubleTap Consulting builds custom websites, product platforms and internal software for streaming, SaaS and commerce teams, working as a senior-only team with direct access between clients and builders. We're hiring a Site Reliability Engineer to help our engagement teams design, build and hand off infrastructure, observability and release processes that clients can own and run long after we're gone.
What you'll do
- Design, build and maintain cloud infrastructure and CI/CD pipelines across client engagements
- Set up monitoring, logging and alerting systems to give teams real visibility into production health
- Define and track SLIs/SLOs, and lead incident response and postmortems when things break
- Automate repetitive operational work through infrastructure-as-code and tooling
- Partner with engineers on each engagement to harden release processes before handoff
- Document infrastructure and runbooks so client teams can operate systems independently after we leave
What we're looking for
- Solid experience running production infrastructure on a major cloud provider (AWS, GCP, or Azure)
- Hands-on experience with infrastructure-as-code tools such as Terraform or CloudFormation
- Strong scripting or programming ability (Python, Go, or similar) for automation work
- Experience with containerization and orchestration (Docker, Kubernetes)
- Track record of building or maintaining observability stacks (metrics, logging, tracing)
- Comfort working directly with client teams and communicating technical tradeoffs clearly
Nice to have
- Experience with Node.js and Postgres environments
- Familiarity with Next.js or React application deployment patterns
- Background supporting high-traffic consumer or e-commerce platforms
- Prior consulting or client-facing engineering experience