LandEarly is looking for a Site Reliability Engineer to help ensure our platform stays fast, available, and resilient as it scales. LandEarly is an auto-apply tool that helps job seekers find fresh job postings, tailor their applications, and submit them quickly, and our infrastructure needs to reliably support a large and growing base of users interacting with the product across web, mobile, and browser extension surfaces. You'll work closely with engineering to build the monitoring, automation, and operational practices that keep our systems running smoothly.
What you'll do
- Design, build, and maintain the monitoring, alerting, and observability systems that give the team visibility into production health
- Improve system reliability, availability, and performance through proactive engineering and capacity planning
- Automate manual operational work, including deployment pipelines, infrastructure provisioning, and incident response tooling
- Participate in an on-call rotation, responding to incidents, driving them to resolution, and leading blameless postmortems
- Partner with software engineers to embed reliability, scalability, and security best practices into the development lifecycle
- Identify and address performance bottlenecks and single points of failure across services and infrastructure
What we're looking for
- Experience operating and supporting production systems in a cloud environment (AWS, GCP, or Azure)
- Strong scripting or programming skills (e.g., Python, Go, or similar) for building automation and tooling
- Hands-on experience with infrastructure-as-code tools such as Terraform, and configuration management or orchestration tools such as Kubernetes
- Solid understanding of monitoring and observability practices, including metrics, logging, and distributed tracing
- Experience with CI/CD pipelines and deployment automation
- Comfort working in a fast-moving environment and participating in on-call responsibilities
Nice to have
- Experience supporting high-traffic consumer-facing web or mobile applications
- Familiarity with database reliability and performance tuning
- Experience with incident management processes and tooling
