Site Reliability Engineer
ExternalPrepare for this interview
EliteAI-generated questions, company research, and talking points tailored to this role
Responsibilities
- Design and improve systems with resilience and graceful degradation in mind. Plan for capacity and possible failure modes.
- Define and measure SLOs and SLIs that reflect customer experience and help teams make better reliability tradeoffs.
- Use observability tools such as Datadog, CloudWatch, logs, metrics, traces, and APM. Build signal-heavy, noise-light visibility into production systems.
- Configure and improve alerting and routing through incident management workflows. Make sure pages are actionable, well-routed, and worth human attention.
- Participate in incident response from detection and triage through communication, resolution, postmortems, and follow-up.
- Continuously improve the incident lifecycle. Focus on better detection, clearer runbooks, stronger postmortems, and concrete remediations.
- Construct or optimize infrastructure, reliability tooling, and automation that eliminate toil and ensure operational consistency.
- Use AI-assisted tools to accelerate coding and documentation. Speed up root-cause exploration, runbook improvement, infrastructure-as-code workflows, and operational tasks.
- Help engineering teams improve production readiness and deployment safety. Support service ownership and operational clarity.
- Communicate reliability concepts clearly across technical and non-technical teams.
- Document operational knowledge to reduce silos. Make it easier for engineers to respond with confidence.
- Contribute to a culture where reliability is shared by SRE and product engineering teams.
Requirements
- Bachelor's or master's degree in Computer Science, Engineering, or a related field, or equivalent industry experience.
- 3+ years of experience in SRE, Software Engineering, Infrastructure Engineering, or a related role.
- Hands-on coding experience in Python, Go, or similar production-oriented programming languages.
- Experience operating production systems and contributing to reliability, observability, incident response, infrastructure, or automation improvements.
- Working knowledge of SLIs, SLOs, error budgets, MTTR, and how reliability d
Benefits
Additional Information
About EarnIn As one of the first pioneers of earned wage access, our passion at EarnIn is building products that deliver real-time financial flexibility for those with the unique needs of living paycheck to paycheck. Our community members access their earnings as they earn them, with options to spend, save, and grow their money without mandatory fees, interest rates, or credit checks. We're fortunate to have an incredibly experienced leadership team, combined with world-class funding partners like A16Z, Matrix Partners, DST, Ribbit Capital, and a very healthy core business with a tremendous runway. We're growing fast and are excited to continue bringing world-class talent onboard to help shape the next chapter of our growth journey. POSITON SUMMARY EarnIn's community members rely on our products to deliver reliability and trust when they need them most. Reliability shapes the product experience, not simply operational concerns. Every noisy alert, unclear runbook, fragile deployment, or repeated incident undermines customer trust and hinders engineering teams. This role enables EarnIn to build and run production systems with greater resilience, clarity, and confidence. As a Site Reliability Engineer II, you will strengthen infrastructure, optimize tooling, deepen observability, streamline incident response, and elevate reliability standards. These actions empower teams to ship quickly and safely. This position will be hybrid, based in our Bengaluru office, with 2 days a week required in the office, as part of our expanding site. EarnIn provides excellent employee benefits, including healthcare, internet/cell phone reimbursement, a learning and development stipend, and opportunities to collaborate with and travel to our Palo Alto HQ and Bangkok Site. Our salary ranges are determined by role, level, and location. We are unable to provide visa sponsorship or immigration support for this position. IMPACT You will operate as a well-rounded SRE practitioner across production operations, observability, incident response, infrastructure-as-code, automation, and software engineering. You will demonstrate growing independence in reliability work. You will not only follow existing playbooks, but also refine them. You will transform production learnings into better alerts, clearer runbooks, safer deployments, stronger observability, and more reliable services. You will harness AI-assisted development and operational workflows to minimize toil, accelerate investigation, enhance documentation, and streamline infrastructure and reliability work. You will meticulously validate AI-generated output before applying it to production systems or operational workflows. You will collaborate with product engineering and platform teams to implement, explain, and support reliability practices, ensuring they are practical, understandable, and actionable.
Your Match
How well this role fits your profile.
Company Intel
What employees say
Worked at earnin? Share your experience