+ Post Job +
Home DevOps & Cloud

Entry Level Remote Site Reliability Engineer

📍 Anywhere 🏷️ DevOps & Cloud 💰 $116,000 / year
This is a full-time, fully remote Site Reliability Engineer opening paying $116,000 USD per year, with the location listed simply as Anywhere.

The role

The team is looking for someone early in their SRE career who's already had some hands-on time keeping a production system alive. It's a lighter entry point than many remote site reliability engineer jobs, most of which require several years of production ownership before they'll even glance at a resume. It isn't one of the site reliability engineer jobs with no-experience postings you'll sometimes see advertised, since 13 months of relevant experience is still required, but it's built for someone at the start of that path rather than someone with a decade of PagerDuty experience behind them.

Responsibilities

  • Monitor uptime and performance across production services, catching problems early enough to act before they turn into outages.
  • Build and maintain automation for repetitive operational work, replacing manual fixes with scripts and tooling where it actually saves time.
  • Lead incident response when something breaks, working the problem until service is restored and downtime is kept to a minimum.
Capacity planning also falls under this seat. That means keeping an eye on where systems are trending toward their limits, not just reacting once they've already been hit. When a service starts creeping toward a resource ceiling, whoever holds this role is expected to flag it, size the fix, and push it through before it becomes a Friday-night page. A fair amount of the job is documentation that other engineers actually use later: runbooks that are followed during a real incident, dashboards that are checked before someone assumes something is broken, and postmortems that name what happened without turning into a blame exercise. None of this shows up as a bullet point on most postings, but it's a big part of how the work actually fills a day.

Qualifications

A bachelor's degree is the baseline education requirement here, typically in computer science or a closely related technical field. Beyond the degree, the role calls for 13 months of hands-on experience supporting production systems, along with solid scripting and automation skills; candidates who've mostly done manual, ticket-driven support work with little automation exposure will likely find the ramp-up steep. Strong fundamentals in Linux systems and basic networking are assumed, even though neither is called out as a separate line item. Thirteen months is short by SRE standards, and the team knows that going in. What they're checking for isn't years on a badge, but whether a candidate has already worked on a live production incident and can talk through what they did, in what order, and why. A candidate who spent that time on a help desk, resetting passwords and closing tickets, is a different applicant than one who spent it on-call for an actual service, even if both resumes say the same job title. The interview process is built around drawing out that difference. Degree requirements are treated as a floor, not a filter for prestige. A computer science degree from a smaller or less well-known school carries the same weight here as one from a name-brand program, and a related engineering degree paired with real production experience will also clear the bar.

Skills

  • System reliability and highly-available architecture
  • Monitoring and alerting tooling
  • Incident response
  • Automation and scripting
  • Cloud platform experience
  • Capacity planning
None of this requires deep specialization yet. Familiarity with infrastructure-as-code tools such as Terraform isn't required, but it will make a candidate stand out; the same goes for prior exposure to AWS, Azure, or GCP specifically rather than to cloud platforms in the abstract. A working knowledge of Python or Go for writing operational scripts is more useful day-to-day than fluency in a broader language, since most of the automation work here is small, targeted tooling rather than application development. On the monitoring side, exposure to tools like Prometheus, Grafana, or Datadog will shorten the learning curve considerably, though the team is realistic about training someone up on their specific stack. What matters more going in is understanding what a good alert looks like versus a noisy one, since a junior engineer who pages the whole team over something that could have waited until morning burns trust fast.

Compensation and benefits

At $116,000 per year, the pay falls comfortably within typical remote site reliability engineer salary ranges reported across the industry for roles at this experience level. On top of base pay, the package includes:
  • Health coverage
  • Paid time off
  • 401(k) matching
  • On-call compensation for after-hours incident work
  • Full remote flexibility, with no requirement to be near an office
  • Regular performance reviews tied to merit-based pay increases

How this team works

Naukri Mitra is handling the hiring process for this opening, though the day-to-day reporting line runs through the SRE function itself, which is small and fully distributed. There's no large NOC absorbing the first response to every incident, so ownership over the systems in this seat comes fairly quickly rather than after years of proving yourself. Because the team isn't tied to one office or region, this sits among site reliability engineer jobs worldwide that expect coverage across time zones rather than a fixed nine-to-five block. On-call rotations are built around that reality, and schedules are set with enough notice to avoid blindsiding anyone. Junior engineers on this kind of team usually spend their first stretch shadowing incidents and building out monitoring coverage before they're handed a rotation of their own. That's a fairly standard on-ramp for this field, and it's part of why the experience bar here is set lower than on more senior listings. There's also a reasonably clear growth path attached to this seat. Engineers who stick around a year or two typically move toward owning a specific service area outright, and from there toward more senior SRE or platform roles, either on this team or elsewhere once they've built up a track record of handling real production incidents on their own. The team is small enough that someone doing good work gets noticed quickly, for better or worse.

How to apply

Applications should include a resume and a short note on the largest production incident you've personally worked on, what broke, and what you did about it. Interviews typically run in two rounds: a technical conversation covering monitoring, automation, and incident handling, followed by a shorter call with the hiring manager about schedule and on-call expectations. Given the remote and worldwide nature of the role, expect at least one interview to be scheduled outside standard US business hours to accommodate time zone differences on the team.
Apply Now