Senior Site Reliability Engineer, Robotics & Cloud Infrastructure

External

Bedrockocean · Worldwide

Full-timeRemote1d ago

AWSBashDockerGrafanaIAMKubernetes

Cover Letter Connect

Prepare for this interview

Elite

AI-generated questions, company research, and talking points tailored to this role

About the role

We're looking for an SRE who is equally comfortable on the robotics side- compute on the vehicle, topside operator machines, field deployments- and the cloud side: data ingestion, processing pipelines, and our customer-facing platform. You'll build the automation, observability, and operational guardrails that let a small team run continuous AUV operations without continuous heroics, turning manual recovery steps into self-healing systems and shrinking the set of failures that only one person knows how to fix. This is a hands-on senior infrastructure role with a strong automation mandate and a shared on-call rotation. You'll set reliability direction across vehicle-side and cloud-side systems, raise the operational bar for the team, and mentor others toward it. You'll be a force multiplier for reliability across the company, not a ticket queue. Reports to: Head of Software. East Coast location is required to support coverage across both European operations and the East Coast during 12-hour on-call shifts. Travel to field deployments and Richmond HQ is expected (approximately 5-15%).

Responsibilities

Own reliability across the full path from vehicle to customer: AUV onboard compute (Jetson-class modules, ROS 2), topside/operator systems, cloud data pipelines, and the platform that delivers data products.
Build and extend infrastructure automation- provisioning, configuration management, deployment, and self-recovery- so that routine field operations and pipeline runs require minimal manual intervention.
Design and improve observability: metrics, logging, tracing, and alerting that give both robotics and data teams early, actionable signal across vehicle fleets and cloud services.
Drive down on-call burden by identifying and eliminating single points of failure, writing runbooks, and automating the manual steps that currently require tribal knowledge.
Participate in a shared on-call rotation covering both robotics-side and cloud-side incidents in 12-hour shifts spanning European and East Coast business hours; lead and contribute to blameless post-incident reviews.
Define and track reliability targets, availability, data yield, recovery time, tied to continuous-operations goals, and partner with robotics and data teams to meet them.
Manage cloud infrastructure on AWS (compute, storage, networking, IaC, cost, and security posture) for data processing and platform workloads.
Improve fleet- and vehicle-level configuration management, deployment safety, and rollback so changes reach the field reliably and predictably.

Requirements

5+ years in an SRE, DevOps, or infrastructure engineering role running production systems with real uptime and on-call responsibilities, including senior-level ownership of reliability outcomes.
Experience implementing a scalable incident management and operational excellence mechanism that treats operators as customers, building processes and tooling that serve the people running operations day to day, not just the engineering team.
Strong automation instincts: comfortable scripting and building tooling in Python and/or Go and Bash, and using infrastructure-as-code (Terraform or equivalent).
Hands-on AWS experience across compute, storage, networking, and IAM, plus containerization and orchestration (Docker, Kubernetes or similar).
Working knowledge of Linux internals, networking, and observability tooling (Prometheus/Grafana or equivalents).
Comfort operating across environments that aren't just cloud: embedded or edge compute, intermittent connectivity, and physical systems that fail in messy ways.
A reliability mindset: you instrument before you guess, you automate the second time you do something manually, and you write things down so the next

Benefits

Vision insurance

Additional Information

About Bedrock Ocean Exploration Bedrock Ocean builds and operates autonomous underwater vehicles (AUVs) that collect georeferenced ocean-floor data at commercial scale. We deliver bathymetric and imagery data products to customers through our own platform, and we're scaling toward continuous, around-the-clock data collection campaigns spanning months at a time. Keeping vehicles in the water and data flowing reliably is a core engineering problem and this role owns the reliability of the systems on both ends of that pipeline. Headquartered in Richmond, California, Bedrock Ocean Exploration is building autonomous ocean intelligence that will enable the ocean economy to solve the world's most pressing challenges in maritime security, infrastructure, energy, and climate. Our modular architecture, driven by Siren (autonomous underwater vehicles), Trident (command and control), and Mosaic (subsea data fusion), delivers entirely new intelligence capabilities for government and commercial partners. Missions mobilize from any vessel of opportunity in 24 to 72 hours, and our automated pipeline returns comprehensive insights in hours, not weeks like the incumbents, keeping crews safe on shore while cutting cost and time.

Your Match

How well this role fits your profile.

Company Intel

What employees say

Worked at bedrockocean? Share your experience

Interested in this role?

Apply on the company's website.

Cover Letter Connect