View jobs

Staff Cloud Engineer

  • Software Development
  • Full-time
  • San Francisco, CA
  • Remote
  • 210K - 260K USD a year

Staff Cloud Engineer - Earthly Lunar

Write production software, design the cloud architecture behind it, and own how it runs in the real world.

About Earthly Lunar

Earthly Lunar is a guardrails engine for the AI era. Engineering organizations use it to turn their standards for quality, security, and compliance into automated checks across their repositories and CI/CD pipelines. Teams get feedback directly in pull requests, leaders can see where standards aren't being met, and compliance evidence is collected as the work happens. The standards apply whether a developer or an AI agent writes the code.

We're a small, AI-native team backed by Innovation Endeavors and Work-Bench, with angels including Jeff Dean, the CEOs of Datadog, Cockroach Labs, and Dash0, the founder of Sentry, and the creators of Elixir, Pandas, and Ruff/uv. We're building for enterprise engineering teams with complex environments and high expectations for reliability.

The role

We're looking for someone who is an excellent software engineer and an excellent cloud engineer. You can design and build a substantial backend feature in Go, work through the architecture it needs, deploy it safely, and debug it when production behaves differently from the test environment.

You'll work on Lunar's core product and the infrastructure that runs it, across our hosted service and customer-managed deployments. That includes backend services, deployment automation, Kubernetes, and the operational work that makes enterprise software dependable.

This is a staff-level individual contributor role. You'll stay hands-on while helping decide what we build, making architectural calls, and leading substantial projects through to production. We want someone whose judgment improves the whole team's work, not just their own code.

What you'll do

  • Write and maintain production Go code in Lunar's backend and supporting services. Own features from design and testing through deployment and ongoing operation.

  • Design and evolve our AWS architecture and Kubernetes infrastructure. Make deliberate trade-offs around reliability, security, performance, cost, and how much complexity a small team can support.

  • Build reusable Terraform modules, Helm charts, and deployment tooling. Make provisioning, upgrades, and recovery repeatable for both our hosted service and customer-managed installations.

  • Improve production reliability: useful metrics and alerts, capacity planning, backups and tested recovery, safe rollouts, and clear rollback paths. Participate in incident response and fix the causes of recurring problems.

  • Debug across the stack. Follow a failure through Go code, a database query, container behavior, Kubernetes networking, or cloud infrastructure rather than stopping at a team boundary.

  • Improve CI/CD and the development environment so engineers can test realistic changes and ship frequently without making production fragile.

  • Work with our forward deployed engineers and customer platform teams on difficult deployment and scale problems. Turn what we learn into product improvements that help every customer.

  • Lead technical design and code reviews, mentor teammates, and document decisions well enough that other engineers can maintain what you build. Keep solutions proportionate to the problem.

  • Build and manage AI-assisted engineering workflows for implementation, review, testing, and operational investigation. Make useful workflows repeatable for the team, with appropriate access controls and checks on their output.

What we're looking for

You need substantial software engineering experience and deep production infrastructure experience, plus extensive use of AI in your engineering work.

  • Strong Go and software engineering fundamentals. You've designed, shipped, and maintained production services, not only infrastructure scripts. You're comfortable with concurrency, APIs, testing, performance work, and debugging distributed systems.

  • Deep AWS experience. You've made and lived with cloud architecture decisions involving networking, IAM, compute, storage, and managed databases. You understand failure modes, isolation boundaries, and cost.

  • Extensive Kubernetes experience. You've built and operated production clusters and workloads, handled upgrades, and debugged real failures. You understand scheduling, networking, storage, RBAC, resource management, and Helm beyond following an installation guide.

  • Extensive Terraform experience. You've designed reusable modules and managed production infrastructure through code, including state, environment separation, drift, and safe changes to existing resources.

  • Strong operational instincts. You're comfortable with Linux, Docker, networking, and troubleshooting under pressure. You've taken responsibility for systems after launch and made them easier to operate over time.

  • Experience with production data and observability systems. You can operate and troubleshoot PostgreSQL or comparable SQL databases, object storage such as S3, and metrics, logs, and traces using tools such as Prometheus, Grafana, or OpenTelemetry.

  • Extensive experience using and managing AI workflows. AI coding tools and agents are part of your daily work. You build your own workflows, supply the context and tools they need, and manage their permissions, cost, and failure modes. You can explain how you verify generated code and infrastructure changes, and you remain responsible for the result.

  • A staff-level track record. You've taken ambiguous problems, chosen a practical approach, brought other engineers along, and delivered substantial changes that held up in production.

  • Clear communication and high ownership. You write useful design notes, can explain a trade-off to another engineer or a customer, and make progress without waiting for a detailed specification. You know when to ask for help.

Nice to have

  • Experience building developer tools, infrastructure products, or other B2B software used by engineering teams.

  • Experience shipping software into customer-managed Kubernetes environments, including restricted networks and enterprise security requirements.

  • Experience with gRPC and Protocol Buffers, queues and background workers, or distributed job execution.

  • Experience with GitOps and progressive delivery, Kubernetes operators, or multi-region systems.

  • Familiarity with software supply-chain security, policy as code, or compliance automation.

  • Experience as an early startup engineer or founder.

You might not be a fit if

  • You want an architecture or management role with little hands-on coding.

  • You enjoy operating infrastructure but don't want to build substantial production software, or want to write features without owning how they run.

  • You use AI only occasionally and aren't interested in building and managing your own engineering workflows.

  • You need a narrowly defined remit and a separate team to handle every issue outside it.

How we work

  • We're a small, senior team. You'll work directly with the founder and the founding engineers, with room to shape both the product and how we build it.

  • We use AI heavily and invest in the tools that make the team effective. Engineering judgment, testing, and review still matter.

  • We favor frequent, tested releases and straightforward systems over elaborate architecture we don't yet need.

  • Production ownership is shared engineering work. Part of your job is to make it less dependent on any one person.

  • We're remote, with overlapping working hours and written communication that lets people work independently.

Compensation & logistics

  • Base salary: $210,000-$260,000 (US base), depending on experience, plus equity.

  • Location: Remote (US & Canada).

  • Type: Full-time, staff-level individual contributor.

  • Production support: Incident response is part of the role.

Benefits

  • Healthcare, dental, and vision, including dependents (US & Canada)

  • Incentive stock plan

  • 401(k) / RRSP with employer matching (US & Canada)

  • Life insurance (US & Canada)

  • Credits for gym or activity equipment (US & Canada)

  • Fully remote work

  • Direct ownership of substantial engineering work in an early-stage company


Remote restrictions

  • Workday must overlap by at least 5 hours with San Francisco, CA, USA