Embed with product and platform teams from early stages to ensure reliability is designed in from the start.
Define production-readiness standards and measurable SLIs/SLOs to guide operational excellence.
Build tooling and infrastructure across AWS, GCP, and Azure using Terraform, and share on-call rotation.
We build WebContainers and Bolt.new, an AI-powered app builder that lets you create, edit, and deploy full-stack apps instantly in your browser. We are a fully remote, globally distributed team of passionate engineers serving over 1 million developers monthly.
Design, develop, and maintain scalable backend APIs supporting authentication, payments, subscriptions, analytics, consumption tracking, and text-to-speech services.
Build and optimize reliable backend systems with a strong focus on scalability, performance, security, and maintainability.
Collaborate with cross-functional teams to align backend architecture with product strategy, customer needs, and user experience goals.
Our partner is a technology company building backend infrastructure that powers products used by millions worldwide, focusing on accessibility and learning experiences. They are a fast-growing global company with a collaborative engineering culture and a fully distributed team.
Build enterprise-scale infrastructure using infrastructure-as-code and Kubernetes-native systems.
Sustain platform health and performance by owning critical systems in production.
Enable teams and customers to move faster with abstractions and tooling for AI/ML workloads.
Cake makes cutting-edge AI accessible to enterprise teams by removing infrastructure barriers, enabling 10x faster and cheaper AI/ML platform deployment. Backed by top investors, they have a small senior team focused on ownership and operational excellence.
Own and operate customer-facing managed infrastructure across multiple AWS accounts and regions.
Serve as the senior technical escalation point for production incidents and complex configurations.
Contribute to OpenTelemetry distributions and maintain open source projects like Refinery.
Honeycomb provides observability for developer tools, helping companies like HelloFresh and Slack understand their software. They have over 200 employees and were named to Forbes' Best Startups in 2022 and 2023, with a culture that values inclusion and autonomy.
Work with a team of DevOps and DBA professionals to improve infrastructure and streamline deployments across countries.
Continuously improve Kubernetes platform stability, efficiency, and GitOps-first environment provisioning.
Monitor cloud infrastructure, own on-call operations, and define SLIs/SLOs for reliability improvements.
Sporty Group is a remote-first company focused on sustainability in the sports and gaming industry. They maintain a competitive, performance-driven culture with a distributed team across EMEA.
Design and operate the infrastructure for a high-throughput messaging platform operating at 500K+ events/sec.
Build guardrails, runbooks, and validation gates that enable AI agents to safely execute deployments and operations.
Lead incident response and encode every fix as a new runbook and regression test.
Postscript is an AI messaging platform trusted by 20,000+ Shopify brands to drive revenue through SMS. The company is fully remote, backed by Greylock and Y Combinator, and has a culture of ownership and innovation.
Build and manage AWS infrastructure, CI/CD pipelines, and ensure reliability, security, and cost optimization without an infrastructure team above you.
Work directly with stakeholders to shape architecture and product direction, shipping fast to learn or slowing down to fix the foundation.
Adopt the latest AI improvements to speed up how we build and ship, while collaborating with a remote-first team that meets regularly.
Cello is an AI-powered referral infrastructure company that helps SaaS companies turn users into a sales channel through sharing and recommendations, with APIs and widgets used by 10M+ end users monthly. Backed by $10.3M in funding from top-tier investors, the small team of serial founders and big-tech operators from Twilio, Wise, and Skype values active learning and innovation.
Design and evolve mission-critical APIs for a modern developer platform.
Build secure, scalable backend services using TypeScript, Node.js, and modern frameworks.
Collaborate with engineers across teams to deliver reliable systems that improve developer experience.
Jobgether is a platform that uses AI-powered matching to connect candidates with job opportunities. They focus on efficient, objective hiring processes and are part of a network of partner companies.
End-to-end ownership of internal orchestration platform built on event-driven architecture with Redpanda, including code, architecture, and roadmap.
Own infrastructure-as-code using Terraform Cloud, manage Kubernetes workloads with Helm, and provide self-service tooling for engineering teams.
Set SLOs, handle production on-call, lead incident response, author design docs, and operate AI-natively using tools like Cursor and Notion AI.
Velora unifies Aplos, Raisely, and Keela into one company with a shared mission to help nonprofit organizations thrive by offering fundraising, donor management, financial tracking, and communications tools. We are a financially solid company with a combined team dedicated to making nonprofit work easier, more impactful, and more sustainable.
Keep user-facing services and production systems reliable, scalable, and efficient with automation and infrastructure-as-code.
Operate and troubleshoot production systems on Kubernetes, and contribute to observability with metrics, logs, and SLOs.
Participate in on-call, incident response, and post-incident reviews to drive improvements in automation and processes.
GitLab is the intelligent orchestration platform for DevSecOps, enabling organizations to increase developer productivity and improve operational efficiency. With more than 50 million registered users and over 50% of the Fortune 100 as customers, GitLab fosters a high-performance, all-remote culture driven by values and continuous knowledge exchange.
Design, implement, and maintain highly available and scalable infrastructure solutions.
Monitor system performance, identify bottlenecks, and resolve reliability issues proactively.
Automate infrastructure deployment, configuration management, and operational workflows.
The company is a technology firm that provides critical authorization solutions to organizations worldwide. It is a remote-first organization with a collaborative culture, offering equity opportunities and a focus on team building.
Design and maintain scalable infrastructure-as-code solutions using Terraform and Kubernetes.
Build and operate observability systems while leading incident response and reliability improvements.
Embed security and compliance practices into infrastructure and optimize system performance and cloud costs.
This partner company builds a next-generation platform enabling AI-driven services across global employment infrastructure. It is a highly distributed, async-first organization where engineers thrive in ownership and autonomy.
Design, build, and run distributed cloud architectures and large-scale production systems.
Ensure reliability, observability, performance, and cost efficiency of the platform.
Collaborate with product and backend teams to design system architecture and optimize resource use.
Tinybird helps developers and data teams unlock the power of real-time data, enabling them to build data pipelines and innovative data products quickly. They are a remote-first company with a culture of ownership, transparency, and clear communication.
Design and scale highly reliable platform systems supporting complex cloud-native workloads across multiple deployment environments.
Build and enhance core platform services while contributing to distributed systems, event-driven architectures, and cloud-native infrastructure.
Optimize cloud resources, networking, storage, compute, and observability to improve system performance, scalability, reliability, and maintainability.
Jobgether uses an AI-powered matching process to connect candidates with hiring companies. They operate as a job platform, processing applications and sharing top candidates with employers.
Build and improve platform services, including CI/CD pipelines and cloud infrastructure.
Collaborate with senior engineers to design scalable solutions and enhance developer experience.
Participate in incident response and retrospectives to drive continuous improvement.
Octopus Energy is a tech-powered energy company focused on renewable energy and customer experience. The company culture emphasizes ownership, collaboration, and making a tangible impact across teams.
Own and evolve our SLI/SLO and error-budget frameworks, using them to influence prioritization and product decisions.
Lead incident response, drive postmortems, and turn findings into systemic fixes rather than one-off patches.
Build and maintain observability across metrics, logs, and traces (Datadog), improving signal and reducing alert fatigue.
MyFitnessPal provides tools, resources and support to enable users to reach their health goals. The company values collaboration, mentorship, and inclusive environments, with a focus on reliability and delivery.
Lead the design, development and operation of large-scale, secure observability systems to keep services online and performant.
Deploy and scale Prometheus, ElasticSearch clusters, and high-throughput Kafka data pipelines for millions of customer devices.
Collaborate with the Observability team to build alerting systems, APIs, and self-service monitoring tools using Terraform and multiple languages.
ItD is a new generation consulting and software development company that blends diversity, innovation, and integrity with real business results. It is a woman- and minority-led firm with a global community, empowering employees and offering benefits like medical, dental, vision, 401(k), and career development.
Partner with product engineering squads to own production reliability for high-SLA customer environments, designing automation and defining per-tenant SLOs.
Serve as a primary escalation point for incidents, leading response, post-incident reviews, and reducing SLO burn to prevent repeats.
Influence feature design for scalability and operability, improve alert quality, and eliminate toil through automation.
Grafana Labs is the company behind the open observability cloud, providing a fully managed observability platform for organizations to see, understand, and act on their data. With over 35 million users, 7,000+ customers, and 1,600+ team members across 40+ countries, we foster a remote, collaborative culture rooted in open-source values.
Architect and manage scalable cloud infrastructure for the 3D rendering platform.
Build and maintain CI/CD pipelines with a focus on deployment reliability.
Monitor performance metrics and proactively address bottlenecks before they become incidents.
Homekynd builds the spatial intelligence layer for enterprise retail, transforming photos into 3D room models for immersive furniture visualization. They are a remote-first team on a fast build timeline, seeking engineers who want real ownership over hard problems.
Drive the definition and adoption of SLIs and SLOs across services, reducing toil through automation and incident response.
Design and architect Infrastructure as Code solutions for large-scale environments using Docker, Kubernetes, and cloud-native services.
Serve as primary SRE liaison for development teams, influencing architecture and conducting training for clients.
Noctua Technology, LLC is a company that drives digital transformation by treating operations as a software engineering challenge, focusing on cloud native systems. They are a dynamic team seeking a Senior SRE to define strategy and bridge development and operations for clients.