Design and scale high-performance backend systems for real-time data processing and critical API services.
Work across Go and Node.js to build reliable, scalable infrastructure.
Take ownership of complex technical problems from investigation to deployment.
The company builds technology to help leading organizations detect and prevent online fraud. It operates with a fully remote, asynchronous, and collaborative engineering culture.
Design and evolve scalable cloud infrastructure on Google Cloud Platform, focusing on reliability and automation.
Strengthen observability platform with metrics, logging, and tracing to improve incident response and reduce recovery time.
Champion reliability practices like SLOs, error budgets, and DORA metrics to drive operational excellence.
They operate at the intersection of geospatial intelligence and environmental technology. They are a growing organization with a collaborative, high-impact engineering culture.
Build and operate streaming and polling infrastructure for feature flag delivery.
Own projects end-to-end from design to production, including debugging and on-call.
Collaborate with cross-functional teams and improve observability and reliability.
LaunchDarkly provides a feature management platform that enables developers to control software releases and target features to specific user segments. The company fosters a collaborative, inclusive culture and values diversity among its team.
You'll contribute to infrastructure scaling to infinitely many apps, improving performance and reliability across backend services.
You'll support observability efforts, help implement SLOs, and build foundational services for next-generation cloud infrastructure.
You'll participate in triage and on-call processes to diagnose issues and implement changes to prevent recurrence.
Bubble is an AI visual development platform that empowers anyone to create software without code, from first-time entrepreneurs to enterprise teams. With over 6 million users in more than 100 countries and a mission to break down barriers to entrepreneurship, the company fosters a collaborative and inclusive culture focused on empowering builders worldwide.
Design, build, and maintain scalable distributed backend systems powering platform and product capabilities.
Own backend services end-to-end, from architecture to monitoring, while collaborating with product and engineering peers.
Build and evolve Web3 platform backend components, including protocol integrations and blockchain network support.
Zerion builds backend systems for Web3 infrastructure, powering API and wallet products trusted by 50+ top crypto brands. They have a fully remote team with strong ownership and high code quality, processing over 3.1B requests and supporting 1.7M+ active funded wallets.
Design and evolve cloud infrastructure on GCP for scale and resilience.
Build internal tooling and automation that promote team autonomy and developer productivity.
Advance observability platform with metrics, logging, tracing, and alerting to reduce recovery time.
The company is a well-funded AI/ML company at the intersection of geospatial intelligence and climate technology, building products on scalable cloud infrastructure. The engineering team fosters a culture of reliability and continuous improvement, operating with a focus on SLOs, error budgets, and DORA metrics.
Design, build, and operate the distributed systems that deliver feature-flag configuration to LaunchDarkly SDKs.
Own and improve the reliability, latency, and scalability of our streaming and polling infrastructure.
Debug and resolve complex production issues while sharing the team's on-call rotation.
LaunchDarkly builds a feature management platform that enables safe and gradual software releases. The company is growing and fosters a humble, open, collaborative culture.
Design and build core backend services for context ingestion, indexing, retrieval APIs, and agent-facing integrations.
Create a scalable multi-tenant SaaS foundation with tenant isolation, usage tracking, and reliable service boundaries.
Work across product and infrastructure to make practical tradeoffs between fast experimentation and long-term reliability.
Grafana Labs is the company behind Grafana Cloud, a fully managed observability platform trusted by over 10,000 organizations. We are a 100% remote company with team members across 40+ countries, backed by leading investors, and we value an open-source legacy, global collaboration, and transparency.
Design and implement a scalable observability platform for Whatnot's growing infrastructure.
Work with core infrastructure, platform, and developer tools teams to redesign data collection to visualization.
Utilize AI agents and open standards to ensure visibility into software stack performance and reliability.
Whatnot is the largest live shopping platform in North America and Europe, enabling sellers to build businesses across hundreds of categories. They are a remote co-located team anchored in hubs across the US, UK, Ireland, Poland, Germany, and Australia, and were recently named the #1 Best Startup Employer in America by Forbes.
Set reliability strategy and SLO culture that scales across engineering teams.
Own platform architecture, event-driven messaging, and observability for a global payments platform.
Lead chaos engineering, incident response, and mentorship for the most complex production challenges.
Yuno is an AI-native operating system for global commerce, connecting merchants to pay-ins, payouts, fraud prevention, and stablecoins via a single API. It powers payment infrastructure for global brands like McDonald's and GoFundMe, with a culture of remote work and AI-driven innovation.
Build the world's fastest APIs delivering globally resilient access to blockchain nodes.
Implement customer management, control, and billing systems.
Scale systems for data ingestion, storage, and query capabilities.
Validation Cloud is an AI platform powering Web3 finance, delivering products across Data x AI, Staking, and Node API, trusted by billions in staked assets. Backed by over $20M in venture funding, it has a world-class team spanning San Francisco, New York, London, and beyond.
Design, build, and maintain robust, scalable, and secure infrastructure systems supporting Laurel's AI-driven platform.
Manage and optimize cloud infrastructure (AWS and Azure), Kubernetes orchestration, and CI/CD pipelines to increase deployment frequency and reliability.
Implement comprehensive observability, monitoring, and alerting to maintain system health and partner with engineering teams to optimize performance and cost-efficiency.
Laurel is an AI Time platform for professional services firms, automating work time capture and connecting time data to business outcomes for clients like EY and Crowell & Moring. The company comprises top AI, product, and engineering talent, is VC-backed by Google Ventures and IVP, and fosters an inclusive, ambitious culture.
Build and operate monitoring, tracing, alerting, and observability infrastructure for system reliability.
Drive platform security initiatives with preventative controls and resilient architecture.
Lead incident response and recovery, including root-cause analysis and preventative measures.
This role is with a partner company managing AI-powered products. They are a growing technology organization with a fully distributed US-based team and a collaborative culture focused on large-scale infrastructure and AI technology.
Lead the discovery, design, and delivery of complex reliability and infrastructure initiatives, translating ambiguous problems into robust technical solutions.
Define and operate reliability practices including SLOs, SLIs, error budgets, alerting strategies, and observability standards.
Manage and scale production Kubernetes environments, build cloud infrastructure on AWS, and mentor less-senior engineers.
Jobgether is an AI-powered recruitment platform that connects candidates with partner companies through automated matching. They are a globally distributed organization focused on fair, efficient hiring processes.
Build and run monitoring, tracing, and alerting infrastructure to ensure platform reliability and security.
Lead incident response and recovery, including root cause analysis, and improve deployment processes for fast, safe code changes.
Collaborate with engineering teams to deliver a stable, scalable platform and handle load for resource-intensive applications.
WellSaid Labs is the leading AI voiceover studio for enterprise and professional use, providing ultra-realistic voices that the world’s biggest brands trust. We are a fully distributed team across the U.S. with a focus on responsible AI and an inclusive culture.
Ship reliable features at scale that deliver the most value to customers by partnering with engineering and product.
Perform thorough and thoughtful code reviews to maintain a high standard of code quality.
Identify key system metrics and ensure adequate monitoring coverage for new and existing data storage and analytical services.
Twilio powers real-time business communications and data solutions that help companies and developers worldwide build better applications and customer experiences. They employ thousands of Twilions worldwide and foster a remote-first, antiracist culture that supports diversity, equity, and inclusion.
Architect and build a robust, scalable, and highly available distributed infrastructure.
Build a cutting-edge cloud-native platform on top of the public cloud and automate cloud resource management.
Work closely with core database development and security teams to produce the SaaS offering.
ClickHouse is a real-time analytics and data warehousing company recognized on the Forbes Cloud 100 list. With over 4,000 customers and rapid growth, the company is a leader in its space.
Design, build, and operate event-driven services for real-time pentest data processing.
Solve distributed systems challenges at scale, including data migration and batch-to-streaming conversion.
Set high standards for operational excellence with SLOs, instrumentation, and on-call.
Horizon3 is a remote cybersecurity company that enables organizations to find, fix, and verify exploitable attack vectors using its NodeZero autonomous pentesting platform. The company is a fusion of former US Special Operations cyber operators and engineers, fostering a culture of respect, collaboration, and ownership.
Own and scale cloud infrastructure including compute, networking, storage, and data systems.
Lead BYOC and private cloud deployments with infrastructure-as-code and GitOps foundations.
Establish reliability through service-level objectives, observability, and incident response processes.
A technology company builds a developer-focused platform with scalable cloud infrastructure. This is a remote-first opportunity with a small, autonomous engineering team operating in North America, LATAM, and Europe, offering high autonomy and ownership.