Build foundational systems across PostgreSQL, GCP, and internal tooling.
Develop shared infrastructure for agentic systems, including frameworks, evals, and harnesses.
Collaborate with engineering and non-engineering teams to improve developer velocity and scalability.
Ambrook rebuilds financial infrastructure for American family-run businesses, modernizing accounting, banking, and spending. We're a Series A startup backed by Thrive Capital, with an early team focused on empowering the stewards of land and labor.
Own database reliability and drive high availability and disaster recovery strategies.
Architect for scale, including partitioning, sharding, and capacity forecasting.
Optimize performance through deep-dive query optimization and proactive bottleneck remediation.
Finom is a European tech startup developing an all-in-one financial B2B platform integrating banking, accounting, and invoicing. With over $346 million in total funding, the company nurtures an innovative and inspiring work environment where bold ideas thrive.
Act as an SME in production database troubleshooting, ServiceNow instance performance, and Linux performance analysis.
Own complex incident root cause analysis across multi-tenant PostgreSQL/MariaDB fleets at scale.
Develop observability tooling and automation to prevent recurrence of critical issues.
ServiceNow provides an AI platform for business reinvention, helping 85% of the Fortune 500 work smarter. The company fosters an AI-native culture and is growing, with a focus on technology and talent working together.
You'll operate production day-to-day, including oncall, incident response, and postmortems.
You'll own reliability practice by defining SLIs/SLOs and error budgets.
You'll ship infrastructure through code in a GitOps workflow for cloud and Kubernetes.
Alpaca is a global leader in agent-first brokerage infrastructure for stocks, ETFs, options, crypto, and more. With a team of 400+ globally distributed members, we foster a culture of curiosity, empathy, and accountability.
Automate the full lifecycle of multi-terabyte databases like MongoDB, PostgreSQL, and Elasticsearch, from provisioning to retirement.
Drive static credentials out of the system, moving toward short-lived, identity-based auth across services and infrastructure.
Harden multi-region disaster recovery and own the telemetry and CI/CD backbone that the entire company runs on.
Close builds an AI-powered, communication-first CRM for small, scaling businesses, designed to eliminate busywork and boost sales. They are a bootstrapped, profitable 120-person, fully remote team that values transparency and long-term thinking.
Collaborate with engineering teams to design scalable, secure systems.
Establish SLOs, manage incident response, and drive reliability improvements.
Leverage expertise in Go, Python, Kubernetes, and cloud platforms.
ClickHouse is a leading real-time analytics company recognized on the 2025 Forbes Cloud 100 list. With over 3,000 customers and rapid growth, the company offers a remote-friendly, globally distributed culture.
Build and improve platform services, including CI/CD pipelines and cloud infrastructure.
Collaborate with senior engineers to design scalable solutions and enhance developer experience.
Participate in incident response and retrospectives to drive continuous improvement.
Octopus Energy is a tech-powered energy company focused on renewable energy and customer experience. The company culture emphasizes ownership, collaboration, and making a tangible impact across teams.
Design and develop a highly available, scalable, and secure ClickHouse Cloud platform for regulated environments.
Build innovative deployment automation across cloud, hybrid, and on-prem systems, including disconnected environments.
Collaborate with Security, Dataplane, and Infrastructure teams to ensure compliance with NIST and FedRAMP frameworks.
ClickHouse is a fast-growing private cloud company offering real-time analytics, data warehousing, observability, and AI workloads. With over 3,000 customers and a $400M Series D, the company fosters a culture of innovation and collaboration.
Investigate and resolve complex technical issues involving databases, cloud environments, and distributed systems.
Provide consultative guidance to users, optimizing workloads and preventing recurring issues.
Mentor junior support team members and contribute to improving support processes and technical standards.
Our partner is a technology company that provides database and backend solutions to developers worldwide. The company values remote work, autonomy, and a collaborative culture, with a globally distributed team.
Run and evolve the Kubernetes landscape (Amazon EKS, on-prem via Rancher) for all deployments.
Automate deployment and scaling, and build observability to spot problems before users notice.
Design abstractions that let backend and data engineers ship without opening tickets.
Yazio is a nutrition app company that helps millions of users in over 150 countries lead healthier lives through diet tracking. The platform engineering team is small and senior, with a remote-first culture that values efficiency and work-life balance.
Design, build, and run distributed cloud architectures and large-scale production systems.
Ensure reliability, observability, performance, and cost efficiency of the platform.
Collaborate with product and backend teams to design system architecture and optimize resource use.
Tinybird helps developers and data teams unlock the power of real-time data, enabling them to build data pipelines and innovative data products quickly. They are a remote-first company with a culture of ownership, transparency, and clear communication.
Own the administration, maintenance, patching, and upgrades of production database environments while ensuring high performance and reliability.
Lead database optimization initiatives, including performance tuning, indexing strategies, query optimization, and architecture improvements for high-throughput systems.
Design and maintain backup, recovery, replication, and disaster recovery strategies to protect critical data and ensure business continuity.
Jobgether is a platform that uses AI-powered matching to connect candidates with hiring companies. They are a technology-driven company focused on innovation and reliability in recruitment.
Partner closely with product engineering squads to own production reliability for high-SLA customer environments.
Design and implement automation to scale reliability practices and ensure customers meet SLO targets.
Lead customer-impacting incident response and post-incident reviews, contributing to design docs and code reviews.
Grafana Labs, the company behind the open observability cloud, is founded on open source principles and offers a fully managed observability platform with actually useful AI. Today, more than 35 million users and 7,000+ customers trust Grafana Labs, and we are a 100% remote company with 1,600+ team members across 40+ countries.
As a Staff Backend Engineer, you will shape the future of large-scale software delivery and deployment for self-managed environments.
You will design reliable backend systems, automation frameworks, and Kubernetes-based solutions to improve operational resilience.
Collaborate with engineering, reliability, security, and product teams to define scalable solutions and technical strategies.
GitLab is a leading DevOps platform providing a complete CI/CD and software development lifecycle tool. The company is open-core with a strong remote culture and globally distributed teams.
Drive rapid delivery and iteration for the Core Services team of DoiT Cloud Intelligence.
Translate VP-level strategy into a sequenced backlog of problems, user stories, and acceptance criteria.
Leverage hands-on DevOps experience to understand cloud workflows and prioritize enhancements that improve speed, reliability, and simplicity.
DoiT is a global technology company that helps cloud-driven organizations leverage the cloud for business growth and innovation. They are an award-winning strategic partner of AWS, Google Cloud, and Microsoft Azure, working with over 4,000 customers worldwide, and pride themselves on a remote-first culture with flexibility and professional development.
Own the infrastructure and platform powering the marketplace, focusing on reliability, observability, security, and automation.
Manage production AWS and EKS clusters, infrastructure as code with Terraform and GitOps, and CI/CD pipelines via GitHub Actions.
Build automation and internal tooling in Python, Bash, Go, and Node.js/TypeScript, and operate PostgreSQL, MongoDB, and Temporal.
Office Hours is an on-demand expert network that connects leading organizations with trusted experts across various knowledge domains. The company is hyper-growth, profitable, and expanding quickly, backed by top marketplace investors.
Break down larger projects into individual tasks and deliver them in multiple phases.
Support peers and stakeholders in the product development lifecycle by collaborating with product management, design & analytics.
Support the operations and availability of your team's artifacts by creating and monitoring metrics.
Affirm is reinventing credit to make it more honest and friendly, offering buy now, pay later without hidden fees or compounding interest. They are a remote-first company with a global engineering team and a culture that values people first.
Provide advanced technical support to enterprise customers across multiple channels.
Troubleshoot complex software, integration, and infrastructure issues, collaborating with engineering teams.
Act as a trusted technical advisor, helping customers adopt and optimize platform capabilities and influencing product roadmap.
They help enterprise customers maximize the value of modern data and AI platforms. They are a collaborative team focused on innovation, continuous learning, and customer success.
Design, build, and scale reliable infrastructure for Klover's fintech platform using modern technologies like Kubernetes, Terraform, and Istio.
Use AI agents as force multipliers to automate manual processes and improve developer experience.
Collaborate with engineering teams to ensure system reliability, performance, and security across production systems.
Attain powers Klover, a fast-growing fintech platform serving over one million active users monthly, processing over $1.5 billion annually. The company emphasizes collaboration, reliability, and innovation, with a culture of automation and AI-driven development.
Work with a team of DevOps and DBA professionals to improve infrastructure and streamline deployments across countries.
Continuously improve Kubernetes platform stability, efficiency, and GitOps-first environment provisioning.
Monitor cloud infrastructure, own on-call operations, and define SLIs/SLOs for reliability improvements.
Sporty Group is a remote-first company focused on sustainability in the sports and gaming industry. They maintain a competitive, performance-driven culture with a distributed team across EMEA.