Source Job

Global Unlimited PTO

  • Own and execute automated reliability, regression, and correctness test workflows within the CI/CD pipeline.
  • Collaborate with core developers to convert engineering specifications into repeatable test constraints.
  • Investigate complex performance data variations and identify architectural bottlenecks before production deployment.

Rust C++ Kubernetes Docker QA

20 jobs similar to Software Engineer, Database Reliability & Performance

Jobs ranked by similarity.

Global

  • Own and modernize testing infrastructure for Kubernetes operators running PostgreSQL, MySQL, and MongoDB.
  • Speed up feedback loops by re-architecting e2e test pipelines and integrating AI for test generation and triage.
  • Build reliable CI/CD pipelines and partner with QA to scale automated testing across the organization.

Percona provides open source database software, support, and services to make databases and applications run better. It is a remote-only company with a globally dispersed workforce of experts across more than 50 countries, fostering a collaborative and highly-engaged culture.

US 6w PTO

  • Take an active role in influencing our roadmap and your own career objectives.
  • Design, build, operate, and maintain critical systems, owning reliability, performance, and availability.
  • Collaborate with your team to deliver new features and iterate based on results.

Grafana Labs is the company behind the open-source observability platform Grafana, providing a fully managed observability cloud. With over 1,600 team members across 40+ countries and 35 million users, the company thrives on a transparent, collaborative, and open-source culture.

Global

  • Own the long-term evolution of the TimescaleDB toolkit, balancing customer needs, technical excellence, and product strategy.
  • Define and prioritize the roadmap by identifying emerging customer needs across IoT, industrial systems, observability, and financial markets.
  • Partner closely with Product, Sales, and customers to translate real-world workflows into intuitive developer experiences.

Tiger Data, formerly Timescale, provides the fastest PostgreSQL platform for transactional, analytical, and agentic workloads. Backed by $180 million and trusted by over 2,000 customers across 25+ countries, the globally distributed, remote-first team values direct communication, accountability, and collaborative excellence.

Europe 6w PTO

  • Partner closely with product engineering squads to own production reliability for high-SLA customer environments.
  • Design and implement automation to scale reliability practices and define per-tenant SLOs and reliability models.
  • Serve as a primary escalation point for incidents, lead response and post-incident reviews, and improve alert quality.

Grafana Labs is the company behind Grafana Cloud, a fully managed observability platform trusted by over 10,000 organizations to ensure reliability and resolve incidents faster. We are a 100% remote company with team members across 40+ countries, backed by leading investors, and we foster a global collaborative culture and a passion for meaningful work.

US UK Ireland Poland Germany Australia

  • Design and implement a scalable observability platform for Whatnot's growing infrastructure.
  • Work with core infrastructure, platform, and developer tools teams to redesign data collection to visualization.
  • Utilize AI agents and open standards to ensure visibility into software stack performance and reliability.

Whatnot is the largest live shopping platform in North America and Europe, enabling sellers to build businesses across hundreds of categories. They are a remote co-located team anchored in hubs across the US, UK, Ireland, Poland, Germany, and Australia, and were recently named the #1 Best Startup Employer in America by Forbes.

US Unlimited PTO

  • Build and deliver high-quality solutions that power Honeycomb's query and data storage infrastructure.
  • Scope and deliver projects independently, breaking down complex storage problems into achievable steps.
  • Support our services in production, participating in on-call rotations and reducing toil.

Honeycomb is a service defining observability and raising expectations for developer tools, working with companies like HelloFresh and Slack. We've scaled past 200 people, closed Series D funding, and were named to Forbes' America's Best Startups in 2022 and 2023.

Europe 6w PTO

  • Lead multi-quarter technical initiatives on Tempo's architecture, including trace aggregation APIs, autoscaling, and query engine improvements.
  • Drive operational excellence by owning SLOs, reducing toil, and ensuring Tempo operates reliably at scale across growing cell counts.
  • Design APIs for humans and agents, partner with product teams, and mentor engineers to raise the bar across the organization.

Grafana Labs is the company behind the open observability cloud, Grafana Cloud, a fully managed observability platform built for scale. With over 1,600 team members across 40+ countries, we are a 100% remote company backed by leading investors, fostering a global collaborative culture and a passion for meaningful work.

US

  • Design, develop, and execute performance testing scenarios including load, stress, spike, and scalability tests.
  • Investigate application performance issues by analyzing application, infrastructure, and database bottlenecks.
  • Integrate performance testing into CI/CD pipelines and apply AI/ML techniques to improve test automation.

The company is a partner organization focused on healthcare technology, providing enterprise-scale applications. It operates with a collaborative, inclusive remote engineering culture and values diversity.

Global

  • Find bottlenecks in live production systems and characterize them precisely enough that the owning team can act on them without you in the room.
  • Partner closely with database teams (e.g. Multigres, OrioleDB) and infra teams to land concrete performance improvements.
  • Build, communicate, and evolve performance methodologies and tooling that turn live production data into actionable insight.

Supabase is a Postgres development platform that provides a complete backend solution including Database, Auth, Storage, Edge Functions, Realtime, and Vector Search. The company has about 400 team members across 60+ countries, operates as a born-remote and open-source-first culture.

$217,000–$303,900/yr
US Unlimited PTO

  • Work collaboratively with a team to create and maintain the foundational platform for Reddit's infrastructure.
  • Deliver software to improve the availability, scalability, latency, and efficiency of observability components.
  • Contribute upstream changes to open source projects and share on-call responsibilities.

Reddit is a community of communities, built on shared interests, passion, and trust. With 100,000+ active communities and approximately 130 million daily active unique visitors, Reddit is one of the internet’s largest sources of information, employing a flexible-first workforce that values open-source contributions.

$128,700–$153,400/yr
US 4w PTO

  • Collaborate cross-functionally with teammates, external teams, and client stakeholders to integrate quality considerations throughout the SDLC.
  • Develop comprehensive test plans, test cases, and documentation to ensure thorough testing of features and functionalities.
  • Establish and enforce quality standards, track metrics, and automate testing processes to support continuous integration and deployment.

Bellese is a mission-driven digital services company pioneering innovative technology solutions in civic healthcare. They are a remote-first company with a collaborative learning environment focused on making a meaningful impact on public health outcomes.

UK

  • Develop deep, hands-on expertise in InfluxDB and the wider platform.
  • Partner with senior Sales Engineers on demos, proof-of-concept environments, and customer integrations.
  • Provide responsive technical coverage and relay customer feedback to Engineering and Product teams.

InfluxData is the creator of InfluxDB, a leading time series platform for collecting, storing, and analyzing time-stamped data at any scale. It is a remote-first company with a globally distributed workforce that values open source, humility, and getting things done.

US Unlimited PTO

  • Design, build, and operate distributed systems that ingest, process, and store telemetry at very high scale.
  • Own the reliability, performance, capacity, and cost-efficiency of telemetry pipelines and storage systems.
  • Participate in the on-call rotation, help resolve production incidents, and drive root-cause fixes through to completion.

ClickHouse provides a real-time analytics database for data warehousing, observability, and AI workloads. Recognized on the 2025 Forbes Cloud 100 list, it has over 4,000 customers and significant year-over-year growth.

$250,000–$285,000/yr
US Unlimited PTO

  • Define architecture and best practices for the platform and infrastructure layer the product is built on.
  • Own the deploy pipeline and lead the move to a GitOps model (Argo) for fast, safe releases.
  • Design and harden multi-tenant isolation and blast-radius protection for top-tier customers, including dedicated deployments.

We are the Engineering Operations Platform - mission control for the AI software factory, providing visibility, governance, and golden paths. We are a group of 80 passionate individuals, backed by $60M Series C from Sequoia, IVP, and others, with a fully remote culture.

US Unlimited PTO

  • Develop new features with a focus on optimizing performance and efficiency.
  • Collaborate with the team to implement scalable solutions and enhance application performance.
  • Identify and act on opportunities to improve the reliability of our services.

New Relic builds an intelligent observability platform that helps companies thrive in an AI-first world by providing insight into complex systems. It is a global company with a diverse and inclusive culture, fostering innovation and collaboration.

Argentina Colombia Costa Rica Mexico

  • Operate and maintain high-availability database systems including Vitess (distributed MySQL) and Cassandra against established runbooks.
  • Develop automation for common operational tasks using tools like Terraform, Ansible, and scripting languages such as Python, Bash, or Go.
  • Participate in on-call rotations, incident response, and post-incident reviews to ensure service reliability and continuous improvement.

Backblaze is the object storage leader in the open cloud movement, helping customers break free from overpriced legacy solutions with cloud storage designed to unlock budgets and unleash innovators. Founded in 2007, the company generates over $136M ARR, manages over three billion gigabytes of data for 500K+ customers in 175+ countries, and values diversity, equity, and inclusion in its workforce.

Global

  • Serve as the quality gate between engineering and production, ensuring every release meets standards.
  • Develop automated testing frameworks and define release validation strategies.
  • Work at the intersection of engineering and product to continuously improve software quality.

Alpen Labs is a New York-based startup founded in 2022 by four MIT alumni, building a scalable, private and programmable Bitcoin ecosystem. Their team consists of engineers and researchers from companies like Blockstream Research, Palantir, and Nethermind.

Global

  • Evolving Supabase Edge Runtime, an open-source Rust-based host that runs Deno isolate and enforces per-request memory and CPU limits.
  • Implementing monitoring, alerting, and OpenTelemetry tracing to drive latency and reliability improvements.
  • Expanding functions for more use cases like AI inference, MCP servers, and improving developer experience with Supabase CLI.

Supabase is the Postgres development platform, providing a complete backend solution including Database, Auth, Storage, Edge Functions, Realtime, and Vector Search. With a globally distributed team of ~400 members across 60+ countries, they are open-source-first and move fast, building in public.

Global

  • Drive QA strategy for Mirantis Secure Registry and Kubernetes environments.
  • Architect test frameworks for microservices in multi-cluster and hybrid-cloud.
  • Spearhead scalable test automation integrated with CI/CD pipelines.

Mirantis is the Kubernetes-native AI infrastructure company, enabling organizations to build and operate scalable, secure infrastructure for AI and data-intensive applications. It has a global, remote-first culture and is a leader in container management.

US Unlimited PTO

  • Build production-like test environments and tooling for engineers to spin up test cells resembling production topology.
  • Enable production-representative workloads including stress/load testing and traffic replay for controlled environments.
  • Drive system-level test coverage and build shared frameworks to make failure-mode testing repeatable across the engineering org.

Temporal is an open source programming model that simplifies code and makes applications more reliable. The company values curiosity, drive, collaboration, genuineness, and humility, and is building a team to be the reliable foundation of every developer's toolbox.