Source Job

$77,000–$106,700/yr
France

  • Design and develop scalable search and indexing systems for an AI search engine.
  • Ensure operational excellence by participating in on-call rotation and maintaining system quality.
  • Collaborate with a global remote team to solve distributed system challenges.

Distributed Systems Kubernetes Go Automated Testing

20 jobs similar to Senior Site Reliability Engineer - Search

Jobs ranked by similarity.

Romania 5w PTO 16w maternity 6w paternity

  • Design and build distributed data systems handling large-scale ingestion and processing.
  • Drive architectural decisions and take end-to-end ownership of critical components.
  • Collaborate with product teams to translate ambiguous requirements into robust technical solutions.

Our partner builds a large-scale, multi-chain data platform that ingests, models, and delivers blockchain data to users and developers. They are a remote-first, distributed team with a strong engineering culture focused on ownership and collaboration.

India

  • Design and build production AI agent systems for diagnosing and remediating infrastructure issues in large-scale GPU environments.
  • Develop distributed services, orchestration frameworks, knowledge graphs, and retrieval systems to power AI agents.
  • Own services end-to-end from architecture through production, collaborating with infrastructure and engineering teams.

The company builds AI agents that operate and automate large-scale GPU infrastructure. The engineering team is highly collaborative and remote, fostering ownership and autonomy.

$151,000–$206,000/yr
US Canada Unlimited PTO

  • Build large-scale real-time services and applications leveraging massive datasets.
  • Develop and maintain data pipelines, messaging systems, databases, and cloud services.
  • Work with Machine Learning Engineers and Security Researchers on security solutions.

Censys provides real-time Internet intelligence and threat insights to global governments and Fortune 500 companies. It is a growing company with a focus on comprehensive internet mapping and security solutions.

$139,200–$235,200/yr
Canada United States Unlimited PTO

  • Design, build, and operate GitLab Orbit backend services, primarily in Rust, within a distributed, cloud-native environment.
  • Improve deployment, monitoring, and operations using Kubernetes, Helm, Terraform, and cloud services from AWS or GCP.
  • Automate operational work, strengthen observability, and manage production issues to reduce single points of failure.

GitLab is the intelligent orchestration platform for DevSecOps, helping organizations increase developer productivity, improve operational efficiency, and accelerate digital transformation. Trusted by more than 50 million registered users and over 50% of the Fortune 100, GitLab fosters a high-performance culture driven by shared values and continuous knowledge exchange.

India

  • Design and build production AI agent systems that diagnose, investigate, and remediate infrastructure issues across one of the world’s largest GPU fleets.
  • Build the distributed services, orchestration framework, knowledge graph, and retrieval systems that power infrastructure agents.
  • Own services end to end, including architecture, implementation, testing, deployment, observability, and production operations.

Together AI is a research-driven artificial intelligence company that builds open and transparent AI systems. The company has contributed to leading open-source research like FlashAttention and RedPajama, and aims to lower the cost of modern AI through co-designed software, hardware, algorithms, and models.

India

  • Own and drive impactful distributed systems problems end-to-end, from inception through production launch.
  • Collaborate with a strong engineering team to prioritize and solve the most important problems for the company.
  • Raise the quality bar while keeping systems reliable and operationally lean, and mentor fellow engineers.

StarTree is a cloud-based software company that enables businesses to derive advanced insights from real-time and historical data using Apache Pinot. The company was founded by the core engineering team behind Apache Pinot, has secured Series B funding, and was named one of The Information's 50 Most Promising Startups.

$232,000–$290,000/yr
United States Canada 18w maternity 12w paternity

  • Set the long-term technical vision and architecture for backend systems powering experimentation, personalization, and analytics.
  • Architect and evolve high-throughput distributed systems and APIs for real-time data processing with low latency and high availability.
  • Lead complex, multi-team technical initiatives end-to-end, from problem framing through implementation and long-term operational ownership.

Webflow is the agentic web marketing platform for modern marketing teams, helping organizations build, manage, and optimize high-performing web experiences. They have a remote-first culture that values craft, quality, and moving fast, with a focus on innovation and collaboration.

$127,008–$152,410/yr
Ireland 6w PTO

  • Take an active role in influencing the roadmap and your own career objectives.
  • Drive projects from initial ideation all the way to operations in customers' hands.
  • Mentor and support team members while collaborating in a remote-first environment.

Grafana Labs builds Grafana, the open-source observability platform, and Grafana Cloud, a fully managed observability service. It's a 100% remote company with 1,600+ team members across 40+ countries, backed by leading investors and known for a collaborative open-source culture.

$0–$150,000/yr
US EU UK

  • Help design, build, and operate the Kubernetes platform used across PulsePoint.
  • Own reliability, observability, and incident response across platform services.
  • Build infrastructure automation and GitOps workflows to reduce operational toil.

PulsePoint sits at the intersection of healthcare and adtech, helping brands interpret health signals using real-world data. With over 300 employees, the company is a post-acquisition profitable leader in the US healthcare ad market, known for a flat hierarchy and high engineering bar.

Europe 6w PTO

  • Lead multi-quarter technical initiatives on Tempo's architecture, including trace aggregation APIs, autoscaling, and query engine improvements.
  • Drive operational excellence by owning SLOs, reducing toil, and ensuring Tempo operates reliably at scale across growing cell counts.
  • Design APIs for humans and agents, partner with product teams, and mentor engineers to raise the bar across the organization.

Grafana Labs is the company behind the open observability cloud, Grafana Cloud, a fully managed observability platform built for scale. With over 1,600 team members across 40+ countries, we are a 100% remote company backed by leading investors, fostering a global collaborative culture and a passion for meaningful work.

US 6w PTO

  • Take an active role in influencing our roadmap and your own career objectives.
  • Design, build, operate, and maintain critical systems, owning reliability, performance, and availability.
  • Collaborate with your team to deliver new features and iterate based on results.

Grafana Labs is the company behind the open-source observability platform Grafana, providing a fully managed observability cloud. With over 1,600 team members across 40+ countries and 35 million users, the company thrives on a transparent, collaborative, and open-source culture.

$217,000–$303,900/yr
US 17w maternity 17w paternity

  • Lead reliability initiatives across multiple Ads domains including ad serving, auctions, targeting, reporting, measurement, and billing.
  • Design and build platforms, tooling, and automation that improve reliability and developer productivity at scale.
  • Participate in on-call rotations, lead complex incident investigations and coordinate cross-functional response efforts during major production events.

Reddit is a community of communities, built on shared interests, passion, and trust. With 100,000+ active communities and approximately 130 million daily active unique visitors, Reddit is one of the internet's largest sources of information.

Global

  • Design and build agent runtime infrastructure with Firecracker, Rust, and Go
  • Define and enforce security boundaries for running untrusted AI agents
  • Architect global scale distributed systems for scheduling and orchestration

We build agent sandboxes—runtime infrastructure that is fast, durable, and secure by default for AI systems. We are a small, globally distributed team based in San Francisco backed by forward-thinking investors.

$150,000–$190,000/yr
US Canada Europe

  • Build and enhance SpiceDB, a distributed permissions database, and contribute to the open-source ecosystem.
  • Drive best practices in software development, testing, and CI/CD to ensure robust and scalable platform.
  • Collaborate with a high-performing engineering team to address complex challenges in distributed systems and authorization.

AuthZed creates and maintains SpiceDB, an authorization infrastructure used by companies globally to simplify permission management. As a Series A company, they have a fully remote, hardworking team with a software-driven culture across the US, Canada, and Europe.

US

  • Provide technical leadership for reliability across a large-scale advertising technology ecosystem
  • Lead reliability initiatives across ad serving, auctions, targeting, reporting, and billing systems
  • Mentor engineers and influence technical decisions to improve system resilience and developer productivity

The company is a partner organization operating a large-scale advertising technology ecosystem. Its size and culture are not detailed, but the role emphasizes reliability and operational excellence in a high-traffic environment.

Switzerland

  • Lead and develop multiple software engineering teams, ensuring high standards of quality and reliability.
  • Own and drive product development roadmaps across commercial systems, aligning with business objectives.
  • Provide technical leadership across billing, contracts, customer operations, and Linux security services.

The company is a global business operating commercial software platforms in a highly distributed environment. It is remote-first, international, and values autonomy, ownership, and continuous learning.

$190,000–$230,000/yr
US

  • Design, implement, and manage scalable cloud infrastructure using Kubernetes and Pub/Sub.
  • Refactor systems for scalability, including transforming stateful components to stateless ones.
  • Drive platform reliability initiatives like alerting, health checking, and incident management.

Syllo is a unified litigation platform that empowers lawyers and paralegals to safely use language models and agentic AI throughout the litigation lifecycle. The company has gained enterprise customers including major law firms and corporations, and is quickly expanding with a focus on reducing litigation costs and improving access to justice.

$180,000–$240,000/yr
US

  • Own large slices of the system end to end, from approach to operation.
  • Turn Beads into a platform and take Gas City to the cloud.
  • Define SLOs, observability, backups, and security baseline for enterprise readiness.

Gas City builds the open-source stack teams use to run coding agents at scale, including the Beads work graph and Gas City agent orchestration. It's a small, flat organization moving toward revenue with a focus on reliability and agent-driven development.

UK

  • Architect and build a robust, scalable, and highly available distributed infrastructure.
  • Build a cutting-edge cloud-native platform on top of the public cloud and automate cloud resource management.
  • Work closely with core database development and security teams to produce the SaaS offering.

ClickHouse is a real-time analytics and data warehousing company recognized on the Forbes Cloud 100 list. With over 4,000 customers and rapid growth, the company is a leader in its space.

Europe

  • Design scalable storage architectures for edge and core GPU deployments using platforms like StorPool, NVMe, VAST Data, and Weka.
  • Optimize storage for AI workloads including distributed training, fine-tuning, and inference with GPU Direct Storage and RDMA.
  • Act as primary storage design authority, influencing platform architecture and mentoring engineers across infrastructure domains.

Radian Arc builds scalable storage architectures powering AI and GPU infrastructure across edge and core environments. They are a growing company with a remote-friendly work model and an inclusive environment focused on next-generation AI infrastructure.