Source Job

US

  • Build and run monitoring, tracing, and alerting infrastructure to ensure platform reliability and security.
  • Lead incident response and recovery, including root cause analysis, and improve deployment processes for fast, safe code changes.
  • Collaborate with engineering teams to deliver a stable, scalable platform and handle load for resource-intensive applications.

Golang Kubernetes AWS Terraform Python

20 jobs similar to Senior Software Engineer, Site Reliability & Security

Jobs ranked by similarity.

$53,300–$119,850/yr
Global Unlimited PTO 16w maternity 16w paternity

  • Lead the discovery, design, and delivery of complex reliability and infrastructure initiatives, translating ambiguous problems into robust technical solutions.
  • Define and operate reliability practices including SLOs, SLIs, error budgets, alerting strategies, and observability standards.
  • Manage and scale production Kubernetes environments, build cloud infrastructure on AWS, and mentor less-senior engineers.

Jobgether is an AI-powered recruitment platform that connects candidates with partner companies through automated matching. They are a globally distributed organization focused on fair, efficient hiring processes.

US

  • Improve system availability, scalability, and resilience across Flowcode's platforms.
  • Manage and scale core AWS infrastructure through Infrastructure as Code (Terraform) and enhance disaster recovery.
  • Oversee monitoring, logging, and alerting infrastructure, and develop high-signal metrics and dashboards.

Flowcode is a technology company specializing in QR code and smart link solutions for offline-to-online engagement. The company is a growth-stage startup seeking high-performing individuals who thrive in a fast-paced, demanding environment.

US

  • Design, build, and maintain robust, scalable, and secure infrastructure systems supporting Laurel's AI-driven platform.
  • Manage and optimize cloud infrastructure (AWS and Azure), Kubernetes orchestration, and CI/CD pipelines to increase deployment frequency and reliability.
  • Implement comprehensive observability, monitoring, and alerting to maintain system health and partner with engineering teams to optimize performance and cost-efficiency.

Laurel is an AI Time platform for professional services firms, automating work time capture and connecting time data to business outcomes for clients like EY and Crowell & Moring. The company comprises top AI, product, and engineering talent, is VC-backed by Google Ventures and IVP, and fosters an inclusive, ambitious culture.

$175,000–$195,000/yr

  • Design and improve monitoring, logging, distributed tracing, dashboards, alerting, SLIs, and SLOs for production health.
  • Build and maintain automation, internal tools, and CI/CD systems to increase engineering efficiency and support reliable deployments.
  • Own complex production incidents from detection to resolution, turning learning into durable improvements and reducing recurring incidents.

Filevine is a Legal AI company delivering Legal Operating Intelligence for the future of legal work. It has earned recognition from Deloitte and Inc. as one of the most innovative and fastest-growing technology companies in the country.

EMEA Unlimited PTO

  • Design and maintain infrastructure-as-code patterns using Terraform and Kubernetes for scalable deployments.
  • Build monitoring, logging, and alerting systems, lead incident response, and drive continuous reliability improvements.
  • Embed security into infrastructure and optimize performance, costs, and automation across the platform.

Remote enables global employment compliantly, allowing businesses to recruit, pay, and manage international teams. With a future-focused culture and fully remote team across six continents, it builds an innovative HR platform with automation and AI.

US Unlimited PTO

  • Own the technical direction of the AWS platform, building cost visibility tooling and managing Kubernetes on EKS.
  • Design and maintain Terraform modules for self-service infrastructure provisioning and standardize CI/CD across services.
  • Harden the platform alongside security, lead incident response, and mentor senior engineers through design and code review.

PlayOn powers high school sports ticketing, streaming, and fundraising through platforms like GoFan, NFHS Network, and MaxPreps. Backed by KKR, the company is a growth-stage leader focused on making high school sports more accessible and connected.

US

  • Design and build sandboxed evaluation environments for AI models to safely execute code and interact with tools.
  • Build backend services and infrastructure supporting large-scale AI and agentic evaluations.
  • Develop agent scaffolding, evaluation harnesses, and systems for provisioning isolated environments using Docker, Kubernetes, and cloud infrastructure.

10a Labs is the safety and threat-intelligence layer trusted by frontier AI labs, AI unicorns, Fortune 10 companies, and leading global technology platforms. The company focuses on adversarial red teaming and model evaluations to help teams deploy AI systems safely.

US

  • Design and implement scalable cloud infrastructure to support growth.
  • Develop monitoring, alerting, and incident response for system reliability.
  • Automate deployment pipelines and ensure high availability and security.

Tekmetric is the all-in-one, cloud-based software helping auto repair shops run smarter, grow faster, and serve customers better. Founded in Houston in 2017, we've grown into an industry-leading team of builders who value transparency, integrity, and a service-first mindset.

$138,700–$173,400/yr
US

  • Design, build, and operate services and automations to manage Kubernetes clusters at scale, partnering with product management and technical leadership.
  • Drive rigorous code reviews and maintain high testing standards across the platform.
  • Manage cloud configurations across AWS and Azure using Terraform, ensuring deep observability and reliability.

Twilio is shaping the future of communications by delivering innovative solutions to hundreds of thousands of businesses and empowering millions of developers. They are a remote-first company with a strong culture of connection and global inclusion, employing a vibrant and diverse team.

Latin America

  • Design and implement reliability strategies for distributed systems across AWS and GCP, defining SLIs and SLOs.
  • Build and enhance observability solutions using monitoring, logging, tracing, and alerting platforms.
  • Lead incident response, root cause analysis, and postmortem processes to improve system reliability.

We specialize in creating high-performing nearshore IT teams to help North American clients innovate faster and more efficiently. We are a people-first, purpose-driven company with a growing team, offering an inclusive culture and real growth opportunities.

$126,290–$190,000/yr
United States 18w maternity 12w paternity

  • Empower engineers on other teams by maintaining monitoring tooling and collaborating on observability best practices.
  • Enhance reliability of Kubernetes applications through resource optimization, streamlined upgrades, and scalability.
  • Participate in on-call and incident response processes, occasionally diving into application code to debug production issues.

Webflow is the agentic web marketing platform for modern marketing teams, helping organizations build, manage, and optimize high-performing web experiences. It serves over 2 million users worldwide across 190 countries, with tens of thousands of projects launched each month, and fosters a culture of grit, speed, and craft.

$180,000–$250,000/yr
US Unlimited PTO

  • Own the technical direction and architecture of critical infrastructure domains, establishing scalable patterns and standards.
  • Lead complex, multi-team infrastructure initiatives from design through implementation and production operation.
  • Design and evolve AWS and Kubernetes infrastructure to enable teams to build and deploy systems reliably at scale.

We provide innovative identity and risk solutions, empowering institutions and individuals to transact with confidence. Our company is backed by world-class investors including Craft Ventures and Andreessen Horowitz, with offices across the US and India, and we are growing extremely quickly.

Global Unlimited PTO

  • Set reliability strategy and SLO culture that scales across engineering teams.
  • Own platform architecture, event-driven messaging, and observability for a global payments platform.
  • Lead chaos engineering, incident response, and mentorship for the most complex production challenges.

Yuno is an AI-native operating system for global commerce, connecting merchants to pay-ins, payouts, fraud prevention, and stablecoins via a single API. It powers payment infrastructure for global brands like McDonald's and GoFundMe, with a culture of remote work and AI-driven innovation.

US

  • Own the infrastructure end-to-end for ScaleOps' self-hosted and SaaS platforms.
  • Manage cloud infrastructure across AWS, GCP, and Azure, including networking, security, and compute.
  • Collaborate with customers and internal teams to ensure rapid feature delivery without compromising reliability.

ScaleOps is redefining autonomous cloud and AI infrastructure, freeing DevOps from manual resource management. Backed by $210M+ in funding, they are trusted by leading enterprises and Fortune 100 companies, with a fast-paced, innovative culture.

UK Unlimited PTO 18w maternity 12w paternity

  • Own the technical strategy for multi-ecosystem scaling, defining architecture for onboarding new language ecosystems.
  • Drive end-to-end remediation automation, leading redesign of CVE workflows to close the loop from detection to verified release.
  • Set platform-wide technical direction spanning package index, build pipelines, and orchestration tooling to serve customers and ecosystem teams.

Chainguard is the trusted source for open source, delivering hardened, secure, and production-ready builds of open source software. They serve Fortune 500 enterprises and global industry leaders, and are venture-backed by leading investors, fostering a culture of customer obsession and intentional action.

Canada Unlimited PTO 18w maternity 12w paternity

  • Build and harden secure CI/CD pipelines with security gates to catch issues before production.
  • Lead security architecture reviews and threat models for Kubernetes-based workloads on GCP and AWS.
  • Harden container images, Kubernetes configurations, and cloud IAM to minimize attack surface.

Chainguard is the trusted source for open source, delivering hardened, secure, and production-ready builds of open source software. The company is venture-backed by leading investors and serves Fortune 500 enterprises, with a culture that values customer obsession, intentional action, and trust.

APAC

  • Design, build, and operate components of the Kubernetes platform and its core subsystems end to end.
  • Write and review Go code for controllers, operators, platform services, and automation.
  • Help operate the platform, including on-call, incident investigation, and follow-up work to prevent recurrence.

ServiceNow provides an AI platform for business reinvention, helping 85% of the Fortune 500 work smarter. The company fosters an AI-native culture where technology and talent collaborate.

$241,000–$270,000/yr
US Unlimited PTO

  • Architect the end-to-end reliability, performance, and resilience of cloud environments, including the SLO framework for critical services.
  • Lead incident response, on-call rotation, root cause analysis, and build a culture of corrective actions.
  • Build observability platforms to detect issues proactively and mentor engineers on reliability standards.

Garner is on a mission to transform the U.S. healthcare system by partnering with employers to steer members to better-performing doctors, resulting in better care and lower costs. With 550+ proprietary clinical metrics, they have helped over 2.5 million people and saved $1B in healthcare costs, recently raising a Series E and doubling five years running.

Ireland

  • Design, build, and deploy production systems with focus on scalability, reliability, and security.
  • Develop and maintain automation to streamline operations and eliminate toil.
  • Proactively monitor systems and implement automated incident response to minimize downtime.

Arista Networks is an industry leader in data-driven networking for large data centers, campus, and routing. With over $8 billion in revenue and a culture valuing diversity, Arista is a Great Place to Work for Best Engineering Team and Best Company for Diversity.

$192,000–$192,000/yr
US Unlimited PTO

  • Design, implement, and maintain scalable and reliable systems.
  • Set up monitoring tools and create incident response plans to quickly identify and resolve issues.
  • Develop and maintain automation tools for deployment, monitoring, and system health checks.

LeoLabs is building the living map of activity in space through a proprietary global radar network and AI-enabled analytics platform. The company collects millions of measurements daily on more than 25,000 objects, protecting billions in assets for commercial and government missions.