Design, deploy, and maintain the reliability, availability, and performance of critical systems and APIs across AWS and GCP.
Build observability frameworks, define SLIs/SLOs, and implement monitoring using Datadog and Kubernetes.
Participate in on-call rotations, incident response, and blameless post-incident reviews to drive systemic improvements.
JumpCloud is an AI-powered unified IT management platform that secures the modern workforce by consolidating identity, device, and access management. The company is remote-first with teams in over 15 countries and values building connections, thinking big, and continuous improvement.
Build and maintain the company's internal platform, driving operational excellence.
Collaborate with engineering squads to ensure applications are safe and reliable.
Take ownership of software infrastructure projects and provide off-hours support.
Loadsmart is a growth-stage logistics technology company valued at over $1 billion, using innovative technology to reinvent the freight industry. With headquarters in Chicago and a globally distributed remote team, it attracts top talent committed to driving meaningful change.
Collaborate with engineering teams to design and implement scalable, secure, and highly available systems for ClickHouse.
Establish and manage SLOs and SLAs for ClickHouse Cloud, ensuring monitoring and alerting are in place for all infrastructure.
Lead incident response, blameless postmortems, and chaos initiatives to continuously improve reliability and performance.
ClickHouse develops an open-source column-oriented database management system and offers a cloud database service. The company is a rapidly scaling, globally distributed startup with employees in over 25 countries, offering a flexible and collaborative culture.
Design, develop, and maintain reliability solutions and SRE utilities using Python in AWS environments to reduce toil and improve platform reliability.
Build observability and monitoring solutions with Grafana and AWS CloudWatch, and implement Infrastructure as Code using Terraform.
Develop CI/CD pipelines, define SRE standards and metrics, and participate in incident management and on-call rotation.
Peraton is a next-generation national security company that delivers mission-critical solutions and transformative IT services to government agencies and the U.S. armed forces. The company operates across land, sea, space, air, and cyberspace, with employees solving the most daunting challenges facing customers worldwide.
Build and maintain scalable cloud infrastructure for high availability.
Enhance observability and monitoring frameworks for accurate alerts.
Support on-call rotations and incident response with post-mortems.
GoGuardian is an award-winning learning solutions company purpose-built for K-12, trusted by educators to promote effective teaching and keep students safe. They are a remote, diverse, and committed team of mission-driven employees focused on improving learning environments.
Architect and scale multi-region microservices, APIs, and authentication infrastructure on AWS/GCP.
Lead SLOs, observability, incident management, and disaster recovery automation to maintain 99.99% availability.
Manage Kubernetes clusters and Terraform IaC while eliminating toil with Python/Go tooling.
JumpCloud is an AI-powered unified IT management platform that secures the modern workforce through identity, device, and access management. The company is remote-first with teams in 15+ countries and values building connections, thinking big, and continuous improvement.
Contribute to platform and harness engineering, including CI/CD and developer tooling.
Build systems to reduce toil and maintain production infrastructure under conversational AI traffic.
Participate in on-call rotation and incident management to ensure platform uptime.
Replicant builds an AI-powered customer service platform that helps contact centers resolve requests and improve agent performance. The company is distributed, with a focus on ownership and collaboration, and serves Fortune 500 companies.
Proactively identify, triage, and resolve performance issues across our Ruby on Rails stack and infrastructure.
Enhance system observability through monitoring metrics, SLOs, and SLIs across Ruby, Rails, and database systems.
Build and maintain AI agents and automations that reduce operational toil across incident response and routine maintenance.
Fleetio is a modern software platform that helps thousands of organizations worldwide manage their fleet operations. We raised $450M in Series D funding in 2025 and maintain a remote-friendly, engineering-driven culture.
Consolidate Terraform and establish conventions for state management, modules, and CI checks.
Improve monitoring, observability, and automation in Datadog and Cloud Monitoring.
Right-size workloads, evaluate Kubernetes architecture, and retire legacy tooling.
Vida is a virtual, personalized obesity care provider that combines evidence-based treatment with advanced technology to help patients improve their health. Trusted by Fortune 100 companies and growing for years, Vida takes a whole-person approach to care and celebrates diversity across its team.
Design, build, and operate shared platform foundations including GCP, Kubernetes, networking, CI/CD, and observability.
Diagnose and troubleshoot complex distributed systems running at high request volume.
Raise the reliability bar through dashboards, alerting, on-call readiness, and automation.
Sanity.io builds an AI-powered content operating system that helps teams model, create, and automate content. The company has 200+ employees and a positive, flexible, trust-based culture that supports growth and work-life balance.
Architect and build the observability platform for metrics, logs, traces, and events across global infrastructure.
Drive instrumentation with OpenTelemetry, building shared libraries and collector deployments for correlated signals.
Run observability as an internal product with published interfaces, versioned clients, and SLOs to ensure adoption.
Smartsheet empowers teams to manage work and scale solutions, uniting human teams with AI agents to automate tasks and uncover insights. With over 20 years of experience, the company fosters a collaborative, innovative culture focused on employee well-being and professional growth.
Design, implement, and evolve cloud platforms with focus on reliability, scalability, and security.
Build and maintain CI/CD pipelines, automate infrastructure using Terraform, Kubernetes, and Docker.
Implement observability, define SLIs/SLOs, and lead incident investigation and root-cause analysis.
Jobgether is an AI-powered job matching platform that connects candidates with hiring companies through a fair, objective review process. The platform ensures applications are quickly evaluated and shortlists are shared with employers, who manage interviews and final decisions.
Keep user-facing services and production systems reliable, scalable, and efficient through automation and infrastructure-as-code.
Build tooling and participate in on-call, incident response, and post-incident reviews to continuously improve reliability.
Contribute to the observability stack using metrics, logs, and SLOs to detect symptoms early and reduce toil.
GitLab is the intelligent orchestration platform for DevSecOps, enabling organizations to increase developer productivity, improve operational efficiency, and reduce security risk. With over 50 million registered users and more than 50% of the Fortune 100 trusting GitLab, we foster a high-performance culture driven by values, AI integration, and continuous knowledge exchange.
Design, implement, and maintain scalable, secure, and highly available cloud infrastructure.
Build and manage CI/CD pipelines, automate operational tasks, and improve deployment processes.
Monitor production systems, participate in incident response, and champion DevOps best practices.
First Due provides transformative end-to-end software solutions for fire and EMS agencies, helping them run safer, smarter, and more effective operations. The company offers a fully remote workplace, comprehensive benefits, and opportunities for advancement, with a culture focused on respect, inclusivity, and equal opportunity.
Champion SRE culture and best practices to improve production reliability and system resilience.
Communicate with stakeholders at all stages and bring fresh ideas to the table.
Participate in on-call rotation, incident response, and blameless post-incident reviews, while writing code and handling alerts.
Megaport is the global leader in Network as a Service (NaaS), transforming how businesses connect to cloud, data centers, and each other. With over 600 employees spread across Asia-Pacific, Europe, and the Americas, we are a collaborative, supportive, and fun team that values curiosity and diversity.
Design, build, and maintain automation and tooling to reduce operational toil.
Actively participate in the incident-management lifecycle, including detection, escalation, mitigation, and post-incident review.
Provide an SRE point of view on capacity planning, resilience testing, and modernization of legacy workloads.
Seismic is the Go-To-Market Performance company, helping organizations turn strategy into revenue through an AI-powered revenue execution platform. Trusted by 2,500 organizations and over 3.5 million users globally, Seismic is headquartered in San Diego with offices across North America, Europe, and Asia-Pacific, fostering an inclusive culture.
Support production systems across Azure, GCP, and datacenters to meet SLA targets.
Automate with Terraform, GitHub Actions, ArgoCD, and Python/Bash/PowerShell.
Participate in on-call rotation and collaborate to resolve incidents and improve reliability.
Kinaxis is a global leader in modern supply chain orchestration, with an AI-infused platform that provides end-to-end visibility. Starting as a team of three in 1984, Kinaxis now has over 2,000 employees globally and is known for its strong culture and technology.
Own the reliability posture of production services, including availability, latency, capacity, and performance.
Define and operate against SLIs and SLOs, using error budgets to drive engineering priorities.
Lead incident response, write post-mortems, and build automation to measurably improve service reliability.
Twilio is a cloud communications platform that delivers innovative solutions to hundreds of thousands of businesses and empowers millions of developers worldwide to create personalized customer experiences. The company is remote-first with a strong culture of connection, global inclusion, and a focus on solving problems and taking initiative.
Embed with product teams to improve operational maturity through on-call, monitoring, and alerting practices.
Run game day exercises and implement reliability techniques in Haskell & TypeScript code.
Champion reliability practices through design reviews and advocate for SLOs tied to customer outcomes.
Mercury is a fintech company that provides banking services for startups. They are building a modern banking platform and value reliability and innovation.
Lead Cloud Platform and SRE teams to scale securely and efficiently.
Drive infrastructure strategy, including Kubernetes (GKE) clusters and developer platform.
Champion SRE culture with SLOs, error budgets, and observability.
Prolific builds human data infrastructure for AI development. The company is a fast-growing, mission-driven organization with cross-functional teams and a strong ownership culture.