Design and evolve cloud infrastructure on GCP for scale and resilience.
Build internal tooling and automation that promote team autonomy and developer productivity.
Advance observability platform with metrics, logging, tracing, and alerting to reduce recovery time.
The company is a well-funded AI/ML company at the intersection of geospatial intelligence and climate technology, building products on scalable cloud infrastructure. The engineering team fosters a culture of reliability and continuous improvement, operating with a focus on SLOs, error budgets, and DORA metrics.
Design and maintain infrastructure-as-code patterns using Terraform and Kubernetes for scalable deployments.
Build monitoring, logging, and alerting systems, lead incident response, and drive continuous reliability improvements.
Embed security into infrastructure and optimize performance, costs, and automation across the platform.
Remote enables global employment compliantly, allowing businesses to recruit, pay, and manage international teams. With a future-focused culture and fully remote team across six continents, it builds an innovative HR platform with automation and AI.
Design and maintain highly available, scalable systems to ensure exceptional customer experiences.
Drive automation and eliminate operational toil through self-service tooling and process improvements.
Lead incident response and mentor engineers to improve reliability practices.
Redzone provides a connected workforce solution for manufacturers to improve plant efficiency and worker productivity. The company is part of QAD Inc. and fosters a collaborative, customer-focused culture with a strong technology team.
Design, build, and operate AWS infrastructure across multiple regions using Kubernetes, Terraform, and Helm.
Own large-scale object storage environments exceeding 15 petabytes, optimizing reliability, performance, scalability, and cost.
Build an internal developer platform for self-service infrastructure, strengthen security practices, and lead incident response and postmortems.
Our partner builds an AI-powered sports media platform handling petabyte-scale media storage and millions of minutes of video monthly. It operates as a remote-first, international, engineering-led team with significant autonomy and a focus on high availability and security.
Design and evolve scalable cloud infrastructure on Google Cloud Platform, focusing on reliability and automation.
Strengthen observability platform with metrics, logging, and tracing to improve incident response and reduce recovery time.
Champion reliability practices like SLOs, error budgets, and DORA metrics to drive operational excellence.
They operate at the intersection of geospatial intelligence and environmental technology. They are a growing organization with a collaborative, high-impact engineering culture.
Set reliability strategy and SLO culture that scales across engineering teams.
Own platform architecture, event-driven messaging, and observability for a global payments platform.
Lead chaos engineering, incident response, and mentorship for the most complex production challenges.
Yuno is an AI-native operating system for global commerce, connecting merchants to pay-ins, payouts, fraud prevention, and stablecoins via a single API. It powers payment infrastructure for global brands like McDonald's and GoFundMe, with a culture of remote work and AI-driven innovation.
Design, build, and operate Kubernetes infrastructure for AI workloads using Terraform and GitOps.
Define SLOs, run incident response, and create runbooks for reliable AI platform operations.
Drive AI-specific observability, FinOps, and security practices across the platform.
We are an AI-native consulting partner working with clients like PayPal, adidas, and NatWest to build digital products and services. Our team of over 600 has scaled quickly, earning Great Place to Work-Certified status multiple years in a row.
Lead the platform engineering function, defining strategy and long-term vision for cloud infrastructure and developer experience.
Partner with senior engineering leadership to establish a platform roadmap focused on reliability, scalability, security, and automation.
Mentor a team of platform engineers, support hiring, and drive operational maturity across global engineering teams.
The company is a fast-growing digital product organization that develops and operates a portfolio of SaaS products. It employs dozens of engineers and fosters a collaborative, remote-first culture focused on ownership and continuous improvement.
Design, build, and operate reliable infrastructure supporting AI-powered products.
Own and improve Kubernetes environments and cloud infrastructure.
Enhance production reliability through observability, automation, and incident response.
The company builds advanced AI-driven products and services. It values engineering excellence, autonomy, and individual contribution, with a global team of skilled engineers.
Ensure production system reliability, scalability, and performance through automation and AI-driven operations.
Design and implement AIOps workflows for event ingestion, anomaly detection, and incident response.
Collaborate with engineering teams to embed reliability into the software lifecycle and improve system resilience.
Jobgether is a platform that uses AI-powered matching to connect candidates with job opportunities. The company is a partner-focused recruitment service with a remote-first, globally distributed team.
Lead technical and managerial direction for the SRE team, defining reliability, observability, and operational excellence strategy.
Coordinate critical incident responses and root cause analysis, collaborating with architecture, development, security, and product teams.
Drive automation, continuous improvement, and adoption of SRE, DevOps, and Platform Engineering best practices.
Experian is a global data and technology company that drives opportunities for people and businesses worldwide. With 25,200 employees in 32 countries, it has a people-centric, inclusive culture recognized by awards such as World's Best Workplaces™ 2025.
Design, build, and operate secure Kubernetes-based infrastructure for AI-assisted development.
Implement GitOps and Infrastructure as Code for reproducible provisioning and operations.
Provide observability, security, and operational discipline for engineering platforms.
Deutsche Telekom IT Solutions is a subsidiary of the Deutsche Telekom Group providing IT and telecommunications services. The company employs over 5300 people and was named Hungary's most attractive employer in 2025.
Lead Cloud Platform and SRE teams, driving infrastructure strategy and ownership including Kubernetes, GCP, and Terraform.
Champion SRE culture, define SLOs, SLAs, and enhance observability and incident management.
Own the developer-facing platform as a product, ensuring self-service infrastructure and security compliance (SOC2, ISO-27001).
Prolific builds human data infrastructure for AI development, connecting researchers with a global participant pool. They foster a culture of high performance, ownership, and cross-functional collaboration.
Drive deployment automation across EKS clusters using GitOps.
Own infrastructure requirements and coordinate maintenance backlog.
Collaborate with engineering teams and integrate AI agents to accelerate workflows.
Ping Identity provides an intelligent cloud identity platform that secures digital experiences. With global offices and serving over half of the Fortune 100, the company fosters a culture of respect and individuality.
Own cloud infrastructure across AWS and GCP, including Kubernetes, networking, databases, and CI/CD pipelines.
Scale single-tenant deployments and build observability, incident response, and compliance practices.
Manage infrastructure cost, improve developer experience, and contribute to backend systems at the infrastructure-application intersection.
Elicit is an AI research assistant that uses language models to help researchers with literature review and evidence synthesis. The company is a ~30-person Public Benefit Corporation with a high-agency, low-bureaucracy culture.
Lead, mentor, and grow a team of SRE/DevOps engineers while partnering with engineering leadership to assess team needs and develop talent.
Oversee the incident management process end to end, including on-call rotations, escalation paths, incident command, postmortems, and root cause analysis.
Define and drive SRE principles like SLIs, SLOs, error budgets, capacity planning, and observability standards, championing a culture of reliability and operational excellence.
Eltropy is a rocket ship FinTech on a mission to disrupt the way people access financial services, enabling community financial institutions to digitally engage in a secure and compliant way through a world-class digital communications platform. Their platform integrates Text, Video, Secure Chat, co-browsing, screen sharing, and chatbot technology, bolstered by AI and contact center capabilities, and they value integrity, transparency, and ownership.
Own the reliability, security, and infrastructure for the AI operations platform running sandboxed agents.
Join a newly formed SRE team to build reliability practice from scratch on real infrastructure.
Manage distributed systems, observability, incident response, and automation with a security-first mindset.
Duvo builds an AI operations platform for retail and CPG enterprises to automate data workflows across systems. They are a fast-moving, humble team focused on solving real customer problems with strong traction.
Lead and mentor a US-based team of Site Reliability Engineers, driving operational excellence across production platforms.
Serve as an escalation point and incident commander for major production incidents, ensuring timely triage and resolution.
Drive automation, reliability, and observability improvements using Datadog, Kubernetes, and CI/CD pipelines.
NationsBenefits is a Healthcare Fintech provider of supplemental benefits, flex cards, and member engagement solutions that partners with managed care organizations. Recognized as one of the fastest-growing companies in America with multiple locations in the US, South America, and India, they offer a fulfilling work environment that encourages associates to contribute to delivering premier service.
Work with teams to define SLIs and SLOs, and create systems for observability.
Analyze failure scenarios, create runbooks, and reduce work that does not add value.
Participate in incident management and facilitate on-call duty to ensure reliable production environments.
Valtech is an experience innovation company that helps brands unlock new value in an increasingly digital world. They are a global team with a values-driven culture that fosters creativity, diversity, and autonomy.
Design and implement scalable cloud infrastructure to support growth.
Develop monitoring, alerting, and incident response for system reliability.
Automate deployment pipelines and ensure high availability and security.
Tekmetric is the all-in-one, cloud-based software helping auto repair shops run smarter, grow faster, and serve customers better. Founded in Houston in 2017, we've grown into an industry-leading team of builders who value transparency, integrity, and a service-first mindset.