Operate and improve Linux infrastructure and Kubernetes clusters across bare-metal, virtualized, and on-premise environments.
Design and maintain complex networking architectures and automation using Ansible, Bash, Python, and GitOps.
Lead incident response, define SLOs, and build observability platforms with Prometheus, Grafana, and ELK.
Jobgether is a platform that connects job seekers with opportunities through an AI-powered matching process. The company fosters a remote-first culture and emphasizes autonomy and ownership for engineers.
Design and build self-service platform capabilities for engineering teams.
Develop and maintain global hybrid infrastructure across bare-metal, Linux, and Kubernetes.
Automate operational tasks and improve observability and reliability.
The partner company builds globally distributed infrastructure and platform capabilities. It is a small, highly autonomous team with a strong focus on reliability and developer experience.
Own and improve production infrastructure reliability and stability.
Prepare, execute, and support deployments and infrastructure changes.
Build and maintain Infrastructure-as-Code solutions using Ansible and Terraform.
Social Discovery Group (SDG) is a group of social discovery companies that solve problems of loneliness, isolation, and disconnection by transforming virtual intimacy into the new normal. Our international team of digital nomads works remotely from all over the world and we are proud to be a two-time 'Great Place to Work' winner (USA & Japan, 2024–2025) and a Top-5 Company for Work-From-Anywhere Jobs (FlexJobs, 2025).
Apply SRE principles to improve reliability, scalability, and performance of production systems.
Design and implement automation to reduce operational toil and improve engineering efficiency.
Lead incident response and develop sustainable solutions for complex production issues.
The hiring company is a technology organization focused on reliability and operational excellence. They offer a fully remote, collaborative environment with opportunities for technical leadership and career growth.
Drive the performance, stability, security, and reliability of production environments with a focus on automation and proactive improvements.
Design and maintain infrastructure using Infrastructure as Code tools like Terraform, and manage Kubernetes and cloud environments.
Lead vulnerability management, incident response, and secure CI/CD practices to ensure resilience and operational excellence.
Jobgether is a platform that uses AI-powered matching to connect candidates with hiring companies. It processes applications and shares shortlists with employers, offering a remote-first and inclusive work environment.
Own critical infrastructure across compute, networking, CI/CD, Kubernetes, and observability.
Manage Kubernetes environments and infrastructure-as-code with Terraform, improving developer experience and reducing operational friction.
Lead production incident response, influence architecture, and integrate AI-powered tools to boost engineering efficiency.
Jobgether is an AI-powered recruitment platform that connects candidates with global hiring companies. This role is with a partner company, a globally distributed technology organization offering a collaborative, informal culture and long-term opportunities.
Lead the discovery, design, and delivery of complex reliability and infrastructure initiatives, translating ambiguous problems into robust technical solutions.
Define and operate reliability practices including SLOs, SLIs, error budgets, alerting strategies, and observability standards.
Manage and scale production Kubernetes environments, build cloud infrastructure on AWS, and mentor less-senior engineers.
Jobgether is an AI-powered recruitment platform that connects candidates with partner companies through automated matching. They are a globally distributed organization focused on fair, efficient hiring processes.
Design, build, and operate AWS infrastructure across multiple regions using Kubernetes, Terraform, and Helm.
Own large-scale object storage environments exceeding 15 petabytes, optimizing reliability, performance, scalability, and cost.
Build an internal developer platform for self-service infrastructure, strengthen security practices, and lead incident response and postmortems.
Our partner builds an AI-powered sports media platform handling petabyte-scale media storage and millions of minutes of video monthly. It operates as a remote-first, international, engineering-led team with significant autonomy and a focus on high availability and security.
Lead Remote's SRE team owning Kubernetes, AWS, PostgreSQL, CI, and observability.
Balance 60% hands-on technical work with 40% people leadership and career growth.
Drive a maturing reliability practice including SLOs, incident response, and on-call.
Remote is a global employment platform that helps companies recruit, pay, and manage international teams. The company is fully remote with a future-focused, async culture and employees across six continents.
Own and scale cloud infrastructure including compute, networking, storage, and data systems.
Lead BYOC and private cloud deployments with infrastructure-as-code and GitOps foundations.
Establish reliability through service-level objectives, observability, and incident response processes.
A technology company builds a developer-focused platform with scalable cloud infrastructure. This is a remote-first opportunity with a small, autonomous engineering team operating in North America, LATAM, and Europe, offering high autonomy and ownership.
Monitor and support global network connectivity across data centers, offices, cloud environments, and private network links.
Troubleshoot routing, switching, firewall, and connectivity issues across first- and second-line support.
Contribute to CI/CD practices and network automation using Python, Ansible, Terraform, and Git.
This company operates in the financial technology sector, providing high-performance network infrastructure for latency-sensitive trading environments. The global team values network reliability and automation, offering a collaborative culture with opportunities for mentorship and career growth.
Operate and enhance the Open Telekom Cloud platform as a system administrator.
Manage network devices, monitor performance, troubleshoot issues, and ensure security.
Automate infrastructure with common frameworks and work in a collaborative specialist team.
Deutsche Telekom IT Solutions, a Deutsche Telekom subsidiary, provides IT and telecom services to large European customers. With 5,300+ employees in Hungary, it was ranked the country's most attractive employer in 2025.
Design and enhance multi-region, multi-cloud infrastructure for reliability and capacity failover.
Author custom Kubernetes tools and infrastructure code using Terraform, Go, Python, or Ruby.
Diagnose and fix complex production issues and performance bottlenecks across cloud environments.
Huntress is a cybersecurity company founded in 2015 by former NSA cyber operators. They provide enterprise-grade security to businesses of all sizes, securing over 5 million endpoints and 11 million identities. Their remote-first team values resilience and collaboration.
Install and configure OpenStack components using Kolla-Ansible in line with approved designs.
Handle second line incidents, service requests, and maintain configuration repositories.
Automate repeatable work with Ansible and CI pipelines, and participate in on-call rotation.
Software Mind develops solutions that make an impact for companies around the globe. The company builds cross-functional engineering teams that take ownership and crave more, with a culture that embraces openness, acts with respect, shows grit & guts, and combines employment with enjoyment.
Drive complex infrastructure migrations and build platform tooling and automation across multiple production environments.
Support development teams by consulting on infrastructure needs and improving observability and incident response.
Provide operational support and maintain platform reliability through structured debugging and on-call rotations.
PENN Entertainment is North America's leading provider of integrated entertainment, sports content, and casino gaming experiences. We operate across numerous locations in North America and foster a culture that cares about career growth and skill expansion.
Build and maintain scalable, reliable, and secure environments on AWS using Infrastructure as Code tools.
Design and manage CI/CD pipelines, oversee Kubernetes clusters, and ensure GitOps practices.
Monitor system health with OpenTelemetry and Grafana, enforce security best practices, and mentor junior engineers.
Deutsche Telekom IT Solutions is a subsidiary of the Deutsche Telekom Group, providing IT and telecommunications services with over 5,300 employees. Recognized as Hungary's most attractive employer, it serves large corporate clients across Europe.
Design and implement reliability strategies for distributed systems across AWS and GCP, defining SLIs and SLOs.
Build and enhance observability solutions using monitoring, logging, tracing, and alerting platforms.
Lead incident response, root cause analysis, and postmortem processes to improve system reliability.
We specialize in creating high-performing nearshore IT teams to help North American clients innovate faster and more efficiently. We are a people-first, purpose-driven company with a growing team, offering an inclusive culture and real growth opportunities.
Troubleshoot infrastructure and application issues, investigate incidents, and drive timely resolution while maintaining clear technical documentation.
Configure and improve monitoring and incident-management processes in collaboration with monitoring specialists and other technical teams.
Design and implement backup strategies, automated deployments, and Infrastructure as Code solutions to improve consistency, scalability, and operational efficiency.
The company operates a modern engineering environment focusing on reliable, secure, and scalable technology infrastructure. It values continuous improvement and collaboration, offering flexible work and extensive learning opportunities.
Lead the design of scalable, fault-tolerant, self-healing systems in a multi-region AWS environment.
Define SLOs and SLIs to drive architectural decisions and error budget policies.
Conduct blameless post-incident reviews and implement long-term preventive measures.
Airalo is the world's first eSIM store, helping travelers access affordable mobile data in 200+ countries. They are a fully remote team of 400+ people across 60+ countries, with a culture of trust, ownership, and freedom.
Manage physical "Metal" environments from bare metal to Kubernetes, including cluster networking and scheduling.
Maintain Crossplane compositions and Terraform modules for cloud service provider resources.
Work with application teams to understand needs and invest in right capabilities.
Grafana Labs is the company behind Grafana Cloud, a fully managed observability platform trusted by over 10,000 organizations. They are a 100% remote company with team members across 40+ countries, backed by leading investors, and known for an open-source legacy and global collaborative culture.