Own and drive key infrastructure modernization initiatives toward container-orchestrated infrastructure.
Design and maintain infrastructure as code across multiple cloud providers.
Provide technical leadership and mentorship across the Systems Engineering team.
Intellum is the leader in corporate education technology, powering large learning programs for brands like Google, Meta, and Amazon. We are a remote-first company with a culture that values curiosity, creativity, perseverance, and kindness, and we invest in our people through personal development budgets and annual retreats.
Improve system availability, scalability, and resilience across Flowcode's platforms.
Manage and scale core AWS infrastructure through Infrastructure as Code (Terraform) and enhance disaster recovery.
Oversee monitoring, logging, and alerting infrastructure, and develop high-signal metrics and dashboards.
Flowcode is a technology company specializing in QR code and smart link solutions for offline-to-online engagement. The company is a growth-stage startup seeking high-performing individuals who thrive in a fast-paced, demanding environment.
Design and build resilient AWS and Kubernetes platforms to improve reliability, scalability, and security.
Define SLOs, build observability, automate operational work, and lead incident response and post-incident reviews.
Partner with engineering, platform, security, and QA teams to establish reliability standards and optimize cost.
Electric Power Engineers (EPE) provides consulting expertise and energy intelligence software solutions for power and energy clients, focusing on renewable energy and grid modernization. With over half a century in the industry, the company fosters innovation and collaboration, working with industry leaders to build a secure and resilient grid.
Support and optimize AWS infrastructure for Amazon Connect and enterprise cloud platforms.
Manage Infrastructure as Code using Terraform and automate tasks with Python and AWS CLI.
Develop dashboards, alerts, and monitoring solutions using CloudWatch, Dynatrace, Splunk, Grafana, and OpenTelemetry.
Miratech is a global IT services and consulting company that helps visionaries change the world by bringing together enterprise and start-up innovation. The company retains nearly 1000 full-time professionals, operates in over 25 countries, and has a culture of Relentless Performance with a 99% project success rate since 1989.
Design, build, and optimize multi-region, high-availability AWS infrastructure.
Drive resiliency and automation using GitOps, modern CI/CD, and Infrastructure as Code.
Build end-to-end telemetry and own incident management to harden reliability.
VGS is the world's leader in payment tokenization, trusted by the most innovative AI and Fortune 500 companies. They are a remote-first company with a culture of ownership, collaboration, and continuous learning.
Design, implement, and evolve cloud platforms with focus on reliability, scalability, and security.
Build and maintain CI/CD pipelines, automate infrastructure using Terraform, Kubernetes, and Docker.
Implement observability, define SLIs/SLOs, and lead incident investigation and root-cause analysis.
Jobgether is an AI-powered job matching platform that connects candidates with hiring companies through a fair, objective review process. The platform ensures applications are quickly evaluated and shortlists are shared with employers, who manage interviews and final decisions.
Lead enterprise-wide reliability and infrastructure projects with high autonomy, architecting scalable solutions and driving SRE best practices.
Partner cross-functionally with Engineering, Product, and Customer Success to align reliability goals with business objectives and communicate complex concepts to diverse audiences.
Provide tier 2/3 technical support to enterprise customers, conduct technical onboarding, and act as a trusted advisor for platform architecture.
Veza is the pioneer in identity security, providing an Access Graph platform that maps identity ecosystems across users, groups, roles, policies, and resources. With over 30 billion access permissions under management and now part of ServiceNow, Veza combines enterprise scale with security innovation.
Actively identify, plan and implement developer tooling and automation.
Participate in on-call rotations and assist with diagnostics and troubleshooting of platform and infrastructure.
Set technical direction for the team's infrastructure decisions and define overall DevOps strategy.
Bluesight creates groundbreaking solutions that increase efficiency, safety and visibility for health systems, hospital pharmacy, and pharmaceutical manufacturers. They are a high-growth healthcare information technology company with over 3,000 customers and a startup culture.
Lead the design of scalable, fault-tolerant, self-healing systems in a multi-region AWS environment.
Define SLOs and SLIs to drive architectural decisions and error budget policies.
Conduct blameless post-incident reviews and implement long-term preventive measures.
Airalo is the world's first eSIM store, helping travelers access affordable mobile data in 200+ countries. They are a fully remote team of 400+ people across 60+ countries, with a culture of trust, ownership, and freedom.
Design and operate scalable AWS infrastructure with containerization and orchestration tools.
Implement monitoring, logging, and infrastructure as code using Terraform.
Improve CI/CD pipelines and troubleshoot production issues in complex SDLC environments.
Sureify builds systems that support millions of users. It is a high-growth, engineering-driven SaaS company with a remote-first culture across the Americas.
Build and maintain the company's internal platform, driving operational excellence.
Collaborate with engineering squads to ensure applications are safe and reliable.
Take ownership of software infrastructure projects and provide off-hours support.
Loadsmart is a growth-stage logistics technology company valued at over $1 billion, using innovative technology to reinvent the freight industry. With headquarters in Chicago and a globally distributed remote team, it attracts top talent committed to driving meaningful change.
Keep the platform reliable and scalable while improving observability and reducing friction in build and release processes.
Collaborate asynchronously across time zones within a remote-first DevOps team.
Use AI tools to accelerate automation and infrastructure work while validating output.
Remote People builds infrastructure to power borderless teams by handling global payroll, benefits, taxes, and compliance. They are a growing, international company committed to building a diverse team.
Champion SRE culture and best practices to improve production reliability and system resilience.
Communicate with stakeholders at all stages and bring fresh ideas to the table.
Participate in on-call rotation, incident response, and blameless post-incident reviews, while writing code and handling alerts.
Megaport is the global leader in Network as a Service (NaaS), transforming how businesses connect to cloud, data centers, and each other. With over 600 employees spread across Asia-Pacific, Europe, and the Americas, we are a collaborative, supportive, and fun team that values curiosity and diversity.
Build, operate, and evolve our cloud infrastructure with reliability, scalability, and security in mind.
Improve observability practices and partner across engineering to strengthen our cloud security posture and operational practices.
Build tooling and automation that reduces manual work and improves the developer experience.
Maze is the user research platform that helps companies build the right products faster by making user insights available at the speed of product development. With less than 150 team members and a global remote workforce, we operate driven by our core values of ownership, curiosity, and pragmatism.
Drive the performance, stability, security, and reliability of production environments with a focus on automation and proactive improvements.
Design and maintain infrastructure using Infrastructure as Code tools like Terraform, and manage Kubernetes and cloud environments.
Lead vulnerability management, incident response, and secure CI/CD practices to ensure resilience and operational excellence.
Jobgether is a platform that uses AI-powered matching to connect candidates with hiring companies. It processes applications and shares shortlists with employers, offering a remote-first and inclusive work environment.
Design, build, and operate shared cloud infrastructure using AWS, Kubernetes, Terraform, Databricks, and Cloudflare.
Deliver SRE and DevOps initiatives to improve reliability, scalability, observability, and deployment safety.
Build reusable infrastructure modules, automation, and self-service workflows to reduce manual work and improve developer experience.
YipitData is the leading market research and analytics firm for the disruptive economy, recently raising up to $475M from The Carlyle Group at a valuation over $1B. We analyze billions of alternative data points daily and have been recognized as one of Inc’s Best Workplaces, cultivating a people-centric culture focused on mastery, ownership, and transparency.
Lead cloud infrastructure strategy for resilient, secure, and cost-efficient multi-account cloud environments across AWS, Azure, and GCP.
Drive Kubernetes excellence as a technical authority for production clusters including EKS and AKS.
Advance AI-enabled operations by introducing LLM-based tooling and agentic workflows to improve infrastructure development and operational efficiency.
Jobgether is a platform that uses AI-powered matching to connect candidates with hiring companies. They are a technology company focused on improving the hiring process through automation and data analysis.
You will lead the reliability and operational evolution of our platform, building and improving system resiliency and establishing SLIs and SLOs.
You will partner with product engineering teams to own and operate their services, evolving observability platforms and strengthening incident practices.
You will contribute to day-to-day cloud infrastructure work alongside reliability specialty, including on-call rotation.
Rocket Money is a financial technology company that empowers people to live their best financial lives by providing insights and services to save time and money. The company runs hundreds of services in production, processing billions of transactions, and has a culture of reliability and innovation.
Build and operate monitoring, tracing, alerting, and observability infrastructure for system reliability.
Drive platform security initiatives with preventative controls and resilient architecture.
Lead incident response and recovery, including root-cause analysis and preventative measures.
This role is with a partner company managing AI-powered products. They are a growing technology organization with a fully distributed US-based team and a collaborative culture focused on large-scale infrastructure and AI technology.
Build and run monitoring, tracing, and alerting infrastructure to ensure platform reliability and security.
Lead incident response and recovery, including root cause analysis, and improve deployment processes for fast, safe code changes.
Collaborate with engineering teams to deliver a stable, scalable platform and handle load for resource-intensive applications.
WellSaid Labs is the leading AI voiceover studio for enterprise and professional use, providing ultra-realistic voices that the world’s biggest brands trust. We are a fully distributed team across the U.S. with a focus on responsible AI and an inclusive culture.