Drive the performance, stability, security, and reliability of production environments with a focus on automation and proactive improvements.
Design and maintain infrastructure using Infrastructure as Code tools like Terraform, and manage Kubernetes and cloud environments.
Lead vulnerability management, incident response, and secure CI/CD practices to ensure resilience and operational excellence.
Jobgether is a platform that uses AI-powered matching to connect candidates with hiring companies. It processes applications and shares shortlists with employers, offering a remote-first and inclusive work environment.
Design and build resilient AWS and Kubernetes platforms to improve reliability, scalability, and security.
Define SLOs, build observability, automate operational work, and lead incident response and post-incident reviews.
Partner with engineering, platform, security, and QA teams to establish reliability standards and optimize cost.
Electric Power Engineers (EPE) provides consulting expertise and energy intelligence software solutions for power and energy clients, focusing on renewable energy and grid modernization. With over half a century in the industry, the company fosters innovation and collaboration, working with industry leaders to build a secure and resilient grid.
Ensure reliability, scalability, and performance of cloud-based systems using Kubernetes and observability tools.
Define and monitor reliability metrics (SLIs, SLOs, MTTR) to continuously improve operational performance.
Automate operational tasks and implement Infrastructure as Code to reduce manual work and enhance efficiency.
Our partner is a technology company focused on building and maintaining reliable, scalable digital environments. They promote a culture of continuous improvement, collaboration, and proactive engineering.
Design, build, and operate shared cloud infrastructure using AWS, Kubernetes, Terraform, Databricks, and Cloudflare.
Deliver SRE and DevOps initiatives to improve reliability, scalability, observability, and deployment safety.
Build reusable infrastructure modules, automation, and self-service workflows to reduce manual work and improve developer experience.
YipitData is the leading market research and analytics firm for the disruptive economy, recently raising up to $475M from The Carlyle Group at a valuation over $1B. We analyze billions of alternative data points daily and have been recognized as one of Inc’s Best Workplaces, cultivating a people-centric culture focused on mastery, ownership, and transparency.
Define and monitor reliability metrics such as SLI, SLO, SLA, MTTR, and MTTD.
Implement observability solutions including monitoring, alerting, dashboards, and APM.
Collaborate with multidisciplinary teams to embed reliability and observability into solutions.
The company is a technology organization focused on building and maintaining reliable digital environments. It fosters a culture of engineering excellence, collaboration, and data-driven decision-making.
Own critical infrastructure across compute, networking, CI/CD, Kubernetes, and observability.
Manage Kubernetes environments and infrastructure-as-code with Terraform, improving developer experience and reducing operational friction.
Lead production incident response, influence architecture, and integrate AI-powered tools to boost engineering efficiency.
Jobgether is an AI-powered recruitment platform that connects candidates with global hiring companies. This role is with a partner company, a globally distributed technology organization offering a collaborative, informal culture and long-term opportunities.
Build and maintain the company's internal platform, driving operational excellence.
Collaborate with engineering squads to ensure applications are safe and reliable.
Take ownership of software infrastructure projects and provide off-hours support.
Loadsmart is a growth-stage logistics technology company valued at over $1 billion, using innovative technology to reinvent the freight industry. With headquarters in Chicago and a globally distributed remote team, it attracts top talent committed to driving meaningful change.
Lead cloud infrastructure strategy for resilient, secure, and cost-efficient multi-account cloud environments across AWS, Azure, and GCP.
Drive Kubernetes excellence as a technical authority for production clusters including EKS and AKS.
Advance AI-enabled operations by introducing LLM-based tooling and agentic workflows to improve infrastructure development and operational efficiency.
Jobgether is a platform that uses AI-powered matching to connect candidates with hiring companies. They are a technology company focused on improving the hiring process through automation and data analysis.
Design, build, and maintain reliable cloud infrastructure on AWS using CI/CD pipelines and IaC tools.
Automate containerized workloads with Docker, implement monitoring and observability solutions, and ensure security best practices.
Collaborate with engineering teams to improve system reliability, troubleshoot issues, and drive operational excellence.
GoFasti is a Talent-as-a-Service company that bridges world-class developers and designers from Latin America with first-class companies globally. They are a remote-first organization focused on matching top talent with international opportunities.
Contribute to infrastructure automation and operational resilience across hybrid cloud and data center operations.
Implement closed-loop auto-remediation systems and SRE tooling to reduce manual intervention and incident resolution time.
Develop and maintain SLO frameworks, alerting policies, and Infrastructure-as-Code pipelines for reproducible deployments.
ServiceNow is the AI control tower for business reinvention, helping 85% of the Fortune 500 work smarter, faster, and better. They foster an AI-native culture where technology and talent are unstoppable together.
Improve system availability, scalability, and resilience across Flowcode's platforms.
Manage and scale core AWS infrastructure through Infrastructure as Code (Terraform) and enhance disaster recovery.
Oversee monitoring, logging, and alerting infrastructure, and develop high-signal metrics and dashboards.
Flowcode is a technology company specializing in QR code and smart link solutions for offline-to-online engagement. The company is a growth-stage startup seeking high-performing individuals who thrive in a fast-paced, demanding environment.
Lead the SRE strategy and execution for a high-growth AI company.
Build and scale a high-performing SRE team while defining reliability standards.
Architect secure, scalable cloud infrastructure and implement observability practices.
This company develops advanced AI products and agentic technology. It operates in a high-growth, international environment with a focus on operational excellence and innovation.
Serve as the technical backbone of cloud infrastructure operations, bridging incident detection and advanced architecture.
Build and maintain CI/CD pipelines, design IaC modules, and optimize cloud resources for performance and cost efficiency.
Lead observability initiatives, integrate DevSecOps practices, and collaborate with cross-functional teams to ensure robust cloud reliability.
CodeRoad provides end-to-end software development services, helping businesses scale with ideal infrastructure solutions. They operate with a nearshore model and focus on empowering businesses through staff augmentation, dedicated teams, and software engineering.
Deliver infrastructure and platform improvements across AWS environments using Terraform and modern cloud practices.
Build secure, reliable cloud solutions and automation using Node.js or Python, supporting CI/CD and testing.
Collaborate with engineers and leaders to drive platform reliability, security, and operational excellence.
They operate a highly transactional, customer-facing platform used by millions of users across Europe. Their culture is collegial, agile, and engineering-led, with shared responsibility and an emphasis on learning from mistakes.
Lead the discovery, design, and delivery of complex reliability and infrastructure initiatives, translating ambiguous problems into robust technical solutions.
Define and operate reliability practices including SLOs, SLIs, error budgets, alerting strategies, and observability standards.
Manage and scale production Kubernetes environments, build cloud infrastructure on AWS, and mentor less-senior engineers.
Jobgether is an AI-powered recruitment platform that connects candidates with partner companies through automated matching. They are a globally distributed organization focused on fair, efficient hiring processes.
Drive automation and modernize cloud infrastructure using AWS, Kubernetes, and infrastructure as code.
Design and implement enterprise-grade service mesh architectures and CI/CD pipelines.
Collaborate with cross-functional teams to align technical solutions with business and regulatory requirements.
Our partner is a technology company driving digital transformation through modern cloud infrastructure. They are a growing organization with a focus on secure, scalable solutions and engineering excellence.
Architect and maintain critical cloud platform components on AWS EKS with high availability and automated resilience.
Establish SRE standards including SLO/SLI tracking, error budget frameworks, and automated operational tooling.
Design and implement OpenTelemetry capture pipelines for telemetry data feeding downstream platforms.
Inflect is a US-based advisory and marketplace that revolutionizes how companies buy and sell digital infrastructure services. They operate with a focus on high-impact consulting and autonomous work.
Design and implement reliability strategies for distributed systems across AWS and GCP, defining SLIs and SLOs.
Build and enhance observability solutions using monitoring, logging, tracing, and alerting platforms.
Lead incident response, root cause analysis, and postmortem processes to improve system reliability.
We specialize in creating high-performing nearshore IT teams to help North American clients innovate faster and more efficiently. We are a people-first, purpose-driven company with a growing team, offering an inclusive culture and real growth opportunities.
Apply SRE principles to improve reliability, scalability, and performance of production systems.
Design and implement automation to reduce operational toil and improve engineering efficiency.
Lead incident response and develop sustainable solutions for complex production issues.
The hiring company is a technology organization focused on reliability and operational excellence. They offer a fully remote, collaborative environment with opportunities for technical leadership and career growth.
Design, build, and operate AWS infrastructure across multiple regions using Kubernetes, Terraform, and Helm.
Own large-scale object storage environments exceeding 15 petabytes, optimizing reliability, performance, scalability, and cost.
Build an internal developer platform for self-service infrastructure, strengthen security practices, and lead incident response and postmortems.
Our partner builds an AI-powered sports media platform handling petabyte-scale media storage and millions of minutes of video monthly. It operates as a remote-first, international, engineering-led team with significant autonomy and a focus on high availability and security.