Improve system availability, scalability, and resilience across Flowcode's platforms.
Manage and scale core AWS infrastructure through Infrastructure as Code (Terraform) and enhance disaster recovery.
Oversee monitoring, logging, and alerting infrastructure, and develop high-signal metrics and dashboards.
Flowcode is a technology company specializing in QR code and smart link solutions for offline-to-online engagement. The company is a growth-stage startup seeking high-performing individuals who thrive in a fast-paced, demanding environment.
Set reliability strategy and SLO culture that scales across engineering teams.
Own platform architecture, event-driven messaging, and observability for a global payments platform.
Lead chaos engineering, incident response, and mentorship for the most complex production challenges.
Yuno is an AI-native operating system for global commerce, connecting merchants to pay-ins, payouts, fraud prevention, and stablecoins via a single API. It powers payment infrastructure for global brands like McDonald's and GoFundMe, with a culture of remote work and AI-driven innovation.
You define architecture and operational standards for managed services across AWS. - You serve as the final technical escalation point for complex customer situations. - You shape Honeycomb's open source strategy in the OpenTelemetry ecosystem and mentor engineers.
Honeycomb is a service for observability, defining developer tools for the near and present future. We are a fully distributed company of over 200 talented and inclusive bees, named to Forbes' America's Best Startups of 2022 and 2023.
Own and scale the cloud infrastructure behind our open-source platform: compute, networking, and the data layer.
Lead BYOC: turn customer-cloud deployments into a real product, with provisioning, upgrades, and observability that scale past bespoke work per deal.
Make reliability a product feature: meaningful SLOs, and an incident process people trust.
Nango is a developer infrastructure company that provides API access for agents and apps, enabling AI applications to connect to the real world through integrations. With over 400 paying customers and a team of 14 from top tech companies like AWS, GitHub, and Okta, they are a YC-backed, multi-million ARR company that values ownership and autonomy.
Take an active role as co-owner of production services to ensure they are built, maintained, and operated in a reliable and scalable way.
Collaborate with Software Engineering to drive operational improvements through metric-driven analysis and help scale AWS and Kubernetes infrastructure.
Participate in a weekly on-call rotation to investigate and resolve potential system issues, and automate routine tasks in at least two programming languages.
Zerohash is the leading crypto and stablecoin infrastructure platform, powering the next generation of financial services for banks, brokerages, fintechs, and payment companies. Founded in 2017, the company has raised over $280 million from top venture firms and strategic investors, and is trusted by global brands like Morgan Stanley and Stripe, operating with a compliance-first approach.
Architect and maintain critical cloud platform components on AWS EKS with high availability and automated resilience.
Establish SRE standards including SLO/SLI tracking, error budget frameworks, and automated operational tooling.
Design and implement OpenTelemetry capture pipelines for telemetry data feeding downstream platforms.
Inflect is a US-based advisory and marketplace that revolutionizes how companies buy and sell digital infrastructure services. They operate with a focus on high-impact consulting and autonomous work.
Design and build resilient AWS and Kubernetes platforms to improve reliability, scalability, and security.
Define SLOs, build observability, automate operational work, and lead incident response and post-incident reviews.
Partner with engineering, platform, security, and QA teams to establish reliability standards and optimize cost.
Electric Power Engineers (EPE) provides consulting expertise and energy intelligence software solutions for power and energy clients, focusing on renewable energy and grid modernization. With over half a century in the industry, the company fosters innovation and collaboration, working with industry leaders to build a secure and resilient grid.
Design and maintain highly available, scalable systems to ensure exceptional customer experiences.
Drive automation and eliminate operational toil through self-service tooling and process improvements.
Lead incident response and mentor engineers to improve reliability practices.
Redzone provides a connected workforce solution for manufacturers to improve plant efficiency and worker productivity. The company is part of QAD Inc. and fosters a collaborative, customer-focused culture with a strong technology team.
Design and maintain infrastructure-as-code patterns using Terraform and Kubernetes for scalable deployments.
Build monitoring, logging, and alerting systems, lead incident response, and drive continuous reliability improvements.
Embed security into infrastructure and optimize performance, costs, and automation across the platform.
Remote enables global employment compliantly, allowing businesses to recruit, pay, and manage international teams. With a future-focused culture and fully remote team across six continents, it builds an innovative HR platform with automation and AI.
Apply SRE principles to improve reliability, scalability, and performance of production systems.
Design and implement automation to reduce operational toil and improve engineering efficiency.
Lead incident response and develop sustainable solutions for complex production issues.
The hiring company is a technology organization focused on reliability and operational excellence. They offer a fully remote, collaborative environment with opportunities for technical leadership and career growth.
Lead the discovery, design, and delivery of complex reliability and infrastructure initiatives, translating ambiguous problems into robust technical solutions.
Define and operate reliability practices including SLOs, SLIs, error budgets, alerting strategies, and observability standards.
Manage and scale production Kubernetes environments, build cloud infrastructure on AWS, and mentor less-senior engineers.
Jobgether is an AI-powered recruitment platform that connects candidates with partner companies through automated matching. They are a globally distributed organization focused on fair, efficient hiring processes.
Lead enterprise-wide reliability and infrastructure projects with high autonomy, architecting scalable solutions and driving SRE best practices.
Partner cross-functionally with Engineering, Product, and Customer Success to align reliability goals with business objectives and communicate complex concepts to diverse audiences.
Provide tier 2/3 technical support to enterprise customers, conduct technical onboarding, and act as a trusted advisor for platform architecture.
Veza is the pioneer in identity security, providing an Access Graph platform that maps identity ecosystems across users, groups, roles, policies, and resources. With over 30 billion access permissions under management and now part of ServiceNow, Veza combines enterprise scale with security innovation.
Design, build, and operate shared cloud infrastructure using AWS, Kubernetes, Terraform, Databricks, and Cloudflare.
Deliver SRE and DevOps initiatives to improve reliability, scalability, observability, and deployment safety.
Build reusable infrastructure modules, automation, and self-service workflows to reduce manual work and improve developer experience.
YipitData is the leading market research and analytics firm for the disruptive economy, recently raising up to $475M from The Carlyle Group at a valuation over $1B. We analyze billions of alternative data points daily and have been recognized as one of Inc’s Best Workplaces, cultivating a people-centric culture focused on mastery, ownership, and transparency.
Own Primer's internal developer platform end to end, including CI/CD pipelines, deployment workflows, and self-service tooling.
Build the human-AI development loop, creating tooling and automation for coding agent workflows.
Treat developer productivity as a measurable system, using frameworks like DORA to identify and fix delivery bottlenecks.
Primer provides a unified infrastructure for global payments, enabling finance and payments teams to reduce complexity and capture revenue. Backed by top investors like Sofina and Accel, they operate as a remote-first, async culture with high autonomy and low bureaucracy.
Design and implement reliability strategies for distributed systems across AWS and GCP, defining SLIs and SLOs.
Build and enhance observability solutions using monitoring, logging, tracing, and alerting platforms.
Lead incident response, root cause analysis, and postmortem processes to improve system reliability.
We specialize in creating high-performing nearshore IT teams to help North American clients innovate faster and more efficiently. We are a people-first, purpose-driven company with a growing team, offering an inclusive culture and real growth opportunities.
You will ensure the reliability and high availability of Tenable's cloud products in cloud environments.
You will respond to support escalations and troubleshoot complex technical problems.
You will develop software, tools, and scripts to automate deployment and monitoring of production systems.
We are the Exposure Management company, trusted by over 40,000 organizations to understand and reduce cyber risk. Our global team supports 65% of the Fortune 500 and 50% of the Global 2000, with a culture of belonging, respect, and excellence.
Empower engineers on other teams by maintaining monitoring tooling and collaborating on observability best practices.
Enhance reliability of Kubernetes applications through resource optimization, streamlined upgrades, and scalability.
Participate in on-call and incident response processes, occasionally diving into application code to debug production issues.
Webflow is the agentic web marketing platform for modern marketing teams, helping organizations build, manage, and optimize high-performing web experiences. It serves over 2 million users worldwide across 190 countries, with tens of thousands of projects launched each month, and fosters a culture of grit, speed, and craft.
Own observability for critical product journeys, defining SLIs/SLOs and building metrics, dashboards, and alerts.
Act as first responder for production incidents, investigating signals and mitigating issues independently.
Work within a cross-functional squad of 6-8 engineers to improve reliability, monitoring, and incident response processes.
Feeld is a dating app creating a safer and more inclusive space for exploring relationships and sexuality. They have a distributed engineering team of around 50 people across Europe and the US, working in small autonomous squads.
Help design, build, and operate the Kubernetes platform used across PulsePoint.
Own reliability, observability, and incident response across platform services.
Build infrastructure automation and GitOps workflows to reduce operational toil.
PulsePoint sits at the intersection of healthcare and adtech, helping brands interpret health signals using real-world data. With over 300 employees, the company is a post-acquisition profitable leader in the US healthcare ad market, known for a flat hierarchy and high engineering bar.
Own and scale cloud infrastructure including compute, networking, storage, and data systems.
Lead BYOC and private cloud deployments with infrastructure-as-code and GitOps foundations.
Establish reliability through service-level objectives, observability, and incident response processes.
A technology company builds a developer-focused platform with scalable cloud infrastructure. This is a remote-first opportunity with a small, autonomous engineering team operating in North America, LATAM, and Europe, offering high autonomy and ownership.