Design, build, and maintain automation and tooling to reduce operational toil.
Actively participate in the incident-management lifecycle, including detection, escalation, mitigation, and post-incident review.
Provide an SRE point of view on capacity planning, resilience testing, and modernization of legacy workloads.
Seismic is the Go-To-Market Performance company, helping organizations turn strategy into revenue through an AI-powered revenue execution platform. Trusted by 2,500 organizations and over 3.5 million users globally, Seismic is headquartered in San Diego with offices across North America, Europe, and Asia-Pacific, fostering an inclusive culture.
Lead Cloud Platform and SRE teams to scale securely and efficiently.
Drive infrastructure strategy, including Kubernetes (GKE) clusters and developer platform.
Champion SRE culture with SLOs, error budgets, and observability.
Prolific builds human data infrastructure for AI development. The company is a fast-growing, mission-driven organization with cross-functional teams and a strong ownership culture.
Own the reliability, performance, and resilience of cloud environments (AWS, Kubernetes) and define SLOs across critical services.
Lead incident response, on-call rotation, and drive root cause analysis to ensure high production quality.
Build and maintain observability systems and automate operational toil using AI tools.
Garner partners with employers to redesign healthcare by using clinical metrics to identify top doctors and incentivize members to better care. The company has helped over 2.5 million people, saved $1B in costs, and doubled annually for five years, fostering a mission-driven, high-performance culture.
Acts as the strategic bridge between Cloud Operations, Product Management, Engineering, and other teams to drive service quality and operational excellence.
Drives large-scale transformation programs, promotes operational best practices, and ensures lessons learned translate into portfolio-wide improvements.
Champions automation, observability, and reliability standards while influencing engineering practices and product roadmaps.
Unit4 is a cloud company redefining ERP for mid-market people-centric organizations with over 40 years of heritage. They are a people-first community focused on trust, accountability, and growth, with a global team and a commitment to sustainability and inclusion.
Lead engineering teams for Porting, Hosting, and A2P Messaging, owning delivery roadmap and operational excellence.
Drive adoption of AI coding assistants and measure impact on team velocity, fostering psychological safety.
Recruit, develop, and retain a world-class engineering team, promoting a culture of candid feedback and career growth.
Twilio is shaping the future of communications, delivering innovative solutions to hundreds of thousands of businesses and empowering millions of developers worldwide. They are a remote-first company with a strong culture of connection and global inclusion, employing a diverse team making a global impact daily.
Define DevOps strategy and lead infrastructure architecture across multi-environment, multi-region cloud systems.
Architect and own scalable Kubernetes platforms, infrastructure as code, and DevSecOps implementation.
Drive platform reliability, performance SLAs, cost optimization, and lead complex migrations and AI/ML platform infrastructure.
Robots & Pencils is an applied AI engineering firm that designs and ships AI co-workers for enterprise operations. Founded in 2009, the company has delivery centers across Canada, the US, Eastern Europe, and Latin America, with teams averaging over 15 years of experience.
Lead both Site Reliability Engineering and Corporate IT, shaping technology operations strategy for a global SaaS environment.
Oversee platform reliability, incident response, compliance, and automation to reduce manual effort and improve efficiency.
Partner with Security and Compliance teams to support FedRAMP authorization and NIST 800-53 compliance frameworks.
Our partner is a high-growth, global SaaS environment seeking a Vice President of Enterprise Technology. They offer a fully remote executive role with approximately 10% travel and a focus on operational excellence.
Lead the architecture and implementation of complex cloud solutions across AWS and GCP.
Drive cloud automation and optimization initiatives to improve scalability and reliability.
Provide technical leadership and mentorship to engineers while collaborating with global teams.
The company focuses on cloud infrastructure and platform engineering. They operate with global teams and emphasize automation, security, and reliability.
Administer HIPAA-compliant AWS landing zone for data, analytics, and AI/ML workloads.
Manage Tableau server deployment on AWS, including patching, upgrades, and performance monitoring.
Lead cloud migration projects, implement IaC and CI/CD, and collaborate with security teams to ensure compliance.
Grady Health System is a leading public health system serving Atlanta and beyond for over 125 years. We are a large organization with a dedicated team of healthcare professionals focused on delivering compassionate, patient-centered care to all, regardless of ability to pay.
Lead the strategy and execution for Upbound's control plane experiences, including Crossplane upstream, enterprise distribution, and marketplace.
Manage multiple engineering teams and managers, scaling the organization to support product velocity and ecosystem expansion.
Drive ecosystem growth by partnering with cloud providers, ISVs, and platform partners to establish Upbound as the standard for programmable infrastructure.
Upbound is the creator and primary maintainer of Crossplane, the open-source project that extends Kubernetes into an AI-native cloud control plane. As a Series B startup backed by GV, Altimeter Capital, and Intel Capital, we serve Fortune 500 companies and platform engineers across 100+ countries, with over 100 million downloads.
Lead a two-month EKS modernization discovery for a high-scale consumer mobile platform and deliver a prioritized roadmap.
Drive execution across blast radius reduction, automated upgrades, compute right-sizing, and Graviton migration.
Act as Pod Leader in the U.S.-Based Virtual Operating Center and be the primary technical contact for customer infrastructure leadership.
EverOps is a premier Embedded Service Provider partnering with customer engineering teams on mission-critical infrastructure and cloud challenges. The company has been fully remote since day one and hires senior engineers who drive meaningful outcomes.
Own the technical relationship and be the trusted advisor for CTOs and platform teams.
Architect real solutions that translate customer constraints into deployable AI infrastructure.
Lead proofs of concept under real conditions to demonstrate operational fit.
Mirantis is a Kubernetes-native AI infrastructure company that helps organizations build scalable and secure infrastructure for modern AI workloads. The company has an installed base of 1,500 enterprise customers and values open source innovation, collaboration, and continuous growth.
Own and drive key infrastructure modernization initiatives toward container-orchestrated infrastructure.
Design and maintain infrastructure as code across multiple cloud providers.
Provide technical leadership and mentorship across the Systems Engineering team.
Intellum is the leader in corporate education technology, powering large learning programs for brands like Google, Meta, and Amazon. We are a remote-first company with a culture that values curiosity, creativity, perseverance, and kindness, and we invest in our people through personal development budgets and annual retreats.
Own the platform including GCP, Kubernetes, Temporal, GPU fleet, and deploy/rollback machinery.
Contribute to AI enablement substrate: GPU capacity, training/inference pipelines, and cost optimization.
Strengthen team practices through tooling, standards, tests, observability, and release processes.
Descript is building a simple, intuitive, fully-powered editing tool for video and audio — an editing tool built for the age of AI. They are a team of 150 backed by top investors like OpenAI and Andreessen Horowitz, with a culture that values collaboration and serendipitous discovery.
Design, build, and ship production services, APIs, and user-facing interfaces.
Build and operate production AI systems including RAG, fine-tuning, and inference optimization.
Architect AWS/GCP environments with Kubernetes and Terraform and control cloud/AI costs.
Motive empowers people who run physical operations with tools to make their work safer, more productive, and more profitable. Serving nearly 100,000 customers across industries, the company values a diverse and inclusive workplace.
Contribute to infrastructure automation and operational resilience across hybrid cloud and data center operations.
Implement closed-loop auto-remediation systems and SRE tooling to reduce manual intervention and incident resolution time.
Develop and maintain SLO frameworks, alerting policies, and Infrastructure-as-Code pipelines for reproducible deployments.
ServiceNow is the AI control tower for business reinvention, helping 85% of the Fortune 500 work smarter, faster, and better. They foster an AI-native culture where technology and talent are unstoppable together.
Provide technical leadership across cloud architecture, engineering, and modernization initiatives.
Design secure, scalable, and cost-effective cloud solutions on AWS, Azure, and multi-cloud environments.
Champion infrastructure as code, DevOps, governance, and reliability while mentoring technical teams.
Jobgether is an AI-powered job matching platform that connects candidates with hiring companies. The platform leverages automation to ensure fair and efficient recruitment processes.
Define and execute cloud platform strategy for Magnet's SaaS products and business objectives.
Lead cloud security, compliance (FedRAMP), and FinOps cost governance initiatives.
Build and develop high-performing cloud engineering teams while enabling AI adoption.
Magnet Forensics develops digital investigative software that acquires, analyzes, and shares evidence from computers and devices. It serves thousands of customers globally and fosters a culture of learning, integrity, and inclusion with employees around the world.
Lead and modernize Sectigo's global infrastructure organization with a focus on reliability and operational maturity.
Develop a measurable operating model using SLAs, SLOs, and key metrics to drive improvement.
Drive automation, AI-enabled operations, and closer collaboration with engineering teams.
Sectigo is an innovative provider of certificate lifecycle management (CLM) solutions, helping large brands simplify digital trust. With over 700,000 customers including 65% of the Fortune 500, they emphasize a culture of support, excellence, and teamwork.
Build and maintain the company's internal platform, driving operational excellence.
Collaborate with engineering squads to ensure applications are safe and reliable.
Take ownership of software infrastructure projects and provide off-hours support.
Loadsmart is a growth-stage logistics technology company valued at over $1 billion, using innovative technology to reinvent the freight industry. With headquarters in Chicago and a globally distributed remote team, it attracts top talent committed to driving meaningful change.