Lead a small Infrastructure Engineering team while balancing hands-on technical contributions and people management.
Drive infrastructure modernization, reliability, and operational excellence in a cloud-native environment.
Champion AI-native engineering practices to improve throughput and reduce operational toil.
The company is a technology partner that builds and operates high-scale production infrastructure platforms. They foster a culture of ownership, experimentation, and cross-functional collaboration, with a strong emphasis on continuous learning and practical problem-solving.
Own the technical direction and architecture of critical infrastructure domains, establishing scalable patterns and standards.
Lead complex, multi-team infrastructure initiatives from design through implementation and production operation.
Design and evolve AWS and Kubernetes infrastructure to enable teams to build and deploy systems reliably at scale.
We provide innovative identity and risk solutions, empowering institutions and individuals to transact with confidence. Our company is backed by world-class investors including Craft Ventures and Andreessen Horowitz, with offices across the US and India, and we are growing extremely quickly.
Lead and grow a team of platform engineers, coaching them on infrastructure and cloud challenges.
Drive the platform roadmap, balancing reliability, cost, security, and developer experience with AWS and Kubernetes.
Partner cross-functionally to align platform priorities with business goals and ensure system reliability.
PerfectServe is a leading provider of clinical communication and physician scheduling solutions in the health IT space. The company has 400+ employees and 30,000+ customers, with over $100 million in annual revenue, and has received multiple Best in KLAS awards.
Lead a distributed team of engineers, mentoring them in technical and professional growth.
Manage technical execution, prioritization, and cross-team collaboration on infrastructure projects.
Foster an inclusive, high-performance team environment with continuous feedback and career development.
Pilot provides small businesses with dedicated finance experts and custom software for accurate bookkeeping and financial management. With over 3,000 customers and $170 million in funding, the company values high trust, ownership, and continuous iteration.
Define and drive the vision and multi-quarter roadmap for the infrastructure foundations team, tying investments to business outcomes.
Lead and mentor a team of infrastructure engineers, fostering ownership, collaboration, and technical excellence.
Own the reliability and safety of the foundational AWS layer, including account provisioning, core networking, and IAM access.
Affirm is reinventing credit to make it more honest and friendly, offering consumers the ability to buy now and pay later without hidden fees or compounding interest. As a publicly traded company, Affirm fosters a culture of thorough technical design review, operational excellence, and capable incident response.
Design, build, and maintain robust, scalable, and secure infrastructure systems supporting Laurel's AI-driven platform.
Manage and optimize cloud infrastructure (AWS and Azure), Kubernetes orchestration, and CI/CD pipelines to increase deployment frequency and reliability.
Implement comprehensive observability, monitoring, and alerting to maintain system health and partner with engineering teams to optimize performance and cost-efficiency.
Laurel is an AI Time platform for professional services firms, automating work time capture and connecting time data to business outcomes for clients like EY and Crowell & Moring. The company comprises top AI, product, and engineering talent, is VC-backed by Google Ventures and IVP, and fosters an inclusive, ambitious culture.
Lead the discovery, design, and delivery of complex reliability and infrastructure initiatives, translating ambiguous problems into robust technical solutions.
Define and operate reliability practices including SLOs, SLIs, error budgets, alerting strategies, and observability standards.
Manage and scale production Kubernetes environments, build cloud infrastructure on AWS, and mentor less-senior engineers.
Jobgether is an AI-powered recruitment platform that connects candidates with partner companies through automated matching. They are a globally distributed organization focused on fair, efficient hiring processes.
Own and scale cloud infrastructure including compute, networking, storage, and data systems.
Lead BYOC and private cloud deployments with infrastructure-as-code and GitOps foundations.
Establish reliability through service-level objectives, observability, and incident response processes.
A technology company builds a developer-focused platform with scalable cloud infrastructure. This is a remote-first opportunity with a small, autonomous engineering team operating in North America, LATAM, and Europe, offering high autonomy and ownership.
Lead Remote's SRE team owning Kubernetes, AWS, PostgreSQL, CI, and observability.
Balance 60% hands-on technical work with 40% people leadership and career growth.
Drive a maturing reliability practice including SLOs, incident response, and on-call.
Remote is a global employment platform that helps companies recruit, pay, and manage international teams. The company is fully remote with a future-focused, async culture and employees across six continents.
Design, build, and optimize multi-region, high-availability AWS infrastructure.
Drive resiliency and automation using GitOps, modern CI/CD, and Infrastructure as Code.
Build end-to-end telemetry and own incident management to harden reliability.
VGS is the world's leader in payment tokenization, trusted by the most innovative AI and Fortune 500 companies. They are a remote-first company with a culture of ownership, collaboration, and continuous learning.
Lead engineering teams for Porting, Hosting, and A2P Messaging, owning delivery roadmap and operational excellence.
Drive adoption of AI coding assistants and measure impact on team velocity, fostering psychological safety.
Recruit, develop, and retain a world-class engineering team, promoting a culture of candid feedback and career growth.
Twilio is shaping the future of communications, delivering innovative solutions to hundreds of thousands of businesses and empowering millions of developers worldwide. They are a remote-first company with a strong culture of connection and global inclusion, employing a diverse team making a global impact daily.
Lead Cloud Platform and SRE teams to scale securely and efficiently.
Drive infrastructure strategy, including Kubernetes (GKE) clusters and developer platform.
Champion SRE culture with SLOs, error budgets, and observability.
Prolific builds human data infrastructure for AI development. The company is a fast-growing, mission-driven organization with cross-functional teams and a strong ownership culture.
Create and test reliable cloud infrastructure services supporting Webflow's product range.
Lead initiatives to reduce triage load, increase reliability, and handle growing customer scale.
Collaborate with product engineering teams to deliver new solutions and improve existing services.
Webflow is an agentic web marketing platform that helps modern marketing teams build, manage, and optimize high-performing web experiences. The company values grit, speed, and craft, fostering a culture of ownership and continuous improvement.
Own and drive key infrastructure modernization initiatives toward container-orchestrated infrastructure.
Design and maintain infrastructure as code across multiple cloud providers.
Provide technical leadership and mentorship across the Systems Engineering team.
Intellum is the leader in corporate education technology, powering large learning programs for brands like Google, Meta, and Amazon. We are a remote-first company with a culture that values curiosity, creativity, perseverance, and kindness, and we invest in our people through personal development budgets and annual retreats.
Lead, coach, and develop a team of SREs focused on developer experience and engineering productivity.
Drive the evolution of CI/CD platforms, deployment automation, and release workflows.
Partner with engineering leaders to enable faster, more reliable software delivery.
Sprout Social is a global leader in social media management and analytics software. Founded in 2010, the company has a hybrid team of 1,400 people across the globe and is consistently recognized as a best place to work.
Improve system availability, scalability, and resilience across Flowcode's platforms.
Manage and scale core AWS infrastructure through Infrastructure as Code (Terraform) and enhance disaster recovery.
Oversee monitoring, logging, and alerting infrastructure, and develop high-signal metrics and dashboards.
Flowcode is a technology company specializing in QR code and smart link solutions for offline-to-online engagement. The company is a growth-stage startup seeking high-performing individuals who thrive in a fast-paced, demanding environment.
Own and scale the cloud infrastructure behind our open-source platform: compute, networking, and the data layer.
Lead BYOC: turn customer-cloud deployments into a real product, with provisioning, upgrades, and observability that scale past bespoke work per deal.
Make reliability a product feature: meaningful SLOs, and an incident process people trust.
Nango is a developer infrastructure company that provides API access for agents and apps, enabling AI applications to connect to the real world through integrations. With over 400 paying customers and a team of 14 from top tech companies like AWS, GitHub, and Okta, they are a YC-backed, multi-million ARR company that values ownership and autonomy.
Drive complex infrastructure migrations and build platform tooling and automation across multiple production environments.
Support development teams by consulting on infrastructure needs and improving observability and incident response.
Provide operational support and maintain platform reliability through structured debugging and on-call rotations.
PENN Entertainment is North America's leading provider of integrated entertainment, sports content, and casino gaming experiences. We operate across numerous locations in North America and foster a culture that cares about career growth and skill expansion.
Lead and mentor a distributed engineering team across US and EU, fostering growth and collaboration.
Drive proactive ownership and engineering excellence for a cloud platform at exabyte scale.
Architect global infrastructure expansions across multi-cloud and compliance environments.
New Relic is an intelligent observability platform that helps companies gain insight into their complex systems and thrive in an AI-first world. As a global team of innovators, we foster a diverse, welcoming, and inclusive environment where employees can be their authentic selves.
Own the technical architecture and evolution of core infrastructure.
Engineer for scale and performance through capacity modeling and bottleneck diagnosis.
Participate in on-call rotation and drive technical recovery during incidents.
Sezzle is a fintech company that revolutionizes shopping through interest-free installment plans, blending cutting-edge technology with financial empowerment. It has a dynamic and innovative team culture focused on shaping the future of fintech and retail.