Implement Astronomer’s software and services in the core of some of the world’s largest businesses and organizations.
Guide our largest and most complex customers in their Apache Airflow journeys.
Act as the core interface between the Customer, Sales, and Product to ensure that the broader solution around the platform is solving pain points and bringing value to customers.
Astronomer empowers data teams with their DataOps platform Astro, powered by Apache Airflow. They are trusted by over 800 enterprises and foster a culture of diversity and equal opportunity.
Own executive-level relationships and strategic engagement for a portfolio of enterprise and high-value customers.
Collaborate with customer teams on Sysdig deployments, architecture, and operational best practices.
Lead Customer Business Reviews and strategic touchpoints with both technical and non-technical stakeholders.
Sysdig is a cloud security company that created Falco, the open standard for cloud threat detection, and leads the cloud security market with runtime insights and open innovation. Trusted by over 60% of the Fortune 500, Sysdig is recognized as a Best Place to Work and one of Deloitte's fastest-growing companies.
Provide Apache Airflow expertise directly to customers, solving challenging problems and optimizing configurations.
Learn and build expertise across software engineering disciplines including Airflow, Kubernetes, and Cloud Engineering.
Own the customer experience, working directly with customers to prioritize and resolve issues, and provide guidance on the path to production.
Astronomer empowers data teams to bring mission-critical software, analytics, and AI to life and is the company behind Astro, the industry-leading unified DataOps platform powered by Apache Airflow. Trusted by more than 800 of the world's leading enterprises, Astronomer lets businesses do more with their data.
Report to the Manager of Customer Reliability and provide senior level technical support on customer issues.
Work with development engineering to track escalations, bugs, and feature requests.
Develop technical troubleshooting sessions, trainings, and maintain internal knowledge base articles.
Sysdig is a cloud security company that stops attacks in real-time using runtime insights and open source Falco. It is a well-funded startup with a large enterprise customer base, recognized as a "Best Places to Work" fostering an inclusive and diverse culture.
Deploy and operate Blitzy's self-hosted platform within a customer-controlled, secure cloud environment.
Own the Kubernetes-based deployment, releases, upgrades, capacity planning, and performance benchmarking.
Serve as the on-account technical presence, partnering with customer infrastructure and security teams.
We are an AI software development platform that autonomously builds custom software for enterprises. Backed by tier 1 investors and led by two co-founders, we are one of the fastest-growing U.S. companies with a culture of speed and customer focus.
You define architecture and operational standards for managed services across AWS. - You serve as the final technical escalation point for complex customer situations. - You shape Honeycomb's open source strategy in the OpenTelemetry ecosystem and mentor engineers.
Honeycomb is a service for observability, defining developer tools for the near and present future. We are a fully distributed company of over 200 talented and inclusive bees, named to Forbes' America's Best Startups of 2022 and 2023.
Own, operate and evolve our Cloud and Kubernetes based Platform.
Collaborate with Product teams to enable and empower them to build and own their services.
Build tooling and AI powered integrations that reduce cognitive load of complex infrastructure operations.
AlphaSense provides AI-driven market intelligence and search to help companies remove uncertainty from decision-making. Founded in 2011, the company has over 2,000 employees globally and is headquartered in New York City.
Define and execute the technical strategy for observability, platform infrastructure, and operational excellence.
Lead the design and evolution of scalable, secure, reliable cloud-native platforms and distributed systems.
Establish reliability best practices including SLIs, SLOs, error budgets, and automation initiatives.
The company is a technology organization that builds and operates large-scale cloud infrastructure. It fosters a collaborative culture centered on innovation, ownership, and impact.
Design and improve monitoring, logging, distributed tracing, dashboards, alerting, SLIs, and SLOs for production health.
Build and maintain automation, internal tools, and CI/CD systems to increase engineering efficiency and support reliable deployments.
Own complex production incidents from detection to resolution, turning learning into durable improvements and reducing recurring incidents.
Filevine is a Legal AI company delivering Legal Operating Intelligence for the future of legal work. It has earned recognition from Deloitte and Inc. as one of the most innovative and fastest-growing technology companies in the country.
Design, build, and maintain robust, scalable, and secure infrastructure systems supporting Laurel's AI-driven platform.
Manage and optimize cloud infrastructure (AWS and Azure), Kubernetes orchestration, and CI/CD pipelines to increase deployment frequency and reliability.
Implement comprehensive observability, monitoring, and alerting to maintain system health and partner with engineering teams to optimize performance and cost-efficiency.
Laurel is an AI Time platform for professional services firms, automating work time capture and connecting time data to business outcomes for clients like EY and Crowell & Moring. The company comprises top AI, product, and engineering talent, is VC-backed by Google Ventures and IVP, and fosters an inclusive, ambitious culture.
Be on an on-call rotation responding to production incidents and support service engineers.
Run infrastructure with Ansible, Puppet, Terraform, and Kubernetes, making monitoring alert on symptoms.
Design and maintain core infrastructure scaling to hundreds of thousands of concurrent users.
Our client's Cloud Operations team is expanding its SRE function, keeping user-facing services and production systems running smoothly. The team specializes in systems like networking, Linux kernel, and distributed systems, blending pragmatic operations with software engineering.
Own cloud infrastructure across AWS and GCP, including Kubernetes, networking, databases, and CI/CD pipelines.
Scale single-tenant deployments and build observability, incident response, and compliance practices.
Manage infrastructure cost, improve developer experience, and contribute to backend systems at the infrastructure-application intersection.
Elicit is an AI research assistant that uses language models to help researchers with literature review and evidence synthesis. The company is a ~30-person Public Benefit Corporation with a high-agency, low-bureaucracy culture.
Support and improve production and development infrastructure for multiple teams handling high traffic products.
Help developers debug intricate issues and architect scalable solutions across cloud and on-premise environments.
Promote CICD strategies, document processes, and mentor junior DevOps engineers.
We are a tech pioneer offering world-class adult entertainment and games on safe platforms. With an international team of dynamic innovators, we have offices in Montreal, Austin, and Nicosia and celebrate diversity and inclusion.
Own the infrastructure end-to-end for ScaleOps' self-hosted and SaaS platforms.
Manage cloud infrastructure across AWS, GCP, and Azure, including networking, security, and compute.
Collaborate with customers and internal teams to ensure rapid feature delivery without compromising reliability.
ScaleOps is redefining autonomous cloud and AI infrastructure, freeing DevOps from manual resource management. Backed by $210M+ in funding, they are trusted by leading enterprises and Fortune 100 companies, with a fast-paced, innovative culture.
Monitor, triage, and resolve customer-reported incidents within defined SLAs.
Serve as primary point of contact for technical issues related to installation, Helm configuration, and integrations.
Troubleshoot Kubernetes-related issues such as ingress, SSO, and other integrations.
ScaleOps redefines autonomous cloud and AI infrastructure, freeing DevOps engineers from manual resource management. The company is backed by over $210M in funding and trusted by leading enterprises including Adobe and Coinbase.
Design, build, and operate reliable infrastructure supporting AI-powered products.
Own and improve Kubernetes environments and cloud infrastructure.
Enhance production reliability through observability, automation, and incident response.
The company builds advanced AI-driven products and services. It values engineering excellence, autonomy, and individual contribution, with a global team of skilled engineers.
Lead technical and managerial direction for the SRE team, defining reliability, observability, and operational excellence strategy.
Coordinate critical incident responses and root cause analysis, collaborating with architecture, development, security, and product teams.
Drive automation, continuous improvement, and adoption of SRE, DevOps, and Platform Engineering best practices.
Experian is a global data and technology company that drives opportunities for people and businesses worldwide. With 25,200 employees in 32 countries, it has a people-centric, inclusive culture recognized by awards such as World's Best Workplaces™ 2025.
Build and operate the control plane for automated cluster deployment from bare metal to customer-ready.
Manage machine lifecycle including joining, wiping, verifying, and rejoining between tenants.
Operate Kubernetes, Postgres, and custom operators across the fleet, scaling from tens to thousands of nodes.
Andromeda provides scaled AI infrastructure for startups, managing compute across numerous capacity providers. The company operates tens of thousands of GPUs for 80+ customers and fosters an inclusive environment.
Design and maintain CI/CD and MLOps pipelines for software and machine learning models.
Build and scale cloud-native infrastructure using Kubernetes, Docker, and GPU clusters.
Champion Infrastructure as Code and observability to ensure high availability and governance.
Bitdeer is a world-leading technology company for AI and Bitcoin mining infrastructure, providing comprehensive solutions and cloud capabilities. Headquartered in Singapore, the company has deployed data centers across multiple countries and fosters a culture of innovation.
Build systems for declarative application and infrastructure lifecycle management, including CI/CD, Kubernetes, and service inventory.
Prioritize and troubleshoot infrastructure issues to minimize downtime and respond to alerts efficiently.
Contribute to setting the SRE team's direction and streamline automation of infrastructure processes.
Counterpart Health develops Counterpart Assistant, an AI-enabled primary care tool that supports physicians in chronic disease management. It is a subsidiary of Clover Health, with a remote-first culture and a focus on value-based care through technology.