Design and build automated reliability and self-healing systems to protect production at scale.
Own and improve incident management tooling and on-call health, reducing alert noise and empowering teams.
Develop and evolve observability infrastructure, including monitoring, alerting, SLOs, and performance regression detection.
Samsara is the pioneer of the Connected Operations Cloud, enabling organizations to harness IoT data for improved safety, efficiency, and sustainability. As a public company with over 2.3 million IoT devices deployed, we foster a culture of growth mindset and collaboration.
Design and maintain Grafana dashboards and telemetry visualizations to monitor system performance and platform health.
Develop and maintain modular Ansible playbooks to automate infrastructure provisioning and configuration.
Configure observability solutions with Prometheus monitoring and alerting, and participate in Agile ceremonies.
Miratech is a global IT services and consulting company that helps visionaries change the world by supporting digital transformation for large enterprises. With nearly 1,000 full-time professionals across 5 continents and 25 countries, the company has a culture of Relentless Performance with a 99% project success rate and over 25% annual growth.
Build and maintain observability across the platform in Datadog, including dashboards, monitors, APM, and log pipelines.
Participate in on-call rotation and incident response, driving blameless post-incident reviews and automating toil.
Leverage AI tools to accelerate debugging, generate runbooks, and build automation for operational efficiency.
IPSY is a beauty subscription platform that connects brands and consumers through curated beauty products. It is a remote-first company with a focus on community and engagement.
Own and evolve observability strategy including monitoring, alerting, dashboards, logging, and distributed tracing.
Define and manage SLIs, SLOs, and reliability metrics, improving MTTD and MTTR through automation.
Build and maintain reliable cloud infrastructure on AWS and Kubernetes while mentoring engineers on SRE best practices.
Filevine is a Legal AI company delivering Legal Operating Intelligence for legal work. Fueled by a team of exceptional collaborators and innovators, Filevine’s rapid growth has earned AI awards and recognition from Deloitte and Inc. as one of the most innovative and fastest-growing technology companies in the country.
Lead operational excellence, reliability, and support of enterprise AI and data platforms, ensuring stability, scalability, and observability.
Design and implement automation, monitoring, and operational tooling for AI/ML platforms including Palantir Foundry, AWS Bedrock, and SageMaker.
Serve as a senior escalation point for complex production issues, driving root cause analysis and improving platform stability.
CSAA Insurance Group, a AAA insurer, offers personal lines of property and casualty insurance to AAA members across 23 states and DC. Founded in 1914, they are one of the top personal lines insurers in the US with over 3,800 employees, known for a values-based culture and recognition in leadership development and community involvement.
Run day-to-day IT help desk for ~75 employees, resolving 75+ tickets per week across software access, account provisioning, and laptop issues.
Own the full laptop lifecycle from vendor pricing to deployment and retrieval.
Build automations and internal tools using LLM APIs to streamline IT workflows.
Frontier is a subsidiary of Fresh Prints that helps companies build full-time, cross-functional teams abroad. The company is fully-remote with around 150 employees, mostly in India and the Philippines.
Own observability end to end and define how we measure reliability.
Own CI/CD pipelines and make shipping fast and safe.
Footprint builds Percy, an AI agent that runs financial crime investigations end to end. The company is backed by QED, Index, and other investors, and its small, senior team ships fast and grew revenue 5x in the past year.
Build core backend services for context ingestion, indexing, and retrieval in a new AI-native data intelligence system.
Design scalable multi-tenant SaaS architecture and agent-facing APIs for reliable context retrieval.
Operate what you build using observability tools, and contribute to technical direction and architecture.
Grafana Labs is the company behind the open observability cloud, Grafana, which helps organizations monitor and act on their data. They are a 100% remote company with over 1,600 team members across 40+ countries, fostering a culture of transparency, collaboration, and open source.
Own production operability by debugging complex issues, improving system visibility, and eliminating recurring problems at the source.
Improve mean time to detect and resolve issues, and reduce recurrence rates through code fixes and architecture improvements.
Partner with product teams to feed production learnings back into design and development, reducing support and incident load.
One Identity helps organizations secure, manage, and analyze information and infrastructure to drive innovation. With a global team, we offer a collaborative environment where employees build products at scale and grow their careers.
Design and operate cloud development environments and execution platforms for engineers and AI agents.
Scale Kubernetes-based systems and improve developer workflows across build, test, and dependency management.
Diagnose production issues and set technical direction for a high-impact infrastructure team.
Airbnb is a global hospitality platform connecting hosts with guests, offering unique stays and experiences. The company has over 5 million hosts and prioritizes inclusion and belonging.
Define and implement SLIs/SLOs for critical services, lead incident response, and conduct blameless postmortems to drive systemic improvements.
Design and improve monitoring and alerting with Prometheus and Grafana, build internal tooling, and automate operational workflows to reduce toil.
Partner with cross-functional engineering teams to improve system resilience, contribute to architectural discussions, and strengthen production readiness standards.
Runpod provides a cloud platform for AI development, used by over one million developers for training, fine-tuning, and deploying AI models. The company is a small, remote-first team that closed a $100M Series A in June 2026, emphasizing ownership, speed, and impact at scale.
Partner directly with the founder on operational priorities and company initiatives.
Create accountability systems, workflows, and processes across the organization.
Develop and maintain dashboards, reporting, and performance tracking systems.
Click and Love is a rapidly growing affiliate marketing and content business that blends organic content, paid advertising, and strategic partnerships. They are a scaling company seeking an operational leader to bring structure and accountability.
Architect and build the systems that encode how GCS operates, including account health observability, intelligent routing, knowledge retrieval, and workflow automation.
Build AI workflows using Anthropic Claude API, Claude Agent SDK, and Claude Code, connecting agents to Slack, Jira, Zendesk, Grafana, Airtable, and internal systems.
Operate within CXO's governance model, maturing versioning, evaluation frameworks, A/B testing, and tying technical performance to GCS outcomes like time to value and resolution rate.
Twilio is a communications platform company delivering innovative solutions to hundreds of thousands of businesses and empowering millions of developers worldwide. They are a remote-first company with a strong culture of connection and global inclusion, committed to fostering a diverse and inclusive workplace.
Design and deploy a modern observability stack including logging, metrics, and distributed tracing across hundreds of services.
Build alerting policies and incident response workflows to reduce manual escalations and improve mean time to detect and resolve.
Automate toil and set SRE standards while mentoring engineers on observability tooling.
WeightWatchers is a global digital health company and the world's #1 doctor-recommended behavioral weight health program. They have a cloud architecture supporting hundreds of services and are expanding their platform engineering team.
Developing standards, infrastructure, and automation for dashboards, alerts, and monitors as code.
Partnering with development teams to establish production readiness and operational readiness.
Building tooling to automate observability and operational workflows, eliminating manual toil.
Glia is the #1 Banking AI platform, empowering community and regional financial institutions with an AI workforce. The company is trusted by over 700 banks and credit unions and has a remote-first culture with offices in Estonia.
Map operational workflows across Sales and broader organization to identify friction and build scalable AI-powered solutions.
Partner with Revenue Operations to design and deploy intelligent workflows, automations, and playbooks.
Drive adoption by creating documentation, guardrails, and training to ensure systems are embraced and trusted.
1mind builds the next generation of revenue teams with a platform that deploys AI-powered Superhumans to serve buyers across the entire journey. As a rapidly growing startup, the company fosters a high-ownership, collaborative culture where every team member drives impact.
Design, operate, and continuously tune platform monitoring across Dynatrace and Splunk.
Own the Single Pane of Glass dashboard and lead all P1/P2 incident responses.
Manage on-call rotation and deliver monthly SLA reports.
We are a growing Service-Disabled Veteran Owned company providing IT solutions to federal clients. We embrace remote work and invest in our employees' growth.
Serve as the escalation point for complex issues, owning them from investigation through resolution.
Write and run SQL queries to investigate transaction discrepancies and data problems.
Use AI tools like Claude to accelerate root-cause analysis and document solutions.
Sezzle is a fintech company revolutionizing shopping with interest-free installment plans. They are building an innovative, dynamic team to shape the future of fintech and retail.
Design and implement enterprise monitoring and observability strategies using AI-driven automation.
Apply machine learning techniques to improve incident detection, prediction, and resolution.
Collaborate with IT teams and stakeholders to optimize event management and operational efficiency.
The partner company specializes in enterprise IT monitoring and observability, leveraging AI and automation. It operates with a global team and offers a fully remote, contract-based work environment.
Architect and manage scalable cloud infrastructure for the 3D rendering platform.
Build and maintain CI/CD pipelines with a focus on deployment reliability.
Monitor performance metrics and proactively address bottlenecks before they become incidents.
Homekynd builds the spatial intelligence layer for enterprise retail, transforming photos into 3D room models for immersive furniture visualization. They are a remote-first team on a fast build timeline, seeking engineers who want real ownership over hard problems.