Define and lead the end-to-end observability strategy covering logging, metrics, tracing, and alerting.
Architect and evolve a unified observability platform ensuring scalability and reliability.
Build and lead a high-performing observability engineering team with strong technical standards.
The company operates a high-scale developer-facing platform focused on reliability and performance. It is a remote-first organization with a globally distributed engineering team committed to building best-in-class developer infrastructure.
Developing standards, infrastructure, and automation for dashboards, alerts, and monitors as code.
Partnering with development teams to establish production readiness and operational readiness.
Building tooling to automate observability and operational workflows, eliminating manual toil.
Glia is the #1 Banking AI platform, empowering community and regional financial institutions with an AI workforce. The company is trusted by over 700 banks and credit unions and has a remote-first culture with offices in Estonia.
Design and implement SOAR playbooks and security automation workflows to streamline SOC operations.
Build integrations between SOAR platforms and security technologies using APIs, scripting, and custom connectors.
Administer SOAR platforms and collaborate with SOC teams to optimize incident response and automation initiatives.
Jobgether uses AI-powered matching to connect candidates with hiring companies. They operate as a platform focused on efficient recruitment, with a remote-first and inclusive culture.
Own and evolve observability strategy including monitoring, alerting, dashboards, logging, and distributed tracing.
Define and manage SLIs, SLOs, and reliability metrics, improving MTTD and MTTR through automation.
Build and maintain reliable cloud infrastructure on AWS and Kubernetes while mentoring engineers on SRE best practices.
Filevine is a Legal AI company delivering Legal Operating Intelligence for legal work. Fueled by a team of exceptional collaborators and innovators, Filevine’s rapid growth has earned AI awards and recognition from Deloitte and Inc. as one of the most innovative and fastest-growing technology companies in the country.
Lead high-priority strategic and operational initiatives across the business.
Partner with leadership teams to identify, analyze, and solve business and product challenges.
Drive cross-functional projects from planning through execution and completion.
The partner company focuses on driving strategic and operational initiatives to influence business growth. It has a collaborative international team culture with diverse perspectives.
Define and implement SLIs/SLOs for critical services, lead incident response, and conduct blameless postmortems to drive systemic improvements.
Design and improve monitoring and alerting with Prometheus and Grafana, build internal tooling, and automate operational workflows to reduce toil.
Partner with cross-functional engineering teams to improve system resilience, contribute to architectural discussions, and strengthen production readiness standards.
Runpod provides a cloud platform for AI development, used by over one million developers for training, fine-tuning, and deploying AI models. The company is a small, remote-first team that closed a $100M Series A in June 2026, emphasizing ownership, speed, and impact at scale.
Design and maintain Grafana dashboards and telemetry visualizations to monitor system performance and platform health.
Develop and maintain modular Ansible playbooks to automate infrastructure provisioning and configuration.
Configure observability solutions with Prometheus monitoring and alerting, and participate in Agile ceremonies.
Miratech is a global IT services and consulting company that helps visionaries change the world by supporting digital transformation for large enterprises. With nearly 1,000 full-time professionals across 5 continents and 25 countries, the company has a culture of Relentless Performance with a 99% project success rate and over 25% annual growth.
Design, scale, and maintain enterprise monitoring and alerting ecosystems across multi-cloud and native systems.
Bridge development and operations to ensure high availability, performance tuning, and deep visibility.
Automate infrastructure and build robust observability pipelines using cloud-native tools like Prometheus, Grafana, and GCP.
Ontrac Solutions is a leading technology consulting firm specializing in cutting-edge solutions that drive business transformation. Their team is committed to innovation, collaboration, and excellence, empowering clients to succeed in an evolving digital landscape.
Monitor, operate, and support production AI infrastructure platforms and resolve incidents.
Investigate performance, availability, and reliability issues across infrastructure and platform components.
Collaborate with engineering teams, hardware vendors, and data center personnel to resolve technical issues.
Mirantis is the Kubernetes-native AI infrastructure company, enabling organizations to build and operate scalable, secure, and sovereign infrastructure for modern AI, machine learning, and data-intensive applications. It is a Silicon Valley leader with passionate, talented colleagues, offering a competitive compensation package and strong benefits.
Own reliability and operational stability of BJAK’s production systems.
Design and improve monitoring, alerting, logging and observability across services.
Lead incident response, troubleshooting and structured root cause analysis.
BJAK's automation systems power end-to-end insurance journeys across quote generation, policy issuance, claims, and more. They are a global engineering team with modern engineering culture, offering fully remote work and a high-ownership environment.
Lead architecture and technical direction for AI tools and platforms.
Guide development of AI-enabled workflows, agents, and automation.
Mentor and lead multiple staff developers while owning code quality and engineering standards.
Apex IT is a global consulting firm providing Salesforce and Oracle enterprise solutions to transform customer, employee, and student experiences. As a remote company, we have top talent across the US and India and offer a flexible work-life balance.
Build and maintain observability across the platform in Datadog, including dashboards, monitors, APM, and log pipelines.
Participate in on-call rotation and incident response, driving blameless post-incident reviews and automating toil.
Leverage AI tools to accelerate debugging, generate runbooks, and build automation for operational efficiency.
IPSY is a beauty subscription platform that connects brands and consumers through curated beauty products. It is a remote-first company with a focus on community and engagement.
Lead client discovery, architecture workshops, and solution design across observability, telemetry, reliability, and operational intelligence initiatives.
Define scalable standards for telemetry onboarding, naming, tagging, RBAC, service ownership, dashboards, alert governance, runbooks, and operational handoff.
AHEAD builds platforms for digital business by weaving together cloud infrastructure, automation, analytics, and software delivery to help enterprises achieve digital transformation. The company prioritizes a culture of belonging and is an equal opportunity employer that values diversity and inclusion.
Co-own the architecture of cloud infrastructure on Azure and Kubernetes clusters for high throughput and availability.
Drive resilience strategy for global scaling, zero-downtime deployments, and disaster recovery.
Evolve observability stack with LGTM (Loki, Grafana, Tempo, Mimir) and lead incident response.
Flip is an AI-powered employee experience platform for frontline workers in retail, manufacturing, and logistics. The company is a young, rapidly growing tech company with a remote-first culture and offices in Berlin and Stuttgart.
Lead the design, development and operation of large-scale, secure observability systems to keep services online and performant.
Deploy and scale Prometheus, ElasticSearch clusters, and high-throughput Kafka data pipelines for millions of customer devices.
Collaborate with the Observability team to build alerting systems, APIs, and self-service monitoring tools using Terraform and multiple languages.
ItD is a new generation consulting and software development company that blends diversity, innovation, and integrity with real business results. It is a woman- and minority-led firm with a global community, empowering employees and offering benefits like medical, dental, vision, 401(k), and career development.
Support customers with installation, configuration, and optimization of AI and database solutions across cloud and on-premise environments.
Troubleshoot complex technical issues, analyze logs, and collaborate with engineering teams to resolve problems.
Provide guidance on AI technologies including LLMs, embeddings, vector search, and RAG to improve customer outcomes.
Our partner company specializes in AI and database technologies, delivering cutting-edge solutions to enterprise customers. They operate fully remotely with a focus on innovation and customer success.
Design, build, and operate reliable infrastructure supporting AI-powered products.
Own and improve Kubernetes environments and cloud infrastructure.
Enhance production reliability through observability, automation, and incident response.
The company builds advanced AI-driven products and services. It values engineering excellence, autonomy, and individual contribution, with a global team of skilled engineers.
Define product strategy and roadmap for an AI-powered observability platform.
Drive cross-functional collaboration with engineering, design, and sales teams.
Conduct customer research and support commercial growth through packaging and pricing.
They build an AI-powered observability platform for modern engineering teams. The company is fast-growing with a collaborative culture focused on ownership and innovation.
Define and implement SLOs, SLIs, and Error Budgets to ensure production system reliability.
Lead incident command during major outages and drive blameless postmortems.
Develop observability strategies, including monitoring, logging, tracing, and alerting.
Oowlish is a rapidly expanding software development company in Latin America. It is certified as a Great Place to Work and offers a nurturing environment with professional development opportunities.
Design and deploy a modern observability stack including logging, metrics, and distributed tracing across hundreds of services.
Build alerting policies and incident response workflows to reduce manual escalations and improve mean time to detect and resolve.
Automate toil and set SRE standards while mentoring engineers on observability tooling.
WeightWatchers is a global digital health company and the world's #1 doctor-recommended behavioral weight health program. They have a cloud architecture supporting hundreds of services and are expanding their platform engineering team.