Operate, maintain, and improve the Dragos cloud fleet across Azure, AWS, and GCP.
Own the full customer environment lifecycle including onboarding, configuration, upgrades, and offboarding.
Drive SRE practices: define SLOs, manage error budgets, and lead post-incident reviews.
Dragos is the global leader in operational technology (OT) cybersecurity, combining technology, threat intelligence, and expert services to protect critical infrastructure. The remote-first team spans North America, Europe, the Middle East, and APAC, built on authenticity, transparency, and trust.
Own the end-to-end major incident lifecycle, acting as Incident Commander for critical events.
Drive reliability metrics and improve MTTR for Yuno's 99.99% uptime target.
Run blameless postmortems and translate findings into actionable reliability improvements.
Yuno builds payment infrastructure for global market participation, enabling companies to integrate over 1,000 payment methods via a single API. They empower high-performing teams and use advanced AI for smart routing and fraud prevention across 80+ countries.
Lead design, implementation, and troubleshooting of enterprise network infrastructure including routing, switching, firewalls, VPN, and wireless.
Own and mature observability strategy using Datadog, Dynatrace, LogicMonitor, and SNMP-based polling to build proactive monitoring and alerting.
Drive infrastructure as code with Terraform and GitHub-based CI/CD pipelines for scalable, secure, and repeatable IT operations.
EverOps is a premier Embedded Service Provider that partners with customer IT teams to address mission-critical delivery and infrastructure challenges. The company operates a U.S.-Based Virtual Operating Center and offers equity, unlimited PTO, and sponsored healthcare.
Lead end-to-end management of major production incidents, coordinating cross-functional teams from detection to resolution.
Drive improvements in operational reliability by reducing MTTD, MTTR, and optimizing on-call programs.
Facilitate blameless postmortems and establish incident management standards including severity frameworks and runbooks.
Jobgether is an AI-powered job matching platform that connects candidates with hiring companies. They focus on innovation and operational excellence within a globally distributed engineering organization.
Manage the Jira support help desk and ensure SLAs are met.
Develop and maintain Python scripts for customer data import and export.
Triage customer issues by checking logs, running database queries, and performing root cause analysis.
Converge fuses cyber insurance, security, and technology to provide businesses with clear, confident cyber protection. Our team is composed of top-tier professionals across underwriting, technology, and claims who share a passion for driving the future of cyber risk management.