Help define and mature Engineering Operations by improving application health visibility, service reliability, and operational analytics.
Build and implement scalable processes for Incident, Problem, and Change Management that engineers actually want to use.
Connect engineering systems, data, and teams to reduce fragmentation and improve operational visibility across the organization.
Turnitin is a recognized innovator in global education, developing learning integrity solutions that help educators and institutions uphold academic integrity. With over 16,000 academic institutions using our services in more than 185 countries, we foster a remote-first culture and a diverse community of colleagues across 35+ countries.
Lead, mentor, and grow a team of SRE/DevOps engineers while partnering with engineering leadership to assess team needs and develop talent.
Oversee the incident management process end to end, including on-call rotations, escalation paths, incident command, postmortems, and root cause analysis.
Define and drive SRE principles like SLIs, SLOs, error budgets, capacity planning, and observability standards, championing a culture of reliability and operational excellence.
Eltropy is a rocket ship FinTech on a mission to disrupt the way people access financial services, enabling community financial institutions to digitally engage in a secure and compliant way through a world-class digital communications platform. Their platform integrates Text, Video, Secure Chat, co-browsing, screen sharing, and chatbot technology, bolstered by AI and contact center capabilities, and they value integrity, transparency, and ownership.
Lead technical and managerial direction for the SRE team, defining reliability, observability, and operational excellence strategy.
Coordinate critical incident responses and root cause analysis, collaborating with architecture, development, security, and product teams.
Drive automation, continuous improvement, and adoption of SRE, DevOps, and Platform Engineering best practices.
Experian is a global data and technology company that drives opportunities for people and businesses worldwide. With 25,200 employees in 32 countries, it has a people-centric, inclusive culture recognized by awards such as World's Best Workplaces™ 2025.
Lead and develop junior analysts through guidance, feedback, and knowledge sharing to support their professional growth.
Take ownership of critical production incidents, coordinating war rooms and driving resolution to prevent recurrence.
Analyze and improve system performance, reliability, and resilience to ensure production stability.
The role is with a partner company of Jobgether, a platform using AI-powered matching for job applications. The environment is highly autonomous, collaborative, and technically demanding, focusing on ownership and continuous learning.
Define and execute the technical strategy for Observability & Alerting, Platform Infrastructure, and operational excellence.
Lead the evolution of reliable, scalable, secure, and efficient cloud platforms and distributed systems.
Champion SLIs, SLOs, error budgets, capacity planning, and automation across the service lifecycle.
Filevine is a Legal AI company delivering a unified platform for legal work, powered by LOIS (Legal Operating Intelligence System). The company is rapidly growing, recognized by Deloitte and Inc. as one of the most innovative and fastest-growing technology companies.
Ensure reliability, scalability, and operational excellence of analytics and data systems.
Provide technical leadership and direction to an offshore contract team.
Drive incident response, automation, and data governance initiatives.
Workiva provides an AI-powered platform that unifies finance, risk, and sustainability for complex organizations. It is a large enterprise with a collaborative and innovative culture centered on data integrity and trust.
Lead and develop the Enterprise Support team, owning service delivery and driving operational excellence for a strategic client.
Oversee technical engagement, governance, and risk management, ensuring platform stability and long-term client trust.
Drive AI and automation adoption to improve support efficiency and scalability, while collaborating with cross-functional teams.
Jobgether is a job platform that uses AI-powered matching to connect candidates with hiring companies. They are a partner company that manages applications and next steps for this role, offering an inclusive and collaborative working environment.
Lead the Site Reliability Engineering team to drive reliability and scalability of XTB's systems.
Define and drive SRE platform strategy, incident management, and observability engineering.
Mentor team members, foster collaboration, and ensure operational excellence.
XTB is a global investment company offering innovative technological solutions for managing finances through an intuitive app, used by over one million users worldwide. It is a certified Great Place to Work with a focus on technical excellence and collaboration.
Lead and manage complex operational projects from conception to completion, ensuring timely delivery and measurable impact.
Analyze existing business processes, identify inefficiencies, and design and implement solutions for optimization and automation.
Collaborate closely with senior leadership across departments to align operational strategies with business objectives.
Instructure creates intuitive products that simplify learning and personal development. They foster a culture of inclusivity, support, and meaningful connection, giving smart, creative, passionate people opportunities to create awesome.
Own the reliability of Supabase's deployment and release systems against clear SLOs and error budgets.
Standardize and instrument pre-production deployment workflows for trustworthy signal.
Drive disaster-recovery readiness and reduce mean-time-to-detect and recover for deploy-related incidents.
Supabase is the Postgres development platform providing a complete backend solution including Database, Auth, Storage, Edge Functions, Realtime, and Vector Search. They are a globally distributed team of ~400 members across 60+ countries, with over $1B raised and 540,000+ community members.
You will own and deliver quarterly goals for your team, leading engineers through ambiguity to solve open-ended problems.
You will proactively identify technical solutions and operational processes that strengthen incident readiness and response.
You will foster a culture of quality and ownership by setting or improving code review and design standards.
Affirm is reinventing credit to make it more honest and friendly, offering consumers the flexibility to buy now and pay later. The company has a strong engineering culture focused on reliability and ownership.
Support cloud-hosted application testing, implementation, maintenance, and validation.
Monitor and troubleshoot Windows, Linux, server, network, and application environments.
Assist with incident response, documentation, and operational process improvements.
Applied Systems builds cloud software and AI-powered solutions that reinvent insurance technology for agencies and brokers worldwide. With over 40 years of experience, the company fosters a people-first culture built on trust, inclusion, and growth, supporting employees to deliver their best work.
Manage end-to-end operations for Edifecs hosted or customer-hosted application instances, including upgrades, configuration changes, and day-to-day submission operations.
Provide L2/L3 technical support to enterprise clients, perform root cause analysis, and implement long-term solutions.
Develop, enhance, and maintain custom monitoring and alerting systems to proactively detect and resolve issues.
Edifecs is a healthcare technology company that provides enrollment management solutions to enterprise customers. The company fosters a collaborative and learning-oriented culture with a focus on operational excellence.
Manage delivery of managed services, building mature client relationships and anticipating their needs.
Track key metrics to ensure SLAs are consistently achieved and lead incident response with clear communication.
Ensure accurate invoices, reports, and team tools while proactively identifying risks and driving improvements.
Thoughtworks is a dynamic and inclusive community of bright and supportive colleagues who are revolutionizing tech. As a leading technology consultancy with over 30 years of experience, we deliver extraordinary impact by helping clients solve complex business problems with technology.
Design and maintain highly available, scalable systems to ensure exceptional customer experiences.
Drive automation and eliminate operational toil through self-service tooling and process improvements.
Lead incident response and mentor engineers to improve reliability practices.
Redzone provides a connected workforce solution for manufacturers to improve plant efficiency and worker productivity. The company is part of QAD Inc. and fosters a collaborative, customer-focused culture with a strong technology team.
Own the technical operations domain covering IT, security, DevOps, and infrastructure for an AWS-native platform.
Build and lead a lean team, leveraging AI agents and tooling to automate DevOps, security, and compliance tasks.
Partner with engineering leadership to ensure reliability, scalability, and HIPAA/CMS compliance in a regulated environment.
Spark Advisors builds healthcare technology for Medicare advisors, helping seniors navigate complex coverage. They are a fast-growing platform backed by Primary Ventures and Viewpoint Ventures, recognized as one of Inc. Magazine's Best Workplaces of 2025.
Define and drive reliability of systems at the scale of millions of clients, strengthening SRE practices. - Develop observability platforms and serve as a strategic partner to product engineering teams. - Enhance proactive resilience through early-warning systems, AI/ML, and incident management.
XTB is a global FinTech company specializing in online trading of financial instruments. As the largest FinTech in Poland and a leader in Central and Eastern Europe, we operate across multiple continents and are a certified Great Place to Work, focusing on employee development and training.
Lead and scale IT operations strategically, including FinOps automation, security resilience, and vendor management.
Handle hands-on tasks like provisioning accounts, troubleshooting Google Workspace and Zoom, and onboarding new hires.
Mentor a team of 4 and drive workflow automation to reduce manual effort and support organizational growth.
Lemnis is a public charity dedicated to harnessing transformative change to expand learning for all. We are a growing, innovative organization with a collaborative and inclusive culture, committed to fostering an environment where every team member can thrive.
Lead and execute on complex technical troubleshooting and incident resolution from investigation to delivery of permanent solutions.
Design and implement comprehensive monitoring processes, including creating detailed playbooks and runbooks for common scenarios, as well as detailed documentation including post-incident reviews and knowledge base articles.
Leverage AI-powered tools and workflows to automate issue detection, diagnosis, and resolution processes.
Alternative Payments is building the financial operating system for SMBs, consolidating the disconnected tech stack that holds service-based businesses back. We're growing fast, thinking big, and building a global team that wants to be part of something that lasts.
Own the operational health and reliability of Trino, Lightdash, Coder, and other platform services across development and production.
Lead the migration from Hive Metastore to Nessie and drive platform upgrades with zero disruption.
Manage and mentor a team of 3-5 platform engineers, setting standards for operational excellence and automation.
ServiceNow is the AI control tower for business reinvention, providing an AI platform that brings together any AI, any data, and any workflow. The company powers 85% of the Fortune 500 and fosters an AI-native culture focused on innovation and continuous learning.