Design and build production-grade systems end-to-end, from problem definition through deployment and operations.
Work across application services, distributed systems, infrastructure, data pipelines, and ML systems, debugging complex issues across multiple layers.
Frame problems correctly, applying ML when needed, and ensure reliability, performance, and cost efficiency.
Own the architecture health of a billion-scale distributed system, including failure modes, capacity limits, and cross-deployment interactions.
Approve critical-path designs and hunt gaps like single points of failure, unbounded queues, and missing idempotency proactively.
Build the parts nobody else can, prototype risky architectural bets, and ship remediations after serious incidents.
HighLevel is an AI-powered business operating system that gives agencies and SMBs the infrastructure to build, automate and scale. With over 2,000 team members across 10+ countries, HighLevel operates as a global, remote-first organization built for speed and ownership.
Own architecture health of large-scale distributed systems, including failure modes, capacity constraints, and consistency guarantees.
Identify and remediate systemic risks such as single points of failure, unbounded queues, and data-loss scenarios.
Work hands-on with Node.js/Go and GCP technologies to prototype solutions and resolve complex failures.
This company operates large-scale distributed systems processing billions of events and messages. Its engineering culture values technical rigor, proactive problem solving, and clear cross-team communication.
Architect and evolve core control and context planes, including service registries, SLO enforcement, and automated canary releases.
Own the service chassis and golden path, maintaining multi-language Java/Python libraries, Helm charts, and deployment pipelines.
Drive reliability engineering practices, mentor engineers, and lead architectural strategy for distributed systems at production scale.
The company builds large-scale web data products and distributed engineering infrastructure for AI-driven workflows. It operates a remote-first, globally distributed engineering culture focused on reliability, autonomy, and technical excellence.
Lead the development of Reddit's Ingestion Platform, designing and delivering reliable software for distributed data movement across streaming and batch workloads.
Own the architecture of the platform's control and data planes, including pipeline APIs, connectors, and sink integrations, expanding beyond Kafka-to-BigQuery to S3/GCS and Apache Iceberg.
Mentor engineers, drive migrations from legacy systems, and establish robust reliability, security, and operational practices for pipelines running on Kubernetes.
Reddit is a community of communities built on shared interests, passion, and trust, hosting authentic conversations across 100,000+ active communities and approximately 130 million daily active unique visitors. The company fosters an open, collaborative culture with a focus on reliability, performance, and efficiency.
Own large slices of the system end-to-end, from approach to shipping and operations.
Maintain the Beads Team Server and Hosted Gas City, focusing on state, concurrency, and reliability.
Be a maintainer of the open-source repos, representing community needs and compatibility with our roadmap.
Gas City builds the open-source stack that teams use to run coding agents at scale, including the work graph Beads and the agent orchestration tool Gas City. We are a small, flat organization of founding engineers who work 1:1 with the CTO and run our own product on itself, using agents to write and review code.
Own large slices of the system end to end, from approach to operation.
Turn Beads into a platform and take Gas City to the cloud.
Define SLOs, observability, backups, and security baseline for enterprise readiness.
Gas City builds the open-source stack teams use to run coding agents at scale, including the Beads work graph and Gas City agent orchestration. It's a small, flat organization moving toward revenue with a focus on reliability and agent-driven development.
Design and build distributed data systems handling large-scale ingestion and processing.
Drive architectural decisions and take end-to-end ownership of critical components.
Collaborate with product teams to translate ambiguous requirements into robust technical solutions.
Our partner builds a large-scale, multi-chain data platform that ingests, models, and delivers blockchain data to users and developers. They are a remote-first, distributed team with a strong engineering culture focused on ownership and collaboration.
Design, implement, test, and operate production services and APIs.
Lead projects or meaningful components of projects from problem definition through deployment and iteration.
Improve the observability and operability of the systems you own, including metrics, logs, traces, alerting, and incident learnings.
LaunchDarkly provides a platform that helps engineering teams release software and AI with speed, safety, and control using feature flags and observability. The company is growing and emphasizes teamwork, humility, openness, curiosity, and inclusive collaboration.
Design, build, and operate GitLab Orbit backend services, primarily in Rust, within a distributed, cloud-native environment.
Improve deployment, monitoring, and operations using Kubernetes, Helm, Terraform, and cloud services from AWS or GCP.
Automate operational work, strengthen observability, and manage production issues to reduce single points of failure.
GitLab is the intelligent orchestration platform for DevSecOps, helping organizations increase developer productivity, improve operational efficiency, and accelerate digital transformation. Trusted by more than 50 million registered users and over 50% of the Fortune 100, GitLab fosters a high-performance culture driven by shared values and continuous knowledge exchange.
Work on the heart of broker-dealer operations: clearing and settlement for US equities and options, ensuring systems are reliable and scalable.
Design and build new systems to enable new product offerings and revenue streams, from ideation to deployment.
Participate in an on-call rotation to maintain system health while prioritizing elimination of issues to protect team free time.
Alpaca is a US-headquartered global leader in agent-first brokerage infrastructure for stocks, ETFs, options, crypto, and more, serving hundreds of financial institutions across 40 countries. The team is a dynamic group of 400+ globally distributed members who thrive working from their favorite places around the world, with a culture valuing curiosity, empathy, and accountability.
Partner with cross-functional teams to design and deliver scalable backend systems for major product initiatives.
Own the full software lifecycle from technical design to rollout, using A/B experiments and data analysis to drive decisions.
Build and maintain high-performance APIs and distributed services using modern languages and tools.
Reddit is a community of communities, built on shared interests, passion, and trust. It is home to the most open and authentic conversations on the internet, with 100,000+ active communities and approximately 130 million daily active unique visitors.
Architect, build, and own full-stack software platforms supporting enterprise infrastructure and operations teams.
Lead technical design of systems that ingest, process, analyze, and expose infrastructure and operational data at scale.
Develop backend services, APIs, and web interfaces with advanced search and usability tailored for technical users.
Cision is the global leader in consumer and media intelligence, engagement, and communication solutions, serving over 75,000 companies including 84% of the Fortune 500. We foster an inclusive culture where diversity, equity, and inclusion drive innovation and long-term success.
Design and implement scalable cloud infrastructure using Kubernetes, Pub/Sub, and distributed systems technologies.
Collaborate with our AI team to optimize data pipelines and integrate AI to remove performance bottlenecks.
Drive platform reliability initiatives including alerting, health checking, and incident management.
Syllo is building a unified litigation platform that helps lawyers and paralegals use AI throughout the litigation life cycle. We are a quickly expanding company with enterprise customers including major law firms and corporations.
Design and develop scalable search and indexing systems for an AI search engine.
Ensure operational excellence by participating in on-call rotation and maintaining system quality.
Collaborate with a global remote team to solve distributed system challenges.
Algolia is a pioneer and market leader in AI Search, empowering over 18,000 businesses to deliver blazing-fast search experiences. With $150 million in Series D funding and a valuation of $2.25 billion, the company fosters a high-trust, flexible culture and values diversity and collaboration.
Design, develop, and launch backend systems at scale using Python or Kotlin.
Collaborate with team and stakeholders to balance speed and quality while protecting system reliability.
Contribute to community through growth and development activities and navigate large codebases.
Affirm is reinventing credit to make it more honest and friendly, giving consumers the flexibility to buy now and pay later without hidden fees. They are a remote-first company that values transparency and employee well-being, offering competitive benefits and a collaborative culture.
Design and build margin and risk systems for multiple asset classes.
Lead real-time risk enforcement and ensure system reliability and correctness.
Work with engineering, product, and compliance teams to establish scalable infrastructure.
Our partner is a global financial platform seeking experienced engineers. They value technical ownership and collaboration in a distributed environment.
Provide senior technical leadership across multiple engineering teams and business domains.
Drive architecture, design, and implementation of scalable, reliable, and maintainable software systems.
Mentor staff and senior engineers and influence technical strategy across the organization.
Oportun is a mission-driven financial services company that provides responsible credit, savings, and budgeting tools to help members build a better financial future. It has provided over $22.7 billion in credit and values speed, high standards, and using AI to work smarter.
Architect and build robust, scalable, and highly available distributed infrastructure.
Build a cutting-edge cloud-native platform on public cloud and automate resource management.
Improve reliability, security, and cost efficiency of cloud services.
ClickHouse builds a real-time analytics database platform and manages ClickHouse Cloud data plane end-to-end with compute, networking, and security. As a rapidly scaling global startup, the company operates across 25+ countries and fosters a flexible, remote-friendly culture with equity and healthcare benefits.
Design and build agent runtime infrastructure with Firecracker, Rust, and Go
Define and enforce security boundaries for running untrusted AI agents
Architect global scale distributed systems for scheduling and orchestration
We build agent sandboxes—runtime infrastructure that is fast, durable, and secure by default for AI systems. We are a small, globally distributed team based in San Francisco backed by forward-thinking investors.
Design, build, and operate high-scale observability pipelines for logs, metrics, traces, and exceptions.
Lead cross-functional initiatives to resolve scaling bottlenecks and evolve production infrastructure safely.
Partner with engineering teams to improve observability tools and provide technical leadership across teams.
Jobgether is an AI-powered job matching platform that connects candidates with hiring companies. It emphasizes learning, professional growth, and an inclusive work environment.