Own end-to-end ML system execution including data pipelines, training workflows, evaluation systems, inference architecture, and deployment.
Fine-tune and adapt models using state-of-the-art methods such as LoRA, QLoRA, SFT, DPO, and distillation.
Architect scalable inference systems, balance latency, cost, and reliability, and deploy production-grade ML solutions.
Gina's Tech Jobs is a recruiting and staffing company that helps firms hire technical talent. They are a small agency focused on IT roles, fostering a high-trust, collaborative environment.
Optimize machine learning inference systems for latency, throughput, and cost-efficiency.
Profile and troubleshoot GPU/CPU bottlenecks, implement advanced techniques like quantization and speculative decoding.
Collaborate with research and engineering teams to productionize new models and improve inference infrastructure.
The company is an AI-focused organization that develops advanced machine learning systems for production environments. It values technical excellence and experimentation, offering a flexible remote work environment.
Develop and operate production-ready AI and machine learning systems for enterprise-scale products.
Build and optimize LLM-powered applications, RAG pipelines, and intelligent agents.
Implement software engineering best practices for AI development including CI/CD and testing.
Our partner is building enterprise-grade AI solutions that deliver measurable business impact. They offer a remote-friendly work environment with a collaborative engineering culture focused on innovation, quality, and continuous learning.
Independently own high-value optimization initiatives across training, inference, or launch-readiness for important Ads ML workloads.
Diagnose bottlenecks in real production systems using profiling, benchmarking, and observability.
Build performance tooling, optimization playbooks, and efficiency primitives that benefit multiple teams.
Reddit is a community of communities built on shared interests and authentic conversations. With 100,000+ active communities and approximately 126 million daily active unique visitors, Reddit has a flexible workforce and values collaboration.
Own end-to-end Machine Learning (ML) system execution including data pipelines, training, and deployment.
Fine-tune and adapt models using state-of-the-art methods like LoRA and DPO.
Architect scalable inference systems and collaborate closely with application engineering.
This company develops advanced production-grade machine learning systems. The team is small and high-trust, with a culture of ownership and pragmatism.
Build and improve ML components across data, training, evaluation, and inference.
Implement evaluation and testing to understand model behavior.
Debug model issues, performance problems, and production incidents.
This company builds core ML components for large-scale production systems. They emphasize real-world learning, iteration, and collaboration with senior engineers.
Design and build production-grade ML inference infrastructure using frameworks like vLLM and Triton.
Optimize GPU utilization, memory efficiency, and model artifact storage for cost-effective performance.
Collaborate with infrastructure and AI teams to establish engineering best practices and scalable platform architecture.
Jobgether is an AI-powered job matching platform that connects candidates with hiring companies. It uses technology to review applications and share top candidate shortlists with employers, operating in a remote-first environment.
Research and implement state-of-the-art techniques to accelerate AI inference: quantization, sparsity, distillation, speculative decoding, and caching.
Partner closely with hardware and compiler teams to ensure algorithmic improvements translate to real gains on custom silicon.
Build profiling tools and comprehensive benchmarking frameworks to measure model quality and efficiency.
EnCharge AI is building the next generation AI platform using novel in-memory-computing architecture. The team consists of experienced AI researchers, silicon & systems engineers, and architects backed by leading investors.
Train, evaluate, and iterate on ML models for customer feedback, including custom fine-tuning pipelines.
Build and maintain LLM-powered features like retrieval pipelines and insight agents.
Design and run robust evaluation frameworks to measure model performance.
Chattermill helps large brands like Uber, Amazon, and Wise put customers at the center using AI. They offer a flexible, trust-based culture with a choice-first environment.
Build and iterate on consumer-facing AI features powered by large language models (LLMs) and generative AI systems
Collaborate with engineers across the AI stack including prompt engineering and agentic workflow optimization
Run structured experiments and monitor production AI systems to optimize latency, cost, and scalability
Quora is a global knowledge sharing platform with over 300M monthly unique visitors, connecting people to share insights and learn. Poe provides a platform for users to chat and build with AI language models. They are a remote-first company with a culture rooted in transparency and experimentation.
Design, develop, and deploy production-grade AI-powered backend systems.
Integrate large language models and machine learning models into scalable architectures.
Optimize system performance and implement strong testing practices.
Our partner company is building advanced AI-powered systems to create meaningful customer value. The team operates in a high-autonomy, fast-moving environment focused on production-ready AI solutions.
Design and scale production ML systems for LLM-based applications.
Build training and evaluation pipelines for continuous model improvement.
Fine-tune foundation models using modern adaptation techniques such as LoRA, QLoRA, SFT and DPO.
A1 is a new AI venture building the next generation of AI-native productivity applications, starting with an email agent that uses autonomous AI. Backed by an initial $100M investment, the company is a small, high-talent founding engineering team focused on solving challenging AI infrastructure problems.
Optimize production LLM serving with vLLM and SGLang to maximize throughput and minimize latency through batching and quantization.
Profile training runs to find bottlenecks and resolve them with attention implementations like FlashAttention on H200 and GB200 hardware.
Deploy and operate multiple models on shared GPU clusters with autoscaling, bin-packing, and efficient handling of mixed workloads.
Egen is a fast-growing technology company with a data-first mindset, partnering with clients on Google Cloud and Salesforce to drive action through data and insights. We are a team of dedicated engineers who thrive on solving tough problems and continually innovate to achieve fast, effective results.
Develop and refine features for deep learning models using large-scale customer and behavioral datasets.
Implement model architecture changes informed by recent academic research from venues like NeurIPS.
Optimize model training pipelines for efficiency and scalability while collaborating with client teams.
OpenTeams builds AI that empowers, offering energy-efficient and cost-effective models with a commitment to open source. The company values freedom, teamwork, accountability, and quality, and reinvests 3% of profits into the open-source community.
Architect and maintain production high-traffic LLM serving systems.
Optimize throughput, latency, and cost for leading open-source LLMs.
Debug and optimize major inference engines like SGLang, vLLM, or TensorRT using PyTorch and CUDA.
We are building decentralized and confidential machine learning infrastructure to enable user-owned AI. Our team is focused on highly scalable and efficient infrastructure for open-source AI at a global scale, with a culture that values innovation and performance.
Build and ship AI features end-to-end, from model to system to user experience.
Design and iterate on prompts, tools, memory, and agent workflows for real-world reliability.
Debug full-stack issues and optimize for latency, cost, and production performance.
A1 builds a proactive smart assistant for everyday users, bringing intelligence to conversations, errands, organizing, and workflows with minimal prompting. The team is small, world-class, and focuses on rapid iteration and shipping high-quality AI products.
Design and maintain scalable ML infrastructure including data pipelines, training workflows, and model deployment systems.
Own end-to-end ML lifecycle operations, ensuring reliable delivery of models into production at scale.
Implement monitoring, telemetry, and feedback loops for ML models running across large-scale device fleets.
Our partner company develops ML systems for connected hardware products used by customers worldwide. They operate in a fast-paced, product-driven environment with a collaborative and technically ambitious culture focused on real-world ML impact.
Lead the design and operation of production machine learning systems for batch and online use cases with a focus on reliability and scalability.
Build and improve ML lifecycle infrastructure including training pipelines, inference workflows, monitoring, and automation.
Partner with cross-functional teams to translate business problems into ML solutions and guide prototypes to robust production systems.
Included Health is a healthcare company delivering integrated virtual care and navigation, aiming to raise the standard of healthcare for everyone. They are a remote-first organization offering comprehensive benefits and fostering a culture of inclusion.
Design, develop, and deploy ML models, including large language models, for various NLP tasks.
Collaborate with cross-functional teams to gather requirements, define architectures, and iterate on model development.
Stay up-to-date with latest research and contribute to best practices for responsible ML development.
Reddit is a community of communities built on shared interests, passion, and trust, hosting the most open and authentic conversations on the internet. With over 100,000 active communities and approximately 126 million daily active users, it is one of the internet's largest sources of information, fostering a culture of authenticity and community.