Design, develop, and deploy AI/ML models including NLP, predictive models, and intelligent automation workflows.
Integrate LLM-powered features using APIs like OpenAI and implement RAG patterns and AI agent workflows.
Collaborate with cross-functional teams to translate requirements into AI use cases and contribute to engineering best practices.
We are a confidential US-based organization dedicated to leveraging AI and technology to drive business transformation. Our team values innovation and collaboration, and we foster a culture that supports employee growth and responsible AI practices.
Build reinforcement learning and agent environments for real customer and Frontier lab use cases, including task specifications, scoring, and evaluation.
Develop benchmarks and evaluation harnesses to measure model and data quality across accuracy, robustness, safety, latency, and cost.
Run fine-tuning, adapter, and other model experiments to evaluate how data and methods influence model behavior, and deploy local or self-hosted models for evaluation and inference.
Appen has been a leader in AI training data for over 30 years, specializing in human-generated data to train, fine-tune, and evaluate models across generative AI, LLMs, computer vision, and speech recognition. They support model development through an AI-assisted data annotation platform and a global crowd of over 1 million contributors in more than 200 countries, fostering a culture of innovation, collaboration, and humility over ego.
Design, implement, and deploy ML/AI models end-to-end, including data pipelines, training workflows, and production optimization.
Collaborate with product, engineering, and data teams to align AI work with business goals and translate technical tradeoffs.
Contribute to AI architecture decisions and raise engineering practices, using AI-forward coding tools like Claude and Cursor.
Robots & Pencils designs AI systems for a human world, pairing engineering with creativity to ship production-ready AI in 30 to 45 days. The company values craft, ownership, and direct feedback, with teams averaging fifteen-plus years of experience.
Develop advanced ML models and agentic workflows to accelerate model development.
Use AI-assisted tools like Claude and Cursor to investigate model behavior and automate analysis.
Set technical direction, mentor engineers, and raise the bar for modeling rigor.
Reddit is a community of communities, built on shared interests, passion, and trust, and is home to the most open and authentic conversations on the internet. With 100,000+ active communities and approximately 130 million daily active unique visitors, Reddit is one of the internet’s largest sources of information.
Lead the development and optimization of machine learning models to detect and prevent AI security risks like prompt injection and jailbreaks.
Build reproducible training and evaluation pipelines on Reddit's ML platform, partnering with platform engineers to improve performance and reliability.
Set the technical vision and multi-quarter modeling roadmap, mentoring engineers and establishing best practices for responsible ML development.
Reddit is a community of communities, built on shared interests and authentic conversations, with 100,000+ active communities and 130 million daily active visitors. It is one of the internet's largest sources of information, fostering a culture of openness and trust.
Design and implement novel training optimization techniques for large-scale neural networks.
Investigate approaches to improve training efficiency, stability, and convergence speed.
Collaborate with infrastructure and inference engineering teams to translate research into production performance.
This partner company focuses on AI research and training optimization for large-scale models. They maintain a small, senior-level team that values deep technical thinking and thoughtful execution.
Design, build, and deploy production ML and LLM-based systems for enterprise clients.
Own technical delivery end-to-end from architecture to deployment and iteration.
Mentor and support other ML engineers through code reviews and technical guidance.
TensorOps is a boutique AI consultancy that designs and ships production-grade AI systems for enterprise clients. The company has shipped AI systems impacting 200M+ end users daily, partnered with 11 unicorns, and operates fully remotely with a supportive, fast-growing culture.
Build and improve computer vision and 3D perception models that detect and assess infrastructure from LiDAR and imagery.
Develop LLM-powered products like agentic workflows, MCP servers, and AI copilots for customer tools.
Own models end-to-end, from data and training through evaluation, deployment, and monitoring in production.
Cyvl is a Physical AI company building purpose-built sensors, computer vision, and AI to turn every drive into current data for infrastructure. We're a fast-moving Boston startup with 500+ cities using our platform, and a culture of ownership, intensity, and care.
Develop and train ML models for learned behavior systems using imitation and reinforcement learning.
Write production-quality ML code for training, evaluation, and inference in the autonomy stack.
Collaborate with teams to test and integrate learned behavior models across diverse driving environments.
Torc is an autonomous vehicle technology company that develops software for automated trucks, with a goal to transform freight movement. Now a part of the Daimler family, the company fosters a collaborative, energetic, and team-focused culture.
Work directly with leading AI labs and enterprises to define research goals and technical requirements.
Build data intelligence systems and implement ML pipelines for data curation, model training, and evaluation.
Develop LLM applications, including multi-agent systems, RAG workflows, and evaluation harnesses.
Our client is a venture-backed AI company building intelligent systems by combining human expertise with machine learning. With over $40 million in funding and a global expert network, they provide critical infrastructure for AI development.
Build and ship features of AI/ML and LLM-powered systems with guidance from senior engineers.
Implement and maintain AI/ML and AI agent pipelines from data ingestion through model deployment.
Contribute to LLM-powered features such as prompts, evaluations, retrieval, and tool integrations, while documenting experiments clearly.
Robots & Pencils designs AI systems for a human world, pairing engineering with creativity. Teams average fifteen-plus years of experience and value ownership, craft, direct feedback, and continuous learning.
Design, build, and improve production agentic AI systems used to solve complex real-world problems.
Develop agent architectures, model integrations, and evaluation frameworks for reliable, production-grade AI.
Build scalable APIs, services, and infrastructure supporting agent execution and AI-powered product experiences.
Air is the leader in Enterprise Readiness, providing an AI-native platform to align development, production, delivery, and sustainment for government agencies and industrial suppliers. The company is a startup with a focus on mission-driven work and innovation.
Design and maintain reliable, low-latency ML APIs to integrate Safety AI model outputs into cloud applications.
Build scalable data pipelines for continuous model iteration, backtesting, and online evaluation.
Optimize model artifacts for production and monitor rollout health, ensuring predictable failure modes.
Samsara builds a Connected Operations Cloud that helps physical operations use IoT data to improve safety, efficiency, and sustainability. Samsara is a recently public company with an employee-led remote culture and a long-term focus.
Own Merlin's foundation and world-model work, including architecture selection, post-training, and capability roadmap.
Lead and mentor a small team of world-model engineers, setting the technical bar and review culture.
Design model interface to the autonomy stack with structured, schema-constrained plan outputs and build evaluation harnesses.
Merlin is a publicly traded aerospace and defense company building a non-human pilot for full-stack aircraft autonomy. Headquartered in Boston, it is expanding its organization to accelerate the deployment of its autonomy platform.
Develop and improve core AI methods and systems for reliable AI agents across the full lifecycle.
Create novel approaches for simulation, evaluation, and optimization of agent behavior in production.
Turn research ideas into working prototypes and production-facing capabilities.
This is an early-stage AI infrastructure company focused on making AI agents reliable in production. The company values innovation and practical deployment, with a small team driving frontier AI research and product development.
Develop and deploy cutting-edge AI/ML solutions to enhance the platform and improve student learning experiences.
Design, develop, and optimize LLM-powered agentic systems and APIs for real product use cases.
Collaborate with senior engineers and contribute to evaluation frameworks and MLOps pipelines.
Interview Kickstart specializes in interview preparation and career transitions into high-demand tech fields like AI, ML, and Data Science. Over 17,000 tech professionals have been guided by current and former hiring managers to land coveted positions at companies like Google and Amazon.
Design, build, and deploy generative AI systems using large language models, RAG, and agentic workflows.
Own the full lifecycle from experimentation to production, including evaluation, infrastructure, and monitoring.
Collaborate with Product, Engineering, and Data teams to turn AI ideas into scalable, secure customer experiences.
Typeform is a form builder that helps over 150,000 businesses collect data with forms, surveys, and quizzes that people enjoy. With 500 million responses annually, we are a diverse team of 500+ employees committed to excellence, respect, and transparency.
Design and implement autonomous agents and multi-agent orchestration systems for complex, open-ended tasks.\n- Build and optimize production-grade LLM applications with a focus on reliability, observability, and low-latency performance.\n- Architect advanced RAG pipelines and vector database strategies to provide agents with accurate, real-time context.
Jobgether is an AI-powered job matching platform that connects candidates with hiring companies through objective evaluation. The company operates with a distributed, global-first team culture, emphasizing fairness and innovation in recruitment.
Develop and evaluate deep learning models for feature detection, matching, depth estimation, and visual localization.
Own data preparation, training workflows, and evaluation of learned components for vision-based navigation.
Partner with state estimation and deployment engineers to integrate and optimize models for onboard systems.
Shield AI is a venture-backed defense-tech company developing intelligent systems to protect service members and civilians, with products like Hivemind autonomy software and V-BAT aircraft. It has offices and facilities across the U.S., Europe, the Middle East, and Asia-Pacific, actively supporting operations worldwide.
Set technical direction for speech-model and decoder strategy across the team, guiding evaluation of state-of-the-art approaches.
Lead research, adaptation, and implementation of ASR/STT, speech-enhancement, and related audio models.
Improve the speech pipeline end to end, including decoding, endpointing, turn detection, and streaming behavior.
Dialpad is an AI platform for customer experience, built to resolve customer problems in real time across voice and digital. The company is backed by major investors like Andreessen Horowitz and serves market-leading brands, fostering a culture of curiosity and ambition.