Source Job

$97,600–$139,000/yr
United States Canada

  • Build and maintain core infrastructure for Quora's ML platform, ensuring high availability, scalability, and performance.
  • Build and improve distributed systems serving ML models in production, from Large Recommendation Models to Large Language Models.
  • Work on GPU model serving, optimizing latency, throughput, and cost to support larger and more capable models.

Python Go C++ PyTorch Kubernetes

20 jobs similar to Software Engineer (New Grad)

Jobs ranked by similarity.

Canada

  • Architect, design, build, deploy, and maintain Model Serving infrastructure for a world-class Detection Engine.
  • Own projects that scale model serving and data processing to handle 10x traffic, including real-time streaming pipelines and online feature serving.
  • Collaborate with MLE and Data Science teams to build the ML Training platform, improving MLE velocity and model precision and recall.

Abnormal protects the humans behind the world's most critical organizations from AI-powered cybercrime. 4,500+ enterprises trust our behavioral AI platform, and we foster a culture of innovation and impact.

Spain

  • Design and develop large-scale platforms for LLM training and AI workloads.
  • Tackle distributed-systems challenges including intelligent job scheduling and resource optimization.
  • Collaborate with international teams to build production-ready AI infrastructure.

This role is with a partner company, an AI-focused R&D team building infrastructure for large language models. They are a fast-moving, highly technical team with a collaborative and innovative culture.

Canada

  • Develop machine learning solutions for supply chain planning, from tech selection to production code.
  • Write Python code for microservice architectures and contribute to the continuous improvement of the AI platform.
  • Participate in agile processes including sprint planning, pair programming, and retrospectives with cross-functional teams.

Kinaxis is a global leader in modern supply chain orchestration, powering complex global supply chains with an AI-infused platform. With over 2,000 employees worldwide, the company fosters a culture of innovation, collaboration, and continuous learning, and has won several Top Employer awards.

Global 6w PTO 26w maternity 26w paternity

  • Design and write high-performing, scalable software for training models.
  • Develop new tools to support and accelerate research and LLM training.
  • Collaborate with engineering and scientific teams to create a strong post-training ecosystem.

Cohere is a security-first enterprise AI company building cutting-edge foundation models and end-to-end products to solve real-world business problems. They are a global team of researchers, engineers, and designers passionate about their craft, with offices in Toronto, London, New York, San Francisco, Montreal, Paris, Berlin, and Seoul.

Global

  • Build and improve the inference layer of the Gcore Inference platform, integrating frameworks like vLLM and TensorRT-LLM.
  • Bring new language and multimodal models into production, optimizing latency, throughput, and cost efficiency.
  • Debug performance issues across model code, GPU execution, and Kubernetes, collaborating with cross-functional teams.

Gcore is a global provider of AI, cloud, network, and security infrastructure and software. They are a team of 550+ professionals with a collaborative culture and partnerships with Intel, NVIDIA, Dell, and Equinix.

$189,507–$274,604/yr
Global

  • Develop and improve ads ranking models, including prediction objectives, feature interactions, user-history modeling, and calibration.
  • Take end-to-end ownership of machine learning systems from data pipelines to production integration.
  • Evaluate and apply advances in deep learning and recommendation modeling to improve ads ranking within production constraints.

Quora operates two knowledge-sharing platforms: Quora, a global Q&A platform, and Poe, a platform for interacting with AI language models. They have a remote-first culture with passionate, collaborative, and high-performing global teams focused on transparency and experimentation.

US Unlimited PTO

  • Build the foundations of the EdgeRunner Research organization, including data pipelines, model evaluation, and efficient parallelization.
  • Own codebases in areas like distributed training, quantization, compression, or compute cluster management.
  • Collaborate with a team of self-starters who operate independently and translate business objectives into technical solutions.

EdgeRunner AI builds state-of-the-art AI models for the tactical edge, enabling warfighters to make faster decisions and interact with robotic platforms. As a Series-A startup, we move quickly, fostering a culture of initiative, ownership, and comfort with ambiguity.

$170,170–$286,000/yr
North America

  • Design and maintain reliable, low-latency ML APIs to integrate Safety AI model outputs into cloud applications.
  • Build scalable data pipelines for continuous model iteration, backtesting, and online evaluation.
  • Optimize model artifacts for production and monitor rollout health, ensuring predictable failure modes.

Samsara builds a Connected Operations Cloud that helps physical operations use IoT data to improve safety, efficiency, and sustainability. Samsara is a recently public company with an employee-led remote culture and a long-term focus.

Europe

  • Design and maintain the MLOps platform for experiment tracking, model registry, and CI/CD practices.\n- Productionize ML models into scalable, low-latency serving infrastructure with monitoring and rollback.\n- Automate retraining, evaluation, and deployment pipelines to reduce manual intervention.

Yuno builds payment infrastructure that connects companies to over 300 payment methods worldwide via a single API, using AI for intelligent routing and fraud prevention. It is a growing company with a global reach, founded by veterans from payments and technology, and emphasizes innovation and remote collaboration.

US

  • Develop advanced ML models and agentic workflows to accelerate model development.
  • Use AI-assisted tools like Claude and Cursor to investigate model behavior and automate analysis.
  • Set technical direction, mentor engineers, and raise the bar for modeling rigor.

Reddit is a community of communities, built on shared interests, passion, and trust, and is home to the most open and authentic conversations on the internet. With 100,000+ active communities and approximately 130 million daily active unique visitors, Reddit is one of the internet’s largest sources of information.

$220,000–$292,000/yr
US Unlimited PTO

  • Own the platform including GCP, Kubernetes, Temporal, GPU fleet, and deploy/rollback machinery.
  • Contribute to AI enablement substrate: GPU capacity, training/inference pipelines, and cost optimization.
  • Strengthen team practices through tooling, standards, tests, observability, and release processes.

Descript is building a simple, intuitive, fully-powered editing tool for video and audio — an editing tool built for the age of AI. They are a team of 150 backed by top investors like OpenAI and Andreessen Horowitz, with a culture that values collaboration and serendipitous discovery.

$166,600–$208,300/yr
US Canada

  • Build and operate the real-time inference service that scores models for the risk decision engine, with low latency and high availability.
  • Own model deployment infrastructure including registry, versioning, CI/CD, and staged rollouts.
  • Build model observability with availability, latency, error monitoring, and drift detection.

Mercury is a fintech company that builds banking services for startups. They are committed to diversity and inclusion, and are an equal opportunity employer.

$190,000–$230,000/yr
US

  • Design and implement scalable cloud infrastructure using Kubernetes, Pub/Sub, and distributed systems technologies.
  • Collaborate with our AI team to optimize data pipelines and integrate AI to remove performance bottlenecks.
  • Drive platform reliability initiatives including alerting, health checking, and incident management.

Syllo is building a unified litigation platform that helps lawyers and paralegals use AI throughout the litigation life cycle. We are a quickly expanding company with enterprise customers including major law firms and corporations.

US Unlimited PTO

  • Build the component layer around our layout synthesis engine, including API contracts, services, and evaluation gates.
  • Turn model retraining into a one-command job with built-in benchmarks and readable results for the whole team.
  • Own latency and cost budgets for learned capabilities and partner with infrastructure engineers on MLOps and deployment.

Higharc is a VC-backed startup that is changing how new homes are designed and built using spatial AI and generative floor plan technology. The company is fully remote, has raised over $175M, and values flexibility, collaboration, and asynchronous deep work.

$164,200–$229,900/yr
US

  • Refine and maintain data infrastructure for ML and analytics workflows on data from hundreds of millions of users.
  • Own the Data Movement Platform enabling batch and stream processing, investing in Spark, Flink, and Airflow technologies.
  • Build automated solutions to minimize toilsome work, providing a declarative, self-service experience for data users.

Reddit is a community of communities built on shared interests, passion, and trust, home to the most open and authentic conversations on the internet. With over 100,000 active communities and approximately 130 million daily active unique visitors, Reddit is one of the internet's largest sources of information, fostering a flexible and inclusive culture.

Switzerland

  • Design and deliver sophisticated ML and data architectures using Linux, Kubernetes, and open-source technologies.
  • Work closely with sales and technical teams to translate customer requirements into scalable infrastructure solutions.
  • Influence product direction by sharing customer insights and technical feedback with engineering teams.

The partner company helps organizations adopt modern AI and machine learning technologies across public and private cloud environments. They operate with a global Field Engineering team and a distributed work environment.

$176,100–$308,200/yr
North America Canada

  • Own a major subsystem of a novel exploitability engine end-to-end, including design, delivery, and quality.
  • Drive design and code reviews, raise the engineering bar, and mentor engineers through influence.
  • Partner with product, security R&D, and SecOps to turn customer problems into subsystem design.

ServiceNow provides an AI platform that automates workflows to free people from busywork, serving as the AI control tower for business reinvention. It serves 85% of the Fortune 500 and fosters an AI-native culture where technology and talent are unstoppable.

Latin America

  • Build and operate model and inference serving infrastructure, managing latency, throughput, autoscaling, and reliability for real-time and batch inference.
  • Own the ML deployment lifecycle: model registry, versioning, promotion workflows, rollout strategies, and safe rollback.
  • Operate agentic and LLM workloads in production, managing inference providers, gateways, quotas, guardrails, and graceful degradation under load.

ReadyOn is an AI-native Labor Operating System that redefines how enterprises manage frontline labor by matching workers to shifts in real time. Headquartered in San Francisco with over 100 employees, it grew revenue 8x year over year in 2025.

$150,000–$220,000/yr
Global Unlimited PTO

  • Lead the effort to make Runpod the fastest and most cost-efficient place for LLM inference, owning performance end to end.
  • Profile and diagnose performance bottlenecks across the serving stack, from scheduling to kernels, and implement fixes.
  • Work closely with product and infrastructure teams to shape how inference is offered, turning improvements into production-ready defaults.

Runpod is the AI Developer Cloud, providing a platform for over one million developers to experiment, train, fine-tune, deploy, and scale AI. We're a small, remote-first team that takes ownership seriously, moves fast, and has processed more than 20 billion inference requests.

$230,000–$322,000/yr
US

  • Lead the development and optimization of machine learning models to detect and prevent AI security risks like prompt injection and jailbreaks.
  • Build reproducible training and evaluation pipelines on Reddit's ML platform, partnering with platform engineers to improve performance and reliability.
  • Set the technical vision and multi-quarter modeling roadmap, mentoring engineers and establishing best practices for responsible ML development.

Reddit is a community of communities, built on shared interests and authentic conversations, with 100,000+ active communities and 130 million daily active visitors. It is one of the internet's largest sources of information, fostering a culture of openness and trust.