Source Job

Europe 6w PTO

  • Own agent quality end-to-end for Bliro's chat and voice agents, including prompting, context, and model evaluation.
  • Build eval sets, LLM judges, and benchmarks to measure quality, latency, and cost with statistical rigor.
  • Establish a repeatable process for improving agent quality and set the long-term technology roadmap.

Python LLM Agents

20 jobs similar to Senior AI Engineer, Agents (Python)

Jobs ranked by similarity.

Global

  • Design and implement autonomous agents and multi-agent orchestration systems for complex, open-ended tasks.\n- Build and optimize production-grade LLM applications with a focus on reliability, observability, and low-latency performance.\n- Architect advanced RAG pipelines and vector database strategies to provide agents with accurate, real-time context.

Jobgether is an AI-powered job matching platform that connects candidates with hiring companies through objective evaluation. The company operates with a distributed, global-first team culture, emphasizing fairness and innovation in recruitment.

North America 3w PTO

  • Develop, configure, deploy, and optimize AI agents using Cresta's AI platform and tools.
  • Build AI agent integrations with external systems such as APIs, databases, and CRMs to ensure seamless workflow integration.
  • Collaborate with customers and internal stakeholders to gather technical requirements and translate business needs into AI agent solutions.

Cresta builds a unified AI platform that combines conversational AI agents, real-time human agent augmentation, and conversation intelligence to drive revenue and efficiency for customer experience channels. Born from Stanford AI Lab, the company has raised over $270 million from top investors and is led by AI industry veterans.

Global

  • Develop and deploy cutting-edge AI/ML solutions to enhance the platform and improve student learning experiences.
  • Design, develop, and optimize LLM-powered agentic systems and APIs for real product use cases.
  • Collaborate with senior engineers and contribute to evaluation frameworks and MLOps pipelines.

Interview Kickstart specializes in interview preparation and career transitions into high-demand tech fields like AI, ML, and Data Science. Over 17,000 tech professionals have been guided by current and former hiring managers to land coveted positions at companies like Google and Amazon.

$180,000–$250,000/yr
Global

  • Work directly with leading AI labs and enterprises to define research goals and technical requirements.
  • Build data intelligence systems and implement ML pipelines for data curation, model training, and evaluation.
  • Develop LLM applications, including multi-agent systems, RAG workflows, and evaluation harnesses.

Our client is a venture-backed AI company building intelligent systems by combining human expertise with machine learning. With over $40 million in funding and a global expert network, they provide critical infrastructure for AI development.

US Unlimited PTO

  • Lead the engineering strategy and execution for evaluations of AI agents, owning the core evaluation platform.
  • Design scalable evaluation infrastructure, APIs, workflows, and production systems across software categories.
  • Mentor and develop a team of engineers as the technical authority on agentic evaluation.

This company focuses on building credible, scalable evaluations of AI agents from software vendors. It operates as a fully remote, inclusive team with a flexible culture and a focus on professional growth.

US

  • Design and deploy AI agents that automate sales, operations, and support workflows.
  • Own technical delivery end-to-end, from architecture to production and post-launch improvements.
  • Work directly with clients to understand workflows and explain technical tradeoffs.

Stello is an AI transformation firm that partners with mid-market and enterprise companies to automate repetitive work with AI agents. They work across sales, operations, and customer support, deploying solutions inside existing client tools.

US

  • Set the technical direction for AI engineering across the team: agent architecture patterns, evaluation methodology, deployment and monitoring strategies
  • Design the AI platform layer, including shared agent frameworks, tool integrations, and evaluation infrastructure
  • Work directly with clients on the most complex engagements, identifying new problem domains and ensuring production quality

Addepar is a global data and AI platform empowering investment professionals to turn complex financial information into actionable intelligence. With over 1,500 firms in 60 countries managing nearly $10 trillion in assets, Addepar fosters a culture of ownership, collaboration, and innovation.

UK

  • Build tooling for capturing and processing data from agents and humans at significant scale.
  • Solve hard problems around compute, orchestration, scaling, security, and reliability.
  • Help develop approaches for training, benchmarking, and evaluating AI agents.

Prolific builds human data infrastructure for AI development, connecting researchers with a global pool of participants to collect high-quality, ethically sourced behavioral data. They are a mission-driven company at the forefront of AI innovation, with a remote culture and a focus on impactful work.

India

  • Build reinforcement learning and agent environments for real customer and Frontier lab use cases, including task specifications, scoring, and evaluation.
  • Develop benchmarks and evaluation harnesses to measure model and data quality across accuracy, robustness, safety, latency, and cost.
  • Run fine-tuning, adapter, and other model experiments to evaluate how data and methods influence model behavior, and deploy local or self-hosted models for evaluation and inference.

Appen has been a leader in AI training data for over 30 years, specializing in human-generated data to train, fine-tune, and evaluate models across generative AI, LLMs, computer vision, and speech recognition. They support model development through an AI-assisted data annotation platform and a global crowd of over 1 million contributors in more than 200 countries, fostering a culture of innovation, collaboration, and humility over ego.

$196,984–$222,000/yr
US

  • Design, build and operate production AI agents, including an autonomous coding agent that works from Jira and Slack, opening PRs and fixing CI failures.
  • Build the retrieval and context layer behind our AI brain: embeddings, semantic search, knowledge-graph modeling, and integration with Snowflake and internal docs.
  • Harden agents against prompt injection and over-broad permissions, and drive cost and latency down using caching and structured outputs.

Weedmaps is a global leader in the cannabis industry, providing a platform for consumers and businesses. The company has a collaborative culture focused on innovation and community, serving the U.S. and worldwide.

Canada Unlimited PTO

  • Design and deliver production-grade agents that investigate, reason, and act on live observability data.
  • Own agent work from rough prototype to production, including evals.
  • Extend the agentic workspace with new canvas capabilities, MCP server, and skills.

Honeycomb provides an observability platform that helps engineers understand and debug complex systems. It is a fully distributed company of over 200 people, with a culture that values impact, autonomy, and inclusivity.

US

  • Design and build autonomous AI agents using modern agentic frameworks to analyze infrastructure and make intelligent decisions.
  • Develop and deploy ML models and MLOps pipelines that learn from infrastructure patterns to optimize resource policies and scaling.
  • Own AI systems end-to-end, from architecture to production, ensuring performance, reliability, and cost-effectiveness.

ScaleOps is redefining autonomous cloud and AI infrastructure by freeing DevOps teams from manual resource management, reducing cloud costs by up to 80%. Backed by over $210M from leading VCs, the company is trusted by Fortune 100 companies and fosters an innovative, mission-driven culture.

North America Europe

  • Build tools for understanding traces, evaluating agent behavior, and turning results into improvements.
  • Run experiments on Mastra's primitives to find where they fall short and improve configuration and behavior.
  • Help turn findings into capabilities developers can use in their own applications.

Mastra is the open-source TypeScript framework for building AI agents. We're a small, fully remote team backed by Y Combinator and Spark Capital with $35M raised.

$150,000–$190,000/yr
Global Unlimited PTO

  • Monitor production health across every graph, catching errors and silent failures proactively.
  • Triage incoming issues, fix small things directly, and route larger problems to the right owner.
  • Track cost, latency, usage, and build business metrics to show the agent's ROI.

LangChain builds the foundation for intelligent agent engineering, helping developers move from prototypes to production-ready AI agents. With over 100M monthly open source downloads and backing from top VCs, the company is at a stage where all team members have meaningful impact.

Latin America

  • Build production Agentforce agents including topics, actions, grounding, guardrails, and evaluation harnesses.
  • Implement actions across Flow, Apex, prompt templates, and external APIs, grounding answers in Data 360.
  • Design guardrails against prompt injection and data leakage, and run evaluation and regression tests before changes.

AspenView Technology Partners builds high-performing nearshore IT teams for North American clients, focusing on innovation and efficiency. They are a people-first, purpose-driven company that believes great culture drives great outcomes.

India

  • Build RL environments, agentic systems, LLM pipelines, and evaluation frameworks for real-world AI use cases.
  • Develop benchmarks and evaluation harnesses to assess model quality across accuracy, safety, latency, and cost.
  • Conduct fine-tuning and model experiments, deploy self-hosted models, and document reproducible methodologies.

The employer is an organization focused on applied AI research, building practical and reusable AI systems for real-world use cases. It values curiosity, accountability, innovation, collaboration, and continuous learning in a remote environment.

Global

  • Build LLM-based agents on the platform's scaffolding, integrating tool calls, internal APIs, and guardrails.
  • Take agents to production on AWS with containers, CI/CD, secrets, permissions, and security controls.
  • Define and run evals, monitor with Langfuse, and document runbooks for independent operation.

Muttdata builds innovative Data Products and Machine Learning solutions to help companies solve complex business challenges. It is a fast-growing, remote-first startup that values collaboration, continuous learning, and a positive, ownership-driven culture.

India

  • Build production AI and agentic workflows end to end, from PRD and architecture through implementation, evaluation, and iteration.
  • Own the retrieval stack, model selection, and evaluation harness to ensure measurable quality, latency, and cost.
  • Design and maintain a reusable component library and set engineering standards for the CoE and partner teams.

QAD | Redzone builds an intelligent manufacturing and supply chain platform that connects people, processes, and data into a single System of Action. It fosters a collaborative culture of smart, hard-working people who support one another and prioritize diversity, equity, and inclusion.

Brazil

  • Design and develop specialized AI agents using Generative AI, LLMs, and agentic frameworks like LangChain and Semantic Kernel.
  • Build MCP servers and integrate AI applications with event-driven architectures, ensuring secure and traceable decisions.
  • Implement Human-in-the-Loop workflows and apply PromptOps practices for scalable, secure, and high-impact AI solutions.

The hiring company is a technology organization in Brazil focused on building AI-powered applications and agentic systems. It fosters a collaborative and international culture, with fully remote work and opportunities to work on innovative AI technologies.

$250,000–$300,000/yr
US Canada Unlimited PTO

  • Own the build out of new agents, skills, and platform capability for teams across TLDR.
  • Build and deploy agents end to end, from design through implementation, evals, and rollout to internal users.
  • Partner with stakeholders across sales, editorial, and people ops to find where an LLM belongs in their process.

TLDR runs the largest network of tech newsletters in the world, with over 8 million subscribers covering startups, software engineering, AI, and more. Our 31-person full-time team is bootstrapped, profitable, and on track for $35M in revenue this year, with a culture of owning functions rather than slices.