The opportunity
HubSpot is an all-in-one marketing, sales, and service software platform that helps businesses grow and succeed. With a user-friendly interface and powerful tools, HubSpot enables businesses to attract, engage, and delight customers, ultimately driving growth and increasing revenue.
What you'll do
Build infrastructure for agent runtime, evaluation, quality measurement, and model optimization.
Create tooling that helps teams understand where their evals are strong,: where coverage is missing, and where users are asking questions the system is not yet prepared to handle.
Design signal pipelines that surface user frustration, agent failure states,: and quality issues before they show up as CSAT drops.
Help define a HubSpot-specific AI benchmark for evaluating model performance on real HubSpot workloads.
Build systems to evaluate new models across agents, improving product: quality, reducing cost, and supporting future model routing.
Make fine-tuning and task-specific optimization more repeatable for HubSpot’s AI use cases.
What they're looking for
- Has deep experience with production ML, LLM, or AI infrastructure.
- Has built systems at scale where reliability, latency, quality, and cost matter.
- Thinks in platforms, not one-off solutions.
- Can work across infrastructure, product, backend, and ML teams.
- Has experience with areas like eval frameworks, model serving, fine-tuning,: signal extraction, experimentation, feedback loops, or model optimization.
- Is excited by ambiguous, foundational problems where many of the answers still need to be invented.
- Can influence senior engineers and leaders through technical credibility and clear judgment.