The opportunity
As a Senior Software Engineer, AI Ops at Scale AI, you will own the long-term technical health, performance, and stability of AI solutions deployed across our strategic public sector partners.
What you'll do
Handover Gate & Onboarding: Act as the technical gatekeeper during the formal transition from Delivery to Maintenance. Conduct deep-dive reviews to ensure baseline code, prompts, and architecture meet strict maintainability and documentation standards before sign-off.
Tiered SLA & Incident Management: Own technical response and resolution targets across multi-tiered service models (from Business-Hours Essential to 24/7 Mission-Critical). Lead Incident Governance, Root Cause Analysis (RCA), and P1/P2 mitigations within strict active support windows.
AI Lifecycle Governance: Monitor production model performance, latency, and data drift. Manage prompt configuration repositories to maintain behavioral consistency and perform regression testing when LLM providers update underlying endpoints.
Request Classification & Technical Scope: Operationalize the boundary between Routine Maintenance (In-Scope) and System Evolution (Out-of-Scope). Assess incoming client requests and run comparative benchmarking on new AI models.
Automation & Reliability Engineering: Eliminate operational toil by engineering self-healing data pipelines, automated RAG indexing syncs, and telemetry tooling. Influence upstream "Delivery" teams to adopt architectural patterns that simplify ongoing maintenance.
Client Technical Interface: Serve as the senior technical point of contact for government and enterprise IT leads. Translate technical AI concepts (data drift, prompt versioning, API deprecation) into clear business impacts for non-technical stakeholders.
What they're looking for
- Background: 5+ years in Software Engineering, MLOps, SRE, or Forward Deployed: Engineering in heavy data or production AI environments.
- Technical Stack: Advanced proficiency in Python, SQL, REST/gRPC APIs, and cloud architecture (AWS, Azure, or GCP). Hands-on experience with MLOps tooling, vector databases, and LLM orchestration frameworks (e.g., LangChain, LlamaIndex).
- AI Governance Expertise: Practical understanding of prompt version control, model benchmarking against evaluation datasets, RAG pipeline mechanics, and data drift detection.
- Engineering Mindset: A drive to build systematic, automated fixes rather than applying temporary patches. Strong grasp of CI/CD for machine learning pipelines.