About the position
ENVIRONMENT:
Our client is seeking highly specialised AI Platform Engineers to design, build, operate and optimise enterprise-grade AI infrastructure within a complex, regulated environment. This is not a general cloud engineering, IT infrastructure or data science role. The successful candidates must have hands-on experience supporting production AI workloads across multi-cloud environments and must be comfortable working across AI infrastructure, model serving, agentic AI, security, observability, infrastructure-as-code and AI cost governance.
DUTIES:
- Design, deploy and optimise scalable multi-cloud AI platform infrastructure.
- Build reusable platform components for AI gateways, model serving, vector databases, data pipelines and GPU workloads.
- Develop and maintain infrastructure-as-code using Terraform, Pulumi, CloudFormation or equivalent technologies.
- Design infrastructure supporting agentic AI, including orchestration environments, tool-calling, agent memory, state management and multi-agent communication.
- Implement cloud-agnostic model-serving patterns that support workload portability.
- Define and manage AI platform SLAs covering availability, inference latency, throughput and reliability.
- Implement platform observability, monitoring, incident management, release management and operational runbooks.
- Design and implement zero-trust security controls for AI platforms.
- Manage AI compute expenditure through cost attribution, chargeback/showback, workload optimisation and usage reporting.
- Maintain technical documentation, architectural decision records and governance evidence.
- Mentor engineers and contribute to platform engineering standards and delivery practices.
REQUIREMENTS:
- Senior level: approximately 5–8 years of relevant cloud and AI platform engineering experience.
- Lead/Principal level: approximately 8–12 years of relevant experience, including technical leadership and responsibility for engineering teams or platform squads.
Mandatory Technical Experience:
- Candidates must demonstrate meaningful production experience in most of the following:
- At least two of the following AI ecosystems:
- AWS Bedrock or SageMaker
- Microsoft Azure AI Foundry or Azure OpenAI
- Databricks AI
- Enterprise Hugging Face deployments
- Kubernetes, Docker, Helm and containerised platform services.
- Terraform, Pulumi, AWS CDK, CloudFormation or equivalent infrastructure-as-code.
- CI/CD and automated deployment of cloud or AI platform components.
- Production model-serving infrastructure, AI gateways or inference endpoints.
- Platform observability using tools such as Prometheus, Grafana, Datadog, OpenTelemetry or Databricks Lakehouse Monitoring.
- Cloud security, identity and access management, including OAuth/OIDC, JWT, RBAC or ABAC.
- Production incident management, SLAs, release management and operational readiness.
- Experience within banking, financial services or another highly regulated enterprise environment.
Specialist AI Experience:
- Candidates should demonstrate practical experience in one or more of the following:
- Agent orchestration frameworks such as LangGraph, AutoGen, AWS Bedrock Agents or Microsoft Foundry Agent Service.
- Model Context Protocol, tool-calling APIs and agent state or memory management.
- Retrieval-augmented generation and vector database infrastructure.
- Cloud-agnostic model serving using tools such as ONNX, BentoML, Triton Inference Server or vLLM.
- MLOps platforms such as MLflow, Kubeflow or Airflow.
- GPU cluster management and inference or training workload optimisation.
- Prompt-injection prevention, output filtering, data-exfiltration controls and AI threat modelling.
AI Finops Experience:
- Candidates should have experience with some combination of:
- AI or cloud cost attribution and tagging.
- Chargeback and showback models.
- Token, GPU, DBU or provisioned-throughput cost management.
- Rightsizing, workload scheduling and reserved or spot-instance optimisation.
- Cost dashboards, anomaly detection and cost-per-use-case reporting.
- Communicating technical cost trade-offs to senior technology, business or finance stakeholders.
Qualifications:
- Postgraduate qualification in Computer Science, Information Technology, Data Science, Mathematics, Statistics, Engineering or a related quantitative field.
- A Master’s degree is preferred and may be required for certain senior appointments.
- Relevant certifications are strongly preferred, including:
- AWS Solutions Architect Professional or AWS Machine Learning
- Microsoft Azure AI Engineer
- FinOps Certified Practitioner
- Certified Cloud Security Professional or equivalent
- HashiCorp Terraform Associate
- Kubernetes certification
Desired Skills: