Start Your Search Here

push notification bell

Would you like to receive notifications about IT & Computers jobs in Glasgow?

push notification bell

You have blocked notifications

Oops! You have blocked notifications. Click here for more info

You have blocked notifications, please check your browser settings.

push notification bell

You're currently subscribed to job notifications

Want to change your notifications for job alerts?

push notification bell

Subscribe to notifications

You will no longer receive notifications

Job Search

Hackajob

Glasgow / Global

Lead Software Engineer - LLM Ops Platform Reliability

Job Description

hackajob is partnering directly with JPMorganChase to hire for this role.

JOB DESCRIPTION

Help shape how AI systems run reliably in production at scale. In this role, you'll build and operate large language model serving infrastructure, bringing strong engineering fundamentals and site reliability practices to cutting-edge AI platforms. You'll work hands-on with cloud and Kubernetes-based deployments, deep observability, and cost-aware performance tuning. If you enjoy solving hard production problems and making platforms measurably better, you'll find meaningful impact and growth here.

As a Lead Software Engineer at JPMorgan Chase in the AI and Machine Learning Platform team, you will build and scale AI infrastructure that modernizes traditional infrastructure management and site reliability engineering through applied AI. You will own the reliability, performance, and cost-efficiency of the LLM inference platform end to end. You will operate large language model serving stacks (such as vLLM and llm-d) in production at scale, with deep instrumentation and strong operational rigor. You will partner across engineering to deliver secure software, improve stability, and lead incident response and continuous improvement.

Job Responsibilities

Design, develop, troubleshoot, and deliver secure, high-quality production software and services for AI infrastructure

Build backend services and APIs that enable reliable operation of AI infrastructure in production

Operate and scale LLM serving infrastructure (such as vLLM and llm-d), including model hosting, request routing, continuous batching, and KV-cache optimization

Deploy, host, and lifecycle-manage open-source and proprietary LLMs on Amazon EKS and Amazon SageMaker, as well as on-prem and local GPU clusters, using reproducible infrastructure as code and continuous delivery pipelines

Implement observability (logs, metrics, traces) with dashboards and actionable alerting, including Prometheus metrics and Grafana/Alertmanager integration for LLM and GPU workloads

Tune GPU and accelerator capacity, autoscaling, and cost efficiency for LLM inference workloads using performance and optimization techniques (e.g., quantization, parallelism, speculative decoding)

Lead reliability engineering for LLM endpoints through capacity planning, load/soak testing, safe rollouts ...

Apply Now