Callosum
Greater London / Global
You have blocked notifications
Oops! You have blocked notifications. Click here for more info
You have blocked notifications, please check your browser settings.
You're currently subscribed to job notifications
Subscribe to notifications
You will no longer receive notifications
Greater London / Global
Callosum in London is hiring for a role owning end-to-end performance for our inference platforms. You will manage KV caches, batching, memory, and scheduling across heterogeneous hardware to scale model serving.
The candidate should have deep LLM inference knowledge, strong distributed system experience, and low-level debugging skills on GPUs, networks, and Linux. Visa sponsorship and relocation are available; on-site in London.
#J-18808-LjbffrGreater London / Global
, United Kingdom / Global
Greater London / Global
Greater London / Global
Greater London / Global
Greater London / Global