Actively recruiting / 7 applicants
We’re here to help you
Jane Cervantes is in direct contact with the company and can answer any questions you may have. Email
Jane Cervantes, RecruiterAbout the Role
We're looking for a Machine Learning Engineer to fine-tune open-source LLMs and deploy them as fast, cost-efficient production services. You'll own the full path from dataset preparation and training through quantisation, serving and performance optimisation.
Key Responsibilities
- Fine-tune open-source LLMs and SLMs (e.g., Llama, Qwen, Mistral, Gemma, DeepSeek) using SFT, LoRA/QLoRA and other PEFT techniques
- Build, clean and validate domain-specific instruction datasets
- Configure and run GPU training, including hyperparameters, mixed precision (BF16/FP16), gradient accumulation and checkpointing
- Evaluate fine-tuned models against baselines using measurable criteria
- Quantise and optimise models for production (AWQ/GPTQ, INT8/INT4, KV cache, continuous batching)
- Deploy and serve self-hosted models with vLLM, TGI, Triton or similar
- Benchmark and improve latency, throughput, GPU utilisation and inference cost
- Expose models through production APIs (FastAPI or similar)
Must-Haves
- Strong Python and PyTorch
- Hands-on experience fine-tuning LLMs with Hugging Face Transformers, PEFT and TRL (or Axolotl/Unsloth)
- Solid grasp of transformer architecture and training fundamentals
- Experience deploying LLMs on GPU infrastructure with vLLM, TGI, Triton or equivalent
- Practical experience with quantisation and inference optimisation
- Experience with Docker and Linux environments
Nice-to-Have
- RAG, vector databases and AI agent development
- Automated evaluation pipelines and LLM metrics (hallucination, faithfulness)
- MLOps tooling: MLflow, W&B, Kubernetes, CI/CD for ML
- AWS: SageMaker, EKS, Bedrock
- Distributed training (DeepSpeed, multi-GPU), knowledge distillation, CUDA fundamentals
- Experience with A100/H100/L40S GPUs