About The Role
The role owns the architecture, development, and scaling of machine learning systems, driving the transition of advanced AI models from research into high-throughput production environments.
The team collaborates closely with applied scientists and backend engineers to ensure models achieve optimal performance, low latency, and robust reliability under heavy enterprise workloads.
Key Responsibilities
- Design and implement scalable machine learning pipelines for model training, validation, and inference using Python, PyTorch, and distributed computing frameworks
- Deploy, monitor, and scale models on cloud platforms like AWS SageMaker or GCP Vertex AI with automated CI/CD pipelines
- Optimize model inference latency, throughput, and memory footprint through quantization, pruning, and hardware acceleration techniques
- Build feature and data ingestion pipelines handling large-scale datasets, ensuring consistency between training and production feature stores
- Implement comprehensive monitoring frameworks to track model performance, data drift, and anomaly detection in real-time production environments
- Write rigorous unit and integration tests, conduct code reviews, and establish engineering best practices for the broader machine learning team
What We Are Looking For
- 3-6 years of professional software and machine learning engineering experience with a track record of deploying models to production
- Strong proficiency in Python and hands-on experience with deep learning frameworks such as PyTorch or TensorFlow
- Solid understanding of MLOps best practices, containerization with Docker, and orchestration using Kubernetes
- Experience with cloud infrastructure (AWS, GCP, or Azure) and modern feature stores or vector databases
- Bachelor's or Master's degree in Computer Science, Machine Learning, Statistics, or a related quantitative field
- Bonus: Experience fine-tuning large language models, contributing to open-source ML projects, or publishing research at top-tier AI conferences