Job Description
Ready to take your AI engineering expertise to the next level? This is an opportunity to join a technically driven team building and deploying production-grade AI and machine learning solutions not simply experimenting with the latest models.
You’ll have the opportunity to fine-tune models such as Llama, Mistral and Qwen , build self-hosted inference environments, optimise AI systems for performance and cost, and develop scalable machine learning solutions using modern cloud, containerisation and MLOps practices.
What we are looking for
• 3+ years of experience building and shipping machine learning or AI-powered systems in production
• 7 plus years’ commercial development experience
• Strong Python skills, with hands-on experience using frameworks such as PyTorch, TensorFlow, or similar
• Practical experience with large language models - prompting, fine-tuning, RAG, or agent frameworks (e.g., LangChain, LlamaIndex, or custom implementations).
• Hands-on experience training and fine-tuning open-source models (e.g., Llama, Mistral, Qwen) full fine-tuning or parameter-efficient methods (LoRA/QLoRA) and deploying them on your own infrastructure rather than relying solely on hosted APIs.
• Experience standing up self-hosted inference serving on your own infra (e.g., vLLM, TGI, Triton, Ray Serve), including GPU provisioning, batching, and cost/latency optimization
• Solid understanding of ML fundamentals: model evaluation, overfitting, data leakage, and experimentation methodology
• Experience with vector databases and embedding-based retrieval (e.g., Pinecone, Weaviate, pgvector, FAISS)
• Comfort working with cloud infrastructure (AWS, GCP, or Azure) and containerized deployments (Docker, Kubernetes), including GPU-backed compute.
• Familiarity with MLOps practices model versioning, CI/CD for ML, monitoring, and rollback strategies for both hosted and self-hosted models
• Strong software engineering fundamentals: clean code, testing, code review, and API design
• Proficiency working with AI coding assistants (e.g., Claude Code, GitHub Copilot, Cursor) to accelerate development, while critically reviewing and validating AI-generated code
Nice to Have
• Experience with model quantization, distillation, or optimization techniques (e.g., GPTQ, AWQ, ONNX) for efficient self-hosted inference
• Background in NLP, computer vision, or recommendation systems.
• Experience with streaming data pipelines (Kafka, Spark) or feature stores.
• Contributions to open-source ML/AI tooling
• Prior experience in a startup or fast-growth environment
Reference Number for this position is GZ61533 which is a Permanent remote position offering a salary of R1000k per annum , negotiable based on experience and ability.
Contact Garth on ***email_hidden*** or call 011 463 3633 to discuss this and other opportunities.
Are you ready for a change of scenery? We are a specialist niche recruitment agency. We offer our candidates a wide range of opportunities, ensuring we successfully match the right technology professionals with the right roles.
Do you know a developer or technology specialist looking for a new opportunity? We pay cash for successful referrals!