Job Description
Introduction
Our client builds intelligent software products that helps customers make better decisions faster. They're a small, international, fast-moving team that cares deeply about shipping reliable AI systems, not just impressive demos. As an AI Engineer, you'll be at the center of that effort — designing, building, and productionizing the machine learning and generative AI capabilities that power the platform.
About the Role
We're looking for an AI Engineer who can move fluidly between research and production: someone who can prototype a new model or prompting approach quickly, then harden it into a reliable, observable, cost-efficient service that runs at scale. You'll work closely with product and full-stack engineering to turn ambiguous problems into shipped AI features.
Duties & Responsibilities
What You'll Do
• Design, build, and deploy machine learning and LLM-based systems, including retrieval-augmented generation (RAG) pipelines, fine-tuned models, and agentic workflows.
• Own the full lifecycle of AI features: data collection and evaluation, model/prompt selection, experimentation, deployment, and monitoring in production.
• Build and maintain data pipelines, embeddings stores, and vector databases that support search, retrieval, and personalization.
• Evaluate and integrate third-party model APIs (OpenAI, Anthropic, etc.) alongside open-source and self-hosted models, balancing quality, latency, and cost.
• Fine-tune, train, and deploy open-source models on our own infrastructure — including data prep, training/fine-tuning runs, quantization, and self-hosted inference serving at scale.
• Establish evaluation frameworks and metrics to measure model quality, drift, hallucination rate, and business impact over time.
• Collaborate with full-stack engineers to expose AI capabilities through clean, well-documented APIs and SDKs.
• Implement guardrails, safety checks, and monitoring for AI systems running in production.
• Stay current with the fast-moving AI/ML landscape and bring back practical recommendations on tools, models, and techniques.
• Write clear technical documentation and communicate trade-offs to both technical and non-technical stakeholders.
Desired Experience & Qualification
What We're Looking For
• 4+ years of experience building and shipping machine learning or AI-powered systems in production.
• Strong Python skills, with hands-on experience using frameworks such as PyTorch, TensorFlow, or similar.
• Practical experience with large language models — prompting, fine-tuning, RAG, or agent frameworks (e.g., LangChain, LlamaIndex, or custom implementations).
• Hands-on experience training and fine-tuning open-source models (e.g., Llama, Mistral, Qwen) — full fine-tuning or parameter-efficient methods (LoRA/QLoRA) — and deploying them on your own infrastructure rather than relying solely on hosted APIs.
• Experience standing up self-hosted inference serving on your own infra (e.g., vLLM, TGI, Triton, Ray Serve), including GPU provisioning, batching, and cost/latency optimization.
• Solid understanding of ML fundamentals: model evaluation, overfitting, data leakage, and experimentation methodology.
• Experience with vector databases and embedding-based retrieval (e.g., Pinecone, Weaviate, pgvector, FAISS).
• Comfort working with cloud infrastructure (AWS, GCP, or Azure) and containerized deployments (Docker, Kubernetes), including GPU-backed compute.
• Familiarity with MLOps practices — model versioning, CI/CD for ML, monitoring, and rollback strategies for both hosted and self-hosted models.
• Strong software engineering fundamentals: clean code, testing, code review, and API design.
• Proficiency working with AI coding assistants (e.g., Claude Code, GitHub Copilot, Cursor) to accelerate development, while critically reviewing and validating AI-generated code.
• Excellent communication skills and comfort working in a remote, async-friendly environment.
Nice to Have
• Experience with model quantization, distillation, or optimization techniques (e.g., GPTQ, AWQ, ONNX) for efficient self-hosted inference.
• Background in NLP, computer vision, or recommendation systems.
• Experience with streaming data pipelines (Kafka, Spark) or feature stores.
• Contributions to open-source ML/AI tooling.
• Prior experience in a startup or fast-growth environment.
Why Join
• Work on AI problems that ship to real customers, not just internal experiments.
• High ownership, high trust — you'll shape how we build and evaluate AI systems from the ground up.
• Will start as a fully remote and will move to a hybrid model, with flexible working hours.
• Competitive salary, equity, and benefits.