We are looking for a Technical Lead – AI Solutions to own the architecture and delivery of production Generative AI systems for enterprise clients. This is a hands-on leadership role: you set the technical direction — model selection, serving strategy, system architecture — and stay close to the code while leading a team of AI engineers. The ideal candidate combines deep engineering fundamentals with practical GenAI depth across fine-tuning, inference and serving, and orchestration, and has the judgment to make sound build-vs-buy, self-hosted vs. API, and cost–quality–latency trade-off decisions.
Key Responsibilities
Technical leadership: own solution architecture, system design, and technical decision-making for AI initiatives; translate business requirements into scalable technical solutions.
Team development: mentor and guide AI engineers, perform design and code reviews, establish engineering best practices, and support hiring.
Design, develop, and ship AI-powered applications — RAG pipelines, agentic workflows, copilots, and conversational AI — remaining hands-on in production code.
AI strategy & architecture decisions: model selection (open-weight vs. API), fine-tune vs. prompt vs. RAG choices, serving strategy, and cost–quality–latency trade-offs, with clear articulation to stakeholders.
Model training & fine-tuning: lead adaptation of open-weight LLMs using SFT, LoRA, and preference tuning, with sound dataset curation and experiment workflows.
Inference & serving: own self-hosted LLM inference using engines such as vLLM, SGLang, or TensorRT-LLM — model onboarding, throughput/latency benchmarking, capacity and cost planning on GPU infrastructure.
Orchestration & MLOps: establish automated pipelines for data processing, training, evaluation, and deployment; model versioning, rollout/rollback, and CI/CD for models and services.
Evaluation & quality ownership: define the evaluation strategy — offline eval harnesses, LLM-as-judge pipelines, golden datasets, tracing, and cost/quality monitoring — and hold releases to it.
Reliability, security & responsible AI: production SLAs and latency budgets, guardrails against prompt injection and data leakage, and compliance-aware deployment for enterprise environments.
Required Skills – Engineering Foundation
7+ years of hands-on experience building and shipping software and AI/ML systems in production, with recent years focused on Generative AI.
Strong Python (async, typing, packaging); working knowledge of Go or Node.js is a plus.
Excellent grounding in Data Structures, Algorithms, and OOP, with proven depth in solution design, system design, and software architecture.
Experience building scalable microservices, REST/streaming APIs, and distributed systems.
SQL/NoSQL databases, caching, and messaging platforms.
Hands-on experience with AWS, Azure, or GCP, including GPU infrastructure and managed AI services.
Docker, Kubernetes, Git, and CI/CD pipelines; comfort owning services in production.
AI / LLM Skills (Mandatory – Must Haves)
Track record of shipping multiple Generative AI / LLM systems to production — RAG, AI agents, fine-tuned models, or conversational systems — including at least one led end to end.
Model training & fine-tuning: SFT with LoRA/QLoRA on open-weight models, preference tuning (DPO), dataset curation, and experiment tracking; familiarity with the Hugging Face ecosystem.
LLM inference & serving: practical experience with at least one of vLLM, SGLang, TensorRT-LLM, TGI, or Ollama/llama.cpp; working knowledge of continuous batching, PagedAttention/KV-cache management, chunked prefill, prefix caching, quantized serving, and prefill vs. decode phases.
Architecture judgment: self-hosted vs. API-based serving, model selection, and cost–quality–latency trade-offs at production scale.
Agentic orchestration: tool/function calling, structured outputs, and multi-step agent workflows using agent frameworks or custom loops; familiarity with MCP is a plus.
Evaluation-driven development: systematic evals, regression guarding, and prompt engineering discipline across a team.
Leadership Skills (Mandatory)
Experience leading technical teams — mentoring engineers, driving design/code reviews, and raising the engineering bar.
Strong stakeholder management: communicating technical trade-offs, timelines, and risks to product and business leadership.
Ability to lead technical initiatives while remaining actively involved in development.
Preferred / Nice to Have
Distributed and large-scale training: DeepSpeed/FSDP, multi-GPU training, MoE fine-tuning, RLHF/GRPO.
Inference environments: GPU memory budgeting, latency/throughput trade-offs, autoscaling model replicas, and cost economics of self-hosted vs. API-based serving.
Quantization workflows: AWQ, GPTQ, FP8/INT8 calibration, and quality–latency trade-off evaluation.