Technical Lead – AI Solutions

Ganit

Bengaluru, Chennai, Mumbai7 yrs expFull TimeHybrid50 - 80 LPA

Posted 27 days ago

Solution DesignPythonArtificial IntelligenceMachine LearningGenerative AICloud PlatformsAWSAzureGCPNLP

About the Role

We are looking for a Technical Lead – AI Solutions to own the architecture and delivery of production Generative AI systems for enterprise clients. This is a hands-on leadership role: you set the technical direction — model selection, serving strategy, system architecture — and stay close to the code while leading a team of AI engineers. The ideal candidate combines deep engineering fundamentals with practical GenAI depth across fine-tuning, inference and serving, and orchestration, and has the judgment to make sound build-vs-buy, self-hosted vs. API, and cost–quality–latency trade-off decisions.


Key Responsibilities

  1. Technical leadership: own solution architecture, system design, and technical decision-making for AI initiatives; translate business requirements into scalable technical solutions.
  2. Team development: mentor and guide AI engineers, perform design and code reviews, establish engineering best practices, and support hiring.
  3. Design, develop, and ship AI-powered applications — RAG pipelines, agentic workflows, copilots, and conversational AI — remaining hands-on in production code.
  4. AI strategy & architecture decisions: model selection (open-weight vs. API), fine-tune vs. prompt vs. RAG choices, serving strategy, and cost–quality–latency trade-offs, with clear articulation to stakeholders.
  5. Model training & fine-tuning: lead adaptation of open-weight LLMs using SFT, LoRA, and preference tuning, with sound dataset curation and experiment workflows.
  6. Inference & serving: own self-hosted LLM inference using engines such as vLLM, SGLang, or TensorRT-LLM — model onboarding, throughput/latency benchmarking, capacity and cost planning on GPU infrastructure.
  7. Orchestration & MLOps: establish automated pipelines for data processing, training, evaluation, and deployment; model versioning, rollout/rollback, and CI/CD for models and services.
  8. Evaluation & quality ownership: define the evaluation strategy — offline eval harnesses, LLM-as-judge pipelines, golden datasets, tracing, and cost/quality monitoring — and hold releases to it.
  9. Reliability, security & responsible AI: production SLAs and latency budgets, guardrails against prompt injection and data leakage, and compliance-aware deployment for enterprise environments.

Required Skills – Engineering Foundation

  1. 7+ years of hands-on experience building and shipping software and AI/ML systems in production, with recent years focused on Generative AI.
  2. Strong Python (async, typing, packaging); working knowledge of Go or Node.js is a plus.
  3. Excellent grounding in Data Structures, Algorithms, and OOP, with proven depth in solution design, system design, and software architecture.
  4. Experience building scalable microservices, REST/streaming APIs, and distributed systems.
  5. SQL/NoSQL databases, caching, and messaging platforms.
  6. Hands-on experience with AWS, Azure, or GCP, including GPU infrastructure and managed AI services.
  7. Docker, Kubernetes, Git, and CI/CD pipelines; comfort owning services in production.


AI / LLM Skills (Mandatory – Must Haves)

  1. Track record of shipping multiple Generative AI / LLM systems to production — RAG, AI agents, fine-tuned models, or conversational systems — including at least one led end to end.
  2. Model training & fine-tuning: SFT with LoRA/QLoRA on open-weight models, preference tuning (DPO), dataset curation, and experiment tracking; familiarity with the Hugging Face ecosystem.
  3. LLM inference & serving: practical experience with at least one of vLLM, SGLang, TensorRT-LLM, TGI, or Ollama/llama.cpp; working knowledge of continuous batching, PagedAttention/KV-cache management, chunked prefill, prefix caching, quantized serving, and prefill vs. decode phases.
  4. Architecture judgment: self-hosted vs. API-based serving, model selection, and cost–quality–latency trade-offs at production scale.
  5. RAG architecture: embedding models, vector databases, hybrid search, chunking strategies, and reranking.
  6. Agentic orchestration: tool/function calling, structured outputs, and multi-step agent workflows using agent frameworks or custom loops; familiarity with MCP is a plus.
  7. Pipeline orchestration & MLOps: workflow orchestrators, model/artifact versioning, automated evaluation gates, and deployment automation.
  8. Evaluation-driven development: systematic evals, regression guarding, and prompt engineering discipline across a team.


Leadership Skills (Mandatory)

  1. Experience leading technical teams — mentoring engineers, driving design/code reviews, and raising the engineering bar.
  2. Strong stakeholder management: communicating technical trade-offs, timelines, and risks to product and business leadership.
  3. Ability to lead technical initiatives while remaining actively involved in development.


Preferred / Nice to Have

  1. Distributed and large-scale training: DeepSpeed/FSDP, multi-GPU training, MoE fine-tuning, RLHF/GRPO.
  2. Inference environments: GPU memory budgeting, latency/throughput trade-offs, autoscaling model replicas, and cost economics of self-hosted vs. API-based serving.
  3. Quantization workflows: AWQ, GPTQ, FP8/INT8 calibration, and quality–latency trade-off evaluation.
  4. Voice/multimodal AI: ASR/TTS integration, streaming pipelines, real-time latency budgets.
  5. Experience with AI governance, responsible AI practices, or regulated-industry deployments.
  6. Bachelor's or Master's degree in Computer Science, IT, or a related field.


Ideal Candidate Profile

  1. 7–12 years of experience building AI/ML and software systems, with strong, clean, hands-on coding skills.
  2. Has led LLM systems from prototype to production, with ownership across training, serving, and application layers.
  3. Strong intuition for inference performance and cost — batching, KV cache, quantization, and GPU utilization.
  4. Earns the team's trust technically while managing stakeholders and delivery cross-functionally