AI Platform Engineer – Observability (Remote)

Fluid AI

Mumbai1 yr - 5 yrs expNot SpecifiedRemote5 - 10 LPA

Posted 27 days ago

Observability PlatformsMonitoring ToolsLogging SolutionsTracing ToolsAI InfrastructureDockerKubernetesAWSGCPData VisualizationAlerting SystemsReliability EngineeringProblem-solvingProactive ThinkingPythonJavascriptTypescriptfastAPILLMGenerative AI

AI Platform Engineer – Observability (Remote)

Location: Remote (India)

Experience: 1–5 Years

About the Company

We are a cutting-edge Generative AI company building enterprise-grade Agentic AI platforms for leading organizations across banking, financial services, manufacturing, and the public sector. Our AI platform enables autonomous AI agents that can reason, collaborate, integrate with enterprise systems, and automate complex business workflows.

We are looking for engineers who are passionate about building scalable AI infrastructure and developer platforms that power the next generation of enterprise AI applications.

About the Role

We are looking for an AI Platform Engineer – Observability to build the monitoring, debugging, tracing, and performance capabilities of our enterprise Agentic AI Platform.

You will work across Python services, JavaScript applications, Kubernetes infrastructure, and AI workflows to provide deep visibility into how the platform operates in production. Your work will help engineers and customers understand system behavior, identify bottlenecks, troubleshoot issues, and improve the reliability of autonomous AI systems.

This is a platform engineering role focused on building developer tools, observability infrastructure, and performance capabilities that support enterprise-scale AI deployments.

Responsibilities

  1. Design and build observability capabilities for an enterprise Agentic AI platform.
  2. Instrument Python backend services, JavaScript applications, and Kubernetes workloads.
  3. Implement distributed tracing, metrics, and logging across AI agents, APIs, databases, and enterprise integrations.
  4. Build dashboards and visualization tools to monitor platform health and AI workflow execution.
  5. Develop debugging, profiling, and diagnostic tools for autonomous AI systems.
  6. Investigate production issues, identify root causes, and improve system reliability and performance.
  7. Integrate and extend observability technologies such as OpenTelemetry, Grafana, Prometheus, Loki, Tempo, and Jaeger.
  8. Collaborate with platform and product engineering teams to improve developer experience and operational excellence.

Required Skills

  1. 1–5 years of software engineering experience.
  2. Strong programming skills in Python and/or JavaScript/TypeScript.
  3. Experience building backend services or distributed applications.
  4. Understanding of REST APIs, Linux, and networking fundamentals.
  5. Strong debugging, analytical, and problem-solving skills.
  6. Interest in performance engineering, monitoring, and distributed systems.

Good to Have

  1. Experience with FastAPI, Node.js, React, or Next.js.
  2. Experience with Docker and Kubernetes.
  3. Experience with OpenTelemetry or distributed tracing.
  4. Experience with Grafana, Prometheus, Loki, Tempo, Jaeger, or similar observability platforms.
  5. Familiarity with cloud platforms such as AWS, Azure, or GCP.
  6. Exposure to LLMs, AI platforms, or Agentic AI systems.

Why Join Us?

  1. Build the observability layer for a cutting-edge enterprise Agentic AI platform.
  2. Work on modern technologies spanning AI, distributed systems, cloud infrastructure, and Kubernetes.
  3. Solve challenging engineering problems involving performance, scalability, and reliability.
  4. Shape the developer experience and operational tooling that powers enterprise AI deployments.
  5. Join a fast-growing engineering team building AI solutions for leading enterprises across banking, financial services, manufacturing, and the public sector.