Design and develop scalable backend services and REST APIs using Python, FastAPI/Flask/Django.
Build and integrate Generative AI/LLM-based applications, including RAG, AI agents, embeddings, and tool/function calling.
Work with Azure OpenAI and other LLM providers.
Use AI frameworks such as LangChain, LlamaIndex, LangGraph, or similar.
Implement LiteLLM for LLM gateway, routing, fallback, and usage/cost management.
Implement rate limiting, throttling, caching, retries, and API quotas for scalable services.
Use Langfuse for LLM observability, tracing, prompt management, token/cost tracking, and evaluation.
Design and work with PostgreSQL/MySQL, MongoDB, and Redis.
Deploy and manage applications using Microsoft Azure, Docker, and CI/CD.
Develop unit/integration tests, troubleshoot production issues, and optimize performance and cost.
Participate in system design, code reviews, and technical architecture discussions.
Required
5–7 years of software development experience with strong Python and backend development skills and hands-on experience in Generative AI/LLM applications.
Strong understanding of REST APIs, microservices, RAG, embeddings, vector databases, prompt engineering, AI agents, rate limiting, and distributed systems.
Hands-on experience with Azure OpenAI, LiteLLM, and Langfuse is preferred, along with exposure to LangChain/LlamaIndex/LangGraph or similar AI frameworks.
Experience with Microsoft Azure, Docker, CI/CD, databases, Redis, API security, and cloud-native application development is required.
Strong problem-solving, communication, ownership, and ability to build production-grade AI solutions are essential.
Innovation isn’t powered by machinery. It’s powered by people. Together, we think bigger. We listen and learn. We explore and solve. We’re a team of many backgrounds and places, working toward one common goal: To help customers build the tangible future, where innovation improves quality of life and the benefits of technology are more accessible to all. Join us.