Hi, I'm Sushant Kulkarni
3+ years of experience architecting production-grade AI systems, autonomous agentic workflows, multi-stage RAG pipelines, and enterprise backend microservices scaled to handle 10,000+ complex real-time conversations.
Architecting Next-Gen AI & High-Performance Systems
Combining deep machine learning orchestration with high-throughput backend infrastructure.
Professional Summary
AI Engineer & Backend Developer with 3+ years of experience building production-grade AI applications and scalable backend systems using Node.js, NestJS, Python, and AWS.
Experienced in designing LLM-powered solutions, RAG pipelines, AI agents, and multimodal workflows while developing high-performance REST APIs, microservices, and cloud-native architectures. Strong expertise in system design, backend optimization, database design, and deploying secure, scalable applications for enterprise environments.
AI Agents & RAG
Multi-agent workflows, vector search (Pinecone, ChromaDB), and function calling via LangChain & LangGraph.
Scalable Backend
Event-driven microservices built on NestJS, FastAPI, WebSockets, Kafka, and Redis caching.
LoRA Fine-Tuning
Specializing domain-specific open-source LLMs & Whisper models to 96% clinical extraction precision.
AWS & Observability
Serverless Lambda, API Gateway, SQS queues, and end-to-end tracing with LangSmith.
Core Engineering Stack
A comprehensive toolkit spanning Artificial Intelligence, System Architecture, and Cloud Infrastructure.
LLMs & Generative AI
Large Language Models (GPT-5, Azure OpenAI, Claude, Llama), AI Agents, Prompt Engineering, Prompt Chaining, Function Calling.
RAG & Vector Search
Retrieval-Augmented Generation (RAG), Semantic Search, Natural Language Processing (NLP), Graph & Vector Databases.
ML Frameworks & Multimodal
Multimodal AI workflows, LLM Evaluation Frameworks, fine-tuning techniques, and ML libraries.
Backend & Microservices
Designing high-performance REST APIs, event-driven architectures, webhooks, and microservices.
Streaming & Caching
WebSocket real-time communication, Kafka event streaming, Redis caching, and Webhooks.
Languages & Frontend
Core programming languages and modern React frontend engineering.
Cloud & DevOps
AWS and Azure cloud infrastructure management, serverless deployment, and Jira project tracking.
Databases
Relational and NoSQL database management systems.
Work Experience
Proven track record of driving AI innovation and technical team leadership.
Software Engineer | Team Lead
- Architected and deployed 3+ production-grade intelligent applications using NestJS, Python, and AWS serverless primitives, scaling backend microservices to process 10,000+ complex healthcare conversations.
- Led 0-to-1 system architecture and project execution for 2 core initiatives, directing PostgreSQL database design, REST API orchestration, and cloud infrastructure from concept to launch.
- Implemented LoRA fine-tuning on open-source LLMs for specialized medical data parsing, substantially improving domain accuracy from 89% to 96% for clinical entity extraction.
- Engineered multi-stage RAG pipelines and autonomous agentic workflows featuring dynamic tool-use and function calling powered by LangChain, LlamaIndex, Pinecone, and ChromaDB.
- Developed a low-latency AI interview agent using WebSockets, Node.js event-driven architecture, and streaming APIs across Anthropic Claude, OpenAI GPT-4o, and Mistral models.
- Integrated computer vision pipelines for live candidate proctoring, leveraging OpenCV and InsightFace for real-time face detection, facial verification and multi-face tracking with more than 90% accuracy.
- Enforced structured payload validation using Pydantic and established dynamic, cost-aware model routing to optimize token usage and minimize API overhead expenses.
- Integrated LangSmith for end-to-end LLM observability, distributed tracing, and prompt evaluation, systematically tracking performance bottlenecks and eliminating output drift.
- Automated end-to-end healthcare intake workflows, reducing manual processing effort by ~80% through structured parsing, Redis caching, and SQS queue management.
- Built candidate onboarding application in Next.js, cutting verification turnaround time from ~1 hour to ~20 minutes.
- Containerized backend services using Docker, AWS Lambda, API Gateway, and S3, maintaining 99.9% system availability.
Co-Founder & Web Developer
- Co-founded and built a client-facing development initiative delivering custom web and backend solutions across multiple business domains.
- Designed and developed end-to-end applications, managing system architecture, database design, backend services, and cloud deployment.
- Built scalable backend systems and REST APIs using Node.js tailored for high reliability in real-world commercial environments.
- Collaborated directly with client leadership to translate complex business requirements into clean, scalable software architecture.
Key Engineering Projects
Production platforms showcasing AI agent workflows, fine-tuning, and robust microservices.
Agent Onboarding System
Architected a production-grade full-stack platform using NestJS and React.js, automating onboarding for 5,000+ agents.
- Reduced end-to-end processing time by 66% (60 → 20 minutes) by engineering automated validation pipelines and streamlined approval workflows.
- Engineered scalable backend services using Prisma ORM, PostgreSQL, and Kafka optimizing schemas and queries for high-performance data retrieval.
- Secured platform with Okta, JWT-based authentication, and Role-Based Access Control (RBAC) for multi-level authorization.
- Improved reliability using Redis for distributed caching & session management, maintaining sub-second API response times under concurrent load.
- Designed modular backend services using Object-Oriented Design and wrote Jest unit tests achieving 90%+ test coverage.
Interview Agent
Architected an AI-powered interview agent using Mistral, LLaMA, and GPT-4, enabling dynamic, context-aware conversations for technical role assessments.
- Engineered a multi-model orchestration pipeline to optimize response latency & cost-efficiency across 1,200+ live interviews.
- Developed a high-performance backend using Python and FastAPI to manage conversation state, prompt chaining, and real-time streaming responses.
- Integrated multimodal stack with Whisper (large-v3) for STT, Deepgram for low-latency TTS, and InsightFace for real-time biometric proctoring.
- Built vision-based AI workflows for automated resume evaluation & candidate scoring, triggering interview pipelines based on profile analysis.
- Orchestrated event-driven workflows for scheduling, evaluation scoring, and automated candidate feedback notifications.
Scribing AI
Developed an AI-powered medical scribe platform, building LLM pipelines and backend services using Python, FastAPI, Node.js, Express.js, and PostgreSQL to transcribe doctor–patient conversations and generate structured clinical summaries.
- Built end-to-end pipelines to extract medical findings and generate SOAP (Subjective, Objective, Assessment, Plan) notes for physicians.
- Processed 10,000+ healthcare conversations with ~96% accuracy in extracting clinically relevant insights.
- Fine-tuned Whisper models with 400K+ trainable parameters using LoRA.
- Engineered scalable APIs, transcript-processing workflows, database integrations, and real-time report generation pipelines.
- Reduced manual documentation effort by automating clinical note generation and standardizing reporting workflows.
Education & Certifications
Academic foundations and verified cloud engineering credentials.
AWS Certified Cloud Practitioner
Amazon Web Services (AWS)
Issued: Apr 2026 – Valid thru: Apr 2029
Bachelor of Engineering - Information Technology
MET BKC College of Engineering, Nashik, India
Graduation: Jul 2020 – Jun 2024
Let's Build Something Intelligent
Interested in AI Agents, RAG architecture, high-scale backends, or engineering leadership? Feel free to reach out.