Back to Developers
Rajasri

Rajasri

Associate Member of Technical Staff

Hyderabad, India 1+ yrs exp 83 · Excellent

About

Associate Member of Technical Staff with experience building production-grade LLM gateways, inference platforms, and cloud-native ML systems. Skilled in Python, FastAPI, Kubernetes, vLLM, LoRAX, OpenAI, Anthropic, Google Gemini, asynchronous inference, multi-tenant APIs, and Kubeflow with expertise in scalable backend and ML infrastructure.

Skills & Expertise (28)

Python Intermediate
7.5/10
1
Years Exp
FastAPI Intermediate
7.0/10
1
Years Exp
Kubernetes Intermediate
7.0/10
1
Years Exp
REST APIs Intermediate
6.5/10
1
Years Exp
Celery Intermediate
6.5/10
1
Years Exp
Docker Intermediate
6.5/10
1
Years Exp
Kubeflow Intermediate
6.5/10
1
Years Exp
SQL Intermediate
6.0/10
1
Years Exp
GitHub Actions Intermediate
6.0/10
1
Years Exp
NumPy Intermediate
6.0/10
1
Years Exp
Pandas Intermediate
6.0/10
1
Years Exp
scikit-learn Intermediate
6.0/10
1
Years Exp
Microservices Intermediate
6.0/10
1
Years Exp
vLLM Ollama Google Gemini API CNN PCA Regression Feast API Design Asynchronous Processing RabbitMQ HashiCorp Vault JWT Authentication Multi-tenant Architecture Circuit Breaker Pattern MySql

Work Experience

Associate Member of Technical Staff

Gaian Solutions (Mobius)

Nov 2025 - Present

Build a production-grade OpenAI-compatible LLM Gateway using Python, FastAPI, and Kubernetes, unifying inference across 4+ backends – Ollama, vLLM, LoRAX, and ONNX – through a declarative API and routing architecture. Integrated 3 cloud LLM providers – Anthropic Claude, Google Gemini, and OpenAI – through a unified CloudLLMService, standardizing provider APIs, responses, and token accounting while supporting prompt caching, extended thinking, Google Search grounding, and asynchronous batch inference with submission, polling, cancellation, JSONL processing, result retrieval, and a configurable circuit breaker for resilient provider throttling and failures. Engineered a secure, configuration-driven LLM inference platform supporting 4+ workloads including chat, completions, embeddings, and LoRA adapters, with JWT-based multi-tenant authentication, HashiCorp Vault credential isolation, and per-tenant/per-product API keys; automated GPU-backed Kubernetes deployments of vLLM and LoRAX using the Kubernetes Python SDK and Jinja2 with namespace validation, secret injection, service-readiness polling, and REST-based health monitoring. Built asynchronous multimodal GenAI services for 3 workloads – image-to-image, image-to-3D, and text-to-image – using Celery and RabbitMQ, enabling scalable background processing for AI workloads. Developed 4 end-to-end ML pipelines using Kubeflow and Elyra, covering PCA, linear regression, logistic regression, and CNN workflows across data preprocessing, model training, and evaluation.

Full Stack Intern – Agent Orchestration Framework

Gaian Solutions (Mobius)

Apr 2025 - Oct 2025

Developed FastAPI routes and LLM provider integrations for the Agent Orchestration Framework, implementing agent creation workflows, API integration, and robust error handling.

Education

B.Tech, Computer Science & Data Science - Kakinada Institute of Engineering and Technology

- 2024 · Afghanistan

Certifications

No certifications added yet

Interested in this developer?

Profile Score Breakdown

📷 Photo 10/10
📄 Resume 10/10
💼 Job Title 10/10
✍️ Bio 10/10
🛠️ Skills 20/20
🎓 Education 10/10
⏱️ Experience 8/15
💰 Rate 0/5
🏆 Certs 0/5
Verified 5/5
Total Score 83/100

Profile Overview

Member sinceAug 2026

Skills (28)

Click a skill to find developers with the same skill