SR

Hire Sravanthi R. - Fine-Tuning Engineer

Machine Learning Engineer

7+ years
san francisco, california, united states
Scale Ai

About sravanthi

I’m an AI/ML Engineer with over 5 years of experience building and deploying large-scale AI systems across e-commerce, media, and accessibility domains. My work focuses on optimizing LLM fine-tuning, RAG pipelines, and real-time personalization using cloud-native and GPU-accelerated architectures. At Scale AI and Meta, I designed Llama-based inference pipelines, integrated FAISS vector search, and developed data streaming workflows with Kafka and PySpark—achieving measurable improvements in inference speed, accuracy, and customer engagement. I’m passionate about scalable model deployment, GPU optimization, and applied research in GenAI. I love transforming complex machine learning models into production-grade solutions that directly impact user experience and business growth. Core Skills: LLM Fine-Tuning (Llama 2, GPT-4), RAG, PyTorch, DeepSpeed, Triton, FAISS, AWS SageMaker, Kubernetes, Docker, PySpark, MLflow, MLOps

Key Skills

apache kafkaaws lambdaaws sagemakerazure databrickscloud computingcomputer visiondata engineeringgpu optimizationllm fine tuningmlopsnlpopencvpysparkpythonpytorch+2 more

Experience

Machine Learning Engineer

Current

Scale Ai

* Spearheaded LLM fine-tuning pipeline on AWS, improving enterprise model inference accuracy by 42%, reducing retraining * latency by 35%, significantly enhancing customer GenAI solution adoption. * Optimized real-time media personalization and recommendation workflows, increasing user engagement by 38% and ad * revenue by 26%, directly contributing to measurable business growth metrics. * Implemented RAG retrieval pipelines for multi-terabyte datasets, boosting query response performance by 48% and * accelerating knowledge delivery for enterprise AI applications across diverse industries. * Designed and deployed large-scale LLM inference endpoints on AWS SageMaker, leveraging DeepSpeed and PEFT for * efficient fine-tuning and GPU utilization across multi-tenant enterprise environments. * Built streaming data pipelines using Apache Kafka and PySpark to feed real-time user interactions into embedding-based * recommendation and personalization layers for GenAI applications. * Integrated vector search and retrieval systems with FAISS and custom embedding pipelines, enabling low-latency RAG * responses and scalable AI-driven content recommendations. * Developed model evaluation frameworks using MLflow, custom metrics, and automated safety checks to monitor * hallucinations, bias, and drift across multiple fine-tuned foundation models. * Implemented Python SDKs and REST APIs to support enterprise consumption of GenAI services, enabling seamless * integration with existing applications and internal tooling on AWS. * Architected GPU cluster autoscaling and container orchestration with AWS EKS, Docker, and Kubernetes, ensuring reliable * real-time model serving and optimized resource utilization. * Built end-to-end media preprocessing pipelines using OpenCV and FFmpeg, integrating annotation workflows through Scale * Rapid, enabling high-quality datasets for supervised and generative AI tasks.

Machine Learning Engineer

Meta

* Spearheaded Llama 2 fine-tuning pipelines on AWS, improving inference throughput by 42% and reducing first-token * latency across enterprise applications, enabling scalable real-time AI-driven content personalization. * Optimized RAG pipelines integrating Llama 2 with media embeddings, increasing customer engagement predictions by 38% * and accelerating e-commerce ad relevance, contributing directly to revenue growth and ROI metrics. * Designed token-level caching and dynamic batching in Llama 2 inference, lowering GPU costs by 27% while maintaining * model accuracy, supporting Meta’s large-scale AdTech and personalization workloads. * Engineered distributed Llama 2 inference architecture on AWS GPU clusters using PyTorch 2, DeepSpeed-MII, and Triton, * ensuring low-latency responses for high-demand real-time ML applications. * Developed multimodal media pipelines combining SAM and ImageBind embeddings with Llama 2 on AWS, enabling scalable * text-image/video integration for advanced AI personalization solutions. * Built vector search pipelines using FAISS integrated with Llama 2 on AWS, supporting retrieval-augmented generation, * semantic search, and cross-modal recommendation systems for e-commerce and content applications. * Implemented quantization and LoRA-based fine-tuning for Llama 2 models, leveraging FlashAttention2 and vLLM * frameworks to optimize inference performance across AWS GPU workloads. * Orchestrated real-time streaming data pipelines via Kafka and AWS S3, integrating Llama 2 with media and transactional * inputs for personalized recommendations and AdTech scoring systems. * Designed monitoring and observability stack on AWS using Prometheus, Grafana, and OpenTelemetry to track Llama 2 * inference performance, model latency, and system resource utilization across production environments.

Software Engineer

Accenture In India

* Improved real-time object detection accuracy by 23% on edge devices, enabling visually impaired users to safely navigate * urban environments with low-latency AI-driven audio guidance. * Optimized OCR and NLP pipelines, increasing scene-text recognition efficiency by 28%, accelerating narration generation, * and significantly enhancing accessibility experience for over 100 pilot users. * Deployed AI personalization models on AWS cloud, reducing response latency by 35% and adapting speech and object * detection dynamically to individual user preferences in real time. * Developed and deployed PyTorch and TensorFlow computer vision models, leveraging TFLite for mobile edge inference and * AWS SageMaker endpoints for large-scale scene understanding. * Engineered asynchronous media streaming pipelines using Python asyncio and WebRTC, integrating OCR, CV, and TTS * modules to deliver real-time multimodal accessibility feedback. * Implemented NLP/NLU and speech synthesis layers using HuggingFace Transformers, spaCy, and AWS Polly, enabling * personalized audio narration with low-latency real-time performance. * Built CI/CD pipelines and model versioning with Docker, Kubernetes, and AWS Lambda, ensuring continuous integration of * AI models, automated deployment, and robust monitoring of inference workflows.

Education

Concordia University - St. Paul

Masters

Concordia University - St. Paul

Masters

Osmania University

Bachelors

Interested in connecting with sravanthi?

Sign up for NinjaHire to send a connection request.

Common Questions

What is sravanthi's expertise?

sravanthi specializes in Fine-Tuning Engineer, with expertise in apache kafka, aws lambda, aws sagemaker, azure databricks, cloud computing.

Where is sravanthi located?

sravanthi is based in san francisco, california, united states.

How much experience does sravanthi have?

sravanthi has 7+ years of professional experience.

How can I contact sravanthi?

You can connect with sravanthi through NinjaHire by signing up for a free account.

Looking for a different Fine-Tuning Engineer?

Describe exactly who you need and NinjaHire will source them for you.

Type a role to try NinjaHire for free