About harshitha
I’m an AI/ML Engineer with 5+ years of experience building production-grade NLP and Generative AI systems, with a focus on making models reliable, measurable, and scalable in real-world environments. At Scale AI, I work on LLM-driven document intelligence platforms, designing systems that combine retrieval and generation to answer questions and generate insights from large-scale enterprise data. A big part of my work involves improving output quality through evaluation frameworks, prompt optimization, and human-in-the-loop feedback, while also keeping systems efficient and cost-effective. Previously at Amazon, I worked on multilingual NLP systems powering Alexa and e-commerce experiences, where I focused on improving intent detection and making models perform better across diverse languages and real-world user inputs. What interests me most is solving the gap between “models that work in notebooks” and “systems that work in production.” I enjoy working across the stack — from data and modeling to APIs and deployment — to build end-to-end AI systems that actually deliver value. Core Areas Retrieval-Augmented Generation (RAG) and semantic search systems LLM evaluation, prompt engineering, and response quality improvement RLHF-style workflows and human feedback integration Multilingual NLP and intent understanding Model fine-tuning (LLaMA, Mistral, BERT, XLM-R, BART) Scalable API development and ML system design Tech Stack Python, PyTorch, Hugging Face, FastAPI, MLflow, Docker, Kubernetes, Spark, AWS, FAISS, Pinecone Selected Impact Improved LLM response quality by 20–25% through evaluation and feedback loops Increased evaluation consistency by 30% across datasets Reduced API latency by 25–30% in production Improved multilingual intent detection by ~20–25% Reduced inference latency to <200ms for real-time systems Achieved 10–15% gains via task-specific fine-tuning
Key Skills
Experience
Ai / Ml Engineer
CurrentScale Ai
Worked on a GenAI-based document Q&A and content generation platform, leveraging external LLM (OpenAI, Anthropic) to generate summaries and answer queries from enterprise documents, supporting 100K+ document queries/month across internal use cases. Designed and contributed to RLHF-style workflows with human-in-the-loop feedback, including annotation processes, pairwise ranking, and preference-based learning, helping improve response quality scores by 20–25% based on human evaluation benchmarks. Built LLM evaluation frameworks using automated metrics (BLEU, ROUGE, F1) along with human preference scoring, improving evaluation coverage and consistency by 30% across test datasets. Applied prompt engineering techniques (few-shot, zero-shot, system prompts, templates) and conducted A/B testing, leading to 15–20% improvement in answer relevance and reduction in hallucinated responses. Implemented a RAG pipeline in Python for document-based queries, including document preprocessing, chunking strategies, tokenization and context window management, embedding generation (Hugging Face), and similarity-based retrieval (cosine similarity, top-k search). Incorporated retrieved document context into prompts to generate grounded and context-aware responses, improving factual accuracy and reducing unsupported answers in evaluation datasets. Contributed to supervised fine-tuning (SFT) workflows by preparing labeled datasets and evaluating performance improvements on open-source and internal task-specific models (LLaMA/Mistral), observing 10–15% task-specific accuracy gains. Performed data curation and dataset optimization, including document cleaning, deduplication, bias detection, and edge-case identification, along with error analysis and failure mode classification to improve model robustness. Leveraged MLflow for experiment tracking and model versioning, improving reproducibility and enabling comparison across multiple prompt and dataset variations.
Machine Learning Engineer
Amazon
* Worked on a multilingual NLP system supporting 10+ languages for Alexa and e-commerce use cases, mainly focusing on improving intent detection and reducing cases where the system failed to understand user queries, which over time led to around 20-25% improvement. * Fine-tuned transformer models such as BERT, mBERT, XLM-R, and BART using PyTorch, trying different hyperparameters and model variations and iterating based on both offline results and production feedback. * Trained models on large multilingual datasets using distributed PyTorch on GPU instances, which helped reduce training time by around 30-35% and made it easier to run multiple experiments. * Built a retrieval and generation pipeline (RAG) using embedding-based search and sequence-to-sequence models like BART and T5 to generate more grounded and context-aware responses. * Worked with text embeddings based on BERT and XLM-R for semantic search and query matching, improving how similar queries across different languages map to the same intent. * Contributed to summarization and response generation features, focusing on handling longer context and improving clarity and usefulness of model outputs. * Used Apache Spark and PySpark to process large multilingual datasets, including data cleaning, tokenization using WordPiece and BPE, and preparing training data. * Worked closely with data and annotation teams on labeling and cleaning pipelines, identifying noisy or inconsistent data that was impacting model performance, especially in low-resource and mixed-language scenarios. * Ran multiple model experiments and comparisons across BERT, XLM-R, and sequence-to-sequence models, performed error analysis on failed predictions, and tracked experiments using MLflow. * Built and maintained machine learning APIs using FastAPI and contributed to CI/CD and MLOps workflows to ensure smooth testing, deployment, and versioning of models.
Education
Bapuji Institute Of Engineering & Technology, Davanagere
Bachelor Of Engineering
The University Of Texas At Arlington
Masters
Interested in connecting with harshitha?
Sign up for NinjaHire to send a connection request.
Common Questions
What is harshitha's expertise?
harshitha specializes in Fine-Tuning Engineer, with expertise in agile methodology, amazon kinesis, annotation pipeline collaboration, application insights, automatic text summarization.
Where is harshitha located?
harshitha is based in san francisco, california, united states.
How much experience does harshitha have?
harshitha has 6+ years of professional experience.
How can I contact harshitha?
You can connect with harshitha through NinjaHire by signing up for a free account.
Other Fine-Tuning Engineers
Looking for a different Fine-Tuning Engineer?
Describe exactly who you need and NinjaHire will source them for you.
Type a role to try NinjaHire for free
