BM

Hire Bharath M. - Fine-Tuning Engineer

Ai / Ml Engineer

6+ years
san francisco, california, united states
Scale Ai

About bharath

I’m an AI/ML engineer with over five years of experience building machine learning systems that run in production. I work across the full stack of ML, from data pipelines and model training to backend services, deployment, and monitoring. My recent work has focused on large language models, retrieval systems, and scalable AI platforms. I’ve fine-tuned models, built RAG pipelines, optimized inference, and deployed services using Python, PyTorch, Docker, and Kubernetes in cloud environments. I’m comfortable working in fast-paced teams, taking ownership of projects, and making sure systems are reliable and maintainable. I focus on building solutions that are practical, scalable, and easy for teams to support in the long term.

Key Skills

adlsaifalgoc++distributed traininggenerative aihyperparameter optimizationknowledge engineeringlarge language modelslimellm fine tuningmeta internal ml platformsml apismlopsmodel evaluation and a/b testing+5 more

Experience

Ai / Ml Engineer

Current

Scale Ai

Architected and owned a human-in-the-loop Generative AI platform supporting LLM training, RLHF workflows, and large-scale evaluation, with a strong focus on data quality, alignment, and enterprise reliability. Designed and implemented Retrieval-Augmented Generation (RAG) pipelines, applying NLP techniques (semantic similarity, text normalization, keyword extraction, and document parsing) while optimizing chunking strategies, embedding selection, and vector retrieval tuning to improve factual grounding and reduce hallucinations in LLM outputs. Fine-tuned and evaluated open-source LLMs (Mistral, LLaMA 2) as base models using PyTorch and Hugging Face, leveraging preference modeling, RLHF signals, and prompt optimization to achieve 25–30% improvement on targeted evaluation benchmarks. Developed automated LLM evaluation frameworks measuring relevance, faithfulness, instruction adherence, and toxicity, combining offline metrics with human preference scoring and controlled A/B testing. Applied model interpretability techniques (SHAP, LIME) on supporting ML components (such as retrieval ranking, scoring, andvevaluation models) to analyze feature influence and assist debugging workflows. Implemented data quality, bias, and safety checks, including inter-annotator agreement analysis, gold-set validation, fairness assessments, toxicity detection, and prompt-injection resistance testing. Deployed and scaled training, inference, and evaluation workloads on AWS EKS using Docker and Kubernetes, optimizing performance and reducing LLM inference costs through batching, caching, and model selection strategies. Established MLOps and CI/CD pipelines using MLflow, GitHub Actions, and Terraform for experiment tracking, dataset and prompt versioning, automated deployments, monitoring, and cross-functional collaboration.

Machine Learning Engineer

Meta

Improved large-scale conversational AI systems by optimizing transformer-based inference pipelines using PyTorch and TorchScript, reducing average response latency by 25–30% while improving stability and overall production efficiency. Built and maintained retrieval-grounded response pipelines using FAISS and BM25, grounding conversations in large internal knowledge bases and improving response relevance, coverage, and factual accuracy by 20% in offline evaluations. Developed and fine-tuned BERT, RoBERTa, and DistilBERT models for intent classification, semantic similarity, ranking, and multilingual language understanding across multiple high-traffic conversational flows. Implemented controlled text generation and summarization workflows using T5 (base), ensuring safety-aware, policy-compliant responses suitable for production conversational systems. Applied model optimization techniques including mixed-precision inference, post-training quantization, and TorchScript graph optimizations to improve inference throughput and reduce infrastructure cost without degrading model quality. infrastructure, reducing training time by 30% and accelerating experimentation cycles. Built scalable data ingestion and preprocessing pipelines using Apache Spark, Hive, PyArrow, and Airflow, processing terabytes of conversational and behavioral data to generate clean, privacy-compliant training datasets. pipelines to automate testing, model validation, container builds, and safe promotion across environments. quality regressions, data drift, and reliability issues in production ML services. Supported responsible AI and analytics efforts through A/B testing, explainability using SHAP and LIME, bias evaluation, largescale analysis with Presto and Python, and development of internal ML APIs using Flask and FastAPI.

Education

Saint Louis University

Masters

Interested in connecting with bharath?

Sign up for NinjaHire to send a connection request.

Common Questions

What is bharath's expertise?

bharath specializes in Fine-Tuning Engineer, with expertise in adls, aif, algo, c++, distributed training.

Where is bharath located?

bharath is based in san francisco, california, united states.

How much experience does bharath have?

bharath has 6+ years of professional experience.

How can I contact bharath?

You can connect with bharath through NinjaHire by signing up for a free account.

Looking for a different Fine-Tuning Engineer?

Describe exactly who you need and NinjaHire will source them for you.

Type a role to try NinjaHire for free