GR

Hire Goutham R. - Vector Database Engineer

Senior Ai / Ml Engineer

10+ years
st. petersburg, florida, united states
Raymond James

About goutham

Experienced AI/ML Engineer with 9+ years driving the design, deployment, and operationalization of scalable machine learning and Generative AI platforms across cloud-native environments (GCP & AWS). Specialized in LLMOps, retrieval-augmented generation (RAG), and multimodal AI using cutting-edge technologies like LangChain, Hugging Face Transformers, Vertex AI, and Stable Diffusion. Proven track record in fine-tuning and deploying advanced LLMs including LLaMA2, Mistral, Gemini, and PaLM API, building secure, high-performance model serving infrastructures with BentoML, KServe, and Ray on Kubernetes. Passionate about bridging innovation and production excellence to deliver intelligent document processing, semantic search, and end-to-end MLOps solutions. Led large-scale LLM deployment projects involving distributed model inference using Ray on Kubernetes clusters / Built centralized prompt orchestration systems using PromptLayer for reusable, optimized prompt engineering / Integrated monitoring and drift detection in production using Evidently AI, Prometheus, and Grafana / Designed secure document retrieval systems using RAG pipelines integrated with enterprise-grade IAM / Delivered serverless scoring and reporting pipelines using Cloud Scheduler, AWS Lambda, and Cloud Run / Developed explainable AI workflows using SHAP and LIME for compliance and stakeholder transparency / Built scalable multi-model serving platforms with support for parallel inference across GPUs / Automated data validation, schema enforcement, and transformation for structured and unstructured data / Architected feature stores using Feast and BigQuery to serve real-time and batch ML models / Migrated legacy ML infrastructure to cloud-native architectures using Terraform and Kubernetes / Standardized deployment workflows using TFX for lifecycle tracking, evaluation, and rollout governance / Maintained compliance-aligned ML documentation and operational playbooks across regulated domains / Developed RESTful and serverless APIs for LLM, NLP, and vision models with FastAPI and Flask

Key Skills

bentomlelasticsearchfaisskservelangchainmilvusnltkpineconepromptlayerpytorchqdrantscikit learntensorflowtensorflow servingtfx+5 more

Experience

Senior Ai / Ml Engineer

Current

Raymond James

* Engineered a scalable, cloud-native Generative AI platform integrating LLMs, vector databases, and multimodal pipelines for enterprise-level document understanding and secure content generation. * Architected modular RAG (Retrieval-Augmented Generation) pipelines using LangChain and Vertex AI for document summarization and knowledge retrieval. * Fine-tuned LLaMA2, Mistral, and Gemini models using Hugging Face Transformers and PaLM API for secure, context-aware content workflows. * Developed high-dimensional vector retrieval layers using FAISS, Pinecone, and Weaviate to query financial and regulatory data efficiently. * Deployed containerized LLMs and diffusion models across GKE using BentoML and KServe to support scalable, parallel inference. * Built distributed content generation workflows with Ray, enabling concurrent LLM inference for large-scale summarization and Q&A.

Senior Ml Engineer

State Of Vermont

* Designed and operationalized an AWS-based MLOps platform for real-time health surveillance, disease forecasting, and public health resource planning. * Architected scalable ML pipelines using SageMaker Pipelines, Argo Workflows, and Apache Airflow for reproducible training, validation, and deployment. * Built a centralized feature store with Feast, integrating data from AWS S3 and PostgreSQL to support both real-time and batch inference for epidemiological models. * Automated model versioning and lifecycle workflows using MLflow, enabling seamless experimentation-to-production transitions across Kubernetes environments. * Deployed containerized model inference services using Docker, Triton Inference Server, and TensorFlow Serving on Amazon EKS clusters. * Standardized infrastructure provisioning and Kubernetes management using Terraform, Helm, Kustomize, and Vault. * Developed low-latency APIs with FastAPI, secured using OAuth2 and AWS Secrets Manager for real-time prediction and monitoring services. * Implemented reproducibility and governance through DVC and GitOps pipelines for dataset versioning and artifact tracking.

Machine Learning Platform Engineer

Mantech

* Developed a secure, reproducible ML platform to support defense-sector machine learning workloads under strict compliance and audit requirements. * Designed robust ML pipelines using DVC and Git for dataset versioning, feature engineering, and checkpoint tracking. * Implemented MLflow to enable experiment tracking, model comparison, and visual lineage for audit readiness. * Deployed models via AWS SageMaker Inference Endpoints, supporting A/B testing and controlled production rollouts. * Automated CI/CD workflows using GitHub Actions, orchestrating build-and-train pipelines for continuous delivery. * Containerized ML workloads with Docker and deployed in Kubernetes clusters for scalable and isolated execution environments. * Built and secured inference APIs using FastAPI, exposing models for downstream secure consumption.

Interested in connecting with goutham?

Sign up for NinjaHire to send a connection request.

Common Questions

What is goutham's expertise?

goutham specializes in Vector Database Engineer, with expertise in bentoml, elasticsearch, faiss, kserve, langchain.

Where is goutham located?

goutham is based in st. petersburg, florida, united states.

How much experience does goutham have?

goutham has 10+ years of professional experience.

How can I contact goutham?

You can connect with goutham through NinjaHire by signing up for a free account.