About guru
AI/ML Ops Engineer specializing in designing and operating high-performance AI platforms across hybrid environments, combining bare metal GPU infrastructure with cloud-native ecosystems such as AWS HyperPod. Deep expertise in building scalable AI factories using OpenShift and Kubernetes, enabling efficient orchestration of distributed training, inference, and end-to-end ML pipelines.Hands-on experience with NVIDIA GPU technologies, including CUDA, Multi-Instance GPU (MIG), and NVIDIA GPU Operators for automated provisioning, monitoring, and lifecycle management of GPU resources. Proficient in architecting GPU-optimized clusters on bare metal, leveraging high-throughput, low-latency interconnects and performance-tuned storage systems (NVMe, parallel file systems) to support large-scale model training workloads.Strong background in cloud AI infrastructure, particularly AWS HyperPod, SageMaker, EKS, and high-performance networking (EFA), enabling distributed training frameworks such as Ray clusters and Kubeflow pipelines. Experienced in integrating Ray, Kubeflow, and containerized ML workflows for scalable experimentation, hyperparameter tuning, and production-grade deployments.Expertise in building resilient and fault-tolerant systems with a focus on hybrid architecture—seamlessly bridging on-prem GPU clusters with cloud environments. Skilled in designing network architectures for high-bandwidth, low-latency communication, critical for distributed AI workloads.Advanced knowledge of MLOps practices, including CI/CD for ML, model governance, lineage tracking, observability, and compliance. Familiar with tools such as MLflow, Prometheus, Grafana, ArgoCD, and Tekton for pipeline automation, monitoring, and lifecycle management.Passionate about optimizing performance at every layer—from GPU utilization and scheduling to storage throughput and network efficiency—while ensuring scalability, reliability, and governance across modern AI platforms.
Key Skills
Experience
Ai / Ml Ops Engineer
CurrentNatwest Group
AI/ML Ops Engineer specializing in designing and operating high-performance AI platforms across hybrid environments, combining bare metal GPU infrastructure with cloud-native ecosystems such as AWS HyperPod. Deep expertise in building scalable AI factories using OpenShift and Kubernetes, enabling efficient orchestration of distributed training, inference, and end-to-end ML pipelines. Hands-on experience with NVIDIA GPU technologies, including CUDA, Multi-Instance GPU (MIG), and NVIDIA GPU Operators for automated provisioning, monitoring, and lifecycle management of GPU resources. Proficient in architecting GPU-optimized clusters on bare metal, leveraging high-throughput, low-latency interconnects and performance-tuned storage systems (NVMe, parallel file systems) to support large-scale model training workloads. Strong background in cloud AI infrastructure, particularly AWS HyperPod, SageMaker, EKS, and high-performance networking (EFA), enabling distributed training frameworks such as Ray clusters and Kubeflow pipelines. Experienced in integrating Ray, Kubeflow, and containerized ML workflows for scalable experimentation, hyperparameter tuning, and production-grade deployments. Expertise in building resilient and fault-tolerant systems with a focus on hybrid architecture—seamlessly bridging on-prem GPU clusters with cloud environments. Skilled in designing network architectures for high-bandwidth, low-latency communication, critical for distributed AI workloads. Advanced knowledge of MLOps practices, including CI/CD for ML, model governance, lineage tracking, observability, and compliance. Familiar with tools such as MLflow, Prometheus, Grafana, ArgoCD, and Tekton for pipeline automation, monitoring, and lifecycle management. Passionate about optimizing performance at every layer—from GPU utilization and scheduling to storage throughput and network efficiency—while ensuring scalability, reliability, and governance across modern AI platforms.
Senior Technical Member
Adp
Db2 Dba Z / Os
Mphasis
* Performance Tuning using IBM PM database, CA insight tools, Data Studio, EZ-Db2 tools. * Worked on OPTHINT applying for performance improvement. * Involved in reviewing 100+ programs for access path changes for every Release. * Worked on EZ-DB2 tool for identifying performance issues and improving wherever necessary. * Worked on Optimize queries based on the query optimization guidelines and used Data studio tool for query optimization. Like * o Indexing all the predicates in JOIN, WHERE, ORDER BY and GROUP BY clauses. * o Avoid using functions in predicates. * o Avoid using wildcard (%) at the beginning of a predicate. * o Avoid unnecessary columns in SELECT clause. * o Use inner join, instead of outer join if possible. * o Using DISTINCT and UNION only if it is necessary. Etc. * Worked on generating performance report for HIGH CPU usage. High GETPAGES, Latches, Timeouts, Deadlocks and fine tuning the quires. * Involved in planning & setting up of jobs and scripts for migrating from SQL Replication to Q Replication for an entire RM Banking Platform. Successfully Implemented and tested in UAT environment * Setting up and implementing Q Replication for large number of tables for each release. * Running EZ-DB2 Traces for analyzing Performance of application programs * Identified lots of worst performing queries in Development itself and improved it. * Supported the bank during the critical situation under Db2. * Database design and its related objects. * Q-Replication setup for new tables and columns. * Setting up new off host connection le using JDBC and ODBC. * Setting up housekeeping jobs for new databases. * Part of scrum meetings. * Maintaining GIT Repository. * Creating and setting up of Utilities like Reorg, Copy, Runstats, Load, Unload, Modify, etc. * Generating bad SQL performance report using query monitoring and path checker. * Verifying the packages after release. * Monitoring and optimizing the performance of the database/health of applications using Omegamon.
Education
K N S Institute Of Technology, Bangalore
Bachelors
Phillipines Jobs And Careers
Bachelors
Interested in connecting with guru?
Sign up for NinjaHire to send a connection request.
Common Questions
What is guru's expertise?
guru specializes in MLOps Engineer, with expertise in amazon web services, application development, bitbucket, bmc remedy, bmc remedy ticketing system.
Where is guru located?
guru is based in bangalore, karnataka, india.
How much experience does guru have?
guru has 12+ years of professional experience.
How can I contact guru?
You can connect with guru through NinjaHire by signing up for a free account.
Other MLOps Engineers
Looking for a different MLOps Engineer?
Describe exactly who you need and NinjaHire will source them for you.
Type a role to try NinjaHire for free
