BB

Hire Brunoo B. - Fine-Tuning Engineer

Ai / Ml Engineer

5+ years
long beach, california, united states
Piper Sandler

About brunoo

I am an AI/ML Engineer with 3+ years of experience building and deploying scalable machine learning and Generative AI systems in production environments. I specialize in: LLMs, RAG pipelines, LangChain, LangGraph, multi-agent systems LLM fine-tuning (LoRA, QLoRA, RLHF) and inference optimization (vLLM, quantization, KV caching) Machine learning & deep learning (PyTorch, TensorFlow, XGBoost, time-series models) MLOps & ML infrastructure (Kubeflow, KServe, MLflow, Kubernetes, ArgoCD) Data engineering & distributed systems (PySpark, Kafka, ETL pipelines, Spark) Cloud platforms (AWS, SageMaker, EKS, Lambda) What I’m good at: Building production-grade ML pipelines and scalable AI systems Designing and optimizing RAG architectures to improve retrieval accuracy and reduce hallucinations Deploying and managing ML models with high availability and observability Optimizing LLM inference for latency, throughput, and cost efficiency Developing end-to-end ML workflows from data processing to deployment What I’m passionate about: Applying Generative AI and machine learning to solve complex real-world problems, improve decision-making, and build reliable, scalable AI systems that deliver measurable business impact. Currently seeking: AI/ML Engineer / Generative AI Engineer / LLM Engineer / MLOps Engineer roles focused on production ML systems, LLM applications, and scalable AI infrastructure.

Key Skills

agile and sdlc methodologiesapi developmentaws lambdaazure sqlbig databusiness planningbusiness processcloudcloud infrastructurecloud securitycloud servicescompliance enablementdatabasesdockergenerative ai+18 more

Experience

Ai / Ml Engineer

Current

Piper Sandler

Designed and developed a multi-agent equity research platform using LangGraph, integrating retrieval, macroeconomic analysis, sentiment analysis, and fact-checking agents to automate complex financial research workflows, reducing analyst effort by 40%. Enhanced RAG pipeline performance by resolving chunking inconsistencies and eliminating noisy data, achieving 30% improvement in retrieval precision and reducing hallucination rates, validated using RAGAS and LLM evaluation frameworks. Architected and implemented MCP servers to standardize communication between LLM agents and external tools, while enabling execution trace observability (branching, retries), improving debugging efficiency and root-cause analysis by 5x. Fine-tuned open-source LLMs using LoRA and RLHF techniques, and deployed scalable inference endpoints with vLLM on GPU infrastructure, delivering cost-efficient, enterprise-grade AI solutions. Optimized LLM inference performance using advanced techniques such as KV caching, speculative decoding, quantization, and GPU memory optimization, resulting in 35% reduction in latency, 40% lower memory utilization, and increased throughput. Built and deployed production-grade ML pipelines on AWS EKS using Kubeflow, KServe, and ArgoCD, ensuring highly available model serving with 99.99% uptime and robust CI/CD integration. Streamlined MLOps workflows by implementing Terraform-based infrastructure and GitOps practices, improving deployment velocity across 10+ ML models while maintaining full experiment tracking and reproducibility using MLflow.

Machine Learning Engineer

Persistent Systems

* Developed an AIOps-based predictive analytics platform to detect system failures and anomalies using Random Forest, XGBoost, and LSTM models, improving incident prediction accuracy by 32% and reducing unplanned outages by 28%. * Built scalable data ingestion and processing pipelines using Python, PySpark, and Apache Kafka to handle high-volume log and monitoring data from servers, applications, and cloud environments. * Performed advanced feature engineering on time-series and log data using Pandas and NumPy to identify patterns related to system performance degradation and incident occurrence. * Designed and deployed anomaly detection models to reduce unplanned downtime by predicting incidents before occurrence, improving system reliability and SLA compliance. * Integrated ML models into production using REST APIs (Flask/FastAPI) and deployed them on AWS (EC2, S3, SageMaker) for real-time monitoring and alerting. * Implemented model monitoring, automated retraining pipelines, and CI/CD workflows using Docker, Git, and Jenkins, improving deployment efficiency by 45% and ensuring consistent model performance. * Collaborated with DevOps, Cloud, and Support teams in an Agile environment using Jira, accelerating issue resolution cycles by 30% and delivering data-driven solutions aligned with business KPIs.

Education

California State University, Long Beach

Master Of Science

Srm Ist Chennai

Bachelor Of Technology

Interested in connecting with brunoo?

Sign up for NinjaHire to send a connection request.

Common Questions

What is brunoo's expertise?

brunoo specializes in Fine-Tuning Engineer, with expertise in agile and sdlc methodologies, api development, aws lambda, azure sql, big data.

Where is brunoo located?

brunoo is based in long beach, california, united states.

How much experience does brunoo have?

brunoo has 5+ years of professional experience.

How can I contact brunoo?

You can connect with brunoo through NinjaHire by signing up for a free account.

Looking for a different Fine-Tuning Engineer?

Describe exactly who you need and NinjaHire will source them for you.

Type a role to try NinjaHire for free