RP

Hire Ruhin P. - Fine-Tuning Engineer

Research Support Specialist, Ml Engineer

5+ years
union, new jersey, united states
Institute For Advanced Computational Science

About ruhin

I ship ML systems end-to-end. Model training, evaluation harnesses, the C++ and Python plumbing underneath, the AWS services that serve it. The work I'm proudest of is usually the part nobody sees. At Yoloh, I built a production OCR + LLM pipeline (Llama-2 + GPT-4) that processes insurance documents at 95% F1 on a 10,000-document holdout. It automated 80% of field extraction and cut manual review from 250 to 100 hours per week. The pipeline that made that possible was the unglamorous part: rule-based validation, cross-field consistency checks, confidence-based human-in-the-loop QA. I also deployed the AWS Lambda microservice and DynamoDB + Neo4j backend that serve the resulting JSON, cutting p95 latency by 50%. At Stony Brook's Institute for Advanced Computational Science, I work on ML for scientific code generation. I fine-tuned a 7B CodeLlama model with QLoRA (rank 16, alpha 32) using a two-stage domain adaptation pipeline, improving generation accuracy on the MADNESS quantum chemistry codebase by 18%. To make that possible, I wrote the systems code around it: production C++ for HDF5 serialization with MPI-safe parallel I/O (merged to master), a Python tool server exposing 13 tools for HPC orchestration over SSH, a Tree-sitter ETL pipeline that cut dataset prep by 98%, and an evaluation harness scoring generated code on compilation, test pass rate, and API correctness. The pattern across both: the model is half the work. The other half is the data pipeline that feeds it, the evaluation that tells you it's actually working, and the infrastructure that lets it run reliably. I like all of it. B.Tech in ICT from PDEU (2022), M.S. Data Science from Stony Brook (2025), with internship stops at Blackcoffer and Yoloh in between. I came into ML through the engineering side — building things that work, then figuring out how to make them work better — and that's still how I operate. What I'm looking for: a role at a team shipping real LLM or applied ML systems, where the engineering is taken as seriously as the modeling. That could be ML Engineering, Software Engineering on an ML infra or platform team, or Data Engineering with ML depth. Open to early-career positions at companies where the work is technically serious. Core stack: Python · C++ · PyTorch · PEFT/QLoRA · AWS · Apache Spark · Docker · SQL Feel free to reach out or connect.

Key Skills

amazon dynamodbamazon web servicesanalytic problem solvinganalytical skillsaws lambdabertbig datac (programming language)c++communicationcritical thinkingdata analysisdata analyticsdata automationdata engineering+59 more

Experience

Research Support Specialist, Ml Engineer

Current

Institute For Advanced Computational Science

● Engineered HDF5 serialization for MADNESS scientific functions in 875 lines of C++ with MPI-safe parallel I/O and roundtrip validation to < 1e-8 error; merged to master with CMake integration and three automated ctest entries. ● Architected an 868-line Python server exposing 13 tools for HDF5 inspection and SLURM job management on SeaWulf HPC, using SSH stdio transport with per-session lifecycle and non-blocking submission for long-running calculations. ● Automated a Python + Tree-sitter pipeline to extract, validate, and label 23,089 C++ functions by section type, cutting dataset preparation time by 98% and enabling rapid iteration across training runs on SeaWulf HPC. ● Implemented a two-stage domain adaptation pipeline with an automated evaluation harness scoring generated C++ on compilation success, test pass rate, and API correctness, improving code-generation accuracy by 18%. ● Delivered tiered CLI commands and a fixture validation framework for an open-source quantum chemistry toolkit with a 46-test contract suite (46/46 passing across 18 test files) and regression comparison tooling; contributed via pull requests.

Data Science Intern

Yoloh

● Engineered a robust Optical Character Recognition solution for insurance management by integrating Llama 2, GPT-3, and GPT-4 models, achieving a 95% accuracy rate in data extraction and automating 80% of policy processing tasks. ● Designed and implemented a scalable database architecture using DynamoDB and Neo4j, streamlining user response storage and enabling personalized communication through AWS services, resulting in a 40% improvement in customer satisfaction. ● Implemented an AWS Lambda function to unify and deliver user responses in JSON format, cutting overall processing latency by 50% and boosting reporting speed for BI pipelines.

Data Scientist Intern

Blackcoffer

* ● Created a Python-based solution integrated with the Twitter API to provide actionable insights, boosting the client’s marketing and social media impact by 20%. ● Automated the conversion of raw data into structured Excel sheets, reducing manual input effort by 70% and accelerating transformation cycles by 35%. ● Delivered high-quality projects ahead of schedule by collaborating with a team of 5 developers, contributing to a 5% increase in company revenue.

Education

Stony Brook University

Master Of Science

Masters

Pandit Deendayal Energy University

Bachelor Of Technology

Interested in connecting with ruhin?

Sign up for NinjaHire to send a connection request.

Common Questions

What is ruhin's expertise?

ruhin specializes in Fine-Tuning Engineer, with expertise in amazon dynamodb, amazon web services, analytic problem solving, analytical skills, aws lambda.

Where is ruhin located?

ruhin is based in union, new jersey, united states.

How much experience does ruhin have?

ruhin has 5+ years of professional experience.

How can I contact ruhin?

You can connect with ruhin through NinjaHire by signing up for a free account.

Looking for a different Fine-Tuning Engineer?

Describe exactly who you need and NinjaHire will source them for you.

Type a role to try NinjaHire for free