About harish
Data Engineer and Analyst.
Key Skills
Experience
Data Engineer
CurrentAon
Built and operated scalable microservices and FastAPI LLM/ML services (summarization, extraction, classification, recommendations), including RAG + vector search pipelines, real-time Kafka/Docker/Kubernetes workflows, CI/CD (Jenkins/GitHub Actions), and Airflow-driven ETL/training/batch inference on Azure ML with Snowflake/SQL Server—plus monitoring in Grafana and cross-functional delivery of production AI solutions.
Data Engineer
Xerox
Developed the ingestion layer for Xerox's managed print fleet, landing meter reads, toner levels, and error codes from DCA agents across 50K+ devices into S3 cataloged in AWS Glue, with contract and CRM data brought in via Fivetran. Cleaned and conformed the manufacturer-specific telemetry in PySpark on Databricks, deduplicating repeated scans and flagging backward meter reads, and landed trusted device-reading tables as Delta in a medallion layout. Modeled billing marts in dbt that join meter reads against contracted per-page rates to compute billable volume per device, serving as the source of truth finance invoiced clients from. Added dbt and Great Expectations tests enforcing billing-critical rules, which cleared a recurring class of invoice disputes, and scheduled the daily run as Airflow DAGs with per-site retries so a late-reporting client no longer corrupted the billing run.
Data Engineer
Hsbc
Set up batch ingestion that pulled transaction and account data from PostgreSQL and semi-structured customer activity from MongoDB into Hive on Hadoop, then wrote PySpark jobs that turned it into the regulatory and AML reporting datasets due each morning. Contributed to Kafka ingestion feeding near-real-time transaction monitoring, and built curated star-schema tables in Redshift for downstream analyst reporting. Helped migrate Hive/HDFS workloads to AWS EMR and S3 (Parquet) during the team's move off the on-prem cluster, rewriting HiveQL into Spark SQL and tuning jobs that cut average batch runtime ~40%. Took ownership of source-to-warehouse reconciliation, automating checks in Python and SQL that caught upstream data drops before they reached regulatory reports.
Education
Kumaraguru College Of Technology
Bachelors
Rochester Institute Of Technology
Masters
Interested in connecting with harish?
Sign up for NinjaHire to send a connection request.
Common Questions
What is harish's expertise?
harish specializes in Vector Database Engineer, with expertise in celery, conditional image generation, confluence, crewai, docker.
Where is harish located?
harish is based in united states.
How much experience does harish have?
harish has 6+ years of professional experience.
How can I contact harish?
You can connect with harish through NinjaHire by signing up for a free account.
Other Vector Database Engineers
Looking for a different Vector Database Engineer?
Describe exactly who you need and NinjaHire will source them for you.
Type a role to try NinjaHire for free
