PN

Hire Praveen N. - Generative AI Engineer

Generative Ai Engineer

8+ years
arlington, texas, united states
Realpage, Inc.

About praveen

Data Engineer building scalable data platforms and applying GenAI to automate data engineering workflows. I’m a Data Engineer and GenAI Engineer focused on building scalable data platforms and applying LLMs to automate real-world data workflows. Over the past 6+ years, I’ve worked across GCP, AWS, and Azure designing end-to-end data systems—from ingestion and transformation to analytics and ML-ready datasets. My work focuses on Python, PySpark, SQL, and cloud-native lakehouse architectures, building pipelines that process large datasets reliably and support downstream analytics and machine learning use cases. Recently, I’ve been working at the intersection of Data Engineering and Generative AI, building systems that use RAG pipelines, agentic workflows, and MCP-based tool orchestration to automate parts of the data engineering lifecycle. This includes using LLMs to assist with schema mapping, data validation, and transformation logic, while integrating retrieval systems, structured tool-calling, and human-in-the-loop review patterns to ensure outputs remain reliable and production-ready. On the data engineering side, I design scalable ETL/ELT pipelines, implement lakehouse architectures, and optimize distributed workloads using Spark, Databricks, Snowflake, and BigQuery. I focus on data quality, schema contracts, observability, and performance tuning so that pipelines remain reliable and maintainable as data volumes grow. On the GenAI side, I work with RAG architectures, embeddings, vector retrieval, MCP-based agent tooling, and prompt engineering to integrate LLM capabilities into production data platforms. My focus is on making these systems practical, measurable, and safe for real workflows, rather than just experimental prototypes. Technologies I work with most often: Python, SQL, PySpark, Apache Spark Airflow / Composer orchestration Databricks, Snowflake, BigQuery Lakehouse architectures and distributed data processing RAG pipelines, vector retrieval, MCP-based agent orchestration Cloud platforms: GCP, AWS, Azure I enjoy solving messy data problems and building systems that turn raw data into reliable data products—and increasingly using AI to accelerate that process. I’m currently open to Data Engineering, GenAI, and LLM platform roles across the United States (H-1B transfer required) and always happy to connect with teams building data platforms, AI infrastructure, or intelligent data systems. #DataEngineering #DataEngineer #PySpark #ETL #GenAI #LLM #RAG #AIEngineering #CloudData #OpenToWork

Key Skills

agile project managementbig databig querycloud composercollaborationdata cleaningdata governancedata modelingdata profilingdata securitydata validationsdatabricksdataflowetlgcp+5 more

Experience

Generative Ai Engineer

Current

Realpage, Inc.

Data Engineer

Globallogic

* Building a GCP-based LLM evaluation data platform that processes high-volume feedback data into structured training and analytics datasets using distributed pipelines and GenAI-driven automation. * Responsibilities * Designed and implemented cloud-native data pipelines on GCP using Dataflow, Dataproc (PySpark), GCS, and BigQuery to ingest and process high-volume LLM feedback datasets from APIs and batch file drops. * Built scalable medallion architecture pipelines to normalize schema-variant datasets and convert them into curated analytics and model training datasets. * Developed data quality frameworks including schema validation, reconciliation checks, deduplication, and quarantine pipelines to prevent invalid data from reaching curated layers. * Implemented Airflow orchestration for end-to-end pipeline scheduling, retries, dependency management, and safe backfills across multiple ingestion workflows. * Optimized BigQuery performance using partitioning, clustering, and query tuning techniques to improve analytical workloads and reduce processing time. * Designed metadata-driven RAG workflows using Vertex AI (Gemini), embeddings, and vector retrieval to assist with schema mapping, transformation logic generation, and data validation tasks. * Built agentic pipelines using MCP-style tool orchestration to coordinate tasks such as schema diffing, transformation generation, validation checks, and execution of PySpark workflows. * Implemented human-in-the-loop (HITL) review patterns and confidence scoring to ensure LLM-generated transformations meet reliability and data quality standards. * Collaborated with platform engineers and ML teams to deliver training-ready datasets and analytics tables supporting model evaluation, experimentation, and reporting. * Technologies: * Python, SQL, PySpark, Apache Spark, Dataflow, Dataproc, BigQuery, GCS, Airflow, Vertex AI, RAG pipelines, embeddings, FAISS/vector retrieval, MCP orchestration, cloud-native data engineering

Data Engineer

Athena Tech

* Worked on a cloud-based insurance claims data platform built on AWS and Snowflake to support reporting, analytics, and operational insights for claims processing and risk analysis. * Responsibilities * Designed and implemented AWS-based ingestion pipelines using S3 landing zones, Glue crawlers, and Glue ETL jobs to process large volumes of claims data from multiple upstream systems. * Developed scalable PySpark transformation pipelines on EMR to enrich and standardize high-volume claims datasets for downstream analytics. * Built Snowflake data warehouse layers using staging tables, automated COPY INTO ingestion, and incremental MERGE-based upsert pipelines. * Implemented data validation and reconciliation logic using Python to detect schema inconsistencies, duplicates, and data quality issues. * Optimized Snowflake performance through warehouse tuning, clustering strategies, and SQL query optimization to improve reporting performance. * Developed automated monitoring and alerting workflows using CloudWatch and pipeline logging to improve reliability and reduce operational incidents. * Automated infrastructure deployment and environment setup using Terraform and CI/CD pipelines with Jenkins, enabling consistent deployments across environments. * Delivered Tableau dashboards that provided business users with interactive insights into claims trends, processing efficiency, and operational KPIs. * Technologies: * Python, SQL, PySpark, Apache Spark, AWS S3, AWS Glue, EMR, Snowflake, Terraform, Jenkins, Tableau, CloudWatch

Education

The University Of Texas At Arlington

Masters

Chaitanya Bharathi Institute Of Technology

Bachelor Of Engineering

Interested in connecting with praveen?

Sign up for NinjaHire to send a connection request.

Common Questions

What is praveen's expertise?

praveen specializes in Generative AI Engineer, with expertise in agile project management, big data, big query, cloud composer, collaboration.

Where is praveen located?

praveen is based in arlington, texas, united states.

How much experience does praveen have?

praveen has 8+ years of professional experience.

How can I contact praveen?

You can connect with praveen through NinjaHire by signing up for a free account.

Looking for a different Generative AI Engineer?

Describe exactly who you need and NinjaHire will source them for you.

Type a role to try NinjaHire for free