About abdullah
Data engineer with 5+ years of experience designing, building, and automating data pipelines, cloud data lakes, and real-time streaming systems across commercial banking, retail banking, wealth management, and automotive analytics. Deep expertise in AWS (S3, EMR, Lambda, Step Functions, SageMaker), Databricks, PySpark, Spark, and Kafka. Track record of cutting pipeline development time through standardized data-quality checks, replacing a batch-oriented fraud-detection setup with a real-time Kafka pipeline, and standardizing risk scoring across automotive programs that had previously been tracked ad hoc. Also builds the trust and security layer that production ML systems need, from explainability pipelines that let risk analysts see why a model made a call to fine-grained access controls that keep sensitive account data locked down by default. Consistent record of working directly with clients and risk analysts to translate reporting and compliance requirements into production systems.
Key Skills
Experience
Data Engineer
CurrentNavy Federal Credit Union
* Enterprise AWS Data Lake — Core Banking & Wealth Management * Designed an enterprise AWS data lake for analytics, storage, and regulatory reporting on high-volume account and transaction data across retail banking and wealth management. * Built a security framework using Lambda and DynamoDB for fine-grained, least-privilege access to sensitive account data in S3. * Designed a PySpark solution standardizing data quality checks across 30+ datasets, cutting pipeline development time by ~40%. * Built Spark/Scala jobs on Hadoop and PySpark/SparkSQL jobs on EMR for portfolio performance, interest-accrual, and regulatory reporting. * Implemented infrastructure as code with Terraform and GitLab to provision the data lake environment. * Real-Time Kafka Streaming — Fraud Detection & Payments * Built real-time pipelines with Kafka and Python to stream card and wire transactions into the fraud-detection system. * Led an architecture assessment of EMR/Redshift/S3, finding the batch setup couldn’t meet real-time fraud-detection latency requirements. * Integrated Step Functions with Lambda, S3, DynamoDB, and SQS for end-to-end transaction enrichment and alerting. * Set up Jenkins CI/CD via Bitbucket webhooks for the fraud-rules codebase. * ML-Driven Credit Risk & AML Scoring * Automated a SageMaker workflow with Step Functions, from feature publishing in S3 through training to real-time deployment. * Processed historical transaction and bureau data on Hadoop/Spark/Hive, orchestrated with Airflow, to build model training features. * Deployed Dockerized model-serving and batch-scoring services on Kubernetes. * Environment: AWS (EMR, S3, RDS, Redshift, Lambda, DynamoDB, Step Functions, SQS, Aurora, SageMaker), Hadoop, Spark, Kafka, Airflow, Kubernetes, Terraform, Jenkins, Snowflake, Python
Data Engineer
Flash Infolabs
* Project: Canara Bank Reporting & Analytics Automation * Canara Bank's account and transaction reporting ran off spreadsheets someone had to reconcile by hand every week. This project got the client off that grind with a Python and SQL pipeline into AWS S3, an automated weekly refresh, and a Power BI dashboard the operations team could trust, plus the documentation to run it themselves after the engagement wrapped. * Worked directly with the client to translate their reporting requirements into a data model, since they didn't have a dedicated data team of their own. * Built a Python and SQL ETL pipeline to pull account and transaction data from multiple spreadsheet-based sources into AWS S3 and a central reporting database. * Added validation checks in Python to catch inconsistencies in the client's incoming data files before they reached the reporting database. * Automated the weekly data refresh with a scheduled Python script, so the client's team no longer had to run the process by hand. * Designed a Power BI dashboard giving the client's operations team a consolidated view of account activity, replacing manual Excel-based reporting. * Documented the pipeline and dashboard so the client's internal team could maintain it after the contract ended. * Environment: Python, SQL, AWS S3, Power BI
Data Engineer
Ltimindtree
HSBC Credit-Risk & Loan Delinquency Forecasting HSBC needed more than a monthly delinquency score — they needed to know why a loan got flagged. Built a Databricks pipeline turning raw loan and account data into monthly risk scores, with a SHAP layer so analysts could see which factors drove each prediction. Built a Databricks ETL framework using PySpark, Spark SQL, Delta Lake, and medallion architecture (Bronze-Silver-Gold) to standardize raw loan and account data from core-banking systems into ML-ready datasets for credit-risk forecasting. Developed parameterized, reusable Databricks Jobs running data cleansing, aggregation, and quality checks across multiple bank portfolios to generate monthly account-level risk features. Designed a Databricks-based SHAP explainability pipeline computing feature-level contributions for monthly credit-risk predictions, so analysts could explain the reasoning behind each score. Connected MLflow model artifacts and Gold Delta prediction tables via PySpark, SQL, and SQLAlchemy, so analysts could pull SHAP values and feature-importance metrics directly into credit decisions. Built ML pipelines that created training datasets, trained and evaluated models, and generated batch risk scores predicting monthly loan delinquency per account. Applied Gaussian Process Regression and probabilistic modeling to estimate prediction uncertainty and confidence intervals for portfolio-level loss forecasts. Secured access to sensitive account and loan data using Databricks Secret Scopes and Azure Key Vault, keeping credentials out of the pipeline code.
Education
Governors State University
Master Of Science
Dayananda Sagar College Of Engineering, Bangalore
Bachelor Of Engineering
Interested in connecting with abdullah?
Sign up for NinjaHire to send a connection request.
Common Questions
What is abdullah's expertise?
abdullah specializes in Vector Database Engineer, with expertise in alembic, c (programming language), cascading style sheets, chatopenai, core java.
Where is abdullah located?
abdullah is based in new york, new york, united states.
How much experience does abdullah have?
abdullah has 6+ years of professional experience.
How can I contact abdullah?
You can connect with abdullah through NinjaHire by signing up for a free account.
Other Vector Database Engineers
Looking for a different Vector Database Engineer?
Describe exactly who you need and NinjaHire will source them for you.
Type a role to try NinjaHire for free
