Designed and developed 15+ production ETL pipelines and 20+ Airflow DAGs using Python, Airflow, Dataproc, and BigQuery, automating ingestion, transformation, orchestration, and publishing of enterprise healthcare datasets.
Implemented SQL transformation pipelines involving schema mapping, validation, enrichment, deduplication, BigQuery views, and business-rule processing across 50+ enterprise tables.
Developed reusable unit and integration testing frameworks for ETL pipelines, validating API responses, schema changes, data quality rules, and business logic before production deployment.
Implemented automated pipeline monitoring, retry logic, and email alerting for failed API calls and workflow executions, improving reliability across enterprise data pipelines.
Created multi-stage GitHub Actions CI/CD pipelines integrating SonarQube, Snyk, and automated artifact deployments to GCS buckets, reducing deployment time from 2 hours to under 15 minutes.
Built Tableau dashboards enabling business stakeholders to monitor pipeline health, operational KPIs, and analytical datasets supporting enterprise reporting.
Led AI-driven modernization initiatives using Vertex AI and Gemini, reverse engineering legacy Java/.NET applications to accelerate migration planning and reduce developer onboarding time by 40%.
Maintained and enhanced enterprise decision-support applications built using Java, C#, .NET, and IBM MQ, supporting internal reporting workflows used by multiple business units.
Developed reporting functionality using MDX queries, SQL Server stored procedures, and DB2 stored procedures, improving report execution performance and reliability.
Developed and maintained 18 production enterprise web applications using Java, Spring Boot, and Angular in Agile teams supporting telecom platforms.
Designed REST APIs and backend services following test-driven development, improving feature delivery while reducing production defects by 20%.
Migrated enterprise applications from on-premise infrastructure to AWS using Docker, AWS Managed Workflows, RDS, and EC2, reducing infrastructure costs by 18%.
Implemented CI/CD automation supporting cloud-native deployments and reducing release effort by 30%.
Developed ETL pipelines using Spark Streaming and Airflow, driving a 26% reduction in data processing time and enhancing overall data pipeline performance.
Implemented automation solutions for extracting and loading data from diverse sources such as PDF, Excel, JSON, PostgreSQL, and MongoDB, streamlining data ingestion processes and ensuring data integrity.
Played a key role in optimizing SQL queries for downstream operations, resulting in a notable 30% improvement in query execution times, enhancing data retrieval efficiency and overall system performance.
Created a consolidated auto insurance management dashboard using Flask and MySQL to enhance operational efficiency and optimize workflow processes.
Collaborated with cross-functional teams to integrate the dashboard with APIs, reducing latency by 40%.
Trained supervised Machine Learning models like RF, SVM, and GBM to detect fraudulent claims and integrated them with the dashboard, reducing detection time by 27% and improving accuracy by 4%.
Developed a Deep Learning model, using Python and TensorFlow, trained on the Twitter dataset to flag inappropriate content.
Utilized Hadoop’s MapReduce architecture with General Purpose GPU and PyCUDA for extensive parallelization, leading to ∼12% reduction in processing time.
Download my CV for my detailed work experience as well as links to my publications and projects!