Data Engineer

Shreyas Shende

Data Engineer at Morgan Stanley working across the SDE, data, and applied-AI stack — production pipelines, agentic AI workflows, and internal tooling. Growing deeper into data science, with machine learning engineering as the long-term goal. MS in Computer Science from NJIT.

Data Engineer @ Morgan Stanley
MS CS, NJIT
New York Metropolitan Area
Shreyas Shende

Beyond the Code

When I'm not building pipelines, this is where I'm usually at.

  • Gaming
  • Anime
  • Football
  • Hiking
  • Swimming

Skills

A snapshot of the toolkit — the resume has the full list.

Data Engineering & Cloud

AirflowAWSDatabricks

Programming & SQL

PythonSQLPySpark

Analytics & BI

Power BITableauA/B Testing

Machine Learning (Applied)

TensorFlowPyTorchScikit-learn

Collaboration & DevOps

Git / GitHubCI/CDJIRA

Experience & Education

Hover a card (or tab to it) to read the details.

ExperienceEducationCurrent → 2019
Arc 03Current

Data Engineer

Morgan Stanley · New York, NY · Dec 2025 – Present

  • Designed and automated production ETL pipelines integrating data from 7+ enterprise APIs (Jira, Rally, OpenPages, etc.), processing ~35K records daily through scheduled AutoSys workflows, reducing pipeline failures by 30% and improving data reliability, resiliency, and quality through validation, structured logging, and fault-tolerant error handling.
  • Built Agentic AI workflows leveraging enterprise GPT services to automate classification of 100+ monthly operational risk records — identifying cloud providers, cloud services, business impacts, and AI-related issues — with human-in-the-loop validation for audit readiness.
  • Developed internal AI agents for Jira issue classification and automated pull request reviews, flagging security risks, coding standard violations, exposed secrets, and performance concerns before deployment.
  • Automated CI/CD by configuring and maintaining 15 Jenkins pipelines and deployment workflows, while building operational Power BI dashboards and database fallback strategies over 2M+ records to improve monitoring and platform reliability.
Arc 02

Graduate Research Assistant (Master's Project)

New Jersey Institute of Technology · Newark, NJ · Jan 2025 – May 2025

  • Built an end-to-end RNA-seq analysis pipeline in Python and PyTorch Geometric, automating data preprocessing, graph construction, model training, and feature selection across Cervical, Kidney, Alzheimer, and Lung cancer datasets; reduced feature space by >90% and improved classification accuracy by 5–10% over DESeq2 and EdgeR.
  • Led a 3-member research team in collaboration with Brown University's Alpert Medical School to source and validate clinical RNA-seq datasets, resulting in a co-authored research paper demonstrating the pipeline's adaptability across diverse genomic datasets.
Arc 01

Platform Engineering & Automation Intern

Vendorpass (FIS Global) · Remote, US · May 2024 – Nov 2024

  • Automated Python and Bash ETL workflows, reducing manual mapping effort by 90% and improving pipeline reliability.
  • Built a Streamlit dashboard integrated with Oracle REST APIs to automate reporting across 20K+ records while implementing CI/CD validation and post-clone automation that reduced environment downtime from 4+ hours to 1 hour.
Education

M.S. Computer Science

New Jersey Institute of Technology · Newark, NJ · Aug 2023 – May 2025

GPA 3.95/4.0

Education

B.E. Computer Engineering

University of Pune · Pune, India · Aug 2019 – Jun 2023

GPA 3.66/4.0

Projects

A few things I've shipped recently.

Ad Analytics Pipeline

Jul 2025 – Aug 2025

Containerized Python + Airflow pipeline to ingest, validate, and orchestrate raw ad campaign data into BigQuery. Databricks (PySpark) transformations clean and aggregate campaign KPIs (CTR, CPC, ROI), delivered through an interactive Power BI dashboard.

Pipeline

  1. Ad Campaign Data
  2. Airflow (Docker)
  3. BigQuery
  4. Databricks / PySpark
  5. Power BI

Spotify Listener Analytics Dashboard

May 2025 – Jun 2025

Star-schema data model transforming nested JSON into Parquet for scalable analytics with Spark SQL. Interactive Power BI dashboard with DAX-based KPIs for engagement, churn, and revenue trends — cut reporting turnaround by 70%.

Pipeline

  1. Nested JSON
  2. Python ETL
  3. Parquet (Star Schema)
  4. Spark SQL
  5. Power BI (DAX)

Reddit Data Ingestion & Analytics Pipeline

Mar 2025 – Apr 2025

Scalable ETL pipeline using Airflow, Docker, and Celery to ingest Reddit data via APIs into Amazon S3, orchestrating transformations with AWS Glue and Athena. Automated ingestion of 100+ daily posts into Redshift for SQL-based trend and sentiment analytics.

Pipeline

  1. Reddit API
  2. Airflow + Celery
  3. Amazon S3
  4. AWS Glue / Athena
  5. Amazon Redshift

NFL Management System

Full-stack app (React frontend, Express backend) with a Python/BeautifulSoup4 scraper gathering three seasons (2021–2023) of NFL data — players, coaches, teams, finances, and stats — stored in MySQL.

Pipeline

  1. NFL Web Pages
  2. BeautifulSoup4 Scraper
  3. MySQL
  4. Express API
  5. React Frontend
View more on GitHub

Certificates & Publications

AWS Certified Cloud Practitioner (CLF-C02)

AWS Certified Cloud Practitioner (CLF-C02)

Amazon Web Services

Microsoft Certified: Azure Data Fundamentals (DP-900)

Microsoft Certified: Azure Data Fundamentals (DP-900)

Microsoft

Oracle Cloud Infrastructure 2025 Generative AI Certified Professional

Oracle Cloud Infrastructure 2025 Generative AI Certified Professional

Oracle

SnowPro Associate: Platform Certification

SnowPro Associate: Platform Certification

Snowflake

13

Citations

12 since 2021

2

h-index

2 since 2021

A Comprehensive Study on Simultaneous Localization and Mapping (SLAM): Types, Challenges, and Applications

A. Khole, A. Thakar, S. Shende, V. Karajkhede

2023 International Conference on Sustainable Computing and Smart Systems (ICSCSS), IEEE · 2023

Cited by 6Read paper

A Compendium on Distributed Systems

A. Khole, A. Thakar, A. Kulkarni, H. Jadhav, S. Shende, V. Karajkhede

arXiv preprint arXiv:2302.03990 · 2023

Cited by 6Read paper

RGE-GCN: Recursive Gene Elimination with Graph Convolutional Networks for RNA-seq based Early Cancer Detection

S. Shende, V. Narayanan, V. Fenn, Y. Huang, D. Goksuluk, G. Choudhary, et al.

arXiv preprint arXiv:2512.04333 · 2025

Cited by 1Read paper

Real-Time Monocular SLAM: Accurate Localization and Mapping Using Point Map and Search-by-Projection Approach

A. Khole, A. Thakar, S. Shende, V. Karajkhede

Authorea Preprints · 2023

Let's build something

Open to Data Engineer roles and interesting problems. Reach out any time.

© 2026 Shreyas Shende. Built with Next.js.