Open to work & collaboration · New York, NY

I build models that turn messy data into decisions people act on.

A data scientist blending generative AI, machine learning, and experimentation to help teams make confident, evidence-based decisions.

abhitayshinde@gmail.com (646) 225-0802 LinkedIn GitHub Google Scholar
Scroll

01 · ExperienceWhere I've built

Feb 2026 – Present

Data Scientist Customer Acquisition

Ask2AI · New York, NY

Feature Engineering Causal Inference Polars
  • Corrected selection bias in the scoring model, by building a feature engineering and causal inference pipeline in Python and Polars that integrated 11 source systems into a unified 30 million customer dataset.
  • Scored the rejected-applicant population at 0.86 held-out AUC, by training gradient-boosted models within a Double Machine Learning framework to correct the selection bias in approved-only training data.
  • Expanded the approvable population by 15% at flat expected default, by handing corrected scores to credit policy.
Jun 2025 – Sep 2025

Data Scientist Generative AI & Personalization

Julius Baer · New York, NY

Agentic LangGraph LLM-as-a-Judge RAG
  • Replaced the analyst-and-copywriter handoff for client-facing investment documents, by engineering an agentic LangGraph system retrieving client holdings, stated preferences, and live market news into a single generated draft.
  • Automated compliance review at 86% agreement with human reviewers, by building an LLM-as-a-judge critic that scored each draft against the firm's rubric and regenerated failures until they passed.
  • Cut content selection time for relationship managers by 35%, by shipping the pipeline as a chat interface with human-in-the-loop revision, letting RMs iterate on a living client document.
  • Surfaced a 25% engagement lift, using exploratory analysis on customer data in Python and SQL against matched controls.
Jan 2025 – May 2025

Data Scientist Co-op Customer Acquisition

Ask2AI · New York, NY

Uplift Modeling DiD Mixed-Integer Optimization
  • Improved cross-channel ROI attribution by 22% and identified optimal credit limit thresholds, by building uplift models and difference-in-differences designs that isolated incremental campaign impact from last-touch bias.
  • Lifted customer LTV 12% and cut churn 25%, by pairing uplift modeling with mixed-integer optimization to reallocate acquisition budget across channels.
Jan 2025 – May 2025

Data Scientist Capstone Customer Analytics & Segmentation

TD Bank · New York, NY

Customer Analytics Fraud Detection XGBoost
  • Raised fraud model precision-recall AUC from 0.20 to 0.67, by training XGBoost with SHAP on 100 million transactions.
  • Validated fraud-detection strategy for senior leadership, by quantifying incremental catch rate against a holdout.
  • Shaped segmentation and anti-money-laundering strategy, by clustering customers on transaction behavior and delivering automated reporting on the resulting risk tiers.
Sep 2024 – Jan 2025

Teaching Assistant Algorithms to Data Science

Columbia University · New York, NY

Student Mentorship Experiment Design
  • Mentored 250+ MS students on ML algorithms, analytical reasoning, experiment design, and inference through structured weekly sessions.
  • Designed case studies and assignments on A/B test design, modeling, and data-driven decision-making for applied data science coursework.

Earlier Experience & Internships

May 2023 – Nov 2023

Data Science Intern Operational Analytics

Navin Fluorine International Limited · Mumbai, IN

Predictive Maintenance Operational Analytics
  • Reduced equipment downtime by 35% and cut ingestion latency by 40% by architecting a predictive pipeline with XGBoost and Airflow.
  • Lifted recall 15% and shortened repair cycles by deploying anomaly detection using XGBoost and Isolation Forest to surface high-risk events.
Dec 2021 – May 2022

Founding IT Intern Acquisition & Retention Analytics

E-Revbay Pvt. Ltd. · Mumbai, IN

Acquisition Analytics Retention Analytics
  • Increased qualified customer leads by 50% per quarter by leading acquisition analytics pipelines using Equifax data, SQL, and Python.
  • Eliminated 25% of manager intervention time and accelerated response by building real-time dashboards and root-cause pipelines in Tableau.
Jun 2021 – Aug 2021

Artificial Intelligence Intern

Verzeo · Remote

Data Preprocessing Computer Vision CNN
  • Processed tabular diamond datasets for classification, improving model accuracy by 17% through data cleaning and feature engineering.
  • Built flower image recognition with CNN techniques, achieving 85% accuracy while reducing error rates by 15%.

02 · Selected ProjectsSide Quests

LLMs · Graph RAG · Retrieval 9 min read

IBM Agentic Library: Graph RAG for Document Q&A

Built a Graph RAG pipeline with Neo4j and FAISS for structured retrieval. Knowledge-graph modeling improved retrieval precision by 30%+ over naive search.

Explore the analysis →
Analytics · Experimentation 12 min suite read

Product Analytics Suite

A combined suite of three analyses covering feature impact, growth allocation, and root-cause diagnosis using PSM, DiD, uplift modeling, and MMM.

Explore the analysis →
NLP · Vision · Detection 7 min read

M(iche)Langelo: Analysis on AI-Generated Art

Built an end-to-end Reddit pipeline for style, caption, and source detection using CLIP, BLIP, and SuSy. Transfer learning to a 3-class SuSy head improved MidJourney and DALL-E detection on real-world art streams.

Explore the analysis →

03 · Publications & MediaOn the record

Media Praise

Grand finale collage for Hackathon on Plastic-Free Rivers with AI at REVA University
Featured in ThePrint (ANI PR)

"Hackathon on Plastic-Free Rivers with AI" Grand Finale at REVA University

11 September 2023 · REVA University + Kyndryl

Team EcoGuards placed 1st out of 750 teams from 19 countries (1,311 participants), winning a INR 1,50,000 prize for a Vision AI system that detects, classifies, and segments river plastic from drone imagery.

Research Publications

Predictive maintenance for metro systems

Google Scholar · Lead Author

Demonstrated a sensor-driven pipeline that predicts equipment failures so operations teams can schedule targeted maintenance and reduce downtime.

Diabetes detection optimized for recall

Google Scholar · Lead Author

Prioritized recall in model design to reduce missed diagnoses, improving early detection reliability for clinical use.

NLP to SQL for mobile learning

Google Scholar

Built an NLP-to-SQL interface that lets non-technical users query student and CSV data directly from mobile devices.

Semi-supervised disease prediction

Google Scholar

Applied semi-supervised methods to leverage unlabeled clinical data and improve prediction robustness in ambiguous diagnostic cases.

Leadership & Impact

04 · EducationThe Paper Chase

M.S. Data Science

Columbia University · New York, NY

Sep 2024 – Dec 2025
GPA 3.7/4.0
Relevant Coursework
Agentic AI Fintech & Data Economy (PhD elective) Big Data Analytics Statistics Applied Machine Learning

B.Tech Honors (Computer Engineering, Data Science/Analytics)

NMIMS University · Mumbai, India

Jun 2020 – May 2024
GPA 3.9/4.0
Focus Areas
Computer Engineering Data Science Analytics

05 · ContactSay hello

Always happy to discuss data science, experimentation, causal inference, and GenAI applications.