I'm Swaminathan Sankaran

I like building ML systems that don't break the moment they hit real data. Currently studying Data Science at UB and looking for internship opportunities.

Swaminathan Sankaran

About

A bit about me

I’m a Master’s student in Data Science at the University at Buffalo, and curiosity has shaped a lot of how I learn and work. I enjoy understanding how things work, whether that is a technical system, a human decision, or a bigger idea in philosophy, psychology, and rationality.

A lot of that mindset comes from the things I’ve always enjoyed outside of work. I used to be a competitive gamer, which taught me how to think strategically, stay composed, and keep improving through practice. I’m also a speedcuber, so I’ve always liked patterns, problem solving, and the challenge of getting better with focus and repetition. I enjoy reading, exploring different perspectives, and learning through food, culture, and the ways people think and live.

That same curiosity is what drew me to data science and machine learning. It also shapes how I approach my work. I enjoy building systems that are thoughtful, useful, and reliable, and I’m especially interested in turning ideas into practical solutions that create real impact.

I’m currently looking for internship opportunities where I can keep learning, contribute meaningfully, and grow as both a problem solver and engineer.

Thanks for stopping by. Don’t hesitate to say hi. I’d love to hear from you.

Projects

Selected Work

Vision-Language / Evaluation

Aug 2026 GitHub

Compositional Grounding Probes for Open-Vocabulary Segmentation

  • Built a 914-item benchmark mined automatically from COCO, LVIS, and Visual Genome to test text-prompted segmentation models
  • Engineered a reproducible evaluation pipeline with checkpointed GPU batch inference, unit-tested geometry utilities, and automated metrics
  • Exposed systematic model failures on compositional prompts: 20.7% accuracy on negation vs 93% on controls, and 0 of 51 correct on paired spatial queries
  • Traced all failures to the text-grounding stage via an oracle ablation that scored 99.6% with ground-truth inputs

MLOps / Production Systems

Feb 2026 GitHub

Drift-Aware MLOps Pipeline

  • Designed a drift-aware MLOps pipeline to detect silent model degradation and trigger automated retraining, improving model reliability under changing data distributions
  • Built and containerized the system using FastAPI, Airflow, MLflow, PostgreSQL, and Evidently AI with Docker Compose, with Kubernetes manifests for production readiness
  • Monitored system health in real time using Prometheus and Grafana dashboards
  • Validated end-to-end performance by simulating covariate and concept drift, triggering automated retraining that recovered shifted-distribution PR-AUC from 0.37 to 0.89 while preserving historical ROC-AUC at 0.75

Multimodal Deep Learning

Jan 2026 GitHub

Molecular Similarity Prediction

  • Modeled molecular similarity for drug discovery to improve candidate selection and ranking of chemical compounds
  • Designed a multi-modal architecture fusing 2D image features (ResNet-18), 3D geometry (SchNet), and fingerprint embeddings, trained with contrastive learning across 291,742 molecules
  • Achieved ~0.92 Pearson correlation on 200 expert-annotated pairs, outperforming the Tanimoto similarity baseline on unseen molecular pairs

Medical Imaging

Dec 2025 GitHub

CT Image Tamper Detection

  • Built to detect tampering at the patch level within CT scans, where localized manipulation is harder to catch than whole-scan forgeries
  • Jointly trained a 3D convolutional compressor with an ImageNet-pretrained ResNet-18 across 169 volumetric lung CT scans to convert 16-slice 3D patches into compact 2D feature maps
  • Achieved 0.95 validation AUC, outperforming 2.5D, 3D, and projection-based baselines at lower computational cost

Experience

Where I've Worked

Machine Learning Engineer, Intern

Zolvit (formerly Vakilsearch)

Feb 2024 — Aug 2024
  • Eliminated ~23 hours of weekly manual data entry by engineering a fully automated ingestion pipeline for 350+ complex documents, combining Google Vision OCR with a fine-tuned T5-large model to replace a fragile legacy process
  • Drove 99.9% service uptime and eliminated recurring outages by deploying Dockerized ML microservices on AWS EC2 with FastAPI, restart policies, and real-time monitoring via Grafana
  • Automated routing of customer-uploaded documents across categories by achieving 92% classification accuracy, benchmarking Logistic Regression, SVM, and XGBoost on TF-IDF and Doc2Vec features extracted via Google Vision OCR
  • Improved legal precedent retrieval precision by ~28% for lawyers by building a hybrid search engine combining Elasticsearch keyword search with Pinecone vector embeddings across 10,000+ documents
  • Maintained reliability of production ML systems by handling incident response, debugging failures, and improving robustness of legacy pipelines

Education

Academic Background

University at Buffalo

MS in Engineering Science (Data Science) GPA 3.81

Aug 2025 — Dec 2026

Relevant Coursework

Statistical Learning I & II Machine Learning Probability Theory Database Data Science Data Models & Query Languages Numerical Methods

Vellore Institute of Technology

B.Tech in Computer Science and Engineering (AI & ML)

Aug 2019 — Jul 2023

Relevant Coursework

Machine Learning Deep Learning Reinforcement Learning Computer Vision Applied Linear Algebra Statistics Data Structures & Algorithms Database Management Systems

Skills

Technical Toolkit

Programming Languages

Python
C+C/C++
SQLSQL
RR

Machine Learning & AI

PyTorch
TVtorchvision
PyGPyTorch Geometric
TensorFlow
Scikit-learn
XGXGBoost
Hugging Face
EMEmbeddings
RGRAG
RLRepresentation Learning
LCLangChain
LGLangGraph
TBTensorBoard

Data Processing & Analysis

Pandas
NumPy
DIMedical Imaging (DICOM)
RDRDKit

Data Visualization & BI

MPMatplotlib
SBSeaborn
Plotly
Streamlit
TATableau
BIPower BI

MLOps & Cloud

Docker
Kubernetes
AWSAWS
Airflow
MLflow
EVEvidently AI
Prometheus
Grafana

Databases & Data Engineering

PostgreSQL
MySQL
Elasticsearch
PNPinecone
Spark
Snowflake

Developer Tools & Systems

Linux
Git/GitHub
PTPytest
APIREST APIs
GPUCUDA / GPU Training
AGCoding Agents

Awards & Certifications

Recognition