SK // Portfolio
About Experience Projects Skills Contact Download Resume
Open to Work
MS Applied Data Science // Syracuse University

Siddhant Kasture

Machine learning and NLP engineer focused on building AI systems that are not just accurate, but trustworthy. Currently exploring LLM reliability, multi-agent architectures, and production forecasting infrastructure.

View Projects Download Resume ↓ Get In Touch
Scroll to explore
GitHub
LinkedIn
Machine Learning
Natural Language Processing
Deep Learning
AWS Cloud
Data Engineering
Generative AI
RAG Systems
Multi-Agent Systems
LLM Evaluation
Forecasting
Machine Learning
Natural Language Processing
Deep Learning
AWS Cloud
Data Engineering
Generative AI
RAG Systems
Multi-Agent Systems
LLM Evaluation
Forecasting
01

About

I'm Siddhant, a Data Scientist and recent MS Applied Data Science graduate from Syracuse University's iSchool, focused on building machine learning and NLP systems that work reliably in high-stakes domains like finance, healthcare, and agriculture.

Over the past year, I've worked across research and industry: conducting sentiment analysis and training BERT models at NEXIS Technology Lab, automating mortgage data pipelines at GenNext.Mortgage, leading an LLM reasoning workstream at HyperQuark Intelligence Labs, and building evaluation harnesses for LLM and VLM outputs at Handshake AI. Along the way, I've developed a strong instinct for where models break, and how to fix them.

That perspective shapes the work I'm most drawn to: building production forecasting infrastructure for agricultural risk, evaluating LLM reliability in financial QA, designing multi-agent systems with built-in fairness constraints, and shipping agentic tools with MCP and LangGraph. I care about building AI that's not just accurate, but trustworthy.

Quick Facts

Location Syracuse, New York
Degree MS Applied Data Science · 4.0/4.0 GPA
University Syracuse University, iSchool
Graduation May 2026
Languages Python, R, SQL, JavaScript
Focus Areas ML · NLP · Cloud · GenAI · Forecasting
Status Open to Opportunities
02

Experience

Dec 2025 – Present

AI Evaluation Specialist

Handshake AI · Remote

Build evaluation harnesses and behavioral test suites for LLM and VLM outputs, contributing metrics that drove an 80% improvement in model accuracy and safety benchmarks. Reduced hallucination rates by 15% by instrumenting drift monitoring across release cycles and translating noisy production telemetry into actionable engineering tickets.

Apr 2026 – Present

Research Team Lead

HyperQuark Intelligence Labs · Remote

Lead the LLM reasoning workstream within the lab's research group, coordinating experimental design and weekly research direction across collaborators. Drove team research on a DAG-based, job-weighted skill-assessment framework, improving ISCO classification accuracy by 25% over keyword-based baselines.

May – Jun 2025

Data Analyst Intern

GenNext.Mortgage · Boston, MA

Built secure API integrations using OAuth2, REST/SOAP protocols, and XML/JSON parsing via Google Apps Script to automate mortgage pricing data extraction. Automated extraction for 200+ daily loan records, reducing manual entry time by 40% with 99% data accuracy.

Sep 2024 – Dec 2025

Research AI Engineer

NEXIS Technology Lab · Syracuse, NY

Predicted 2024 U.S. Election outcomes with 84% accuracy by building a sentiment analysis pipeline on 50,000 tweets using NLP techniques and BERT transformers. Fine-tuned a BERT model via Hugging Face, achieving 12% performance gain over baseline through systematic hyperparameter optimization.

All Projects

01

Professional Ag Marketing AI Tool

Runner-up (2nd), CCDS/ProAg 2026 · production cattle/hog market-intelligence tool with daily CI/CD

FastAPI XGBoost Next.js
02

OpenAg Risk Twin

Solo calibrated corn risk system: conformal 80% interval, 7 regimes, precompute-at-refresh

FastAPI DuckDB Conformal
03

Betting Edge

Multi-Agent DQN-based sports betting system with risk profiling and ethical guardrails

Python DQN LangChain
04

Financial LLM Hallucination Study

Benchmarking hallucination rates across 6 LLMs on financial QA with RAG and probing tests

Streamlit RAG LLM-as-Judge
05

RAG-TechDocHelper

Retrieval-Augmented Generation system using LangChain, GPT-4, and FAISS vector store

LangChain FAISS GPT-4
06

Genomic Sequence Analysis

Deep learning pipeline on AWS for promoter classification across 100K+ genomic samples

AWS CNN Dask
07

Diabetes Risk Assessment

Predicting Type II Diabetes from socio-economic factors using TensorFlow neural networks

TensorFlow SMOTE XGBoost
08

MCP Deep Researcher

LangGraph-based research agent with MCP server, multi-turn memory, and SHA-256 caching

LangGraph MCP Tavily
09

Global Soccer Ranking DBMS

Relational database with SQL, Azure Data Studio, Power BI dashboard, and MS Power App for team rankings

SQL Azure Power BI
10

Energy Consumption Analysis

IST 687 · Statistical and ML analysis of residential energy patterns across 5,000+ homes with Random Forest modeling

Python Random Forest EDA
04

Skills

Honest self-assessment. Dot scale reflects depth of experience, not just exposure. Hover for project context.

Core Can whiteboard it, ship it, debug it at 2am
Languages
Python
All Projects
R
Energy Analysis
SQL
Soccer DBMS
ML & Data
Scikit-learn
Diabetes · Hallucination Study
XGBoost
ProAg · OpenAg · Diabetes
pandas
All Projects
NumPy
All Projects
NLTK
Hallucination Study · NEXIS Lab
Dask
Genomic Analysis
Visualization
Matplotlib
Diabetes · Genomic · Energy
seaborn
Diabetes · Energy Analysis
Power BI
Soccer DBMS
Streamlit
Hallucination Study · Betting Edge
Tableau
NEXIS Lab
Tools
Git
All Projects
Jupyter
All Projects
VS Code
All Projects
Prompt Engineering
Hallucination Study · Handshake AI
MySQL
Soccer DBMS
Proficient Shipped projects with these, comfortable in production
ML & AI Frameworks
TensorFlow
Genomic · Diabetes
Keras
Genomic · Diabetes
Hugging Face
NEXIS Lab · Hallucination Study
LangChain
RAG-TechDocHelper · Betting Edge
LangGraph
MCP Deep Researcher
MCP Server
MCP Deep Researcher
DQN / RL
Betting Edge
Backend & Deployment
FastAPI
ProAg · OpenAg
Next.js
ProAg · OpenAg
Render / Vercel
ProAg · OpenAg
GitHub Actions
ProAg · OpenAg · RAG-TechDocHelper
NLP & GenAI
RAG
RAG-TechDocHelper · Hallucination Study
ChromaDB
Hallucination Study
TextBlob
NEXIS Lab
BERT
NEXIS Lab · Hallucination Study
FAISS
RAG-TechDocHelper · Hallucination Study
Cloud & Infrastructure
AWS SageMaker
Genomic Analysis
Amazon S3
Genomic Analysis
Azure
Soccer DBMS
AWS Bedrock
Hallucination Study
Databases & Tools
DuckDB
OpenAg Risk Twin
MongoDB
SQLite
REST APIs
GenNext.Mortgage
Databricks
RLHF
Handshake AI
Familiar Have used, would ramp up quickly
HTML/CSS
Google Apps Script
GenNext.Mortgage
statsmodels
OpenAg Risk Twin
MS Access
JavaScript
Tavily
MCP Deep Researcher
Docker
C
05

Education

MS Applied Data Science
Syracuse University, School of Information Studies
Aug 2024 – May 2026 GPA: 4.0/4.0
Phi Kappa Phi Honor Society (2026)
BE Electronics & Telecom
University of Mumbai, Thakur College of Engineering
Aug 2019 – May 2023 GPA: 8.97
Get in touch

Let's build
something great.

Production Pipeline
Professional Ag Marketing AI Tool
INGESTION + DAILY CI/CD DATA SOURCES Yahoo Finance futures LE=F · HE=F · ZC=F · ZM=F USDA AMS PDFs CI/CD Pipeline GitHub Actions · 14 steps ~4 min · auto-deploy Feature Panel 2019–2025 XGBOOST MODELS time-split · 2025 gated holdout · macro-F1 Hog Price-Direction 78.5% high-conviction (2024) vs 55.29% baseline · down vs up Cattle Basis-Direction 70.25% high-conviction (2024) vs 54.11% baseline · narrow vs widen SERVING (LIVE) FastAPI · Render model-serving API Next.js · Vercel confidence-tier dashboard Baselines: majority-class + logistic regression on training data (2019–2023), benchmarked against the XGBoost models.
Automated pipeline
Models
Live serving
Data / baselines
Precompute-at-Refresh
OpenAg Risk Twin
DAILY GITHUB ACTION: COMPUTE, VALIDATE, COMMIT Corn ZC=F free public data COMPUTE DuckDB warehouse Regime detection (7) structural-break Conformal calibration 80% → ~81% coverage Point forecast = naive random walk A/B test: regime features hurt (reported) Calibrated 80% interval per-regime coverage + LLM decision-readiness layer Contract-validated forecast committed to repo (versioned artifact) THIN SERVE (LIVE) FastAPI Render Next.js Vercel Honest spine: no model beats the random walk on 30-day point forecast. Value is calibrated uncertainty, not accuracy
Compute / validate
Calibrated output
Live serving
Data / honest spine
System Architecture
Betting Edge
MULTI-AGENT PIPELINE User Input Streamlit UI Risk Profiler User segmentation DQN Engine Prediction agent Ethical Net Bias guardrails Value-Betting Algorithm +EV identification AI Assistant + Dashboard Deep query | Recommendations DATA LAYER Sports APIs Historical Odds Match Features User Profiles
Core pipeline
Ethics layer
Data / support
Evaluation Pipeline
Financial LLM Hallucination Study
EVALUATION FRAMEWORK FinQA Multi-step reasoning FiQA Factual lookup SEC 10-K RAG context 6 LLMS Claude Sonnet 4.5 Claude Haiku 4.5 GPT-5 Gemini 2.5 Flash Llama 3.3 70B Qwen3 32B Semantic Sim Embedding cosine ROUGE-L Token overlap LLM-as-Judge Llama 3.1 8B Ensemble Score > 0.5 = hallucination PROBING TESTS Paraphrase Temporal Counterfactual Unanswerable
Core pipeline
Detection methods
Datasets / probes
Cloud ML Pipeline
Genomic Sequence Analysis
AWS PIPELINE FASTA Files 100K+ samples Dask Parallel chunk Amazon S3 Storage layer SageMaker ml.g5.2xl NVIDIA A10G GPU TensorFlow CNN Classification Output OPTIMIZATION Stratified split | Dropout | Early stopping | 60% faster
AWS services
Preprocessing
Output
Model Comparison Pipeline
Diabetes Risk Assessment
ML PIPELINE CDC BRFSS 2023 dataset SMOTE Oversampling + feature eng. Logistic Regression Random Forest XGBoost Dense NN (best) Results F1: 0.57 DNN winner BatchNorm + Dropout Socio-economic predictors: income, education, BMI, physical activity, smoking, healthcare access
Best model / output
Preprocessing
Baseline models
Agentic Pipeline
MCP Deep Researcher
LANGGRAPH AGENT User Query Streamlit / MCP LangGraph Router Search depth config Breadth params OpenAI GPT-4o Reasoning engine Tavily Search Web search API Response + source scores + citations PERSISTENT MEMORY 8-message context | Multi-turn state CACHING + CREDIBILITY SHA-256 hash | Source scoring | Dedup MCP SERVER (server.py) Deployable as tool in Claude Desktop or Cursor IDE
Core agent flow
Router / orchestrator
MCP integration