Machine learning and NLP engineer focused on building AI systems that are not just accurate, but trustworthy. Currently exploring LLM reliability, multi-agent architectures, and production forecasting infrastructure.
I'm Siddhant, a Data Scientist and recent MS Applied Data Science graduate from Syracuse University's iSchool, focused on building machine learning and NLP systems that work reliably in high-stakes domains like finance, healthcare, and agriculture.
Over the past year, I've worked across research and industry: conducting sentiment analysis and training BERT models at NEXIS Technology Lab, automating mortgage data pipelines at GenNext.Mortgage, leading an LLM reasoning workstream at HyperQuark Intelligence Labs, and building evaluation harnesses for LLM and VLM outputs at Handshake AI. Along the way, I've developed a strong instinct for where models break, and how to fix them.
That perspective shapes the work I'm most drawn to: building production forecasting infrastructure for agricultural risk, evaluating LLM reliability in financial QA, designing multi-agent systems with built-in fairness constraints, and shipping agentic tools with MCP and LangGraph. I care about building AI that's not just accurate, but trustworthy.
Build evaluation harnesses and behavioral test suites for LLM and VLM outputs, contributing metrics that drove an 80% improvement in model accuracy and safety benchmarks. Reduced hallucination rates by 15% by instrumenting drift monitoring across release cycles and translating noisy production telemetry into actionable engineering tickets.
Lead the LLM reasoning workstream within the lab's research group, coordinating experimental design and weekly research direction across collaborators. Drove team research on a DAG-based, job-weighted skill-assessment framework, improving ISCO classification accuracy by 25% over keyword-based baselines.
Built secure API integrations using OAuth2, REST/SOAP protocols, and XML/JSON parsing via Google Apps Script to automate mortgage pricing data extraction. Automated extraction for 200+ daily loan records, reducing manual entry time by 40% with 99% data accuracy.
Predicted 2024 U.S. Election outcomes with 84% accuracy by building a sentiment analysis pipeline on 50,000 tweets using NLP techniques and BERT transformers. Fine-tuned a BERT model via Hugging Face, achieving 12% performance gain over baseline through systematic hyperparameter optimization.
A production market-intelligence tool that predicts livestock basis and price direction with a confidence score, built for the CCDS / ProAg Data Analytics Contest. Two XGBoost models (hog price-direction, cattle basis-direction) were validated on a time-based split with 2025 held out as a gated holdout, using macro-F1 as the headline metric. I led production and deployment: the FastAPI/Render backend, the Next.js/Vercel frontend integration, and a fully automated daily CI/CD pipeline (GitHub Actions, ~14 steps) that pulls futures data, parses USDA AMS PDFs, regenerates the panel, and redeploys end-to-end in roughly 4 minutes. Built with Gautam Balgi and Ameya Bhalerao.
A public, solo corn (ZC=F) probabilistic risk system framed as risk-management infrastructure under regime uncertainty, deliberately not a price predictor. Built entirely on free public data with a precompute-at-refresh architecture: a daily GitHub Action computes a contract-validated forecast and a thin FastAPI service serves it. A 2020–2026 walk-forward evaluation found that no model beats a naive random walk on the 30-day point forecast, and a pre-registered A/B test showed regime features actually hurt, reported as a negative result rather than buried. The value lives in calibrated honesty: a conformal 80% interval, seven validated market regimes, and an LLM layer that narrates the band in plain producer language.
A full-stack AI application combining a Multi-Agentic DQN-based predictive engine with user-facing interfaces for risk segmentation and responsible decision support. Led system integration: merged individual agent modules (prediction, DQN, ethical classifier) into a unified agentic pipeline, built the AI Assistant with deep query functionality, enhanced the dashboard, and managed API ingestion across the full stack.
View on GitHubBuilt a Streamlit evaluation framework benchmarking GPT, Claude, and Gemini on financial QA tasks using F1, ROUGE-L, and semantic similarity metrics across 100+ query-answer pairs. Developed a hallucination detection ensemble combining entailment classifiers and factual consistency checks, and designed adversarial probing tests to stress-test LLM reliability under misleading financial prompts.
View on GitHubA scalable genomic classification pipeline built on AWS infrastructure, designed to analyze over 100,000 genomic samples for promoter sequence identification. The project demonstrates end-to-end ML engineering, from distributed data preprocessing with Dask to GPU-accelerated model training on SageMaker.
View on GitHubA comprehensive machine learning study exploring whether non-clinical, socio-economic factors can predict Type II Diabetes risk. Using the 2023 CDC BRFSS dataset, the project benchmarks multiple classifiers and a custom deep neural network to demonstrate the predictive power of lifestyle variables.
View on GitHubAn agentic deep research system built on LangGraph and OpenAI that uses Tavily for web search, featuring multi-turn conversation with persistent memory, configurable search depth/breadth parameters, source credibility scoring, and intelligent caching with SHA-256 hashing. Deployable as an MCP server via server.py, enabling seamless integration as a research tool within Claude Desktop or Cursor IDE.
View on GitHubRunner-up (2nd), CCDS/ProAg 2026 · production cattle/hog market-intelligence tool with daily CI/CD
Solo calibrated corn risk system: conformal 80% interval, 7 regimes, precompute-at-refresh
Multi-Agent DQN-based sports betting system with risk profiling and ethical guardrails
Benchmarking hallucination rates across 6 LLMs on financial QA with RAG and probing tests
Retrieval-Augmented Generation system using LangChain, GPT-4, and FAISS vector store
Deep learning pipeline on AWS for promoter classification across 100K+ genomic samples
Predicting Type II Diabetes from socio-economic factors using TensorFlow neural networks
LangGraph-based research agent with MCP server, multi-turn memory, and SHA-256 caching
Relational database with SQL, Azure Data Studio, Power BI dashboard, and MS Power App for team rankings
IST 687 · Statistical and ML analysis of residential energy patterns across 5,000+ homes with Random Forest modeling
Honest self-assessment. Dot scale reflects depth of experience, not just exposure. Hover for project context.