Open to ML roles & internships

Machine Learning Engineer_

Hi, I'm ABC XYZ — I turn messy, million-row data into models that actually ship. NLP, RAG pipelines and big-data ML, built end-to-end, evaluated hard, deploy-ready.

Gwalior, India M.Tech IT @ IIITM Gwalior B.Tech CGPA 9.34
25M+Records processed
34M+Candidate pairs generated
0.82Macro F0.5 — held-out set
FINALISTSmart India Hackathon '24

01 // About

Research in. Production out.

I'm a results-driven Machine Learning Engineer and M.Tech student specializing in end-to-end ML pipelines, NLP, and large-scale data processing. I've engineered high-performance classification models over 25M+ record datasets and deployed robust retrieval-augmented generation systems — with rigorous evaluation and a documented allergy to hallucinating models.

I live in the gap most teams complain about: taking something that works in a notebook and making it work in the real world — clean data in, honest metrics out, pipelines that don't fall over at scale.

Rapid prototyping

Notebook → working pipeline in days, not quarters. Ship, measure, iterate.

Pipeline optimization

DuckDB, blocking keys, caching — I make 25M rows feel like 25.

Rigorous evaluation

Faithfulness checks, held-out splits, real metrics. If it isn't measured, it didn't happen.

02 // Experience

Where I've shipped.

Machine Learning Intern

Zeronsec India Pvt. Ltd. May 2024 — July 2024
  • Architected an Elastic Stack ingestion pipeline transforming high-volume, unstructured logs into queryable, structured data for downstream analysis.
  • Developed automated Python preprocessing with pandas & regex to normalize datasets — significantly improving data quality and consistency.
  • Collaborated with cross-functional engineering teams to integrate processed pipelines into core analytics workflows and real-time dashboards.
PythonpandasElastic StackRegexData Pipelines

// next entry: your team's pipeline problems →

Let's write it ↗

03 // Projects

Selected work.

Three systems, three different flavors of hard: scale, semantics, and accessibility.

01

Business Entity Resolution

Amazon ML Challenge 2026

Offline entity-resolution pipeline that reconciles ~25 million records across three disparate data sources — no cluster harmed.

  • Disk-backed DuckDB + multi-key blocking strategies to generate 34M+ candidate pairs at 77.7% recall.
  • Engineered pairwise similarity features; trained an XGBoost classifier to a 0.82 macro F0.5 on held-out validation.
View Code
02

PaperTrail

RAG · Citation-grounded Q&A

Q&A over ML research corpora — answers with precise inline source citations you can actually verify.

  • Engineered a RAG pipeline fusing dense embeddings (sentence-transformers + FAISS) with BM25 keyword search — capturing semantic concepts and exact terms.
  • Tuned chunk sizing & overlap to maximize faithfulness; rigorous evaluation against retrieved passages to eliminate unsupported hallucinations.
View Code
03

Voice Bridge

Smart India Hackathon '24 · Finalist

Accessible mobile app + speech pipeline giving the deaf & mute community a voice — in real time, in their language.

  • Curated and processed diverse multilingual audio datasets; built speech-to-text transformation pipelines for real-time inference.
  • Integrated the ML backend into a highly accessible Flutter mobile application.
View Code

04 // Stack

The toolbox.

ML & AI

×10
PythonNumPypandasscikit-learnXGBoostHugging FaceLangChainNLPRAGFeature Engineering

Data & Databases

×7
SQLDuckDBMySQLPostgreSQLMongoDBElastic StackETL Pipelines

Languages & Web

×6
C/C++JavaJavaScriptReactNext.jsExpress

DevOps & Tools

×7
GitGitHubDockerLinuxAWSFlaskVercel

05 // Education

Credentials, hard-earned.

IIITM Gwalior

M.Tech — Information Technology 2026 — 2028

Relevant coursework:

Machine Learning TechniquesParallel & Distributed ComputingMathematical Foundations of Data ScienceAlgorithms

SRM University

B.Tech — CSE (Big Data Analytics) 2022 — 2026

Undergraduate record:

Big Data AnalyticsComputer Science ★ CGPA 9.34

06 // Contact

Got data?
Let's talk.

Open to ML engineering roles, internships, research collaborations, and any dataset that keeps you up at night. Average reply time: faster than my training loops.

Open Mail App
Email copied — talk soon