Body Measurement Dataset Collection
Collected 50+ image samples with body measurements and metadata for an AI body-measurement estimation model. Delivered with a full data dictionary and zero quality rejections from the client.
Available for new projects
ML Dataset Engineer / Data Science
I build datasets that ML models actually train well on. If your model is underperforming, it's probably a data problem. I handle the full pipeline — collection, cleaning, structuring, and validation — so your team can focus on modeling, not fixing CSVs.
Most ML problems are data problems. I fix the data.
I specialize in collecting, cleaning, and preparing data at scale. Every Upwork project I've completed has a 5★ rating, and my job success score is 100%.
I cover the whole pipeline: collecting from scratch, deduplication, normalization, feature engineering and quality validation. You get a dataset that's documented, reproducible and ready to train on.
I'm currently studying Computer Science with an AI focus at NUST and hold two IBM certifications in Python data analysis and machine learning.
Manual, API-based, research, or client-provided sources. Delivered at scale with full documentation.
Deduplication, normalization, missing-value handling and format standardization for production.
Distribution checks, outlier detection and consistency audits. Zero rejections from clients.
Collected 50+ image samples with body measurements and metadata for an AI body-measurement estimation model. Delivered with a full data dictionary and zero quality rejections from the client.
Built a comprehensive indoor plants dataset with a full QA validation pipeline. Structured, labeled and model-ready as a CSV export.
Wrote a Python extraction pipeline that pulled training data from PDF documents and delivered labeled, model-ready CSV files with documentation.
End-to-end data cleaning and logistic regression model: preprocessing, feature scaling and model evaluation. Shipped production-ready with a full README.
Jan 2026 — Present
AI data collection, cleaning and preprocessing. Delivered model-ready datasets for a range of ML applications, with a 100% job success score and a 5★ rating on every completed project.
Jan 2026 — Present
Python-based extraction from PDFs, websites and documents into clean, structured, model-ready data.
Feb 2026
Python for data analysis, statistical methods, visualization and business insights.
Jan 2026
Covered ML algorithms, model development and optimization techniques.
2023 — 2027
National University of Sciences & Technology, focusing on algorithms, data structures and AI.
Fixed-price packages with clear deliverables. Need something else? Ask for a custom quote.
Extract and structure training data from PDF documents. Delivered as labeled, model-ready CSV files.
A clean dataset and a logistic regression model in Python. Production-ready with full documentation.
Collect, organize and validate image datasets with metadata mapping. Zero quality rejections guaranteed.
"Very Proficient and always successfully delivers positive results. Excellent work on the Indoor Plants Dataset - perfectly structured and validated."
"Zero quality rejections on the body measurement dataset. Daniyal delivered exactly what was needed - 50+ image samples with perfect metadata mapping. Highly professional."
"Exceptional data extraction and structuring from PDFs. Delivered clean CSV files ready for training immediately. Great communication throughout the project."
Have a dataset that needs building or fixing? Tell me about the project and I'll get back to you within a few hours.