Available for new projects

Daniyal Ali Dana

ML Dataset Engineer / Data Science

I build datasets that ML models actually train well on. If your model is underperforming, it's probably a data problem. I handle the full pipeline — collection, cleaning, structuring, and validation — so your team can focus on modeling, not fixing CSVs.

  • Pakistan · GMT+5
  • Replies within 4 hrs
Portrait of Daniyal Ali Dana
Job success
100%
Client rating
5.0/5
Projects delivered
4
IBM certifications
2
01

About

Most ML problems are data problems. I fix the data.

I specialize in collecting, cleaning, and preparing data at scale. Every Upwork project I've completed has a 5★ rating, and my job success score is 100%.

I cover the whole pipeline: collecting from scratch, deduplication, normalization, feature engineering and quality validation. You get a dataset that's documented, reproducible and ready to train on.

I'm currently studying Computer Science with an AI focus at NUST and hold two IBM certifications in Python data analysis and machine learning.

  • Data collection

    Manual, API-based, research, or client-provided sources. Delivered at scale with full documentation.

  • Data cleaning

    Deduplication, normalization, missing-value handling and format standardization for production.

  • QA & validation

    Distribution checks, outlier detection and consistency audits. Zero rejections from clients.

02

Selected work

Apr 2026 5.0

Body Measurement Dataset Collection

Collected 50+ image samples with body measurements and metadata for an AI body-measurement estimation model. Delivered with a full data dictionary and zero quality rejections from the client.

  • Image collection
  • Metadata
  • QA validated
Apr 2026 5.0

Verified International Indoor Plants Dataset

Built a comprehensive indoor plants dataset with a full QA validation pipeline. Structured, labeled and model-ready as a CSV export.

  • Data collection
  • Validation
  • CSV export
Apr 2026 Expert level

PDF Data Extraction & Structuring

Wrote a Python extraction pipeline that pulled training data from PDF documents and delivered labeled, model-ready CSV files with documentation.

  • Python
  • PDF extraction
  • CSV output
Mar 2026 5.0

Customer Churn Prediction Model

End-to-end data cleaning and logistic regression model: preprocessing, feature scaling and model evaluation. Shipped production-ready with a full README.

  • Python
  • Scikit-learn
  • Logistic regression
  • Data cleaning
03

Stack

Languages
  • Python
  • Java
  • C++
  • R
  • Bash
Data tools
  • Pandas
  • NumPy
  • Scikit-learn
  • Jupyter
  • Excel
  • Google Sheets
ML techniques
  • K-Means
  • Logistic Regression
  • SVM
  • Decision Trees
  • Naive Bayes
  • Hyperparameter Tuning
Vision & extraction
  • YOLO
  • Tesseract OCR
  • Scrapy
  • MATLAB
Visualization
  • Matplotlib
  • Tableau
  • EDA
  • Distribution analysis
Pipeline
  • Collection
  • Cleaning
  • Preprocessing
  • Feature engineering
  • Validation
  • Train/test splits
04

Experience & education

  1. Jan 2026 — Present

    ML Dataset Engineer · Freelance, Upwork

    AI data collection, cleaning and preprocessing. Delivered model-ready datasets for a range of ML applications, with a 100% job success score and a 5★ rating on every completed project.

  2. Jan 2026 — Present

    PDF Data Extraction Specialist · Freelance, Upwork

    Python-based extraction from PDFs, websites and documents into clean, structured, model-ready data.

  3. Feb 2026

    Data Analysis with Python · IBM Certification

    Python for data analysis, statistical methods, visualization and business insights.

  4. Jan 2026

    Machine Learning with Python · IBM Certification

    Covered ML algorithms, model development and optimization techniques.

  5. 2023 — 2027

    BS Computer Science (AI) · NUST

    National University of Sciences & Technology, focusing on algorithms, data structures and AI.

05

Services

Fixed-price packages with clear deliverables. Need something else? Ask for a custom quote.

PDF Data Extraction

Extract and structure training data from PDF documents. Delivered as labeled, model-ready CSV files.

  • Extraction
  • Structuring
  • Labeling
  • Quality check
From $40 2-day delivery

Clean Dataset + Regression Model

A clean dataset and a logistic regression model in Python. Production-ready with full documentation.

  • Cleaning
  • Model
  • Testing
  • README
From $5 1-day delivery

Image Dataset Organization

Collect, organize and validate image datasets with metadata mapping. Zero quality rejections guaranteed.

  • Collection
  • Organization
  • Metadata
  • QA report
From $50 Custom timeline
06

What clients say

"Very Proficient and always successfully delivers positive results. Excellent work on the Indoor Plants Dataset - perfectly structured and validated."
Verified Upwork ClientIndoor Plants Dataset · Apr 2026
"Zero quality rejections on the body measurement dataset. Daniyal delivered exactly what was needed - 50+ image samples with perfect metadata mapping. Highly professional."
AI Model DeveloperBody Measurement Dataset · Apr 2026
"Exceptional data extraction and structuring from PDFs. Delivered clean CSV files ready for training immediately. Great communication throughout the project."
ML Pipeline LeadPDF Data Extraction · Apr 2026
07

Let's work together

Have a dataset that needs building or fixing? Tell me about the project and I'll get back to you within a few hours.