Pikachu The Story and Mystery of Pokémon

A Comprehensive Data Science Exploration

P8105 – Data Science I | Columbia University | Fall 2025

Executive Summary

This project analyzes 742 Pokémon across nine generations using machine learning and statistical modeling. We achieve 99.5% accuracy in legendary prediction and develop a comprehensive dual-rating system for competitive battling.

Key Results: Legendary Pokémon are statistically distinct (BST 680 vs 420), concentrated in Dragon/Psychic/Steel types, and accurately classified through Random Forest modeling.

Dataset Overview

742
Pokémon Species
101
Legendary Pokémon
18
Elemental Types
519
Average Base Stats

Project Components

Pikachu
Exploratory Analysis
In-depth exploration of Pokémon attributes, type distributions, and stat correlations.
Eevee
Machine Learning
Random Forest classification achieving 99.5% AUC for legendary prediction.
Mew
Statistical Modeling
PCA analysis revealing 4 distinct clusters: Sweepers, Tanks, Balanced, Early-Game.
Jigglypuff
Type Effectiveness
Analysis of type strengths and optimal combinations for competitive play.
Togepi
Interactive Explorer
Dynamic filtering and visualization for real-time data exploration.
Pichu
Dual Rating System
Separate offensive and defensive ratings (100-point scale each) with role classification and tier assignment.

Research Questions

Question Status Method
Can we identify legendary Pokémon from their stats? Complete Random Forest (99.5% AUC)
How do physical attributes correlate with battle stats? Complete Correlation Analysis
Which type combinations are strongest/weakest? Complete Type Effectiveness Matrix
Which types are most likely to be legendary? Complete Chi-Square Test
Can we develop a dual-rating system for competitive battling? Complete Dual Rating Algorithm
Does our optimal team match competitive standards? Complete Competitive Meta Analysis

Key Findings

Top Types

Discoveries

Legendary Characteristics

  • Statistical Distinction: BST 680 (legendary) vs 420 (regular), p < 0.001
  • 99.5% Accuracy: Random Forest classification
  • Type Bias: Dragon (30%), Psychic (25%), Steel (20%)

Pattern Analysis

Most Common Types: Water (18%), Normal (15%), Grass (12%) Rarest Types: Flying (4%), Fairy (6%), Ice (6%) Dual-Type Rate: 47% have secondary typing

Cluster Profiles

Profile BST Legendary % Examples
High Sweepers 580 35% Mewtwo, Rayquaza
Def Tanks 480 8% Snorlax, Steelix
Balanced 450 2% Arcanine, Gyarados
Early-Game 320 0% Pidgey, Rattata

Type Distribution

Methodology

Analysis Pipeline

1.Data Collection

Web scraping from PokémonDB using rvest, handling 1000+ species and alternate forms

2.EDA

Feature engineering with tidyverse: one-hot encoding, correlation matrices, distribution analysis

3.Clustering

PCA + K-Means (k=4) using factoextra, explaining 70% variance in first 3 components

4.Classification

Random Forest tuning with caret: 500 trees, mtry=3, achieving 99.5% AUC

5.Dual Rating System

Separate offensive and defensive ratings using S-curve BST mapping, effective stat calculations, and type matchup analysis

Tech Stack

Languages & Tools • R (≥ 4.0), RStudio, Git

Core Packages • tidyverse, plotly, kableExtra • caret, randomForest, pROC • factoextra, cluster, FactoMineR

Visualization • ggplot2, plotly (interactive) • crosstalk, DT (tables)

Interactive Tools

Explore the Analysis

Bulbasaur Charmander Squirtle Pikachu Eevee Mew Jigglypuff Togepi

EDA ML Analysis Rating System Interactive Report Team

Team

P8105 – Data Science I | Fall 2025

  • Ruipeng Li (rl3616) - Data Cleaning, EDA, Rating System
  • Xuange Liang (xl3493) - Modeling, Rating System
  • Yiwen Zhang (yz4994) - Visualization, Documentation
  • Leah Li (yl5828) - Visualization, Documentation

Columbia University | Mailman School of Public Health

View on GitHub