Self-directed
PokéTyper

multi-modal ML
serving
Project overview
Multi-modal Pokémon type classifier
Predicts a Pokémon's 1–2 elemental types from four modalities, fusing sprite vision, stats, Pokédex text, and hand-engineered color/shape features with a transformer that decides which modality to trust per creature.
Date
2026
Field
AI / ML
Stack
Python, PyTorch, ViT, MiniLM, FastAPI, ONNX
4 fused modalities
Gen-9 test macro-F1 ≈ 0.48
cross-gen split
PokéTyper
Context
An 18-way multi-label classifier evaluated under a deliberate cross-generational split to measure real generalization, not interpolation.

Most classifiers hide which signal drove a prediction.
The hard part
PokéTyper exposes its own learned trust, you can read out that it distrusts the sprite for a serpent like Gyarados and leans on text and stats instead.


What it took
- Fused a frozen ViT-B/16, a residual-MLP tabular branch, frozen MiniLM text embeddings, and a classic-CV visual prior as tokens through a transformer with a learnable [FUSION] token and gate.
- Trained with asymmetric focal loss for severe class imbalance, EMA, and cosine LR with warm restarts, then made predictions meaningful with post-hoc temperature scaling and learned per-type thresholds.
- Backed claims with an Optuna HPO sweep and a controlled ablation matrix, surfacing that text is the most valuable modality and frozen ImageNet vision the least transferable across the art-style shift.

Outcome






