Machine Learning All-in-one
What is machine learning?
- Arthur Samuel 1949 (1901 -1990): Machine Learning is a field of study that gives Computers the ability to learn without being explicitly programmed.
- Tom Mitchell 1997 (1951 -): Well-posed Learning Problem: A computer program is said to learn from experience E with respect to some task T and some performance measure P, if it is performance on T, as measured by P, improves with experience E.
Full Cycle of a ML project
- define project
- define and collect data (Data Sampling) + Data Preprocessing & Feature Engineering
- train model: training, Error analysis & iterative improvement -> loop between 2 and 3 until your model is done
- deploy in production (Machine Learning Systems Design): deploy, monitor, and maintain system -> back to 3 and/or 2 if needed
ML Algorithms Cheat Sheet
Learning Paradigms
Where does feedback come from?
Supervised Learning
Classification
- linear
- Logistic Regression
- Support Vector Machine SVM: Support Vector X#Support Vector Machine SVM
- non-linear
- Kernel SVM: Support Vector X#Kernels SVM
- K-Nearest Neighbor (k-NN)
- Naive Bayes
- Decision Tree Classification: Decision Tree & Random Forest#Decision Tree
- Random Forest Classification: Decision Tree & Random Forest#Random Forest
- Pros and Cons

- Multi-class vs. Multi-label Classification
Regression
- Types
- Linear Regression
- Polynomial Regression
- Regularized regression
- Lasso regression: Regression#Lasso Regression
- Ridge regression: Regression#Ridge Regression
- Support Vector Regression (SVR) Support Vector X#^46e7f9
- Decision Tree Regression: Decision Tree & Random Forest#Decision Tree
- Random Forest Regression: Decision Tree & Random Forest#Random Forest
- Pros and Cons

Unsupervised Learning
Clustering
- Types
- Centroid-based Clustering: K-Means Clustering
- Connectivity-based Clustering: Hierarchical Clustering
- Density-based Clustering: DBSCAN
- Graph-based Clustering: Affinity Propagation
- Distribution-based Clustering: Gaussian Mixture Model
- Compression-based Clustering: Spectral Clustering
- Pros and Cons
| Clustering Model | Pros | Cons | |
|---|---|---|---|
| K-Means | interpretability; works well on even-sized and globular-shaped data | need to predefine the number of clusters; not appropriate for outliers; low computation efficiency | |
| Hierarchical Clustering | no need to predefine the number of cluster; high computation efficiency; works well on high dimensional data | not appropriate for large data | |
| DBSCAN | no need to predefine the number of cluster; can determine arbitrarily-shaped clusters; can detect outlier | unstable performance (sensitive to density units parameter) | |
| Affinity Propagation | no need to predefine the number of cluster | low computation efficiency | low space effciency |
| Gaussian Mixture Model | highest computation efficiency; ensure clusters to follow Gaussian distributions | not appropriate when insufficient data in each cluster | |
| Spectral Clustering | high computaion efficiency | need to predefine the number of clusters |
Dimensionality Reduction
- Dimensionality Reduction#Feature Selection
- Dimensionality Reduction#Feature Extraction
- Principal Component Analysis (PCA)
- Linear Discriminant Analysis (LDA)
- Kernel PCA
- Quadratic Discriminant Analysis (QDA)
- T-Distributed Stochastic Neighbor Embedding (t-SNE)
- Uniform manifold approximation and projection (UMAP)
- Autoencoders
Density Estimation
Anomaly Detection
see Outlier & Anomaly Detection
Association Rule Learning
- Apriori: Association Rule Learning#Apriori
- Eclat: Association Rule Learning#Eclat
Semi-Supervised Learning
- dataset: partially labeled, with some data points having labels and others being unlabeled
- how it works
- use the labeled data to learn patterns and then generalize those patterns to the unlabeled data
- minimizes the difference in predictions between similar training examples
- approaches
- self-training
- co-training
- multi-view learning
Self-Supervised Learning
- algorithm learns from the data without explicit human annotations
Note
The distinction between unsupervised versus self-supervised learning can be blurry sometimes. Roughly:
- Unsupervised learning attempts to learn representations without labels by not using any targets of any sort during training, e.g. by using correlations in activity between units.
- Self-supervised learning attempts to learn representations without labels by using the data itself to generate targets, e.g. generating targets using the next word in a sentence
- Put another way, self-supervised learning looks a lot like supervised learning in code, but there is a big difference related to the following question: do you as a machine learning researcher have to actually ask someone to label the data or not.
- approaches
- Contrastive Learning
- Pretext-task Learning
Reinforcement Learning
Decision making
- Q-Learning: Reinforcement Learning#Q-Learning
- R Learning
- TD Learning
Upper Confidence Bound
see Reinforcement Learning#Upper Confidence Bound
Thompson Sampling
Reinforcement Learning#Thompson Sampling
Learning Settings
How is data received or labeled?
- Batch Learning
- Online Learning
- Active Learning
- Continual Learning
Generalization Strategies
How does knowledge move across tasks?
- Inductive Learning
- Transductive Learning
- Deductive Inference
- Transfer Learning
- Multi-Task Learning
- Few-Shot Learning
- Zero-Shot Learning
- Multi-Instance Learning
Modeling Frameworks
- Discriminative Learning
- Generative Learning
- Representation Learning
- Bayesian Learning
- Hebbian Learning
Model Families
Classical ML Models
- Linear models
- Trees
- SVM
- Naive Bayes
- k-NN
Ensemble Methods
- Bagging: Ensemble Learning#Bagging (Bootstrap Aggregating)
- Boosting: Ensemble Learning#Boosting
- Random Forest: Decision Tree & Random Forest#Random Forest
- Gradient Boosting
Deep Learning
Important
- deep learning = training large neural network
- deep learning is most powerful in supervised learning
- applications: Advertisement, Images vision, Audio to Text, Machine translation, Autonomous Driving
- Overview: Key Components of Deep Learning
- Artificial Neural Networks
- Convolutional Neural Networks (CNN)
- Recurrent Neural Networks (RNN)
- Transformer
- Autoencoders
- Generative Adversarial Networks (GANs)
- Variational Autoencoders (VAE)
Model Selection & Improving
- Cross-Validation
- Hyperparameter Tuning
- randomized search
- grid search
- Error analysis
- Error Metrics
- Bias-Variance Tradeoff
- Underfitting vs. Overfitting
- Data Leakage
ML Model into Production
see Machine Learning Systems Design