Cost Functions
Cost functions
- A cost function determines the "cost" (or penalty) of estimating
when the true or correct quantity is really . - This is essentially the cost of the error between the true stimulus value
and our estimate . - cost function formula
where
Cost vs. Loss:
loss applies to a single training sample; cost is the mean of summed loss.
Forms of cost functions
Note that the error can be defined in different ways:
- Find more types of error in Error Metrics.
- In ML, Mean Squared Error is commonly used as the cost function, but with an extra division by 2, which "is just meant to make later partial derivation in gradient descent neater" :
Cost function with regularization
-
When you choose Regularization, a regularization term will be added to the cost function, in order to add penalty and avoid overfitting.
-
Different types of regularization terms can be added, so
Loss = Original Loss + regularization term- L1 regularization: regularization term =
- L2 regularization: regularization term =
- Elastic Net$$
ElasticNet \ Penalty=λ_1 \sum_{j=1}^p∣ w_j∣+ \lambda_2 \sum_{j=1}^p w_j^2
- L1 regularization: regularization term =
Loss and cost for different functions
Loss and cost for linear regression -> Analytic solution
in matrix form, with
The solution will only be unique when the matrix
Loss and cost for logistic regression
MSE is not proper because the cost function would not be convex.
loss function
Combined the above formula together, we get the simplified cost function for logistic regression:
then the cost function with full form (also used in Maximum likelihood estimation for logistic regression):
Loss and cost for Softmax
What is Softmax: Artificial Neural Networks#^6ef895
The loss function associated with Softmax, the cross-entropy loss, is:
Only the line that corresponds to the target contributes to the loss, other lines are zero:
$$\mathbf{1}{y == n} = =\begin{cases}
1, & \text{if
0, & \text{otherwise}.
\end{cases}$$
Cost function:
where
Cross-entropy takes the full distribution into account.
Expected loss function
A posterior distribution tells us about the confidence or credibility we assign to different choices. A cost function describes the penalty we incur when choosing an incorrect option. These concepts can be combined into an expected loss function.
Expected loss is defined as:
where
- The posterior's mean minimizes the mean-squared error.
- The posterior's median minimizes the absolute error.
- The posterior's mode minimizes the zero-one loss.
Good Practice in minimizing loss function
- be aware of Clever Hans Effect: model learns to minimize loss functions, but it does not learn the thing it should learn...