Normalizing Inputs (C2W1L09)

TL;DR
Normalizing neural network inputs can speed up training by making the cost function easier for gradient descent to optimize. The process has two steps: subtract the training-set mean and scale the feature variances, ideally to variance 1. The same training-set values must also transform the test set. Read on to see why feature ranges such as 1–1000 versus 0–1 can otherwise slow optimization.
Transcript
when training a neural network one of the techniques to speed up your training is if you normalize your inputs let's see what that means let's see the training sets with two input features so the input features X are two-dimensional and here's a scatterplot of your training set normalizing your inputs corresponds to two steps the first is to subtra... Read More
Key Insights
- 🔠 Normalizing input features involves subtracting the mean and normalizing the variance of each feature.
- 😫 Using the same normalization values for both training and test sets ensures consistency.
- 🚱 Unnormalized input features can result in a non-symmetric cost function and optimization difficulties.
- 🛝 Normalizing input features helps achieve a more round and easier-to-optimize cost function.
- 💨 Similar scales for input features aid faster learning and optimization.
- 🧡 Normalization is particularly important when input features have dramatically different ranges.
- 🧡 Normalizing input features is a good practice that usually improves training speed, even when ranges are similar.
Install to Summarize YouTube Videos and Get Transcripts
Explore YouTube Video Summarizer or Get YouTube Transcript Extractor
Questions & Answers
Q: How do you normalize inputs when training a neural network?
First, calculate the mean vector μ across the training examples and subtract it from every example so the features have zero mean. Then calculate the feature variances and scale each feature so that features such as x1 and x2 have variance 1.
Q: Why does input normalization speed up neural network training?
Normalization puts the input features on similar scales, which tends to make the cost function more symmetric and easier to optimize. Gradient descent can then take larger, more direct steps toward the minimum instead of repeatedly oscillating across an elongated cost function.
Q: What values should be used to normalize the test set?
Use the same μ and variance values calculated from the training set to normalize the test set. Do not estimate them separately on the test set, because training and test examples should undergo the same transformation.
Q: What happens when neural network inputs are not normalized?
Unnormalized features can produce a very elongated or squished cost function. Gradient descent may then require a small learning rate and many back-and-forth steps before reaching the minimum.
Q: When is input normalization especially important?
Normalization is especially important when features have dramatically different ranges. The transcript gives the example of x1 ranging from 1 to 1000 while x2 ranges from 0 to 1, a mismatch that can hurt the optimization algorithm.
Q: Is normalization necessary when input features already have similar ranges?
It is less important when the ranges are already similar, such as 0 to 1, −1 to 1, and 1 to 2. Even then, the transcript says this type of normalization pretty much never does harm and will usually help the learning algorithm run faster.
Q: How does normalization change the shape of the cost function?
With features on different scales, the cost function can have elongated contours. Normalized features tend to produce more spherical contours, allowing gradient descent to move more directly toward the minimum.
Q: What scale should normalized input features have?
The features should have zero mean and similar variances, such as variance 1. The transcript also describes inputs as being mostly around −1 to 1 or otherwise having comparable variances, which makes the cost function easier and faster to optimize.
Summary & Key Takeaways
-
Normalizing inputs in neural network training involves subtracting the mean and normalizing the variances of the input features.
-
Normalizing input features ensures that the cost function is more symmetric and easier to optimize.
-
Using normalized input features allows for faster training and better convergence of the gradient descent algorithm.
Read in Other Languages (beta)
Share This Summary 📚
Summarize YouTube Videos and Get Video Transcripts with 1-Click
Try YouTube Summary with ChatGPT & Claude or YouTube Transcript Generator
Explore More Summaries from DeepLearningAI 📚






Summarize YouTube Videos and Get Video Transcripts with 1-Click
Try YouTube Summary with ChatGPT & Claude or YouTube Transcript Generator