#11 Why Is High Test Set Accuracy Not Enough for Production? Machine Learning Engineering for Production (MLOps) Specialization [Course 1, Week 2, Lesson 3]

TL;DR
High test set accuracy is not enough for production because averages can conceal failures on critical queries, protected groups, business categories, and rare classes. A disease classifier can reach 99 percent accuracy by always predicting that the disease is absent when only 1 percent of examples are positive. Read on to see why deployment decisions require analysis beyond a single average metric.
Transcript
the job of a machine learning engineer would be much simpler if the only thing we ever had to do was do well on the holdout test set as hard as it is to do well in the holdout test set unfortunately sometimes that isn't enough let's take a look at some of the other things we sometimes need to accomplish in order to make a project successful we've a... Read More
Key Insights
- Deployment has broader requirements: Doing well on the holdout test set is difficult, but it is only one part of making a production machine learning project successful. A model can achieve low average error and still fail requirements that matter to users, affected groups, retailers, or the application itself. Production acceptability depends on where errors occur, not only how many occur overall.
- Average metrics flatten importance: Average test set accuracy generally gives equal weight to every example. Real applications may not value every example equally, so this calculation can hide severe mistakes on a small but important subset. An improvement in the overall average can therefore accompany a decline on cases whose failure makes the system unacceptable for deployment.
- Search intent changes tolerance: Users searching for apple pie recipes, latest movies, wireless data plans, or information about the Diwali festival may accept several relevant answers. They may forgive the best result appearing second or third. This flexibility means errors on informational or transactional queries can have a different practical effect from errors on queries with one clearly intended destination.
- Navigation demands the top result: Searches for Stanford, Reddit, or YouTube express a clear navigational intent. Users expect the corresponding destination, such as stanford.edu, reddit.com, or youtube.com, to be ranked first. A search engine that fails this expectation can quickly lose user trust, making even a small group of navigational queries disproportionately important during evaluation.
- Higher weighting has limits: Giving disproportionately important examples a higher weight is one possible response to unequal business importance. The lesson notes that this can work for some applications. However, changing example weights does not always solve the entire problem, so teams still need to inspect actual behavior on important examples and data slices.
- Fairness can block deployment: A loan approval algorithm may have strong average test accuracy while showing an unacceptable level of bias or discrimination. In that situation, the average result does not justify production use. The model must be checked for its treatment of applicants across ethnicity, gender, location, language, and other protected attributes relevant to the system.
- Legal requirements affect evaluation: Fairness in loan approval is not presented only as an ethical preference. Many countries have laws or regulations requiring financial systems and loan approval processes not to discriminate using certain protected attributes. This makes performance on protected slices a concrete deployment requirement that aggregate test accuracy cannot replace.
- Recommendation quality can be uneven: An e-commerce system might recommend better products on average while consistently giving irrelevant results to all users of one ethnicity. The overall gain would not make that behavior acceptable. Slice analysis reveals whether apparently strong recommendation quality is shared across major user categories or concentrated in ways that conceal harm.
- Retailer exposure affects the platform: A recommender that continually pushes products from large retailers while ignoring smaller brands may drive small retailers away. Even when average relevance remains high, that outcome can be harmful to the business and feel unfair. Evaluation should therefore include how recommendations are distributed across major retailer categories, not just user-level relevance.
- Product categories require scrutiny: A recommender could offer highly relevant suggestions yet never recommend electronics products. Retailers selling electronics would reasonably be upset, and the pattern might conflict with the platform's long-term health. A slight improvement in average relevance does not automatically outweigh the exclusion of an entire major product category.
- Slice analysis exposes hidden problems: Examining key slices of a dataset helps teams spot patterns that disappear inside an average score. Relevant slices may represent protected groups, major user categories, retailer sizes, or product categories. The purpose is to determine whether the system performs acceptably for each important segment before deployment, rather than relying on one aggregate result.
- Class imbalance can reward uselessness: With 99 percent negative disease examples and 1 percent positive examples, always predicting zero produces 99 percent accuracy. That result requires only a one-line program and no learning algorithm. Its strong score is misleading because it does not provide useful disease diagnosis, demonstrating why rare positive classes need separate attention.
Install to Summarize YouTube Videos and Get Transcripts
Explore YouTube Video Summarizer or Get YouTube Transcript Extractor
Questions & Answers
Q: Why is high test set accuracy not enough for production machine learning deployment?
High test set accuracy averages performance across examples, even though some examples matter much more than others. It can conceal failures on navigational searches, protected groups, retailer or product categories, and rare positive classes. Those failures can damage user trust, create unacceptable bias, harm marketplace participants, or make a diagnostic system useless. Production evaluation must therefore inspect important examples and key slices instead of treating the average score as the complete deployment criterion.
Q: What are disproportionately important examples in machine learning?
Disproportionately important examples are cases whose correct handling matters more to the application than their share of the test set suggests. Navigational web searches such as Stanford, Reddit, or YouTube are examples because users have one clear destination in mind. If the correct site is not ranked first, users may quickly lose trust even when only a small number of queries are affected. These examples require direct attention because average accuracy weights them like ordinary queries and can hide their practical importance.
Q: How do navigational queries differ from informational or transactional queries?
Informational or transactional queries seek knowledge or an action, as with apple pie recipes, latest movies, a wireless data plan, or the Diwali festival. Several results may be useful, so users might accept the best result in the second or third position. Navigational queries such as Stanford, Reddit, or YouTube instead express a clear desire to reach a particular site. The search engine must rank that destination first because users are much less forgiving when their intended site is not immediately returned.
Q: Can weighting important examples solve the average-accuracy problem?
Giving important examples a higher weight is one possible approach. The lesson says this could work for some applications because it makes those examples contribute more strongly to evaluation or learning. However, merely changing the weights does not always solve the entire problem. Teams still need to examine performance on the important examples and relevant data slices to determine whether the system's actual behavior is acceptable.
Q: Why must loan approval models be evaluated on protected attributes?
A loan approval model can achieve high average test accuracy while treating some applicants unfairly. Evaluation should examine performance across ethnicity, gender, location, language, and other protected attributes because aggregate accuracy can conceal discrimination within those groups. Many countries also have laws or regulations requiring financial systems and loan approval processes not to discriminate based on specified attributes. A model with an unacceptable level of bias is therefore not suitable for production merely because its overall score is high.
Q: How can an accurate recommendation system still harm an e-commerce business?
A recommender may improve average relevance while consistently favoring large retailers and ignoring smaller brands. That behavior can cause the platform to lose small retailers and can feel unfair to the businesses participating in the marketplace. The system might also exclude a major category, such as electronics, even while producing relevant recommendations elsewhere. Such patterns can harm retailers and the platform's long-term health, so performance must be checked across major retailer and product categories.
Q: What is key-slice analysis in production machine learning?
Key-slice analysis examines model performance separately for important segments of the dataset. In the examples given, those slices include protected applicant groups, major user categories, retailer groups, and product categories. This analysis shows whether a strong average is masking irrelevant recommendations, discrimination, or the exclusion of particular businesses and products. It is useful because production success requires acceptable behavior within important slices, not only a favorable aggregate metric.
Q: Why is accuracy misleading for rare classes and skewed data distributions?
Accuracy can be dominated by the common class when one outcome is much more frequent than another. In the medical example, 99 percent of the population does not have the disease and only 1 percent does. A one-line program that always prints zero therefore reaches 99 percent test accuracy without using a learning algorithm. The score looks strong, but the program is not useful for disease diagnosis because the rare positive class is precisely the case the system needs to identify.
Summary & Key Takeaways
-
Looking beyond test accuracy: A machine learning engineer cannot judge production readiness solely by performance on a holdout test set. Even low average test set error may be insufficient when a system performs poorly on examples that matter disproportionately to users or the application. Concept drift and data drift were discussed earlier, but production projects also face challenges involving important examples, key data slices, bias, discrimination, rare classes, and skewed distributions. A successful system must satisfy the actual needs of its application.
-
Prioritizing navigational search queries: Informational or transactional searches include apple pie recipes, latest movies, wireless data plan, and learning about the Diwali festival. Users may forgive a search engine for placing a useful result second or third because several results could satisfy them. Navigational searches such as Stanford, Reddit, or YouTube are different. The user has a clear destination, so failing to rank stanford.edu, reddit.com, or youtube.com first can quickly damage trust, even if average search accuracy improves.
-
Evaluating protected data slices: A loan approval model may predict who is likely to repay a loan, but high average test accuracy does not make it acceptable if it unfairly discriminates. Performance should be examined across attributes such as ethnicity, gender, location, and language. Some countries have laws or regulations requiring financial systems and loan approval processes not to discriminate on specified protected attributes. Production evaluation therefore needs to identify unacceptable bias within key slices rather than allowing strong overall results to conceal it.
-
Protecting marketplace participants: An e-commerce recommendation system should treat major user, retailer, and product categories fairly. A model may have high average accuracy while giving irrelevant recommendations to users of one ethnicity, continually favoring large retailers, ignoring smaller brands, or never recommending electronics. These patterns can harm affected businesses, upset retailers, and damage the platform's long-term health. Analysis of key data slices is needed to expose such failures when a single aggregate score makes the recommendations appear successful.
-
Detecting skewed-class failures: Medical diagnosis illustrates why accuracy can become misleading when classes are rare. If 99 percent of examples are negative because 99 percent of the population does not have a disease, while 1 percent are positive, a program that always prints zero obtains 99 percent test accuracy. No learning algorithm is required for that result, yet the program is clearly not useful for diagnosing the disease. Rare classes and skewed data distributions therefore demand analysis beyond the headline accuracy figure.
Read in Other Languages (beta)
Share This Summary 📚
Summarize YouTube Videos and Get Video Transcripts with 1-Click
Try YouTube Summary with ChatGPT & Claude or YouTube Transcript Generator
Explore More Summaries from DeepLearningAI 📚






Summarize YouTube Videos and Get Video Transcripts with 1-Click
Try YouTube Summary with ChatGPT & Claude or YouTube Transcript Generator