How Do Big Data and Machine Learning Work?

1.1K views
•
January 2, 2019
by
a16z
YouTube video player
How Do Big Data and Machine Learning Work?

TL;DR

Big data becomes valuable when machine learning extracts patterns from enough information to predict unknown outcomes from known ones. The amount required depends on the task, while affordable parallel storage systems such as Hadoop made large-scale data collection practical for more companies. This shifts business intelligence from reviewing past results toward building models that can project into the future.

Transcript

Hi, everyone. Welcome to the A6&Z podcast. This is Sonal, and I'm here today with Christopher Nguyen from Adatao, which is a big data company, and its mission is to democratize data intelligence and help people collaborate across the enterprise. And the best way to describe him is as an entrepreneurial scientist. He got his PhD from Stanford in dev... Read More

Key Insights

  • Big data is defined more usefully by whether there is enough information to learn from than by an absolute quantity. A few samples may be sufficient for a simple lesson, while millions may still be insufficient for a complex classification task.
  • The purpose of big data is to give machine learning algorithms enough material to detect patterns automatically. Volume, variety, velocity, veracity, and similar characteristics describe the difficulties of handling data, but they do not explain why organizations accept those difficulties.
  • Machine learning is analogous to human learning from accumulated life experience. Human cognitive capacity may remain broadly stable after a certain age, yet wisdom can continue growing as the brain incorporates more experiences and recognizes recurring patterns or unusual situations.
  • Rule-based computing is limited because a finite collection of rules cannot anticipate every exception. Machine learning addresses this limitation by learning from examples, including corner cases, and developing model parameters that function like experience-based intuition about when ordinary rules should not apply.
  • Traditional business intelligence is primarily backward-looking because it aggregates completed transactions. It can answer questions about past revenue or regional performance, but its historical capabilities were constrained by the data and computing technologies available at the time.
  • Modern business intelligence can use machine learning models to move from describing known history toward predicting unknown outcomes. The shift depends on collecting enough relevant experience and applying algorithms that can turn patterns in that experience into forward-looking projections.
  • Big data existed before organizations routinely collected it. The important change was not the sudden invention of data or machine learning algorithms, but the falling cost and increasing availability of technologies needed to acquire, store, and process large datasets.
  • Hadoop is primarily a highly parallel storage layer built around the Hadoop Distributed File System. Its replication and resilience allow data to be stored reliably across commodity hardware, making terabyte-scale storage affordable for a much broader range of companies.

Explore YouTube Video Summarizer or Get YouTube Transcript Extractor

Questions & Answers

Q: What is the most useful definition of big data?

Big data is information collected at a scale sufficient for a machine learning system to learn useful patterns. The decisive issue is not a universal number of records, because the required amount depends on the task. A simple lesson may require only a few examples, while classifying internet images may require more than two million and still need additional data.

Q: Why is machine learning the purpose of big data?

Machine learning provides the benefit that justifies accepting the storage, processing, variety, velocity, and reliability problems associated with large datasets. Algorithms can examine accumulated information, identify patterns automatically, and construct models from those patterns. Without a way to learn from the collected information, the familiar challenges represented by the various V terms describe costs without explaining the larger reason to incur them.

Q: How is machine learning similar to human experience?

Machine learning resembles the way people develop judgment by accumulating life experiences. A person does not gain wisdom solely by memorizing explicit rules. Repeated examples and unusual cases help that person recognize when a rule works and when it does not. Similarly, a learning algorithm incorporates data into model parameters that support pattern recognition and decisions in situations not covered by simple instructions.

Q: Why can rule-based systems struggle with exceptions?

Rule-based systems depend on people specifying instructions in advance, but ten, twenty, or thirty rules cannot reliably cover every possible exception. Machine learning can use previously observed corner cases to develop a more flexible response. What appears to be intuition in a person can be understood, by analogy, as learned parameters that help a model recognize when an ordinary rule should not apply.

Q: How does machine learning change business intelligence?

Machine learning expands business intelligence beyond aggregating historical transactions and reporting what already happened. Traditional systems might calculate how much revenue a particular region generated yesterday. When an organization has enough relevant data, it can build a model that projects forward. Business intelligence can then address future-oriented questions by predicting unknown outcomes from patterns found in known historical information.

Q: What made large-scale data collection affordable?

Large-scale data became more accessible because the technologies required to collect and store it became cheaper and more widely available. Machine learning algorithms and potentially useful data had already existed, but organizations often did not collect the information. The Hadoop project, followed by companies such as Cloudera and MapR in 2009, helped more businesses afford the acquisition and storage of large datasets.

Q: What does Hadoop do in a big data system?

Hadoop primarily provides a storage layer through the Hadoop Distributed File System. Its architecture is highly parallel and includes storage replication, which supports resilience when information is distributed across multiple machines. It can also operate on commodity hardware. These properties allowed organizations to store terabytes reliably while paying less than they would have under earlier approaches to large-scale storage.

Q: How much data is enough for machine learning?

Enough data means enough examples for the specific system to learn the pattern required by its task. There is no single threshold that applies everywhere. Learning that hitting a head against a brick wall hurts might require about five samples, while classifying images on the internet might not be adequately supported by two million. Big data is therefore qualitatively task-dependent, not merely quantitatively large.

Summary & Key Takeaways

  • Big data is better understood by its purpose than by familiar descriptions involving volume, variety, velocity, and related challenges. Its value comes from supplying machine learning algorithms with enough examples to detect useful patterns. The threshold for having enough data varies substantially according to the learning task being attempted.

  • Machine learning resembles human learning from life experience. Rules provide a foundation, but accumulated examples reveal exceptions and corner cases, producing something analogous to intuition or wisdom. Within a machine learning model, learned parameters can represent this experience-based judgment and help determine when a general rule does not apply.

  • Traditional business intelligence primarily aggregates historical transactions and answers backward-looking questions, such as how much revenue a region generated yesterday. With enough collected data, organizations can build models that project forward and predict unknowns from knowns. Hadoop helped enable this transition by making resilient, parallel data storage affordable on commodity hardware.


Read in Other Languages (beta)

Share This Summary 📚

Explore More Summaries from a16z 📚