A Few Things I Believe About AI: What Makes Data Valuable?
Hatched by Glasp
Sep 07, 2023
6 min read
13 views
A Few Things I Believe About AI: What Makes Data Valuable?
In the world of artificial intelligence (AI), there are two crucial components to intelligence: reasoning and knowledge. While AI models like GPT-4 have made great strides in reasoning capabilities, their knowledge of the world is still limited. This poses a significant challenge for AI builders, as their models' performance is reliant on the right knowledge being provided at the right time. This challenge, known as knowledge orchestration, is one of the biggest unsolved problems in AI. It involves how we store, index, and retrieve the knowledge necessary for large language model tasks.
Fortunately, progress is being made to address this issue. One way developers are tackling knowledge orchestration is by building models with larger context window sizes. The larger the context window, the more knowledge can be included in the prompt, leading to better performance. GPT-4, for example, boasts a 32,000 token context window, which is a significant improvement from previous models.
Additionally, there are companies working on developer tools and infrastructure to facilitate knowledge storage and retrieval. LlamaIndex and Langchain are two examples of platforms that make it easier for developers to chunk, store, and retrieve knowledge from various databases with minimal code. Vector database providers like Pinecone, Weaviate, and Chroma are also competing in this space, with Pinecone currently leading the way.
While knowledge orchestration focuses on the process surrounding knowledge, it's equally important to consider the type of knowledge that is most valuable. One type that excites me is end-to-end interaction data about the lifecycle of a process. This data captures the entire process from beginning to end, including iterations and editing steps. Having access to this data allows models to be steered through techniques like reinforcement learning and fine-tuning, ultimately improving the process over time.
Startups that are horizontally integrated over a process have a significant advantage in this AI-first world. By replacing elements of a process that customers would typically pay other companies for, startups can gain more visibility and control over the entire process. Replit, for example, provides a developer platform that enables customers to write and deploy apps all from a browser window. By integrating over the process of turning ideas into software, Replit has positioned itself as a dominant player in its field.
Integration is key when it comes to improving AI models. OpenAI, for instance, initially built ChatGPT as an API-only product but realized the importance of integrating forward to access end-user data and gather human feedback. Integration through earlier parts of the value chain can also be beneficial. By having access to the editing process that led to the final output of a process, models can learn to recreate or improve that process more effectively.
However, there are challenges to consider when it comes to data sharing and integration. Incumbent companies may face privacy concerns and internal resistance when sharing data between apps in their ecosystems. Startups that prioritize owning an entire end-to-end process and centrally storing the data from that process will have a competitive advantage in this regard.
Moving on to the topic of data network effects, it's essential to understand what makes data valuable. A real data network effect, which is a highly valuable type of defensibility, is much rarer than most people realize. Waze, a popular navigation app, serves as an excellent example of a consumer product with a real data network effect. There are six key elements that contribute to its success:
- Automatic data capture from customer usage.
- Automatic increase in product value as more data is added.
- A high minimum threshold for the amount of data needed before the product provides value, creating a scale defensibility against competitors.
- The value of incremental data doesn't reach a saturation point quickly due to the real-time nature of the service, making it difficult for competitors to replicate the same value without a large user base.
- The customer perceives the value created by the data and uses it to make decisions.
- Constantly changing data ensures the continuous need for new data, providing defensibility against competitors.
Reducing cycle time and marginal cost of data are crucial strategies to maximize the value of data. By reducing cycle time, startups can outpace potential competitors and stay ahead of the curve. Avoiding manual data collection and instead automating the process helps reduce the operational complexities and costs associated with data collection.
However, capturing data alone is not enough. The data must result in improved value for existing users to complete the positive feedback loop that drives a data network effect. While product improvements are typically a manual and periodic process, they play a vital role in enhancing the overall user experience. As usage increases, finding ways to improve the product becomes more challenging, but it's necessary to sustain the network effect.
It's worth noting that many founders overestimate the uniqueness and value of their datasets. In some cases, a smaller amount of data can provide the majority of the product experience value. Real-time data network effects, like the one observed in Waze, require a constant stream of data, making larger networks of users more advantageous.
Furthermore, data scale can offer some defensibility for startups. Yelp, for example, benefits from its extensive coverage of restaurants and local businesses, which provides valuable insights to users. However, the value derived from a scaled-up dataset tends to diminish over time, and differentiation based solely on data quantity is becoming increasingly challenging in a world of abundant data.
In conclusion, AI and data play integral roles in shaping our present and future. Knowledge orchestration remains a significant challenge in the field of AI, as models require the right knowledge to reason effectively. Startups that integrate over an entire process and own the data associated with it have a competitive advantage. Real data network effects are rare but highly valuable, as seen in the case of Waze. Understanding the factors that contribute to data value and leveraging data scale can provide defensibility for startups. However, it's important to remember that data alone is not enough; it's the improvements and value derived from the data that truly matter.
Actionable Advice:
-
Prioritize knowledge orchestration by exploring tools and infrastructure that facilitate storing, indexing, and retrieving knowledge for AI models. Look for platforms that simplify the process and enable seamless integration with different databases.
-
Aim for horizontal integration over a process to gain visibility and control. Identify elements of a process that customers pay other companies for and develop your solution to replace those elements. This will allow you to have a comprehensive view of the process and improve it over time.
-
Focus on data quality and the value it brings to users. Avoid overestimating the uniqueness and value of your dataset. Instead, strive to provide incremental value to users based on their data usage. Constantly seek ways to improve the user experience through periodic enhancements and automation.
Sources
Hatch New Ideas with Glasp AI 🐣
Glasp AI allows you to hatch new ideas based on your curated content. Let's curate and create with Glasp AI :)
Start Hatching 🐣