The Anatomy of a Search Engine: Search, Discovery, and Marketing
Hatched by Kazuki Nakayashiki
Sep 01, 2023
4 min read
5 views
The Anatomy of a Search Engine: Search, Discovery, and Marketing
In 1994, the World Wide Web Worm (WWWW) emerged as one of the first web search engines. At that time, it had an index of 110,000 web pages and web accessible documents. Fast forward to November 1997, and the top search engines claimed to index anywhere from 2 million to 100 million web documents. The growth of search engines has been exponential, with Altavista claiming to handle roughly 20 million queries per day.
The main goal of search engines, such as Google, is to improve the quality of search results. Google was designed to create an environment where researchers can process large amounts of web data and produce interesting and valuable results. However, one resource that has been largely overlooked in traditional search engines is the citation (link) graph of the web. By analyzing the links between web pages, we can gain insights into a page's importance or quality.
PageRank, a model of user behavior, takes into account the citation graph when determining a page's importance. It does not count all links equally, but rather normalizes them by the number of links on a page. The probability that a random surfer visits a page is its PageRank. The damping factor, represented by 'd', is the probability that the surfer will get bored and request another random page. This factor allows for personalization and prevents deliberate manipulation of search rankings.
The predominant business model for commercial search engines is advertising. However, the goals of the advertising business model do not always align with providing quality search results to users. This raises the question of whether search engines should focus on giving users what they are looking for or on suggesting what they might like.
The trade-off between these two problems becomes apparent when looking at the scalability of search engines. Yahoo's directory worked well when there were only 20,000 websites, just as a bookshop with 20,000 titles works well. But as the number of websites increased to 3.2 million, the directory became unusable. Hierarchical directories simply do not scale.
Google excels at giving users what they are looking for, but falls short when it comes to suggesting what users might like or finding things they didn't know they wanted. This is where recommendation and discovery platforms come into play. Amazon, for example, has spent 20 years building a searchable index of everything, but still only has a fraction of the print books market. Physical bookshops act as filters and recommendation platforms, offering a curated experience that online retailers struggle to replicate.
The challenge lies in finding the right balance between curated content and broad coverage. Should crowdsourcing be filtered to ensure quality, or should editorial efforts be scaled up to provide comprehensive coverage? Alternatively, is it possible to create a purely curated product that sacrifices coverage for quality?
In the realm of search engines, there are three distinct categories: giving users what they already know they want (Amazon, Google), working out what users want (Amazon and Google's aspiration), and suggesting what users might like (e.g. Heywood Hill). Each category serves a different purpose and addresses different user needs.
It is interesting to note that despite the failures of Yahoo, many people continue to try and rebuild its model. However, the statement "We've made it to where we are with absolutely no sales or marketing spend" does not always signal success. The existence of search engine marketing (SEM) proves that if PageRank were the ultimate solution, there would be no need for search advertising.
In conclusion, search engines have come a long way since their inception. The focus has shifted from simply indexing web pages to providing high-quality search results. However, the challenges of discovery, recommendation, and marketing still persist. To navigate these challenges, here are three actionable pieces of advice:
-
Embrace the power of the citation graph: Incorporate link analysis and PageRank algorithms to gain insights into the importance and quality of web pages. This will enhance search results and improve the user experience.
-
Explore the balance between curation and coverage: Consider how to provide users with both a curated experience and comprehensive coverage. Find innovative ways to filter crowdsourcing or scale up editorial efforts to strike the right balance.
-
Foster collaboration between search engines and recommendation platforms: Search engines can learn from the success of recommendation platforms like Amazon in terms of personalized suggestions. By incorporating recommendation algorithms, search engines can bridge the gap between giving users what they want and suggesting what they might like.
By addressing these challenges and implementing these recommendations, search engines can continue to evolve and provide users with an even more personalized and comprehensive search experience.
Sources
Hatch New Ideas with Glasp AI 🐣
Glasp AI allows you to hatch new ideas based on your curated content. Let's curate and create with Glasp AI :)
Start Hatching 🐣