Semantic search is retrieval that matches the meaning of a query to the meaning of a document, using entities, context, and intent rather than exact word overlap. This article defines what semantic search is and then follows Google’s sequence: the Knowledge Graph in 2012, Hummingbird in 2013, RankBrain in 2015, neural matching in 2018, BERT in 2019, and MUM in 2021, followed by today’s ranking systems. It then sets out what each change meant for content in one table, and a worked example of semantic search on holisticradar.com. It closes with semantic search and query network research, and with semantic search compared to keyword search, passage ranking, and generative search.
What semantic search is
Semantic search is retrieval that matches the meaning of a query to the meaning of a document, so a page can be found for a query that shares none of its exact words.
Lexical search counts and weights shared words. It works when the searcher and the writer use the same terms, and it fails when they do not, when one word has several meanings, or when the small words in a query change what it asks. Semantic search addresses all three failures by representing queries and documents as meaning: the entities they mention, the relations between those entities, and the intent behind the request.
Google did not switch from lexical to semantic search in one release. It added systems in sequence, and each one understood something the previous ones could not. Keyword matching was never removed. It became one input among several, and a smaller one with each change.
| Problem | Lexical search | Semantic search |
|---|---|---|
| Different words, same meaning | Misses the match | Matches the concept |
| Same word, different meanings | Mixes the results | Resolves the entity |
| Small words that change intent | Ignores them | Reads them in context |
The Knowledge Graph and entities (2012)
The Knowledge Graph, introduced by Google on 16 May 2012, gave Search a database of entities and their relations, so a query could be interpreted as a request about things, not strings.
Google’s announcement, Introducing the Knowledge Graph, described more than 500 million objects and more than 3.5 billion facts and relationships at launch. Its first visible effect was the knowledge panel. Its deeper effect was disambiguation: a query for a name with several meanings could be mapped to distinct entities rather than to a single string.
The Knowledge Graph is the foundation every later system builds on. It supplies the entities that queries and documents are resolved to.
Hummingbird and query meaning (2013)
Hummingbird was a rewrite of Google’s core search algorithm, announced on 26 September 2013 after running for about a month, that interpreted the meaning of the whole query rather than its individual words.
Search Engine Land’s FAQ: the Google Hummingbird algorithm reported Google’s statement that the change affected around 90 percent of searches. The focus was conversational queries, which were becoming common as voice search grew. A query such as “what is the closest place to buy an iPhone to my home” contains words a lexical system weights poorly; Hummingbird let Google read it as a request for a store location near the searcher.
Google’s ranking systems guide now records Hummingbird among the systems noted for historical purposes, absorbed into successor systems. Its principle, whole-query meaning, remains in everything that followed.
RankBrain and unfamiliar queries (2015)
RankBrain is a machine learning system, disclosed by Bloomberg on 26 October 2015, that helps Google understand how words relate to concepts, especially in queries it has never seen before.
Bloomberg’s report, Google turning its lucrative web search over to AI machines, described RankBrain as the third most important signal contributing to a search result. Google stated at the time that about 15 percent of daily queries were new to it, and that RankBrain handled a very large fraction of those. By 2016 Google confirmed that RankBrain was involved in every query.
RankBrain was the first time Google publicly credited machine learning with a ranking role. It let the engine guess the concept behind an unfamiliar phrasing by relating it to familiar ones.
Neural matching and concept matching (2018)
Neural matching is an AI system, announced by Google in September 2018, that understands representations of concepts in queries and pages and matches them to one another.
Google said at launch that neural matching affected 30 percent of queries, and Google’s public search liaison described it as a kind of “super synonyms”. RankBrain relates words to concepts. Neural matching relates the concept of a query to the concept of a page, so a page can be retrieved for a query whose words it never uses.
The consequence for writers is direct: a page is matched to the concepts it covers, not only to the phrases it repeats. Repeating a phrase adds little once the concept is clear.
BERT and word context (2019)
BERT, short for Bidirectional Encoder Representations from Transformers, is an AI system Google applied to Search on 25 October 2019 to understand how combinations of words express different meanings and intent.
Google’s announcement, Understanding searches better than ever before, said BERT would affect one in ten searches in US English. BERT reads each word in relation to the words before and after it, which is what “bidirectional” means. Google’s example was the query “2019 brazil traveler to usa need a visa”: before BERT, the word “to” was underweighted and results described US citizens travelling to Brazil; with BERT, Google understood the traveller was going from Brazil to the United States.
Google expanded BERT to more than 70 languages in December 2019. BERT made prepositions, negations, and word order matter, which is why exact phrasing in a heading now affects what the heading can answer.
MUM and multimodal understanding (2021)
MUM, the Multitask Unified Model, is an AI system Google announced in May 2021 that can both understand and generate language, trained across 75 languages and able to process text and images.
Google described MUM as 1,000 times more powerful than BERT. Its first public application identified more than 800 names for COVID-19 vaccines across more than 50 languages, so that searches for any of them returned current vaccine information. Google’s ranking systems guide states that MUM is not used for general ranking and is applied to specific tasks instead.
MUM matters less as a ranking system than as a marker. It showed that Google’s language models could connect information across languages and formats, which is the capability later used in generative search features.
Google’s current ranking systems
Google’s current ranking systems are listed in A guide to Google Search ranking systems, a Google Search Central page that names each active system and separates retired ones.
The active list includes BERT, MUM, neural matching, RankBrain, and the passage ranking system alongside systems for freshness, link analysis and PageRank, original content, reliable information, reviews, deduplication, site diversity, and spam detection. The retired list records Hummingbird, the Panda and Penguin systems, and the helpful content system, which became part of the core ranking systems in March 2024.
The guide confirms that semantic understanding is not one system. Several language systems run together, each contributing a different kind of understanding, and they operate alongside quality, freshness, and spam systems that are not semantic at all.
What each semantic search change meant for content
Each semantic search change raised one specific requirement for content, and the requirements accumulate rather than replace each other.
| System | Year | What it understands | What content must do |
|---|---|---|---|
| Knowledge Graph | 2012 | Entities and their relations | Name entities in full and state their attributes |
| Hummingbird | 2013 | The meaning of the whole query | Answer the question asked, not the keywords in it |
| RankBrain | 2015 | How words relate to concepts | Cover the concept, including the phrasings searchers use |
| Neural matching | 2018 | Concepts in queries and pages | Make the page’s concept clear without phrase repetition |
| BERT | 2019 | Words in context, including small words | Write headings and answers whose exact wording carries the intent |
| MUM | 2021 | Information across languages and formats | Keep facts consistent across text, images, and versions |
| Passage ranking | 2021 | Individual sections of a page | Make each section answer its heading on its own |
Semantic search on holisticradar.com
Holistic Radar writes holisticradar.com for semantic search by stating each entity once and consistently, answering every heading in its first sentence, and naming each concept on the page that owns it.
One concept shows the full sequence. A searcher may ask “how long does a free SEO audit take”. The site’s free audit is a named method, the Radar Scan, and the words “Radar Scan” appear nowhere in that query.
- Entity. The Radar Scan is stated as a named method of one Organization entity, with fixed attributes: free, five business days, and a report that scores the Holistic Authority Score and names the one layer capping growth.
- Concept. The Radar Scan page describes it as a free audit in plain words, so the concept in the query and the concept on the page can match without the brand name.
- Context. The answer sentence reads “The Radar Scan takes five business days”, with the subject, the predicate, and the value in one sentence, so word order carries the meaning.
- Passage. The sentence sits directly under the heading it answers, so the section stands alone if it is retrieved as a passage.
- Consistency. Every page that mentions the Radar Scan gives the same duration, so no page contradicts the passage that is retrieved.
The same rules apply to every page. Distilled slugs such as /learn/semantic-seo/ name the concept rather than the question form, and one contextual vector per URL keeps each page’s concept unambiguous.
Semantic search and query network research
Semantic search models how queries relate in meaning, and query network research is the Holistic Radar service that maps those same relationships for one market before any page is planned.
Because Google matches by meaning, a site that plans pages from a keyword list plans against the wrong unit. Queries that share a concept belong together, queries that follow each other belong on linked pages, and query paths that end in an action define where commercial pages belong.
Query network research records those relationships as data at L2 of the Six-Layer Pyramid, query network and attributes, so the topical map above it reflects how the market searches rather than how a keyword tool groups phrases.
→ Query network research: the query relationships semantic search already models, mapped for your market
Semantic search vs keyword search, passage ranking, and generative search
Semantic search is most often compared with keyword search and most often extended by passage ranking and generative search.
How does semantic search differ from keyword search?
Semantic search differs from keyword search by matching meaning instead of shared words. Keyword search can only return a page that contains the query’s terms; semantic search can return a page that answers the query in different words and can exclude a page that contains the terms but means something else.
What is passage ranking?
Passage ranking is a Google system that identifies individual sections of a page to judge how relevant the page is to a query. Google announced it in October 2020, saying it would affect 7 percent of search queries across languages once fully rolled out, and it went live in US English in February 2021. It rewards sections that answer their heading on their own.
Is generative search a form of semantic search?
Generative search depends on semantic search: it retrieves passages by meaning, then generates an answer from them and cites some of its sources. Google’s AI Overviews launched in the United States in May 2024. A page that is not retrievable by meaning is not available to be summarized or cited.
Meaning-based retrieval is measured with the same instruments that plan the map. → Semantic analysis instruments