Imagine a map where every word occupies a position. Words with similar meanings are neighbours: cat right next to dog, doctor next to physician, and tractor at the other end of the map. Except this map doesn't have two dimensions—it has several hundred. An embedding is simply the coordinates of a word or piece of text on that map. And once you have coordinates, you can measure distances.
This is one of the most useful concepts to grasp for understanding what AIs actually do, and it's rarely explained clearly. It's also at the heart of nearly every document assistant deployed in the enterprise. Let's break it down.
The simple definition
An embedding (sometimes rendered as plongement in French) is the transformation of a text into a list of numbers—usually several hundred—that encodes its meaning. Each number is a coordinate in a mathematical space, and the overall position summarises what the text is saying.
The remarkable property is that this position isn't arbitrary: it's learned in such a way that texts with similar meanings end up with similar coordinates. The word doctor and the word physician, which share no letters, end up as neighbours. The sentence "my computer won't boot" and the sentence "can't turn on my PC" do too.
This changes everything, because you move from comparing characters to comparing meaning. A classic search engine looks for identical words. An embedding-based engine looks for nearby meanings.
How it's actually used
The dominant use case is called similarity search, and it works in three steps.
First, you split all your documents into chunks and compute the embedding of each one. You store these coordinates in a specialised database, called a vector database. Then, when a user asks a question, you compute the embedding of their question. Finally, you search the database for chunks whose coordinates are closest to those of the question, and you return them.
This is exactly the mechanism powering RAG, the technique that lets an AI answer based on your documents rather than on its training memory. The embeddings find the relevant passages, and the language model drafts the answer from those passages.
Take a concrete case. An employee is looking for the expense reimbursement procedure and types "how do I get reimbursed for my train". The internal document is titled "Expense claims: transport and business mobility". No shared words, apart from a preposition. A classic search finds nothing. An embedding-based search surfaces the document immediately, because the two texts occupy nearby positions on the meaning map. It's this difference that makes document assistants genuinely usable.
The limitations, which you should know about
Proximity isn't relevance. Two texts can be semantically close, yet only one of them answers the question. An embedding-based search returns what resembles, not necessarily what answers. That's why serious systems combine embeddings with other methods, including good old keyword search, which remains unbeatable for finding an exact reference like a contract number.
Chunking is decisive. If you cut your documents mid-argument, each chunk loses its meaning and its embedding becomes misleading. The quality of such a system often depends more on how the documents were chunked than on the model used.
Embeddings inherit biases. The meaning map is learned from human texts, with their implicit associations. The relative positions of words therefore also reflect the corpus's prejudices, which ties into the issue we explored in our article on AI neutrality.
What to take away
An embedding turns meaning into geometry. It's a rare elegant idea: by assigning each text a position in a space, you make meaning measurable, and therefore computable by a machine.
If you hear about vector databases, semantic search, or cosine similarity, you'll now know it all rests on the same intuition: coordinates, and the distances between them. It's also the concept that explains why an AI can find you the right document without you having found the right word. It isn't searching for your word—it's searching for your meaning.