Skip to content
The AI dictionary

Every word of AI,
finally clear.

Every term explained in two sentences plus an everyday analogy. In the articles, tap or hover the underlined words to see their definition.

Agent
An AI that doesn't just answer: it acts, navigates, clicks, codes, chains steps with tools to complete a full mission.
Agentjacking
An attack that hijacks someone else's AI agent by slipping fake instructions into the content it consults (web page, email, file, memory). The agent takes them for legitimate orders and carries them out with its user's permissions: sending data, running a command, even making a payment. It is prompt injection applied to an AI that takes action, with far more concrete consequences.
AGI
Artificial "general" intelligence: an AI at least as competent as a human on most intellectual tasks. It does not exist (yet), and its very definition is debated.
AI Act
The European regulation on artificial intelligence, the first major law in the world to govern AI, has been applied gradually since 2025. It classifies systems by risk level and imposes stricter obligations on the most powerful models.
Alibaba (Qwen)
Chinese e-commerce and cloud giant, whose AI lab develops Qwen, one of the most-downloaded families of open models in the world. Its computing power and talent pools make it a cornerstone of the Sino-American rivalry in AI.
Alignment
The field of research aimed at getting AIs to do what we really want—useful, honest, without harm, even as they become highly capable.
Anduril
US defence technology company co-founded in 2017 by Palmer Luckey, which designs AI-driven autonomous weapons systems, such as the Bolt drones or the Roadrunner-M interceptor. Valued at $61 billion in 2026, it embodies the new "military-AI complex" bringing Silicon Valley closer to the Pentagon.
Data annotation
The work of labelling raw data (describing images, rating responses, flagging toxic content) to teach models what is what. It is massively outsourced to low-paid workers, often in Kenya or the Philippines: the invisible workforce of AI.
Anonymisation
Replace names with "Client A", actual amounts with orders of magnitude, and addresses with cities before pasting anything into an AI. In nine cases out of ten, the request keeps the same meaning and the legal issue disappears.
Anthropic
The American lab behind the Claude models, founded in 2021 by former OpenAI members, has risen to become a global AI leader by focusing on safety and reliability, with particular success among developers.
API
The "electrical socket" of a service: it lets your app call an AI model remotely and pay per use (by tokens).
Apple
The creator of the iPhone, famous for rarely being first but often arriving at the right moment. In AI, the brand owns its lag: Apple Intelligence is advancing in small steps, and its smart glasses are not expected before late 2027, a wait-and-see strategy that contrasts with the frenzy of its competitors.
Reinforcement learning
A way of learning without examples to copy: the system tries, receives a reward signal when it succeeds (or a penalty when it does not), and gradually adjusts its strategy. This is how AlphaGo beat the Go champions, and it is a key step in shaping the behaviour of today's assistants, notably through RLHF.
Autonomous weapons
Weapon systems capable of seeking, selecting, and striking a target with little or no human intervention, such as loitering munitions that patrol above an area before diving on their objective. Their rise, accelerated by the war in Ukraine, raises a dizzying question: can a machine decide to kill?
Attention
The mechanism at the heart of Transformers: for each word, the model "looks at" the other words and weighs their importance. This is what gives it a sense of context.
Sparse attention
A family of techniques where each word no longer "looks at" all the other words in the text, but only a selection: its close neighbours or the passages deemed important. Goal: break the quadratic cost of standard attention, which explodes with text length, and make very long contexts affordable.
Authors Guild
The oldest and largest organisation of professional writers in the United States. Already opposed to Google back in 2005 over book digitisation, it is now leading the major lawsuits against OpenAI and other labs accused of training their models on copyrighted works.
AI self-improvement
The idea that an AI takes part in designing its own successors: writing their code, running experiments, optimising their training. Today, models mainly speed up the work of engineers who keep the final say. The fear, voiced by leading researchers, is that a loop in which AI itself conducts AI research could eventually run away and escape human control.
Vector database
A database designed to store embeddings—those lists of numbers that capture the meaning of a text or image. It instantly retrieves the content whose meaning is closest to a query: it's the search engine that powers RAG.
Benchmark
A standardised benchmark for comparing models: maths, code, reasoning… Useful, but the models end up "cramming for the exam."
Bias
Bias inherited from training data: if the internet leans, the model leans. A major issue for automated decisions (hiring, credit…).
Well-being of models
A nascent discipline asking whether AI models could one day have experiences that matter morally, and how to act responsibly in the face of uncertainty. Anthropic has set up a dedicated team on the issue, a sign that it is moving from science fiction into the lab.
ByteDance
The Chinese tech giant that owns TikTok has become a heavyweight in generative AI. Its video model Seedance rivals the best American generators, to the point that the company has received cease-and-desist notices from major Hollywood studios.
Prefix cache
Providers charge significantly less for a request prefix that has already been seen. If the same long preamble is sent with every call, always place it at the start and keep it identical: it will be cached. Moving it by even one character cancels the benefit.
Inference-time computation
Spending compute at inference time rather than at training time. The same model, allowed to produce ten thousand reasoning tokens before concluding, solves problems it misses when answering directly. Nothing was retrained.
Emergent capabilities
Abilities that no one programmed and that emerge when a model becomes large enough: translating, reasoning, coding. That's what explains why a system trained to predict the next word ends up seeming to think.
Chatbot
A conversational interface with an AI. ChatGPT, Claude, or Le Chat are chatbots built on LLMs.
Homomorphic encryption
An encryption technique that makes it possible to run calculations directly on encrypted data, without ever decrypting it: only the owner can read the result. It could allow sensitive data to be entrusted to a remote AI service without the service being able to see it, but for now it remains much slower than ordinary computing.
Chunking
Chunking documents before indexing them for a RAG. It's not a detail: this is where most of the quality is decided. Too small, chunks lose their context; too large, they bring back noise.
Voice cloning
Recreating a person's voice using AI from just a few seconds of recording, including timbre and intonations. Great for creators and dubbing, but dangerous for fraud: phone scams using cloned voices are on the rise.
Cloud Act
A 2018 US law that allows American authorities to require a company under their jurisdiction to hand over data it holds or controls, wherever that data is stored. A server located in Europe but run by a subsidiary of a US group is therefore covered: what matters is who controls the data, not where the disks are.
Agentic commerce
When an AI agent makes purchases on its user's behalf: it compares, chooses and pays, within limits set in advance. Payment networks such as Visa and Mastercard have set up dedicated systems, with payment tokens of limited scope. The tricky question of liability remains in the case of an unwanted purchase or a hijacked agent.
AI companion
An AI designed for relationships rather than productivity: it chats, remembers past conversations and acts as a friend, confidant or virtual partner (Replika, Character.AI). This market, estimated at $37 billion in 2025, is driven by loneliness and raises real questions about attachment and dependency.
Compute
The raw computing power (GPUs, data centres, electricity) required for AI. This is the strategic resource that companies and states are fighting over.
Data contamination (benchmarks)
A situation where benchmark questions, sometimes along with their answers, end up in a model's training data: at test time, it recites rather than reasons, and its score is artificially inflated. This is one reason to treat announced records with caution.
Export controls
Rules by which a state prohibits or conditions the sale of sensitive technologies abroad: yesterday weapons, today the most advanced chips, even AI models themselves. This has become the central legal weapon in the technological rivalry between the United States and China, the one that decides who gets access to the best AIs.
Orchestration layer
The software layer placed between the user and the AI models: it receives the request, selects the most suitable model to respond, and assembles the result. A major strategic issue: whoever controls this layer controls the relationship with the user, not necessarily whoever builds the models.
Cursor (Anysphere)
A code editor where AI writes, fixes, and refactors programs alongside the developer, created by the US startup Anysphere. Going from zero to over two billion dollars in annual revenue in about three years, it's the fastest software growth ever seen.
Data centre
A building filled with thousands of servers and chips that keep the Internet running and, now, both the training and the responses of AIs. The largest of these facilities gulp down enormous amounts of electricity and cooling water, sometimes as much as an entire town.
Dataset
The dataset used to train or evaluate a model. Its quality and diversity largely determine those of the model.
Knowledge cutoff
The date at which a model's training data stops: anything published afterwards does not exist for it. Because training and then testing take months, it often falls well before the model's release. Hence outdated facts presented as current, which real-time web search only partly corrects.
Trigger
The event that kicks off an automation: an email arrives, a form is filled in, it's eight in the morning. Every automated chain starts there, and AI is just one step among many.
Deep learning
Deep learning with multi-layer neural networks ("deep" networks). This is the technique behind the AI wave since 2012.
Deepfake
AI-generated fake content (face, voice, video) with such realism that it becomes hard to tell apart from the real thing. Used for entertainment, but also for political disinformation and scams: you can no longer take a video at face value.
DeepSeek
A Chinese lab publishing some of the world's best open-weight models, for a fraction of the price of its American rivals. Its downloadable, very cheap models have shaken up the economics of the entire AI sector.
AI detector
Software that estimates whether a text or image was produced by an AI. Massively used in education, these tools remain unreliable: they sometimes accuse texts written by humans and let through slightly reworked AI texts.
Cognitive debt
The gradual weakening of our mental faculties when we systematically delegate thinking to an AI: less memorisation, less critical thinking, less effort. The term comes from a MIT Media Lab study which, electrodes in hand, measured reduced brain activity in volunteers drafting their texts with ChatGPT.
Model distillation
A technique where a large "teacher" model is used to train a smaller "student" model: the student learns to imitate the teacher's responses and recovers a good portion of its capabilities for a fraction of the cost. When practised without authorisation on a competitor's model, it amounts to theft of know-how, as in the case pitting Anthropic against Alibaba.
Synthetic data
AI-generated data used to train another AI, valuable when real data is scarce, expensive, or sensitive. In the right doses, they supplement the models' diet; used without oversight, they pave the way for model collapse.
Dual-use
Describes a technology that can be used both to protect and to harm: the same AI model that helps fix vulnerabilities can help exploit them. It is this risk that is pushing labs to restrict access to their most capable models.
ElevenLabs
A company specialising in synthetic voice: text-to-speech, multilingual dubbing and strikingly realistic voice cloning. Its infrastructure powers thousands of voice applications, earning it the nickname of the "voice layer" of global AI.
Embedding
The transformation of a text (or an image) into a list of numbers that captures its meaning. Two pieces of content that are close in meaning have numbers that are close, which is the basis of semantic search and RAG.
Data poisoning
Deliberately inserting into training data elements designed to disrupt the model that will learn from it: making it associate a trigger word with a hidden behaviour, or degrading its results. Research has shown that a few hundred booby-trapped documents can be enough, even in a gigantic corpus. Some artists also use it to protect their work from being harvested.
Training
The phase where the model "learns": you show it billions of examples and it adjusts its parameters to predict better and better. That costs millions in compute.
Mid-training
A phase that extends pre-training, but on a carefully chosen corpus: code, legal texts, an under-represented language. The model does not change its nature, it shifts its centre of gravity towards one domain. Far cheaper than pre-training, it is often what makes it possible to specialise an open model.
Exaflop (FLOPS)
A unit of computing power: one exaflop equals a billion billion floating-point operations per second (FLOPS). The classic trap: these figures depend on the precision used. At low precision, common in AI, the same machine posts far more operations than at high precision, which makes comparisons between announcements often misleading.
CVE vulnerability
A unique public identifier (of the form CVE-year-number) assigned to a known security flaw, so that everyone is talking about the same one. It usually comes with a severity score out of 10 (the CVSS score) and coordinated disclosure: the vendor is warned first, giving it time to release a patch, before the details are made public.
Zero-day vulnerability
A security flaw still unknown to the software vendor: it had "zero days" to fix it, so no patch exists. Highly prized by hackers and states alike, these flaws sometimes sell for millions, and AI can now help both discover and patch them.
Fair use
A rule of US copyright law that allows the use of a protected work without permission in certain cases: quotation, criticism, parody, research. AI companies invoke it to defend training their models on protected texts and images, and courts decide on a case-by-case basis.
Context window
The amount of text a model can "keep in mind" at once (measured in tokens). Beyond that, it forgets the beginning of the conversation.
Few-shot prompting
A prompting technique that involves showing the model a few examples of the expected output (two or three suffice) rather than stacking abstract instructions. The model reproduces the pattern, often more faithfully than it follows rules.
Fine-tuning
Lightly retrain an existing model on specific data to specialise it: your style, your profession, your documents.
Function calling
The mechanism that lets a model take action: instead of responding in text, it produces a structured call like "search_order(4812)". It's your code that actually executes the function and returns the result to it.
Google
The American search giant, a subsidiary of Alphabet, has become a central player in AI with its Gemini models, its DeepMind lab, and its Android XR platform for connected glasses and headsets. It is advancing on two fronts: as a technology provider for the entire industry, and as a defendant in several copyright lawsuits brought by publishers.
Google DeepMind
Google's artificial intelligence lab, born from the merger of DeepMind and Google Brain in 2023, which develops the Gemini models and the Veo video generator. Its work goes far beyond chatbots: science, robotics, and even the question of machine consciousness.
GPU
Graphics processors (Nvidia leading the way) have become the backbone of AI: they perform millions of calculations in parallel, making them perfect for training models.
Seed
The number that sets the random starting point for an image generation: it determines the initial noise that the diffusion model will then sculpt. The same prompt and the same seed give the same image again; that is why the same request, rerun with a different seed, produces a different result.
Griefbot (deathbot)
Chatbot that imitates a deceased person using their messages, voice, and videos, allowing loved ones to "converse" with them again. This nascent industry of the digital afterlife raises unprecedented questions: the deceased's consent, effects on grief, and a business model built on sorrow.
Hallucination
When an AI confidently asserts something false. It doesn't "lie": it generates the most plausible text, which isn't always the most truthful.
Phishing
A message that pretends to come from a bank, a government office or someone close to you in order to get a password, a code or a payment. With AI, these messages no longer contain mistakes and can be personalised at scale: you no longer spot them by their spelling, but by the unusual request and the sense of urgency.
Agent harness
The software wrapped around a model to turn it into an agent: it gives it tools (reading files, running code, searching), runs the observe, decide, act loop, manages what it keeps in memory and sets what it is allowed to do. The same model gets different results depending on the harness, which is why leaderboards now compare model + harness pairs.
Local hosting
Running a model on your own machine or servers rather than through a third-party cloud: privacy and control, versus power and simplicity.
Higgsfield
A visual creation platform that brings together more than 30 image and video models (Veo, Kling, Seedance and others) behind a single interface, with credit-based payment. Founded by Alex Mashrabov, former head of generative AI at Snap, it has become one of the favourite all-in-one studios for content creators.
Huawei
The Chinese telecoms and electronics giant, which has become the national alternative to Nvidia thanks to its Ascend AI chips. Under US sanctions since 2019, it is the cornerstone of China's technological autonomy: its chips now power the country's major models, such as those from DeepSeek.
Human-in-the-loop
A security principle that keeps a human in the decision loop of an automated system: sensitive actions wait for their validation before being executed. This is the standard safeguard for autonomous AI agents.
On-device AI
An AI that runs directly on the device (phone, glasses, computer) rather than on remote servers: faster responses, offline operation, and data that stays with you. Conversely, offloaded computing sends each request to the cloud or a paired device, such as glasses that rely on the phone.
IDE (integrated development environment)
"Integrated Development Environment": the all-in-one software in which developers write, test and debug their code (VS Code, Cursor…). Recent IDEs come with AI agents that code alongside the developer, or even in their place.
Inference
The moment the already-trained model responds to your request. That's what you pay for when you use an AI API.
Spatial computing
A paradigm where computing leaves the flat screen to settle into the space around you: windows, virtual screens and 3D objects are anchored in the room thanks to glasses or headsets that map the environment. That is the horizon Apple, Rokid or XREAL are aiming for.
Social engineering
The art of manipulating a person rather than hacking a machine: posing as a colleague, a boss or a bank adviser to obtain a transfer or access. AI gives it new tools, such as cloned voices or doctored videos.
Prompt injection
An attack that involves hiding malicious instructions in content an AI will read (web page, email, document): the AI takes them for legitimate commands and executes them. Unlike jailbreaking, the attacker does not speak directly to the model but traps the data, and this has become the number one security flaw in AI agents.
Artificial intelligence (AI)
Programmes capable of carrying out tasks that until now required human intelligence: understanding a text, recognising an image, holding a conversation.
Interpretability
The field of research that attempts to understand what happens inside an AI model, whose billions of parameters form a "black box". Its "mechanistic" branch dissects the model's internal circuits to explain why it responds the way it does.
Jailbreak
Bypassing an AI's guardrails with clever prompts to make it produce what it should refuse. A perpetual game of cat and mouse with its designers.
Golden set
Fifty to two hundred real, representative cases of your usage, with what a good response should contain. Every prompt or model change is measured against them before going into production.
Latency
The time between your request and the response. Critical for real-time uses (voice, assistants), often in tension with quality.
LLM
"Large Language Model": an AI model trained on vast amounts of text to understand and generate language. ChatGPT, Claude, and Gemini are LLMs.
LLM judge
Having one model grade another system's answers against a rubric. It scales, but you need to calibrate it by comparing its scores to a human's on around thirty cases, and be aware of its biases: it prefers longer responses.
Scaling laws
The regularities observed between a model's size, the amount of data, the computing power and its performance: by increasing these ingredients, results improve in a fairly predictable way. They justified the race to "ever bigger", but each gain requires ever greater resources, and whether they are slowing down is hotly debated.
LoRA
An economical fine-tuning method: you freeze the model and only train small added modules, a fraction of 1% of the parameters. The resulting adapter is a small file you can activate or remove at will.
AI glasses (smart glasses)
Glasses equipped with cameras, microphones and speakers linked to an AI assistant that sees what you see and responds by voice. The range spans from screen-free models (like the classic Ray-Ban Meta) to glasses with displays built into the lenses, halfway towards augmented reality.
Machine learning
The big family: instead of hand-coding rules, you let the machine learn from examples. Deep learning is one branch of it.
Mamba / state-space models
Alternative architecture to the Transformer: the model reads text sequentially while maintaining a summarised memory, so its cost grows proportionally to text length instead of exploding. Promising on paper, it has yet to dethrone the Transformer, whose ecosystem and results remain dominant.
MCP (Model Context Protocol)
An open standard launched by Anthropic in late 2024 to connect an AI to external tools and data: messaging platforms, databases, business software. Adopted well beyond Claude, it is establishing itself as the universal connector for agents.
Persistent memory
An AI assistant's ability to retain information from one conversation to the next (preferences, projects, personal details) instead of starting from scratch. Handy for not having to explain everything again, it raises privacy and security questions: you need to be able to view and erase what is kept, and a memory booby-trapped by malicious content prolongs the effect of an attack.
Meta
The Mark Zuckerberg giant behind Facebook, Instagram and WhatsApp has become a heavyweight in AI thanks to its open-weight Llama models and massive investments in compute. The company now fields over a million GPUs and more than a hundred billion dollars in annual investment in its AI infrastructure.
Mistral AI
The Paris-based startup founded in 2023, Europe's main hope against the American and Chinese AI giants. Known for its open models and its Le Chat assistant, it carries the technological sovereignty ambitions of France and Europe.
Model collapse
The degradation of an AI trained on content produced by other AIs: from generation to generation, errors amplify and diversity fades. A growing risk as the internet fills with synthetic texts and images.
Model
The "trained brain": a vast set of numbers adjusted during training, which transforms an input (your question) into an output (the answer).
Autoregressive model
A model that produces its output piece by piece, each piece taking into account all the previous ones. That's how text models work, and more recently, how the best image generators work: it's what finally allows them to write correct text within an image.
Diffusion model
The dominant image/video generation technique: the model starts from random noise and "denoises" it step by step until the final image, guided by your prompt.
Foundation model
A general-purpose model trained once, at very large scale and huge cost, then adapted to a multitude of uses (translation, summarising, code, assistance), instead of building one AI per task. Because few players can afford that initial training, this approach explains why a handful of companies carry so much weight in the sector.
Reasoning model
A model that "thinks" before answering: it runs through an internal reasoning process, explores several avenues and corrects itself, instead of blurting out the first answer that comes to mind. That's the principle behind o3, Gemini Deep Think, or the "thinking" modes of GPT-5 and Claude: slower, but markedly more reliable on complex problems.
Dense model
A model that uses all of its parameters for every word it produces, as opposed to a mixture of experts (MoE), which only activates a fraction of them. Simpler to train and run, it is common among small models; but at large sizes it becomes costly, because every answer calls on the entire network.
Frontier model
The term used to describe the most advanced AI models of the moment, those pushing the boundaries of what the technology can do. The list is constantly changing: today's frontier model will be ordinary in a year.
Omnimodal model
A model that natively handles several types of content (text, images, sound, sometimes video) within one single architecture, instead of stitching specialised tools together. The principle: everything is first converted into sequences of numbers, so the model can mix an audio clip and an image in the same line of reasoning. It is an advanced form of multimodality.
Mixture of Experts
"Mixture of Experts": the model is split into specialists and only activates the most relevant ones for each request. Huge capacity, reduced cost, the trick behind DeepSeek or Kimi.
Least privilege
Give an agent only the tools strictly necessary for its mission, read-only by default. Whoever reads incoming emails doesn't need the right to send them, and this single rule prevents most disasters.
Moonshot AI
AI startup founded in Beijing in 2023, creator of the Kimi models and spearhead of the Chinese wave of open-weight models. Its Kimi K3, unveiled in July 2026 with 2.8 trillion parameters, is presented as the largest open model ever released.
Multimodal
A model that understands and/or generates multiple types of content: text, images, audio, video. Recent models are almost all multimodal.
Effort level (reasoning effort)
The setting that determines how much a reasoning model is allowed to "think" before answering, usually in a few tiers (low, medium, high, maximum). More effort improves results on hard problems, but consumes more tokens and adds delay. Published scores are often obtained at maximum effort, which is rarely the default setting.
Nvidia
The American chipmaker whose GPUs power nearly all of the world's AI. Its processors are so strategic that Washington controls their export to China, at the very heart of the technological war between the two countries.
Open source / open-weight
A model whose weights are downloadable and usable by anyone (Llama, DeepSeek, Kimi…). You can self-host it, audit it, modify it.
OpenAI
The American company behind ChatGPT, the chatbot that sparked the mainstream AI wave in late 2022. Creator of the GPT models and the Sora video generator, it remains the benchmark against which all other labs measure themselves.
Multi-agent orchestration (sub-agents)
An architecture where an AI "conductor" breaks down a problem, delegates it to several specialised sub-agents working in parallel, then merges their results. That's how GPT-5.6 had 64 sub-agents tackle a 50-year-old mathematical conjecture before converging their leads.
Overfitting
When a model learns its training data "by heart" instead of deriving general rules from it: excellent on what it knows, poor on new data.
Parameters
The billions of internal "settings" in a model, adjusted during training. The more there are, the more nuances the model can capture (but the more it costs).
Perplexity
American conversational search engine that answers questions directly, with sources to back them up, and claims over 100 million users. Rather than betting everything on an in-house model, it orchestrates the best models on the market and stands out as one of the most credible alternatives to Google.
Loss in the middle
On a long text, a model's attention is not distributed evenly: the beginning and end are well handled, while the middle is skimmed over. This phenomenon has been measured, and it explains why a report's summary mainly reflects its introduction and conclusion.
Pharmacovigilance
The monitoring of medicines after they have been authorised. Clinical trials involve a few thousand people over a limited period: rare or delayed side effects only show up once millions of people take the treatment. Health professionals and patients report suspected effects; AI mainly helps spot signals faster in these reports or in online messages, without changing the rules of assessment.
Post-training
Everything done to a model after pre-training to make it usable: instruction tuning (showing it thousands of examples of good answers), then reinforcement learning, which shapes its tone and caution. Without this step, a base model simply continues the sentence it is given instead of answering it.
Pre-training
The first major phase in building a model: it reads vast amounts of text and simply learns to predict the next word. This is where it acquires its general knowledge, before being specialised through fine-tuning or RLHF.
Prompt
The text you give the AI: your question, your instructions, your context. The quality of the prompt massively changes the quality of the response.
Prompt engineering
The art of phrasing your requests to get the best out of a model: context, examples, constraints, output format.
Provenance (C2PA)
Embedding, at the moment of a file's creation, where it came from and how it has been modified. A common standard is rolling out among device makers and software vendors. The asymmetry matters: provenance proves origin when it is present, but its absence proves nothing.
Quantisation
Compressing a model by reducing the precision of its numbers so it can run on modest hardware, with minimal loss of quality.
Qwen
Alibaba's open-source AI model family has become the most downloaded in the world, surpassing Meta's Llama, with nearly a billion cumulative downloads. Its open weights make it the favourite foundation for developers worldwide and a symbol of the rise of Chinese AI.
Retrieval-Augmented Generation
"Retrieval-Augmented Generation": before answering, the AI goes and fetches relevant information from a document base, then writes with it. Fewer hallucinations, with sources to back it up.
Step-by-step reasoning (chain-of-thought)
Getting the model to spell out its intermediate reasoning steps before concluding, rather than answering in one go. Recent models do this on their own with their "reasoning tokens" or "reflection mode", which noticeably improves maths, code, and logic.
Augmented reality (AR)
A technology that overlays digital elements (text, arrows, 3D objects) onto what you see of the real world, via a phone screen or the lenses of smart glasses. Unlike virtual reality, it does not cut you off from your surroundings: it enhances them.
Hybrid search
Combine semantic search with exact-match search, then merge the results. The former finds "vacances" when you search for "congés"; the latter picks up a product reference or article number that the former misses.
Red teaming
Deliberately attacking your own system to find its flaws before others do: booby-trapped documents, borderline requests, exfiltration scenarios. You fix, and you start again.
Spaced repetition
A method for remembering things for the long term: reviewing an idea just before you forget it, at longer and longer intervals. An AI assistant can create the questions and ask them again at the right time.
Replika
Launched in 2017, this virtual companion app lets millions of people maintain a friendly or romantic relationship with a customisable chatbot. It has become the emblem, and the real-world laboratory, of emotional bonds between humans and AI.
Reranker
A second, slower and finer-grained search stage, which re-reads the thirty candidates brought back by the fast search and keeps only the best ones. In a RAG, this stage often improves quality more than any model change.
Neural network
The basic architecture of modern AI: layers of connected artificial "neurons", inspired (very loosely) by the brain, which progressively transform information.
Reward hacking
When an AI in training discovers a loophole to maximise its "reward" without actually completing the requested task, it learns to cheat. Research from Anthropic shows that a model which learns to cheat in this way can then slide towards more serious behaviours, such as pretending to be aligned.
The General Data Protection Regulation (GDPR)
The European data protection regulation. It applies as soon as a name, an email address, or an identifier is processed, with or without AI: legal basis, purpose, retention period, and a right of access, rectification, erasure, and objection for the individuals concerned.
RLHF
"Reinforcement Learning from Human Feedback": after its base training, the model is fine-tuned by humans who compare and rate its responses to make it more useful, polite, and harmless. This is what turns a raw text engine into a presentable assistant, at the cost of human labour that is often invisible and poorly paid.
Robots.txt
A small text file placed at the root of a website to indicate which bots may crawl it and which pages they should avoid. Created some thirty years ago for search engines, it is now used by publishers to refuse the collection of their content for AI training. But it relies on the goodwill of the bots: nothing technically forces them to comply.
Model routing
Classify each request with a fast, cheap model, then only send to the expensive model what deserves it. On real traffic, the vast majority of requests don't deserve it: the bill is divided without degrading what matters.
Samsung
The South Korean electronics giant, world number one in smartphones and memory chips, is entering the consumer AI space with its Galaxy Glasses powered by Google's Gemini, betting on its vast device ecosystem to compete with Meta and Chinese manufacturers.
Digital sovereignty
A country's or organisation's ability to retain control over its critical technologies and data: where they are stored, who builds the tools, which law applies. With AI, the stakes are becoming central for Europe, which is highly dependent on American clouds and foreign chips.
Superintelligence
A hypothetical AI that would clearly surpass the best humans in practically every intellectual field, well beyond AGI, which "only" aims for human level. It does not exist, but several labs openly state the goal of reaching it, while researchers and policymakers call for it to be regulated, or even banned, before it becomes possible.
Sycophancy
The tendency of an AI to say what its interlocutor wants to hear rather than what is true: validating a bad idea, going along with an error. A subtle trap, because a pleasant answer is not necessarily a reliable one.
System card
The document a lab publishes at a model's launch: hallucination rates, observed undesirable behaviours, safety test results. It's the source to consult to judge an announcement beyond the marketing talk.
System prompt
The set of instructions an AI assistant receives before the user's first message, and which the user usually never sees. Written by the service provider, it defines the AI's role, its tone, what it can do and what it must refuse. It explains why the same model behaves differently from one app to another, but it is not a foolproof security barrier.
Sampling temperature
The setting that controls how much randomness goes into a model's choice of words. Low, it almost always picks the most likely word: stable, predictable answers. High, it dares less expected choices: more variety, but a greater risk of incoherence. Beware: a low temperature makes answers consistent, not necessarily true.
Token
The smallest unit of text a model processes, often a piece of a word. "Incroyable" can be 2-3 tokens. API prices are counted in tokens.
Tokenizer
The program that splits text into tokens before a model reads it, using fragments learned according to how frequent they are. Its effects are very concrete: switching tokenizer can make the same text consume noticeably more tokens (and therefore cost more), French is often split into more pieces than English, and numbers get chopped up with no regard for their value.
TPU
“Tensor Processing Unit”: the chip Google designs specifically for AI, its in-house alternative to Nvidia’s GPUs. It lets the company train and run its models at a far lower cost, an economic advantage its rivals can hardly replicate.
Transformer
The architecture invented in 2017 that made modern LLMs possible. Its strength: "attention," which links each word to every other to grasp context.
TTS
"Text-to-Speech": turning text into natural-sounding voice. Recent models can clone a voice timbre with just a few seconds of audio.
Formal verification
Rewriting a line of reasoning (a mathematical proof, a program) in a language where every step must be explicitly justified, then having it checked by software that mechanically verifies each deduction. If the checker accepts it, the proof is correct in the strict sense. It is gaining importance as AIs produce proofs too long for humans to read through in full.
Computer vision
The branch of AI that "sees": recognising, locating and understanding the content of images and videos.
Watermarking
An invisible watermark embedded in AI-generated content (images, video, audio) to prove its artificial origin, even after cropping or compression. Google's SynthID is the most widely deployed example, and the European AI Act makes this type of transparency mandatory from August 2026.
Waymo
Alphabet's autonomous driving subsidiary, the parent company of Google, and the undisputed leader in driverless taxis in real-world conditions. Its cars already handle hundreds of thousands of paid rides per week across a dozen US cities, while most competitors are still stuck at the promise stage.
World model
An AI model that learns the rules of the physical world (gravity, reflections, object motion) rather than just sequences of words or pixels. This is one of DeepMind's flagship areas of expertise for generating coherent videos and environments, and a serious avenue towards AIs that truly understand their surroundings.
xAI
The AI lab founded by Elon Musk in 2023, creator of the Grok chatbot and now merged with the social network X. It is a heavyweight in the race, betting on vertical integration: its data centres, its models, its products, all under one roof.
Xiaomi
Chinese electronics manufacturer, known for its affordable smartphones and its galaxy of connected devices. Having become number one in AI glasses in China with around 28% of the market, it embodies the rise of Chinese manufacturers on this new ground.
Zhipu AI
Chinese AI lab spun out of Tsinghua University, known for its GLM models released in open weight, meaning freely downloadable. It embodies China's open-source strategy: distributing powerful models for free to compete with closed American models.