Open-source AI concepts
Plain definitions of the terms that come up in open-source AI projects, each linked to the projects in our catalogue that implement them.
What is AI coding agent?An AI coding agent is a program that uses a large language model to inspect a codebase, edit files and run commands toward a goal you state in natural language. The term also appears as coding agent or agentic coding.What is an AI agent?An AI agent is a program that uses a large language model (LLM) to decide on actions, call tools, observe the results and repeat until a task is finished. This page explains the mechanism, when it is worth using, and how the term appears in open-source projects.What is Computer use?Computer use is the practice of letting an AI model drive a real browser or desktop session, clicking, typing and reading the screen on your behalf. A browser agent applies the same idea to web pages only, usually through a browser automation layer.What is Context window?A context window, also called context length, is the maximum number of tokens a language model can attend to in one request, covering the system prompt, conversation history, tool output and the reply. Tokens beyond that limit are truncated or the request fails, so the window sets a hard ceiling on how much an agent can see at once.What is Embeddings?Embeddings are dense numeric vectors that represent text, images or other content so that similar meanings land close together in vector space. A text embedding model turns a piece of content into such a vector, which is why search, clustering and recommendation systems depend on them.What is GGUF?GGUF is a binary file format that stores a machine learning model's weights, metadata and tokenizer information in one file so inference software can load it without a separate configuration step. It is the format used by llama.cpp and by many other local inference tools.What is Hallucination in LLMs?Hallucination (also called AI hallucination) is when a large language model produces fluent, confident output that is not grounded in its training data, the prompt, or any real source. The model is not lying; it is filling gaps with plausible text, because its objective is to predict the next token, not to check facts.What is KV cache?A KV cache (key-value cache) stores the key and value tensors a transformer has already computed for earlier tokens, so each new token is generated without recomputing the whole prefix. It trades GPU memory for speed, and it is the reason long conversations and long documents get expensive to serve.What is LLM evaluation?LLM evaluation (evals) is the practice of measuring how a language model or an LLM-based system performs on defined tasks, using datasets, metrics and scoring rules. It turns vague impressions about output quality into numbers and pass/fail results you can compare over time.What is LLM gateway?An LLM gateway (also called an LLM proxy or model router) is a service that sits between your application and one or more model providers, giving you a single endpoint for routing, credentials, retries and usage tracking. Instead of calling OpenAI, Anthropic or a self-hosted model directly, your code calls the gateway.What is LLM inference server?An LLM inference server is a process that loads a language model into memory and answers generation requests over an API, usually HTTP. It sits between an application and the model weights, handling batching, memory and scheduling so callers do not have to.What is Local LLM?A Local LLM is a large language model whose weights and inference run on hardware you own, such as a laptop, desktop or home server, instead of a provider's API. Running LLMs locally means prompts, documents and outputs stay on that machine, and the model keeps working when the network does not.What is LoRA? Low-Rank Adaptation ExplainedLoRA (low-rank adaptation) is a fine-tuning method that freezes a model's original weights and trains a small pair of low-rank matrices alongside them, so adapting a large model costs a fraction of full fine-tuning. QLoRA combines LoRA with quantized base weights to cut memory use further.What is Model Context Protocol?Model Context Protocol (MCP) is an open specification that standardises how an AI application connects to external tools, data sources and prompts. It replaces one-off integrations with a single client-server interface.What is Model quantization?Model quantization is the process of storing a model's weights and sometimes its activations in lower-precision numbers, such as 8-bit or 4-bit integers instead of 16-bit floats. The aim is to cut memory use and speed up inference, usually at some cost to accuracy.What is Retrieval-augmented generation?Retrieval-augmented generation (RAG) is a pattern in which a language model's answer is built from documents fetched at question time instead of from the model's weights alone. The model writes the answer, but a retrieval step decides what it gets to read.What is Structured output?Structured output (often called JSON mode) is a way of making a language model return data in a fixed, machine-readable shape instead of free prose. The model is constrained by a schema or format so the response can be parsed by code without guessing.What is Tokenizer?A tokenizer is the component that splits raw text or audio into the small units (tokens) a model can process, mapping each unit to an integer ID. Tokenization is the step that turns human-readable input into the sequence a neural network actually consumes.What is Vector database?A vector database (also called a vector store) stores data as high-dimensional vectors and returns the nearest neighbours to a query vector, usually by approximate search. It is the retrieval layer behind semantic search, recommendation and retrieval-augmented generation.